The phenomenon of high coincidence rate and paper plagiarism were illustrated from the author and content, combined with the academic research for professional titles evaluation materials in universities. The mysteries of coincidence rate were unraveled, which provided a reference to avoid high content coincidence rate for the reader.【结合高校职称评审材料学术检索工作,从作者和内容两方面说明了论文高内容重合率的种种现象以及判定为抄袭的各种情况,为读者解开内容重合率神秘的面纱,并为读者在论文写作中避免高内容重合率提供借鉴作用。】
A new model for plagiarism-identification of scientific papers based on sentence similarity is presented.Large-scale texts are quickly detected with Local Word-Frequency Fingerprint(LWFF) to find suspected plagiarism ones.Sentence similarity is computed according to the Longest Sorted Common Subsequence(LSCS) between source texts and destination texts.The algorithm can mark plagiarism details,and show evidence. 【提出一种基于句子相似度的论文抄袭检测模型。利用局部词频指纹算法对大规模文档进行快速检测,找出疑似抄袭文档。根据最长有序公共子序列算法计算句子间的相似度,并标注抄袭细节,给出抄袭依据。在标准中文数据集SOGOU-T上进行的实验表明,该模型具有较强的局部信息挖掘能力,在一定程度上克服了现有的论文抄袭检测算法精度不高的缺点。】