Thinking :如何进行文本抄袭自动检测:
预测文章风格是否和自己一致 => 分类算法根据模型预测的结果来对全量文本进行比对,如果数量很大,=> 可以先聚类降维,比如将全部文档自动聚成 k=25 类文本特征提取 => 计算TF-IDFTopN相似 => TF-IDF相似度矩阵中TopN文档编辑距离editdistance => 计算句子或文章之间的编辑距离
一、自然语言的分词与关键词处理
在数据分析时,不免接触到自然语言(string),要想对其进行分析,需要以下几个步骤:
(一)对句子进行分词
涉及到 jieba,NLTK的使用
1. import jieba : 主要对中文词句进行分类
import jieba
s
= str()
jieba
.cut
(s
)
jieba
.cut
(s
, cut_all
= True)
jieba
.lcut
(s
)
jieba
.lcut
(s
, cut_all
= True)
jieba
.cut_for_search
(s
)
jieba
.lcut_for_search
(s
)
jieba
.add_word
(w
)
jieba
.del_word
(w
)
import jieba
import jieba
.posseg
as pseg
sentence
= '美国新冠肺炎确诊超57万,纽约州死亡人数过万。当地时间4月13日,世卫组织发布最新一期新冠肺炎每日疫情报告。截至欧洲中部时间4月13日10时(北京时间4月13日16时),全球确诊新冠肺炎1773084例(新增76498例),死亡111652例(新增5702例)。其中,疫情最为严重的欧洲区域已确诊913349例(新增33243例),死亡77419例(新增3183例)。'
words2
= jieba
.cut
(sentence
)
print(' '.join
(words2
))
words
= pseg
.lcut
(sentence
)
temp
= [(word
,flag
) for word
,flag
in words
]
temp
= [(word
,flag
) for word
,flag
in words
if flag
== 'ns']
print(temp
)
2.import NLTK
from nltk
.tokenize
import word_tokenize
import pandas
as pd
data
= pd
.read_csv
('./movies.csv')
sentence
= ' '.join
(data
.title
) + ' '.join
(data
.genres
)
cut_word
= word_tokenize
(sentence
)
cut_text
= ' '.join
(cut_word
)
cut_text
from wordcloud
import WordCloud
import matplotlib
.pyplot
as plt
wc
= WordCloud
(
max_words
= 100,
width
= 2000,
height
= 1200,
stopwords
= ['The', 'Movie'],
collocations
= False,
)
wordcloud
= wc
.generate
(cut_text
)
wordcloud
.to_file
("wordcloud.jpg")
plt
.imshow
(wordcloud
)
plt
.axis
("off")
plt
.show
()
NLTK 使用时如何解决缺少分词数据问题
(二) TF-IDF
当句子中的单词依据词性被分好后,若想比较两个句子的相似度,需要将单词转化为数字形式,因此使用 TF-IDF ,来衡量一个句子与另一个句子的词向量相似度。
T
F
(
t
e
r
m
f
r
e
q
u
e
n
c
y
)
TF(term frequency)
TF(termfrequency):词频,用以衡量一个单词在文档中的重要性。
T
F
=
单
词
出
现
次
数
全
语
料
中
单
词
总
数
TF = \frac{单词出现次数}{全语料中单词总数}
TF=全语料中单词总数单词出现次数
I
D
F
(
i
n
v
e
r
s
e
d
o
c
u
m
e
n
t
f
r
e
q
u
e
n
c
y
)
IDF(inverse document frequency)
IDF(inversedocumentfrequency):逆向文档转换率,用于评价单词的区分度。
I
D
F
=
总
文
档
数
出
现
文
档
数
+
1
IDF = \frac{总文档数}{出现文档数+1}
IDF=出现文档数+1总文档数 工具:
S
k
l
e
a
r
n
,
G
e
n
s
i
m
Sklearn, Gensim
Sklearn,Gensim
TfidfTransformer() 会默认对得到的TF-IDF矩阵进行标准化处理
from sklearn
.feature_extraction
.text
import CountVectorizer
,TfidfTransformer
,TfidfVectorizer
import numpy
as np
corpus
= [
'我 非常 喜欢 看 电视剧',
'我 非常 喜欢 旅行' ,
'我 非常 喜欢 吃 苹果' ,
'我 非常 喜欢 跑步' ,
'王者荣耀 KPL 春季赛 开战啦'
]
cv
= CountVectorizer
(stop_words
= None, token_pattern
='(?u)\\b\\w\\w*\\b')
count
= cv
.fit_transform
(corpus
)
print(cv
.get_feature_names
())
print(count
.toarray
())
print(np
.sum(count
,0))
transformer
= TfidfTransformer
()
tfidf_matrix
= transformer
.fit_transform
(count
)
print(tfidf_matrix
.toarray
())
tfidf_vec
= TfidfVectorizer
()
tfidf_matrix
= tfidf_vec
.fit_transform
(corpus
)
print(tfidf_matrix
.toarray
())
(三)计算相似度
T
F
−
I
D
F
TF-IDF
TF−IDF 获得的是一个
T
F
−
I
D
F
TF-IDF
TF−IDF 向量矩阵,因此使用余弦相似度进行计算相似度
from sklearn
.metrics
.pairwise
import cosine_similarity
print(cosine_similarity
(tfidf_matrix
,tfidf_matrix
))
这样就可以得到所有文章的相似度。
二、项目流程
(一)数据加载与预处理
首先加载停用词,将 content 中的停用词去掉将空值去掉。将换行符等去掉使用 jieba 进行分词,得到后续可使用的语料
import numpy
as np
import pandas
as pd
import jieba
import re
with open('chinese_stopwords.txt', 'r', encoding
= 'utf-8') as file:
stopwords
= [i
[:-1] for i
in file.readlines
()]
sentences
= pd
.read_csv
('sqlResult.csv', encoding
= 'gb18030')
sources
= sentences
.source
contents
= sentences
.content
news
= pd
.concat
([sources
, contents
], ignore_index
= False, axis
= 1)
news
= news
.dropna
(subset
= ['content'])
def split_text(content
):
content
= content
.replace
(' ', '')
content
= content
.replace
('\r', '')
content
= content
.replace
('\n', '')
words
= jieba
.lcut
(content
.strip
())
content
= ' '.join
([word
for word
in words
if word
not in stopwords
])
return content
corpus
= list(map(split_text
, [str(content
) for content
in news
.content
]))
corpus
[:3]
(二)获取 TF-IDF 词向量,并对数据集进行分类预测
使用TF-IDF对语料提取特征语料词向量化
from sklearn
.feature_extraction
.text
import CountVectorizer
, TfidfTransformer
, TfidfVectorizer
cv
= CountVectorizer
(encoding
= 'gb18030', min_df
= 0.015, stop_words
= None)
word_count
= cv
.fit_transform
(corpus
)
transformer
= TfidfTransformer
()
tfidf_matrix
= transformer
.fit_transform
(word_count
)
print(tfidf_matrix
.shape
)
设定 Label,确定是否为新华社发表的文章。使用多项式朴素贝叶斯对数据集进行拟合,得到总体准确率与新华社文章的预测准确率,理论上新华社准确率越高,总体准确率越小越能找到可疑文章。对语料进行预测,得到预测的 tag ,若 tag 为新华社,label为非新华社,则为抄袭文章
labels
= list(map(lambda source
: 1 if '新华' in str(source
) else 0, news
.source
))
from sklearn
.model_selection
import train_test_split
from sklearn
.naive_bayes
import MultinomialNB
, BernoulliNB
from sklearn
.metrics
import roc_auc_score
as AUC
train_x
, test_x
, train_y
, test_y
= train_test_split
(tfidf_matrix
.toarray
(), labels
, test_size
= 0.3, random_state
= 33)
model
= MultinomialNB
(alpha
=1.0, fit_prior
=True, class_prior
=None)
model
.fit
(train_x
, train_y
)
pre_y
= model
.predict
(test_x
)
print("AUC准确率为:", AUC
(test_y
, pre_y
))
tag
= model
.predict
(tfidf_matrix
.toarray
())
accr
= list(map(lambda x
,y
: 1 if (x
== 1 and y
== 1) else 0, labels
, tag
)).count
(1) / labels
.count
(1)
print('预测新华社准确率为:', accr
)
此准确率仍可通过其他方法提升,略
copy_label
= []
copy_label
= list(map(lambda x
,y
: 1 if (x
!= 1 and y
== 1) else 0, labels
, tag
))
print('可疑文章数:',copy_label
.count
(1))
(三)使用可疑文章列表,查找可能抄袭的新华社文章
由于新华社文章较多,占75%,因此使用可疑文章列表(2828个)为基础,搜索可能抄袭的新华社文章 由于文章众多,先进行聚类分析
from sklearn
.cluster
import KMeans
from sklearn
import preprocessing
from sklearn
.preprocessing
import Normalizer
normalizer
= Normalizer
()
scaled_array
= normalizer
.fit_transform
(tfidf_matrix
.toarray
())
kmeans
= KMeans
(n_clusters
= 25, random_state
= 42, n_jobs
= -1)
k_lables
= kmeans
.fit_predict
(scaled_array
)
通过 id:class;class: 新华id 的方式,构建两个字典,作为后续再查找的基础
id_class
= {index
: class_
for index
, class_
in enumerate(k_lables
)}
compare
= pd
.DataFrame
({'labels':labels
, 'tag': tag
})
copy_index
= compare
[(compare
.labels
== 0) & (compare
.tag
== 1)].index
xinhua_index
= compare
[compare
.labels
== 1].index
class_id
= dict()
for i
in range(25):
class_id
[i
] = [key
for key
, value
in id_class
.items
() if value
== i
]
xinhua_class_id
= dict()
for i
in range(25):
xinhua_class_id
[i
] = [value
for value
in class_id
[i
] if value
in xinhua_index
.tolist
()]
from sklearn
.metrics
.pairwise
import cosine_similarity
def find_copy_source_top(copy_index
, top
):
copy_id_list
= copy_index
.tolist
()
copy_news_similar
= dict()
copy_top
= dict()
for news_id
in copy_id_list
:
class_list1
= class_id
[id_class
[news_id
]].copy
()
class_list1
.remove
(news_id
)
copy_news_similar
[news_id
] = {i
:cosine_similarity
(tfidf_matrix
[news_id
], tfidf_matrix
[i
]).tolist
() for i
in class_list1
}
copy_top
[news_id
] = sorted(copy_news_similar
[news_id
].items
(), key
= lambda x
: x
[1], reverse
= True)[:top
]
return copy_top
copy_news_list
= find_copy_source_top
(copy_index
[1010:1015], 5)
(四)查找copy_index列表与新华社相似率大于similar的文章
def find_xinhua_copy_source_top(copy_index
, similar
, top
):
copy_id_list
= copy_index
.tolist
()
from collections
import defaultdict
copy_news_similar
= defaultdict
()
copy_top
= dict()
for news_id
in copy_id_list
:
class_list1
= xinhua_class_id
[id_class
[news_id
]].copy
()
copy_news_similar
[news_id
] = {i
: cosine_similarity
(tfidf_matrix
[news_id
], tfidf_matrix
[i
]).tolist
() for i
in class_list1
if cosine_similarity
(tfidf_matrix
[news_id
], tfidf_matrix
[i
]) > similar
}
copy_top
[news_id
] = sorted(copy_news_similar
[news_id
].items
(), key
= lambda x
: x
[1][0], reverse
= True)[:5]
return copy_top
copy_xinhua_news_list
= find_xinhua_copy_source_top
(copy_index
[1010:1015], 0.85, 5)
copy_xinhua_news_list
(五)查看抄袭句
import editdistance
def print_similar(copy_xinhua_news_list
, copy_news_id
):
for xinhua_id
in copy_xinhua_news_list
[copy_news_id
]:
print('-------------------------------------------------------文章{}-------------------------------------------------'.format(copy_news_id
))
print('怀疑文章抄袭,原文为:\n', news
.iloc
[copy_news_id
,0],':\n', news
.iloc
[copy_news_id
,1])
print('新华社文章为:\n', news
.iloc
[xinhua_id
[0], 0], news
.iloc
[xinhua_id
[0],1])
print('相似度为:\n', xinhua_id
[1])
print('编辑距离:\n',editdistance
.eval(corpus
[copy_news_id
], corpus
[xinhua_id
[0]]))
find_similar_sentence
(news
.iloc
[copy_news_id
].content
, news
.iloc
[xinhua_id
[0]].content
)
print('--------------------------------------------------------------------------------------------------------------')
def find_similar_sentence(candidate
, raw
):
similist
= []
cl
= candidate
.strip
().split
('。')
ra
= raw
.strip
().split
('。')
for c
in cl
:
for r
in ra
:
similist
.append
([c
,r
,editdistance
.eval(c
,r
)])
sort
=sorted(similist
,key
=lambda x
:x
[2])[:5]
for c
,r
,ed
in sort
:
if c
!='' and r
!='':
print('怀疑抄袭句:{0}\n相似原句:{1}\n 编辑距离:{2}\n'.format(c
,r
,ed
))
print_similar
(copy_xinhua_news_list
, 3352)