인지야공/ADP/9번째 글
ADP 실기 — 텍스트 마이닝, 한글이 오히려 안전하다
18회에 “텍스트 마이닝 — 문자열 전처리, 워드클라우드”가 나왔다. 준비하면서 알게 된 건
시험 환경에 konlpy 는 있고 nltk 와 wordcloud 는 목록에 없다는 것이었다.
그래서 영어 문서가 나와도 한글 도구 쪽 흐름으로 푸는 편이 안전했다.
한글 전처리
from konlpy.tag import Okt
import re
okt = Okt()
# 1. 한글·영문만 남기고 전부 지운다
def remove_special_chars(sentence):
sentence = re.sub('[^가-힣ㄱ-ㅎㅏ-ㅣa-zA-Z]', ' ', sentence)
sentence = re.sub(r'\s+', ' ', sentence)
return sentence.strip()
# 2. 형태소 분석 후 의미 있는 품사만
def tokenize(sentence):
tokens = okt.pos(sentence)
return [t[0] for t in tokens if t[1] in ['Noun', 'Verb', 'Adjective']]
# 3. 불용어 제거
def remove_stopwords(tokens, stop_words):
return [t for t in tokens if t not in stop_words]
[^가-힣ㄱ-ㅎㅏ-ㅣa-zA-Z] 이 정규식의 전부다. 가-힣 이 완성형 음절이고, ㄱ-ㅎ 과
ㅏ-ㅣ 는 자모만 남은 경우(ㅋㅋ, ㅠㅠ)까지 살린다. 숫자를 살려야 하면 0-9 를 넣는다.
품사 태거는 셋 다 쓸 수 있다.
| 특징 | |
|---|---|
| Okt | pos, nouns, morphs. 구어체·신조어에 무난하다 |
| Komoran | 오탈자에 강하다. 문서 분석에 쓴다 |
| Mecab | 가장 빠르다. 문서가 크면 이쪽 |
한글 불용어 사전은 없다. 직접 만든다.
stopwords = ['제', '을', '월', '일', '조', '수', '로', '은', '는']
tokens = [t for t in tokens if t not in stopwords and len(t) > 1]
len(t) > 1 한 줄이 조사와 한 글자 잡음을 대부분 걷어 낸다. 불용어 리스트를 아무리 잘 짜도
이 한 줄만 못하다.
영어 전처리
nltk 가 있다면 이 흐름이다.
import nltk
nltk.download('punkt'); nltk.download('stopwords'); nltk.download('wordnet')
from nltk.tokenize import sent_tokenize, word_tokenize
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer, WordNetLemmatizer
sentences = sent_tokenize(text) # 문장 토큰화
words = word_tokenize(text) # 단어 토큰화
stop_words = set(stopwords.words('english'))
filtered = [w for w in words if w.casefold() not in stop_words]
alpha = [w for w in filtered if re.match('^[a-zA-Z0-9]+', w)]
stemmed = [PorterStemmer().stem(w) for w in alpha] # 어간 추출
lemmatized = [WordNetLemmatizer().lemmatize(w) for w in alpha] # 원형 복원
lowercase = [w.lower() for w in alpha]
nltk.download 는 인터넷을 쓴다. 시험장에서는 이 줄부터 막힌다. 그래서 대비책을
따로 준비했다.
corpus = re.sub(r'[^\w\s]', '', text) # 문장부호
corpus = re.sub(r'\d+', '', corpus) # 숫자
tokens = corpus.lower().split()
stop_words = {'the', 'a', 'an', 'and', 'or', 'of', 'to', 'in', 'is', 'are',
'be', 'for', 'on', 'that', 'this', 'it', 'as', 'by', 'with'}
tokens = [t for t in tokens if t not in stop_words and len(t) > 2]
어간 추출과 원형 복원의 차이는 개념으로 물을 수 있다 — 스테밍은 규칙으로 어미를 잘라
사전에 없는 단어가 나올 수 있고(studies → studi), 표제어 추출은 사전을 참조해
실제 단어를 낸다(studies → study). 정확한 대신 느리다.
TF-IDF
from konlpy.corpus import kolaw
from konlpy.tag import Okt
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
corpus = kolaw.open('constitution.txt').read()
corpus = [' '.join(Okt().nouns(corpus))] # 명사만 남겨 공백으로 이어 붙인다
cv = CountVectorizer(max_features=1000)
tdm = cv.fit_transform(corpus)
tfidf = TfidfVectorizer(max_features=1000)
tfidf_matrix = tfidf.fit_transform(corpus)
tdm_df = pd.DataFrame(tdm.A, columns=cv.get_feature_names())
tfidf_df = pd.DataFrame(tfidf_matrix.A, columns=tfidf.get_feature_names())
sklearn 0.23.2 는 get_feature_names() 다. 최신 버전의 get_feature_names_out() 을 치면
없는 메서드라고 나온다. 반대 방향의 실수라 헷갈린다.
konlpy.corpus.kolaw 는 대한민국 헌법 전문이 들어 있는 내장 코퍼스다. 인터넷 없이
연습할 수 있어서 준비할 때 이걸로 계속 돌렸다.
여러 문서에 두루 나오는 단어는 IDF 가 작아져 점수가 깎인다. “그 문서에만 유난히 많이
나오는 단어”를 찾는 것이 TF-IDF 다. 문서가 하나뿐이면 IDF 가 모두 같아져서 그냥 빈도와
다를 게 없어진다 — 위 코드처럼 corpus 가 리스트 원소 하나라면 그 점을 알고 써야 한다.
상위 단어를 뽑아 그린다.
top_tfidf = tfidf_df.sum().sort_values(ascending=False)[:20]
plt.figure(figsize=(8, 6))
plt.barh(top_tfidf.index, top_tfidf.values)
plt.title('Top 20 words by TF-IDF')
plt.xlabel('TF-IDF score'); plt.ylabel('Word')
plt.show()
워드클라우드 — 없을 때를 대비한다
준비할 때 쓴 코드는 이것이다.
from wordcloud import WordCloud
wordcloud = WordCloud(font_path='C:/Windows/Fonts/malgun.ttf',
background_color='white',
max_font_size=100,
width=800, height=800).generate_from_frequencies(word_count)
plt.figure(figsize=(8, 8))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis("off")
plt.show()
한글이면 font_path 를 반드시 준다. 안 주면 전부 네모로 나온다.
그런데 wordcloud 는 시험 환경 패키지 목록에 없었다. import 가 실패하면 빈도 상위
막대그래프로 대체하고, 답안에 “라이브러리 제약으로 막대그래프로 대체했다”고 한 줄 쓴다.
빈도를 보여 준다는 목적은 같다.
from collections import Counter
word_count = Counter(tokens)
top = pd.Series(dict(word_count.most_common(20)))
top.sort_values().plot(kind='barh', figsize=(8, 6))
plt.title('상위 20개 단어 빈도')
plt.show()
빈도를 셀 때 Counter 를 쓰는 이유가 하나 더 있다. 예전 메모에 이런 코드가 있었다.
tokens = list(set(tokens)) # 중복 제거
...
word_count[word] = tokens.count(word)
중복을 제거한 뒤에 빈도를 세면 전부 1 이 나온다. 어휘 목록을 만들 때와 빈도를 셀 때 쓰는 리스트는 달라야 한다. 정리하다 발견하고 고친 부분이다.
Word2Vec 과 단어 네트워크
from gensim.models import Word2Vec
w2v = Word2Vec(window=5, min_count=1, workers=4, sg=1)
w2v.build_vocab([tokens])
w2v.train([tokens], total_examples=w2v.corpus_count, epochs=w2v.epochs)
print(w2v.wv.similarity('국민', '정부'))
print(w2v.wv.similarity('국민', '국회'))
| 파라미터 | |
|---|---|
window | 앞뒤 몇 단어를 문맥으로 볼지 |
min_count | 이보다 적게 나온 단어는 버린다 |
sg | 1 이면 Skip-gram, 0 이면 CBOW |
Skip-gram 은 중심 단어로 주변을 맞히고, CBOW 는 주변으로 중심을 맞힌다. 데이터가 적으면 Skip-gram 쪽이 낫다.
유사도가 임계값을 넘는 쌍만 남기면 단어 네트워크가 된다.
import networkx as nx
from networkx.algorithms.community import greedy_modularity_communities
edges = []
for i, w1 in enumerate(vocab):
for j, w2 in enumerate(vocab):
if i >= j:
continue
sim = w2v.wv.similarity(w1, w2)
if sim >= 0.35:
edges.append((w1, w2, sim))
G = nx.Graph()
G.add_weighted_edges_from(edges)
communities = greedy_modularity_communities(G)
pos = nx.spring_layout(G, seed=42)
nx.draw(G, pos=pos, with_labels=True, font_size=10, font_family=font_name)
plt.show()
i >= j: continue 가 같은 쌍을 두 번 세지 않기 위한 것이다. 이게 없으면 간선이 두 배가
되고 자기 자신과의 유사도(1.0)까지 들어간다.
이 이중 루프는 단어 수의 제곱으로 늘어난다. 어휘가 2000개면 200만 번이라 한참 걸린다. 상위 빈도 단어 200개 정도로 자르고 시작한다.
nx.draw 에 font_family=font_name 을 안 주면 노드 라벨의 한글이 깨진다. 그래프 축 폰트를
설정해도 이건 별도다.
군집별로 무엇이 뭉쳤는지 보기
커뮤니티를 찾았으면 각 커뮤니티에 어떤 단어가 있는지를 보여 줘야 답이 된다.
for i, community in enumerate(communities):
counts = {w: word_count[w] for w in community}
plt.figure(figsize=[5, 5])
plt.pie(counts.values(), labels=counts.keys(), autopct='%1.1f%%', startangle=90)
plt.title(f'Community {i+1}')
plt.show()
커뮤니티 안에서 가중치 합이 큰 상위 노드에만 라벨을 다는 것이 읽기에 낫다. 전부 달면 글자가 겹쳐 아무것도 안 보인다.
node_weights = {n: sum(w for _, _, w in subgraph.edges(n, data='weight'))
for n in subgraph.nodes()}
top_nodes = [n for n, _ in sorted(node_weights.items(), key=lambda x: x[1], reverse=True)[:10]]
nx.draw_networkx_labels(subgraph, pos=pos,
labels={n: n for n in top_nodes},
font_size=10, font_family=font_name)
정리
텍스트 문제가 나오면 이 순서로 간다.
- 정규식으로 잡음 제거 → 형태소 분석 → 불용어와 한 글자 제거
- 빈도 상위 단어를 표와 그림으로 (워드클라우드가 안 되면 막대그래프)
- TF-IDF 로 문서 특징어 추출
- 필요하면 Word2Vec → 유사도 → 네트워크 → 커뮤니티
1과 2까지만 해도 배점의 상당 부분이 나온다. 3과 4는 시간이 남을 때다.