인지야공

인지야공/ADP/9번째 글

ADP 실기 — 텍스트 마이닝, 한글이 오히려 안전하다

18회에 “텍스트 마이닝 — 문자열 전처리, 워드클라우드”가 나왔다. 준비하면서 알게 된 건 시험 환경에 konlpy 는 있고 nltk 와 wordcloud 는 목록에 없다는 것이었다. 그래서 영어 문서가 나와도 한글 도구 쪽 흐름으로 푸는 편이 안전했다.

한글 전처리

from konlpy.tag import Okt
import re

okt = Okt()

# 1. 한글·영문만 남기고 전부 지운다
def remove_special_chars(sentence):
    sentence = re.sub('[^가-힣ㄱ-ㅎㅏ-ㅣa-zA-Z]', ' ', sentence)
    sentence = re.sub(r'\s+', ' ', sentence)
    return sentence.strip()

# 2. 형태소 분석 후 의미 있는 품사만
def tokenize(sentence):
    tokens = okt.pos(sentence)
    return [t[0] for t in tokens if t[1] in ['Noun', 'Verb', 'Adjective']]

# 3. 불용어 제거
def remove_stopwords(tokens, stop_words):
    return [t for t in tokens if t not in stop_words]

[^가-힣ㄱ-ㅎㅏ-ㅣa-zA-Z] 이 정규식의 전부다. 가-힣 이 완성형 음절이고, ㄱ-ㅎ 과 ㅏ-ㅣ 는 자모만 남은 경우(ㅋㅋ, ㅠㅠ)까지 살린다. 숫자를 살려야 하면 0-9 를 넣는다.

품사 태거는 셋 다 쓸 수 있다.

특징
Oktpos, nouns, morphs. 구어체·신조어에 무난하다
Komoran오탈자에 강하다. 문서 분석에 쓴다
Mecab가장 빠르다. 문서가 크면 이쪽

한글 불용어 사전은 없다. 직접 만든다.

stopwords = ['제', '을', '월', '일', '조', '수', '로', '은', '는']
tokens = [t for t in tokens if t not in stopwords and len(t) > 1]

len(t) > 1 한 줄이 조사와 한 글자 잡음을 대부분 걷어 낸다. 불용어 리스트를 아무리 잘 짜도 이 한 줄만 못하다.

영어 전처리

nltk 가 있다면 이 흐름이다.

import nltk
nltk.download('punkt'); nltk.download('stopwords'); nltk.download('wordnet')

from nltk.tokenize import sent_tokenize, word_tokenize
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer, WordNetLemmatizer

sentences = sent_tokenize(text)                     # 문장 토큰화
words     = word_tokenize(text)                     # 단어 토큰화

stop_words = set(stopwords.words('english'))
filtered = [w for w in words if w.casefold() not in stop_words]
alpha    = [w for w in filtered if re.match('^[a-zA-Z0-9]+', w)]

stemmed    = [PorterStemmer().stem(w) for w in alpha]          # 어간 추출
lemmatized = [WordNetLemmatizer().lemmatize(w) for w in alpha] # 원형 복원
lowercase  = [w.lower() for w in alpha]

nltk.download 는 인터넷을 쓴다. 시험장에서는 이 줄부터 막힌다. 그래서 대비책을 따로 준비했다.

corpus = re.sub(r'[^\w\s]', '', text)      # 문장부호
corpus = re.sub(r'\d+', '', corpus)        # 숫자
tokens = corpus.lower().split()

stop_words = {'the', 'a', 'an', 'and', 'or', 'of', 'to', 'in', 'is', 'are',
              'be', 'for', 'on', 'that', 'this', 'it', 'as', 'by', 'with'}
tokens = [t for t in tokens if t not in stop_words and len(t) > 2]

어간 추출과 원형 복원의 차이는 개념으로 물을 수 있다 — 스테밍은 규칙으로 어미를 잘라 사전에 없는 단어가 나올 수 있고(studies → studi), 표제어 추출은 사전을 참조해 실제 단어를 낸다(studies → study). 정확한 대신 느리다.

TF-IDF

from konlpy.corpus import kolaw
from konlpy.tag import Okt
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer

corpus = kolaw.open('constitution.txt').read()
corpus = [' '.join(Okt().nouns(corpus))]        # 명사만 남겨 공백으로 이어 붙인다

cv = CountVectorizer(max_features=1000)
tdm = cv.fit_transform(corpus)

tfidf = TfidfVectorizer(max_features=1000)
tfidf_matrix = tfidf.fit_transform(corpus)

tdm_df   = pd.DataFrame(tdm.A, columns=cv.get_feature_names())
tfidf_df = pd.DataFrame(tfidf_matrix.A, columns=tfidf.get_feature_names())

sklearn 0.23.2 는 get_feature_names() 다. 최신 버전의 get_feature_names_out() 을 치면 없는 메서드라고 나온다. 반대 방향의 실수라 헷갈린다.

konlpy.corpus.kolaw 는 대한민국 헌법 전문이 들어 있는 내장 코퍼스다. 인터넷 없이 연습할 수 있어서 준비할 때 이걸로 계속 돌렸다.

TF-IDF(t,d)=TF(t,d)×log⁡NDF(t)\text{TF-IDF}(t, d) = \text{TF}(t, d) \times \log\frac{N}{\text{DF}(t)}

여러 문서에 두루 나오는 단어는 IDF 가 작아져 점수가 깎인다. “그 문서에만 유난히 많이 나오는 단어”를 찾는 것이 TF-IDF 다. 문서가 하나뿐이면 IDF 가 모두 같아져서 그냥 빈도와 다를 게 없어진다 — 위 코드처럼 corpus 가 리스트 원소 하나라면 그 점을 알고 써야 한다.

상위 단어를 뽑아 그린다.

top_tfidf = tfidf_df.sum().sort_values(ascending=False)[:20]

plt.figure(figsize=(8, 6))
plt.barh(top_tfidf.index, top_tfidf.values)
plt.title('Top 20 words by TF-IDF')
plt.xlabel('TF-IDF score'); plt.ylabel('Word')
plt.show()

워드클라우드 — 없을 때를 대비한다

준비할 때 쓴 코드는 이것이다.

from wordcloud import WordCloud

wordcloud = WordCloud(font_path='C:/Windows/Fonts/malgun.ttf',
                      background_color='white',
                      max_font_size=100,
                      width=800, height=800).generate_from_frequencies(word_count)

plt.figure(figsize=(8, 8))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis("off")
plt.show()

한글이면 font_path 를 반드시 준다. 안 주면 전부 네모로 나온다.

그런데 wordcloud 는 시험 환경 패키지 목록에 없었다. import 가 실패하면 빈도 상위 막대그래프로 대체하고, 답안에 “라이브러리 제약으로 막대그래프로 대체했다”고 한 줄 쓴다. 빈도를 보여 준다는 목적은 같다.

from collections import Counter

word_count = Counter(tokens)
top = pd.Series(dict(word_count.most_common(20)))
top.sort_values().plot(kind='barh', figsize=(8, 6))
plt.title('상위 20개 단어 빈도')
plt.show()

빈도를 셀 때 Counter 를 쓰는 이유가 하나 더 있다. 예전 메모에 이런 코드가 있었다.

tokens = list(set(tokens))          # 중복 제거
...
word_count[word] = tokens.count(word)

중복을 제거한 뒤에 빈도를 세면 전부 1 이 나온다. 어휘 목록을 만들 때와 빈도를 셀 때 쓰는 리스트는 달라야 한다. 정리하다 발견하고 고친 부분이다.

Word2Vec 과 단어 네트워크

from gensim.models import Word2Vec

w2v = Word2Vec(window=5, min_count=1, workers=4, sg=1)
w2v.build_vocab([tokens])
w2v.train([tokens], total_examples=w2v.corpus_count, epochs=w2v.epochs)

print(w2v.wv.similarity('국민', '정부'))
print(w2v.wv.similarity('국민', '국회'))
파라미터
window앞뒤 몇 단어를 문맥으로 볼지
min_count이보다 적게 나온 단어는 버린다
sg1 이면 Skip-gram, 0 이면 CBOW

Skip-gram 은 중심 단어로 주변을 맞히고, CBOW 는 주변으로 중심을 맞힌다. 데이터가 적으면 Skip-gram 쪽이 낫다.

유사도가 임계값을 넘는 쌍만 남기면 단어 네트워크가 된다.

import networkx as nx
from networkx.algorithms.community import greedy_modularity_communities

edges = []
for i, w1 in enumerate(vocab):
    for j, w2 in enumerate(vocab):
        if i >= j:
            continue
        sim = w2v.wv.similarity(w1, w2)
        if sim >= 0.35:
            edges.append((w1, w2, sim))

G = nx.Graph()
G.add_weighted_edges_from(edges)

communities = greedy_modularity_communities(G)

pos = nx.spring_layout(G, seed=42)
nx.draw(G, pos=pos, with_labels=True, font_size=10, font_family=font_name)
plt.show()

i >= j: continue 가 같은 쌍을 두 번 세지 않기 위한 것이다. 이게 없으면 간선이 두 배가 되고 자기 자신과의 유사도(1.0)까지 들어간다.

이 이중 루프는 단어 수의 제곱으로 늘어난다. 어휘가 2000개면 200만 번이라 한참 걸린다. 상위 빈도 단어 200개 정도로 자르고 시작한다.

nx.draw 에 font_family=font_name 을 안 주면 노드 라벨의 한글이 깨진다. 그래프 축 폰트를 설정해도 이건 별도다.

군집별로 무엇이 뭉쳤는지 보기

커뮤니티를 찾았으면 각 커뮤니티에 어떤 단어가 있는지를 보여 줘야 답이 된다.

for i, community in enumerate(communities):
    counts = {w: word_count[w] for w in community}
    plt.figure(figsize=[5, 5])
    plt.pie(counts.values(), labels=counts.keys(), autopct='%1.1f%%', startangle=90)
    plt.title(f'Community {i+1}')
    plt.show()

커뮤니티 안에서 가중치 합이 큰 상위 노드에만 라벨을 다는 것이 읽기에 낫다. 전부 달면 글자가 겹쳐 아무것도 안 보인다.

node_weights = {n: sum(w for _, _, w in subgraph.edges(n, data='weight'))
                for n in subgraph.nodes()}
top_nodes = [n for n, _ in sorted(node_weights.items(), key=lambda x: x[1], reverse=True)[:10]]

nx.draw_networkx_labels(subgraph, pos=pos,
                        labels={n: n for n in top_nodes},
                        font_size=10, font_family=font_name)

정리

텍스트 문제가 나오면 이 순서로 간다.

  1. 정규식으로 잡음 제거 → 형태소 분석 → 불용어와 한 글자 제거
  2. 빈도 상위 단어를 표와 그림으로 (워드클라우드가 안 되면 막대그래프)
  3. TF-IDF 로 문서 특징어 추출
  4. 필요하면 Word2Vec → 유사도 → 네트워크 → 커뮤니티

1과 2까지만 해도 배점의 상당 부분이 나온다. 3과 4는 시간이 남을 때다.

표시는 이 브라우저에만 남는다. 서버로 가는 것은 없다.