인지야공

인지야공/ADP/6번째 글

ADP 실기 — 분류 모델 여덟 개와 평가

기출을 보면 “알고리즘 3개를 제시하고 그중 2개를 선택하고 이유를 설명”이 반복된다 (21회, 24회). 그래서 모델을 깊게 아는 것보다 여러 개를 같은 틀로 빠르게 돌리고 비교하는 것이 실기에서는 낫다. 여기 있는 코드는 전부 같은 뼈대다.

데이터 로드 → 분할 → 모델 생성 → fit → predict → 평가 → GridSearch → 시각화

공통 뼈대

from sklearn.model_selection import train_test_split, GridSearchCV, cross_val_score
from sklearn.metrics import accuracy_score, confusion_matrix, classification_report

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.3, random_state=42, stratify=y)

clf = 모델()
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)

print("Accuracy:", accuracy_score(y_test, y_pred))
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred))

classification_report 하나면 정밀도·재현율·F1 이 클래스별로 다 나온다. 불균형 데이터에서 정확도만 쓰면 안 된다 — 88 대 12 인 데이터는 전부 다수 클래스로 찍어도 정확도가 88% 다. 23회가 정확히 그런 데이터였다.

의사결정나무

from sklearn.tree import DecisionTreeClassifier, plot_tree

clf = DecisionTreeClassifier(max_depth=3, random_state=42).fit(X_train, y_train)

plt.figure(figsize=(12, 8))
plot_tree(clf, filled=True, rounded=True,
          feature_names=iris.feature_names, class_names=iris.target_names)
plt.show()

# 가지치기 — 깊이를 줄여 과적합을 막는다
clf_pruned = DecisionTreeClassifier(max_depth=2, random_state=42).fit(X_train, y_train)

plot_tree 는 해석 가능성을 그림 한 장으로 보여 주는 유일한 모델이라서, 답안에 “현업에서 설명이 필요하면 이 모델”이라고 쓸 근거가 된다. 가지치기 전후 정확도를 나란히 찍어 과적합 이야기를 붙인다.

로지스틱 회귀

from sklearn.linear_model import LogisticRegression

lr = LogisticRegression().fit(X_train, y_train)
print("Accuracy: {:.2f}%".format(accuracy_score(y_test, lr.predict(X_test)) * 100))

결정경계를 그리는 코드는 어느 모델에나 똑같이 쓰인다. 특징 2개짜리로 줄여 놓고 그린다.

x_min, x_max = X[:, 0].min()-0.5, X[:, 0].max()+0.5
y_min, y_max = X[:, 1].min()-0.5, X[:, 1].max()+0.5
xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.01), np.arange(y_min, y_max, 0.01))

Z = lr.predict(np.c_[xx.ravel(), yy.ravel()]).reshape(xx.shape)

plt.contourf(xx, yy, Z, cmap=plt.cm.Set1, alpha=0.5)
plt.scatter(X[:, 0], X[:, 1], c=y, cmap=plt.cm.Set1, edgecolor='k')
plt.show()

np.c_[xx.ravel(), yy.ravel()] 이 격자를 2열짜리 입력으로 펴는 부분이고, .reshape(xx.shape) 이 예측을 다시 격자로 접는 부분이다. 이 두 줄만 외우면 된다.

그래디언트 부스팅과 GridSearchCV

from sklearn.ensemble import GradientBoostingClassifier

clf = GradientBoostingClassifier(n_estimators=100, learning_rate=0.1,
                                 max_depth=3, random_state=42).fit(X_train, y_train)

parameters = {
    'n_estimators': [50, 100, 200],
    'learning_rate': [0.05, 0.1, 0.2],
    'max_depth': [2, 3, 4],
}
grid = GridSearchCV(clf, parameters, cv=5).fit(X_train, y_train)
print("Best parameters:", grid.best_params_)
print("Best score:", grid.best_score_)

변수 중요도 시각화는 문제에 대놓고 나온다(20회 — “컬럼별 중요도 시각화”).

def plot_feature_importance(importance, names):
    idx = np.argsort(importance)
    plt.barh(range(len(importance)), [importance[i] for i in idx], align='center')
    plt.yticks(range(len(importance)), [names[i] for i in idx])
    plt.xlabel('Feature Importance')
    plt.show()

plot_feature_importance(clf.feature_importances_, iris.feature_names)

np.argsort 로 오름차순 정렬해서 barh 로 그리면 중요한 변수가 위에 온다. 정렬을 안 하면 읽을 수 없는 그림이 된다.

SVM — SVC 와 SVR

from sklearn.svm import SVC

param_grid = {'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10], 'kernel': ['linear', 'rbf']}
grid = GridSearchCV(SVC(), param_grid, cv=5).fit(X_train, y_train)

best = grid.best_estimator_
print("Best Parameters:", grid.best_params_)
print("Train score:", best.score(X_train, y_train))
print("Test score:",  best.score(X_test, y_test))

회귀는 SVR 이다. 채점 지표에 주의한다.

from sklearn.svm import SVR

svr = SVR()
scores = cross_val_score(svr, X, y, cv=5, scoring='neg_mean_squared_error')

grid = GridSearchCV(svr, param_grid={'C': [0.1, 1, 10, 100],
                                     'gamma': [0.001, 0.01, 0.1, 1],
                                     'epsilon': [0.01, 0.1, 1]},
                    cv=5, scoring='neg_mean_squared_error', n_jobs=-1).fit(X, y)

scoring 에 'mean_squared_error' 는 없다. 'neg_mean_squared_error' 다. sklearn 은 “점수는 클수록 좋다”로 통일해서, 오차 지표에는 전부 음수를 붙인다. 그래서 best_score_ 가 음수로 나오고, 그건 정상이다.

scoring 에 넣을 수 있는 것들 — accuracy, f1, precision, recall, roc_auc, average_precision, r2, neg_mean_squared_error, neg_log_loss.

SVM 은 스케일링이 필수다. 거리를 쓰기 때문에 단위가 큰 변수가 커널을 지배한다.

나이브 베이즈

from sklearn.naive_bayes import GaussianNB

nb = GaussianNB()
scores = cross_val_score(nb, X_train, y_train, cv=5, scoring='accuracy')
print("Mean CV score: {:.2f}".format(scores.mean()))

grid = GridSearchCV(nb, {'var_smoothing': [1e-5, 1e-4, 1e-3, 0.01, 0.1, 1, 10, 100]},
                    cv=5, scoring='accuracy', n_jobs=-1).fit(X_train, y_train)

빠르다. “속도 측면의 모델 하나, 정확도 측면의 모델 하나를 선정하라”(22회·23회)는 문제에서 속도 쪽 답으로 쓰기 좋다. 변수 간 독립을 가정한다는 한계를 같이 쓴다.

KNN

from sklearn.neighbors import KNeighborsClassifier

knn = KNeighborsClassifier(n_neighbors=5).fit(X_train, y_train)
print("Train: {:.2f}".format(knn.score(X_train, y_train)))
print("Test:  {:.2f}".format(knn.score(X_test, y_test)))

학습이 없고 예측이 느린 모델이다(게으른 학습). 스케일링이 필수고, 차원이 높으면 무너진다. kk 가 작으면 과적합, 크면 과소적합이라 kk 를 바꿔 가며 train/test 점수를 같이 그리면 설명이 쉬워진다.

파이프라인 — 전처리까지 묶는다

전처리를 따로 하면 test 에 실수로 fit 하는 사고가 난다. 파이프라인은 그걸 구조적으로 막는다.

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder

numeric_features = ['Age', 'Fare']
numeric_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='median')),
    ('scaler', StandardScaler()),
])

categorical_features = ['Sex', 'Embarked']
categorical_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='most_frequent')),
    ('onehot', OneHotEncoder(handle_unknown='ignore')),
])

preprocessor = ColumnTransformer(transformers=[
    ('num', numeric_transformer, numeric_features),
    ('cat', categorical_transformer, categorical_features),
])

OneHotEncoder(handle_unknown='ignore') 가 중요하다. train 에 없던 범주가 test 에 나오면 기본값은 에러다. 시험 중에 여기서 멈추면 손해가 크다.

보팅 앙상블

19회 — “분류 분석 3개를 앙상블하여 예측하고 result.csv 로 제출”. 파이프라인에 얹는다.

from sklearn.ensemble import (BaggingClassifier, GradientBoostingClassifier,
                              RandomForestClassifier, VotingClassifier)

base_estimators = [
    ('logistic_bagging', BaggingClassifier(base_estimator=LogisticRegression(),
                                           n_estimators=50, random_state=42)),
    ('logistic_boosting', GradientBoostingClassifier(n_estimators=50, max_depth=4,
                                                     learning_rate=0.1, random_state=42)),
    ('svm', SVC(kernel='rbf', C=1.0, probability=True, random_state=42)),
    ('RandomForest', RandomForestClassifier(random_state=42)),
]

pipeline = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('ensemble', VotingClassifier(estimators=base_estimators, voting='soft')),
])
pipeline.fit(X_train, y_train)

소프트 보팅을 쓰려면 모든 모델이 predict_proba 를 가져야 한다. SVC 는 기본적으로 확률을 안 내므로 probability=True 를 준다. 이거 하나 빠뜨려서 학습 다 끝난 뒤에 에러를 본 적이 있다.

여덟 개를 한꺼번에 태우는 것도 된다.

from lightgbm import LGBMClassifier
from xgboost import XGBClassifier
from sklearn.neural_network import MLPClassifier

estimators = [('rf', RandomForestClassifier()), ('lgbm', LGBMClassifier()),
              ('xgb', XGBClassifier()), ('ann', MLPClassifier()),
              ('svm', SVC(probability=True)), ('nb', GaussianNB()),
              ('lr', LogisticRegression()), ('knn', KNeighborsClassifier())]

soft = VotingClassifier(estimators=estimators, voting='soft').fit(X_train, y_train)
hard = VotingClassifier(estimators=estimators, voting='hard').fit(X_train, y_train)

print("Soft:", accuracy_score(y_test, soft.predict(X_test)))
print("Hard:", accuracy_score(y_test, hard.predict(X_test)))
방식
하드 보팅각 모델의 예측 클래스를 다수결
소프트 보팅각 모델의 확률을 평균내고 큰 쪽. 대개 더 좋다

평가

혼동행렬에서 손으로 지표 뽑기

cm = confusion_matrix(actual, predicted)

accuracy  = np.trace(cm) / np.sum(cm)
precision = cm[1, 1] / np.sum(cm[:, 1])     # 예측이 1인 것 중 실제 1
recall    = cm[1, 1] / np.sum(cm[1, :])     # 실제 1인 것 중 맞힌 것
f1        = 2 * precision * recall / (precision + recall)

sns.heatmap(cm, annot=True, cmap='Blues')
plt.xlabel('Predicted'); plt.ylabel('Actual')
plt.show()

행이 실제, 열이 예측이다. 정밀도는 열로 나누고 재현율은 행으로 나눈다. 이걸 뒤집는 게 가장 흔한 실수다. 지표 계산식을 직접 쓰라는 문제(20회 — 정확도 공식이 주어짐)도 나오므로 classification_report 만 믿지 말고 손으로도 낼 수 있어야 한다.

지표 선택의 근거는 이렇게 쓴다 — 놓치면 큰일 나는 문제(질병·이탈·사기)는 재현율, 잘못 잡으면 곤란한 문제는 정밀도, 둘 다면 F1, 불균형이면 ROC-AUC 또는 PR-AUC.

ROC 와 AUC

from sklearn.metrics import roc_curve, roc_auc_score

y_score = clf.predict_proba(X_test)[:, 1]        # 확률이어야 한다. 예측 라벨이 아니다
fpr, tpr, thresholds = roc_curve(y_test, y_score)
auc = roc_auc_score(y_test, y_score)

plt.plot(fpr, tpr, label=f"AUROC = {auc:.3f}")
plt.plot([0, 1], [0, 1], linestyle='--', label='Random Guessing')
plt.xlabel('False Positive Rate'); plt.ylabel('True Positive Rate')
plt.legend(); plt.show()

xx 축이 위양성률, yy 축이 민감도다. 대각선이 찍기와 같고, 왼쪽 위로 붙을수록 좋다. AUC 는 0.5 가 무작위, 1 이 완벽이다.

두 모델의 차이가 유의한가 — McNemar

정확도가 0.83 과 0.85 로 나왔을 때 “이게 진짜 차이인가”에 답하는 검정이다.

from statsmodels.stats.contingency_tables import mcnemar

# 두 모델의 정오를 교차한 2×2 표
b = np.sum((model1_pred == y_true) & (model2_pred != y_true))
c = np.sum((model1_pred != y_true) & (model2_pred == y_true))
table = [[0, b], [c, 0]]

print(mcnemar(table, exact=False, correction=True))

χ2=(∣b−c∣−1)2b+c\chi^2 = \frac{(\lvert b - c \rvert - 1)^2}{b + c}

예전에 이 검정을 직접 짜 두었는데, 전체 오분류 개수를 넣는 방식이라 틀린 코드였다. McNemar 는 불일치 쌍만 본다 — 두 모델이 똑같이 맞힌 것과 똑같이 틀린 것은 계산에 들어가지 않는다. 그게 이 검정의 전부다. 직접 짜지 말고 statsmodels 를 쓴다.

신경망

ADP 실기에 딥러닝이 정면으로 나온 적은 드물지만, 준비는 해 뒀다. 시험 환경의 tensorflow 는 1.13.1 이다.

from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Dropout

model = Sequential()
model.add(Dense(units=16, activation='relu', input_dim=30))
model.add(Dropout(0.2))
model.add(Dense(units=8, activation='relu'))
model.add(Dropout(0.2))
model.add(Dense(units=1, activation='sigmoid'))

model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
history = model.fit(X_train, y_train, batch_size=32, epochs=100,
                    validation_data=(X_test, y_test))
문제출력층손실함수
이진 분류Dense(1, activation='sigmoid')binary_crossentropy
다중 분류 (원-핫)Dense(k, activation='softmax')categorical_crossentropy
다중 분류 (정수 라벨)Dense(k, activation='softmax')sparse_categorical_crossentropy
회귀Dense(1) — 활성화 없음mse

학습 곡선을 그리는 것이 사실상 필수다. 훈련 정확도만 오르고 검증 정확도가 멈추면 과적합이고, 그 문장 하나가 “모델의 한계와 개선 방안”에 그대로 답이 된다.

plt.plot(history.history['accuracy'], label='train')
plt.plot(history.history['val_accuracy'], label='val')
plt.plot(history.history['loss'], label='loss')
plt.plot(history.history['val_loss'], label='val_loss')
plt.legend(); plt.show()

CNN 은 구조만 바뀐다.

from keras.layers import Conv2D, MaxPooling2D, Flatten, Dense
from keras.callbacks import ModelCheckpoint

model = Sequential()
model.add(Conv2D(32, (3, 3), padding='same', activation='relu', input_shape=X_train.shape[1:]))
model.add(Conv2D(32, (3, 3), activation='relu'))
model.add(MaxPooling2D(pool_size=(2, 2)))
model.add(Conv2D(64, (3, 3), padding='same', activation='relu'))
model.add(Conv2D(64, (3, 3), activation='relu'))
model.add(MaxPooling2D(pool_size=(2, 2)))
model.add(Flatten())
model.add(Dense(512, activation='relu'))
model.add(Dense(num_classes, activation='softmax'))

checkpoint = ModelCheckpoint('cnn_model.h5', monitor='val_accuracy',
                             verbose=1, save_best_only=True, mode='max')
history = model.fit(X_train, y_train, batch_size=128, epochs=50,
                    validation_data=(X_test, y_test), callbacks=[checkpoint])
model.load_weights('cnn_model.h5')

ModelCheckpoint 로 검증 정확도가 가장 높았던 시점의 가중치를 저장해 두고 마지막에 불러온다. 마지막 에폭이 최선인 경우는 별로 없다.

이미지는 /255 로 스케일링하고, Flatten 뒤에 Dense 를 붙인다. 이 두 가지가 구조의 전부다.

표시는 이 브라우저에만 남는다. 서버로 가는 것은 없다.