← 논문 노트 목록
AI

CatBoost: unbiased boosting with categorical features

Liudmila Prokhorenkova(제1저자), Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, Andrey Gulin

NeurIPS 2018 (arXiv:1706.09516) (2018) · DOI: 10.48550/arXiv.1706.09516 · 넣은 날 2026-08-05

1. 서지정보

Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., & Gulin, A. (2018). CatBoost: unbiased boosting with categorical features. Advances in Neural Information Processing Systems (NeurIPS), 31. arXiv:1706.09516.

2. 연구문제

기존 그래디언트 부스팅 구현들은 prediction shift(예측 이동)라는 공통 통계적 결함을 안고 있다 — 훈련 과정에서 모델이 훈련 데이터의 타깃(정답)에 은근히 의존하게 되어, 훈련 시 예측 분포와 실제(테스트) 예측 분포가 어긋난다. 범주형 변수를 수치로 바꾸는 표준 전처리(target statistics)도 같은 종류의 **타깃 누출(target leakage)**을 일으킨다. 이 두 문제를 하나의 원리로 함께 해결할 수 있는가?

3. 방법

4. 데이터

공개 데이터셋 9개로 실험했다(전체는 4/5를 학습·튜닝, 1/5을 테스트에 사용).

데이터셋샘플 수피처 수설명
Adult48,842151994년 인구조사 데이터 — 연소득 5만 달러 초과 여부 예측
Amazon32,76910Kaggle Amazon Employee Access 챌린지
Click Prediction399,482122012 KDD Cup 파생, 클릭=0을 1%로 서브샘플링(5:1 비율)
Epsilon400,0002,000PASCAL Challenge 2008
KDD Appetency50,000231KDD 2009 Cup 축소판
KDD Churn50,000231KDD 2009 Cup 축소판
KDD Internet10,10869다중클래스를 이진(P/N)으로 변환
KDD Upselling50,000231KDD 2009 Cup 축소판
Kick Prediction72,98336Kaggle “Don’t Get Kicked!” 챌린지

5. 결과

6. 한계

저자가 밝힌 한계: 논문 맨 아래에 “Preprint. Work in progress.”라고 스스로 명시했다(arXiv 버전 기준) — 저자 자신도 이 버전이 완성된 최종본이 아니라고 인정하는 셈이다.

읽으면서 느낀 한계:

7. 원문 인용

Two critical algorithmic advances introduced in CatBoost are the implementation of ordered boosting, a permutation-driven alternative to the classic algorithm, and an innovative algorithm for processing categorical features. — Abstract

Both techniques were created to fight a prediction shift caused by a special kind of target leakage present in all currently existing implementations of gradient boosting algorithms. — Abstract

관련 개념

[[gradient-boosting]] · [[target-leakage]]

(아직 만들어진 개념 페이지는 없다. SCHEMA.md 규칙대로 “나중에 쓸 것” 표시로 남겨둔다. 이 논문은 오늘 넣은 TEM 계열 논문들과 주제적으로 안 겹쳐 TEM 개념과는 연결하지 않았다.)