← 논문 노트 목록
AI

The performance of a deep learning system in assisting junior ophthalmologists in diagnosing 13 major fundus diseases: a prospective multi-center clinical trial

Bing Li(제1저자), Huan Chen, Weihong Yu, Ming Zhang, Fang Lu, Jingxue Ma, Yuhua Hao, Xiaorong Li, Bojie Hu, Lijun Shen, Jianbo Mao, Xixi He, Hao Wang, Dayong Ding, Xirong Li, Youxin Chen(교신저자)

npj Digital Medicine (2024) · DOI: 10.1038/s41746-023-00991-9 · 넣은 날 2026-08-09

1. 서지정보

Li, B., Chen, H., Yu, W., Zhang, M., Lu, F., Ma, J., Hao, Y., Li, X., Hu, B., Shen, L., Mao, J., He, X., Wang, H., Ding, D., Li, X., & Chen, Y. (2024). The performance of a deep learning system in assisting junior ophthalmologists in diagnosing 13 major fundus diseases: a prospective multi-center clinical trial. npj Digital Medicine, 7, 8. https://doi.org/10.1038/s41746-023-00991-9

2. 연구문제

안저(眼底, fundus) 사진 한 장으로 13종의 주요 안저질환(당뇨망막병증·망막정맥폐쇄·망막동맥폐쇄·병적근시·망막박리·망막색소변성·황반변성 등)을 자동으로 찾아내는 딥러닝 시스템(DLS, deep learning system)이, 실제 임상 현장에서 경력이 짧은 안과 전공의(junior ophthalmologist)의 진단을 얼마나 도와줄 수 있는가를 전향적(prospective, 결과를 미리 정해두지 않고 앞으로 데이터를 모아가며 검증하는 방식) 다기관 임상시험으로 검증하는 것이 목적이다. 특히 “AI 단독 진단”이 아니라 “AI가 옆에서 제안을 주고 사람이 최종 판단하는” 협업 모델의 효과를 확인하려 했다.

3. 방법

중국 베이징협화의원(Peking Union Medical College Hospital) 등 5개 3차병원이 참여한 전향적 자기대조(self-controlled) 임상시험이다(ClinicalTrials.gov 등록번호 NCT04723160). 안저 사진 1장씩을 세 그룹으로 나눠 비교했다.

정답지(“표준 주석”, standard annotation)는 5년 이상 경력의 망막 전문의 5명이 각자 라벨을 붙이고, 3명 이상이 일치하지 않으면 6번째 중재 전문의(arbitrator)가 포함된 전체 패널 토의로 확정했다.

DLS 자체는 CNN(합성곱신경망) 계열 모델인 seResNext50을 기본 구조로 쓰고, 이미지 화질을 먼저 판정하는 화질평가모델(ResNet-34 기반 회귀모델)과 질병을 판정하는 진단모델 두 단계로 구성했다. 진단모델은 seResNext50 세 개를 병렬로 학습시켜 늦은 융합(late fusion)으로 예측을 안정화했다.

평가지표로는 이 논문이 새로 제안한 “진단 일치도(diagnostic consistency)“를 1차 평가지표로 썼다. 기존의 정확도(accuracy, 라벨 전부가 정답과 일치해야 인정)와 달리, 진단 일치도는 여러 개 라벨 중 표준 진단과 일부라도 겹치면 일치로 인정하는 지표다(한 이미지에 최대 3개 질병까지 라벨을 붙일 수 있었기 때문). 통계 검정은 Cochran–Mantel–Haenszel(CMH-χ2) 검정을 썼고 P<0.05를 유의수준으로 삼았다.

4. 데이터

2020년 8월2021년 1월, 5개 병원에서 외래 환자 750명을 전향적으로 모집했고 748명이 전 과정을 완료했다. 표준 주석 과정에서 화질 불량 3장이 제외되어 최종 1493장의 안저 사진이 분석 대상이 되었다. 평균 연령 51.7±14.7세(1875세), 남성 43.3%, 전원 중국 한족(Han) 환자다. 정상 안저 477장(32.0%), 질병 있는 안저 1016장(68.1%, 그중 92.8%는 질병 1개만 라벨). DLS 모델 자체의 학습 데이터는 별도로 안저 사진 81,395장(학습 77,181 / 검증 1,087 / 테스트 3,127)이 쓰였고, 화질평가모델은 31,082장(학습 25,082 / 검증·테스트 각 3,000)으로 학습했다.

5. 결과

6. 한계

저자가 밝힌 한계:

읽으면서 느낀 한계:

7. 원문 인용

The diagnostic consistency was 84.9% (95%CI, 83.0%~86.9%), 72.9% (95%CI, 70.3%~75.6%) and 85.5% (95%CI, 83.5%~87.4%) in the test group, control group and DLS group, respectively. — Abstract

With the help of the proposed DLS, the diagnostic consistency of junior ophthalmologists improved by approximately 12% (95% CI, 9.1%~14.9%) with statistical significance (P < 0.001). — Abstract

For the detection of 13 diseases, the test group achieved significant higher sensitivities (72.2%~100.0%) and comparable specificities (90.8%~98.7%) comparing with the control group (sensitivities, 50%~100%; specificities 96.7%~99.8%). — Abstract

The diagnostic model used in this system is an extension of the model in our previous work. The fundus disease diagnosis model is based on a CNN, seResNext50, as the main structure. — Methods, Autonomous AI diagnostic system

First, although the dataset represented the true spectrum of the selected fundus diseases, some of the categories contained a few images and might result in bias in the results. — Discussion, 한계 서술 부분

Second, some diseases selected in this study started from peripheral retinal area which is beyond the scope of the fundus image such as RD and RP. Therefore, the DLS could not detect them at the initial stage. — Discussion, 한계 서술 부분

관련 개념

[[fundus-image-diagnosis]] · [[clinical-ai-validation]]

(이 논문은 안저 사진을 딥러닝으로 진단하는 임상 AI 연구라, TEM 등 소재 관찰 개념과는 관찰 대상·스케일이 달라 연결하지 않았다. 위 두 개념은 아직 페이지가 없어 “나중에 쓸 것” 표시로만 남긴다.)