Repository logo

Bias in Developing AI Medical Diagnosis: Impact of Sequential CT Scan Slices on Data Leakage in Lung Cancer

Loading...
Thumbnail Image

Files

Size: 2.21 MB, File format: Adobe PDF

Authors

Rojas, Fredy

Madanian, Samaneh

Supervisor

Degree name

Journal Title

Journal ISSN

Volume Title

Publisher

Elsevier BV

Abstract

Lung cancer is highly heterogeneous, making reliable evaluation of AI-based computer-aided diagnosis systems challenging. Public CT datasets such as IQ-OTH/NCCD are valuable for research, but the released dataset lacks subject identifiers, preventing verified subject-wise train–test partitioning and increasing the risk that slices from the same patient may appear in both training and test sets. This study re-examines IQ-OTH/NCCD to assess how partitioning affects performance and performance-overestimation risk. We evaluated 14 deep-learning pipeline configurations that combined preprocessing, CNN architectures, partitioning, and training strategies. Performance was compared under two conditions: With Order, which preserved the released CT-slice sequence, and Without Order, which shuffled images before fold assignment. Under the null hypothesis, the two conditions were expected to produce comparable performance estimates. Systematically higher performance under Without Order was interpreted as evidence that this partitioning condition may produce more optimistic performance estimates. Performance was evaluated using accuracy, precision, recall, and F1-score Results showed consistently higher performance estimates under the Without Order condition across the evaluated DL pipelines. In 100 repetitions of stratified 5KfoldCV, mean accuracy was 24.12 percentage points higher for the 2DCNN pipeline (95% CI: 23.52–24.74) and 12.42 percentage points higher for the ResNet pipeline (95% CI: 12.19–12.67). Both differences were statistically significant (p<0.001). These findings suggest that the Without Order condition can produce more optimistic performance estimates when applied to IQ-OTH/NCCD. This study raises awareness of the risk of performance overestimation without discouraging the use of IQ-OTH/NCCD. Rather, the findings show that partitioning choices can substantially influence performance estimates when subject identifiers are unavailable. We therefore recommend explicit reporting of data partitioning strategies and the provision of reproducible code within a containerised environment to support transparent evaluation.

Description

Keywords

4203 Health Services and Systems, 42 Health Sciences, Networking and Information Technology R&D (NITRD), Prevention, Cancer, Machine Learning and Artificial Intelligence, Lung Cancer, Lung, 4203 Health services and systems, Digital health, Model evaluation protocol, Bias, Data leakage, Medical imaging, AI in healthcare, Reproducibility

Source

Informatics in Medicine Unlocked, ISSN: 2352-9148 (Print); 2352-9148 (Online), Elsevier BV, 65, 101799-101799. doi: 10.1016/j.imu.2026.101799

Rights statement

© 2026 The Authors. Published by Elsevier Ltd. Open access.

Endorsement

Review

Supplemented By

Referenced By

Creative Commons license

Except where otherwise noted, this item's license is described as Attribution-NonCommercial-NoDerivatives 4.0 International