Bias in Developing AI Medical Diagnosis: Impact of Sequential CT Scan Slices on Data Leakage in Lung Cancer
Loading...
Files
Size: 2.21 MB, File format: Adobe PDF
Date
Authors
Rojas, Fredy
Madanian, Samaneh
Supervisor
Item type
Degree name
Journal Title
Journal ISSN
Volume Title
Publisher
Elsevier BV
Abstract
Lung cancer is highly heterogeneous, making reliable evaluation of AI-based computer-aided diagnosis systems challenging. Public CT datasets such as IQ-OTH/NCCD are valuable for research, but the released dataset lacks subject identifiers, preventing verified subject-wise train–test partitioning and increasing the risk that slices from the same patient may appear in both training and test sets. This study re-examines IQ-OTH/NCCD to assess how partitioning affects performance and performance-overestimation risk. We evaluated 14 deep-learning pipeline configurations that combined preprocessing, CNN architectures, partitioning, and training strategies. Performance was compared under two conditions: With Order, which preserved the released CT-slice sequence, and Without Order, which shuffled images before fold assignment. Under the null hypothesis, the two conditions were expected to produce comparable performance estimates. Systematically higher performance under Without Order was interpreted as evidence that this partitioning condition may produce more optimistic performance estimates. Performance was evaluated using accuracy, precision, recall, and F1-score Results showed consistently higher performance estimates under the Without Order condition across the evaluated DL pipelines. In 100 repetitions of stratified 5KfoldCV, mean accuracy was 24.12 percentage points higher for the 2DCNN pipeline (95% CI: 23.52–24.74) and 12.42 percentage points higher for the ResNet pipeline (95% CI: 12.19–12.67). Both differences were statistically significant (p<0.001). These findings suggest that the Without Order condition can produce more optimistic performance estimates when applied to IQ-OTH/NCCD. This study raises awareness of the risk of performance overestimation without discouraging the use of IQ-OTH/NCCD. Rather, the findings show that partitioning choices can substantially influence performance estimates when subject identifiers are unavailable. We therefore recommend explicit reporting of data partitioning strategies and the provision of reproducible code within a containerised environment to support transparent evaluation.
Description
Keywords
4203 Health Services and Systems, 42 Health Sciences, Networking and Information Technology R&D (NITRD), Prevention, Cancer, Machine Learning and Artificial Intelligence, Lung Cancer, Lung, 4203 Health services and systems, Digital health, Model evaluation protocol, Bias, Data leakage, Medical imaging, AI in healthcare, Reproducibility
Source
Informatics in Medicine Unlocked, ISSN: 2352-9148 (Print); 2352-9148 (Online), Elsevier BV, 65, 101799-101799. doi: 10.1016/j.imu.2026.101799
Publisher's version
Rights statement
© 2026 The Authors. Published by Elsevier Ltd. Open access.
Permanent link
Endorsement
Review
Supplemented By
Referenced By
Creative Commons license
Except where otherwise noted, this item's license is described as Attribution-NonCommercial-NoDerivatives 4.0 International

