Title : AI-enabled non-invasive colorectal cancer prescreening using iris imaging and self-supervised deep learning
Abstract:
Background: Colorectal cancer (CRC) remains a leading cause of cancer-related mortality worldwide, and outcomes are strongly stage-dependent. Screening participation nevertheless remains suboptimal because conventional modalities—colonoscopy and stool-based testing—are invasive, costly, or logistically constrained, particularly in resource-limited settings. Scalable, non-invasive risk-stratification tools that can complement established screening pathways are therefore a pressing clinical need. To our knowledge, population-scale studies investigating iris-image biomarkers in CRC have not previously been reported.
Methods: We conducted an exploratory machine-learning study evaluating a high-potential digital biomarker for CRC risk stratification. Adults with CRC confirmed by colonoscopy and histopathology, together with non-CRC controls, were enrolled under a standardized acquisition protocol and inference was performed on a cloud-based platform designed for population-scale deployment. A self-supervised vision model was pretrained on approximately 500,000 unlabeled iris images to learn generalizable ocular representations and subsequently fine-tuned on 15,000 clinically labeled iris images (as of 10 March 2026), comprising 7,000 CRC-positive cases and 8,000 healthy controls. Model development incorporated preprocessing, augmentation, and supervised optimization with patient-level train/validation/test splits to prevent data leakage. Performance was evaluated using the area under the receiver operating characteristic curve (AUROC), the area under the precision–recall curve (AUPRC), sensitivity at fixed specificity thresholds, and calibration metrics; uncertainty was estimated using bootstrapped confidence intervals. Interpretability was assessed via gradient-based activation mapping to identify image regions contributing to predictions.
Results: In preliminary retrospective evaluation across multiple CRC stages, the model achieved 85% sensitivity and 88% specificity in distinguishing CRC cases from healthy controls, with above-chance discrimination sustained in bootstrapped analyses. These findings support the hypothesis that disease-associated signals may be reflected in iris patterns. Given the exploratory design, however, results may in part reflect confounding factors or dataset-specific characteristics and should be interpreted accordingly. The system is intended for prescreening rather than diagnosis, generating a risk-stratified output to guide referral for confirmatory clinical testing.
Conclusions: These findings support the feasibility of iris-image-based representation learning as a hypothesis-generating approach to AI-enabled CRC prescreening. A rapid, non-invasive, and scalable iris-based tool could meaningfully enhance participation in CRC screening programs and facilitate earlier detection, particularly in underserved settings. Prospective multicenter validation, calibration across diverse populations and imaging devices, and comprehensive evaluation of robustness, subgroup performance, fairness, and real-world clinical utility are required before clinical deployment.

