Abstract
Computational pathology holds substantial promise for improving diagnosis and guiding treatment decisions. Recent pathology foundation models enable the extraction of rich patch-level representations from large-scale whole-slide images (WSIs), but current approaches for aggregating these features into slide-level predictions remain constrained by design limitations that hinder generalizability and reliability. Here we present nnMIL, a simple yet broadly applicable multiple-instance learning framework that connects patch-level foundation models to robust slide-level clinical prediction. nnMIL introduces random sampling at both the patch and feature levels, enabling large-batch optimization, task-aware sampling strategies, and efficient and scalable training across datasets and model architectures. A lightweight aggregator performs sliding-window inference to generate ensemble slide-level predictions and supports principled uncertainty estimation. Across 40,000 WSIs encompassing 35 clinical tasks and 4 pathology foundation models, nnMIL consistently outperformed existing MIL methods for disease diagnosis, histologic subtyping, molecular biomarker detection and pan-cancer prognosis prediction. It further demonstrated strong cross-model generalization, reliable uncertainty quantification and robust survival stratification in multiple external cohorts. In conclusion, nnMIL offers a practical and generalizable solution for translating pathology foundation models into clinically meaningful predictions, advancing the development and deployment of reliable AI systems in real-world settings.
This is a preview of subscription content, access
Access options
- Purchase on SpringerLink
- Instant access to the full article PDF.
Prices may be subject to local taxes which are calculated during checkout
Subjects
- Data processing
- Machine learning
- Prognosis
Data availability
All histopathology images and clinical annotations used for model development and evaluation are publicly available from the following sources: BCCC (https://datahub.aida.scilifelab.se/10.23698/aida/bccc), BRACS (https://www.bracs.icar.cnr.it/), EBRAINS (https://doi.org/10.25493/WQ48-ZGX), IMP-CRC2024 (https://rdm.inesctec.pt/dataset/nis-2023-008), PANDA (https://www.kaggle.com/c/prostate-cancer-grade-assessment/data), BCNB (https://bupt-ai-cz.github.io/BCNB), MCO (https://www.sredhconsortium.org/sredh-datasets/mco-study-whole-slide-image-dataset), SURGEN (https://www.ebi.ac.uk/biostudies/studies/S-BIAD1285), PLCO (https://cdas.cancer.gov/plco/), NLST (https://www.cancerimagingarchive.net/collection/nlst/), TCGA (https://portal.gdc.cancer.gov) and CAMELYON16 and 17 (https://camelyon17.grand-challenge.org/). All datasets are accessible to the research community under their respective data use agreements or institutional licenses. Source data are provided with this paper.
Code availability
The implementation of this project, including training, inference and evaluation pipelines as well as a usage tutorial, is publicly available on GitHub at https://github.com/Luoxd1996/nnMIL (ref. 69).
References
-
Bera, K., Schalper, K. A., Rimm, D. L., Velcheti, V. & Madabhushi, A. Artificial intelligence in digital pathology—new tools for diagnosis and precision oncology. Nat. Rev. Clin. Oncol.16, 703–715 (2019).
-
Yates, J. & Van Allen, E. M. New horizons at the interface of artificial intelligence and translational cancer research. Cancer Cell43, 708–727 (2025).
-
Bhinder, B., Gilvary, C., Madhukar, N. S. & Elemento, O. Artificial intelligence in cancer research and precision medicine. Cancer Discov.11, 900–915 (2021).
-
van der Laak, J., Litjens, G. & Ciompi, F. Deep learning in histopathology: the path to the clinic. Nat. Med.27, 775–784 (2021).
-
Xu, H. et al. A whole-slide foundation model for digital pathology from real-world data. Nature630, 181–188 (2024).
-
Wang, X. et al. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature634, 970–978 (2024).
-
Xiang, J. et al. A vision–language foundation model for precision oncology. Nature638, 769–778 (2025).
-
Saillard, C. et al. H-optimus-0. GitHubhttps://github.com/bioptimus/releases/tree/main/models/h-optimus/v0 (2024).
-
Chen, R. J. et al. Towards a general-purpose foundation model for computational pathology. Nat. Med.30, 850–862 (2024).
-
Lu, M. Y. et al. A visual–language foundation model for computational pathology. Nat. Med.30, 863–874 (2024).
-
Vorontsov, E. et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nat. Med.30, 2924–2935 (2024).
-
Ma, J. et al. A generalizable pathology foundation model using a unified knowledge distillation pretraining framework. Nat. Biomed. Eng.10, 545–564 (2026).
-
Huang, Z., Bianchi, F., Yuksekgonul, M., Montine, T. J. & Zou, J. A visual–language foundation model for pathology image analysis using medical twitter. Nat. Med.29, 2307–2316 (2023).
-
Wang, X. et al. Foundation model for predicting prognosis and adjuvant therapy benefit from digital pathology in GI cancers. J. Clin. Oncol.43, 3468–3481 (2025).
-
Campanella, G. et al. Real-world deployment of a fine-tuned pathology foundation model for lung cancer biomarker detection. Nat. Med.31, 3002–3010 (2025).
-
Kondepudi, A. et al. Foundation models for fast, label-free detection of glioma infiltration. Nature637, 439–445 (2025).
-
Shao, D. et al. Do multiple instance learning models transfer? In Proc. 42nd International Conference on Machine Learning 54219–54238 (PMLR, 2025).
-
Ilse, M., Tomczak, J. & Welling, M. Attention-based deep multiple instance learning. In International Conference on Machine Learning 2127–2136 (PMLR, 2018).
-
Lu, M. Y. et al. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat. Biomed. Eng.5, 555–570 (2021).
-
Zhang, H. et al. DTFD-MIL: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition 18802–18812 (IEEE, 2022).
-
Li, B., Li, Y. & Eliceiri, K. W. Dual-stream multiple instance learning network for whole slide image classification with self-supervised contrastive learning. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition 14318–14328 (IEEE, 2021).
-
Xiang, J. & Zhang, J. Exploring low-rank property in multiple instance learning for whole slide image classification. In The Eleventh International Conference on Learning Representations Paper 774 (ICLR, 2023).
-
Shao, Z. et al. TransMIL: transformer based correlated multiple instance learning for whole slide image classification. Adv. Neural Inf. Process. Syst.34, 2136–2147 (2021).
-
Li, J. et al. Dynamic graph representation with knowledge-aware attention for histopathology whole slide image analysis. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition 11323–11332 (IEEE, 2024).
-
Neidlinger, P. et al. Benchmarking foundation models as feature extractors for weakly supervised computational pathology. Nat. Biomed. Eng.10, 1113–1123 (2026).
-
Ma, J. et al. Pathbench: a comprehensive comparison benchmark for pathology foundation models towards precision oncology. Preprint at https://doi.org/10.48550/arXiv.2505.20202 (2025).
-
El Nahhas, O. S. et al. From whole-slide image to biomarker prediction: end-to-end weakly supervised deep learning in computational pathology. Nat. Protoc.20, 293–316 (2025).
-
Chen, R. J. et al. Pan-cancer integrative histology-genomic analysis
-
Vaidya, A. et al. Molecular-driven foundation model for oncologic pathology. Preprint at https://doi.org/10.48550/arXiv.2501.16652 (2025).
-
Ding, T. et al. A multimodal whole-slide foundation model for pathology. Nat. Med.31, 3749–3761 (2025).
-
He, T. et al. Bag of tricks for image classification with convolutional neural networks. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition 558–567 (IEEE, 2019).
-
Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J. & Maier-Hein, K. H. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat. Methods18, 203–211 (2021).
-
Lakshminarayanan, B., Pritzel, A. & Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. Adv. Neural Inf. Process. Syst.30, 6402–6413 (2017).
-
Gal, Y. & Ghahramani, Z. Dropout as a Bayesian approximation: representing model uncertainty in deep learning. In International Conference on Machine Learning 1050–1059 (PMLR, 2016).
-
Zimmermann, E. et al. Virchow2: scaling self-supervised mixed magnification models in pathology. Preprint at https://doi.org/10.48550/arXiv.2408.00738 (2024).
-
Yacob, F. et al. Weakly supervised detection and classification of basal cell carcinoma using graph-transformer on whole slide images. Sci. Rep.13, 7555 (2023).
-
Brancati, N. et al. BRACS: a dataset for breast carcinoma subtyping in H&E histology images. Database2022, baac093 (2022).
-
Roetzer-Pejrimovsky, T. et al. The digital brain tumour atlas, an open histopathology resource. Sci. Data9, 55 (2022).
-
Neto, P. C. et al. An interpretable machine learning system for colorectal cancer diagnosis from pathology slides. NPJ Precis. Oncol.8, 56 (2024).
-
Bulten, W. et al. Artificial intelligence for diagnosis and gleason grading of prostate cancer: the panda challenge. Nat. Med.28, 154–163 (2022).
-
Xu, F. et al. Predicting axillary lymph node metastasis in early breast cancer using deep learning on primary tumor biopsy slides. Front. Oncol.11, 759007 (2021).
-
Myles, C., Um, I. H., Marshall, C., Harris-Birtill, D. & Harrison, D. J. SurGen: 1020 H&E-stained whole slide images with survival and genetic markers. GigaScience14, giaf086 (2025).
-
Hawkins, N. & Ward, R. MCO study whole slide image collection [Dataset]. UNSWhttps://doi.org/10.4225/53/555921d09f76b (2015).
-
Jonnagaddala, J. et al. in Nursing Informatics 2016 387–391 (IOS Press, 2016).
-
Weinstein, J. N. et al. The cancer genome atlas pan-cancer analysis project. Nat. Genet.45, 1113–1120 (2013).
-
Taylor, A. M. et al. Genomic and functional approaches to understanding cancer aneuploidy. Cancer Cell33, 676–689 (2018).
-
Bejnordi, B. E. et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. JAMA318, 2199–2210 (2017).
-
Bandi, P. et al. From detection of individual metastases to classification of lymph node status at the patient level: the camelyon17 challenge. IEEE Trans. Med. Imaging38, 550–560 (2018).
-
Ling, X. et al. Comprehensive benchmark dataset for pathological lymph node metastasis in breast cancer sections. Sci. Data12, 1381 (2025).
-
Cai, L., Huang, S., Zhang, Y., Lu, J. & Zhang, Y. AttriMIL: revisiting attention-based multiple instance learning for whole-slide pathological image classification from a perspective of instance attributes. Med. Image Anal.103, 103631 (2025).
-
Gohagan, J. K. et al. The prostate, lung, colorectal and ovarian (PLCO) cancer screening trial of the National Cancer Institute: history, organization, and status. Control. Clin. Trials21, 251S–272S (2000).
-
Team, N. L. S. T. R. The National Lung Screening Trial: overview and study design. Radiology258, 243–253 (2011).
-
Hoffer, E., Hubara, I. & Soudry, D. Train longer, generalize better: closing the generalization gap in large batch training of neural networks. Adv. Neural Inf. Process. Syst.30, 1729–1739 (2017).
-
Yang, Y. & Xu, Z. Rethinking the value of labels for improving class-imbalanced learning. Adv. Neural Inf. Process. Syst.33, 19290–19301 (2020).
-
Luo, X. et al. Ensemble learning of foundation models for precision oncology. Preprint at https://doi.org/10.48550/arXiv.2508.16085 (2025).
-
De Fauw, J. et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat. Med.24, 1342–1350 (2018).
-
Sali, R., Jiang, Y., Attaranzadeh, A., Holmes, B. & Li, R. Morphological diversity of cancer cells predicts prognosis across tumor types. J. Natl Cancer Inst.116, 555–564 (2024).
-
de La Torre, J., Puig, D. & Valls, A. Weighted kappa loss function for multi-class classification of ordinal data in deep learning. Pattern Recognit. Lett.105, 144–154 (2018).
-
Yuan, Z., Yan, Y., Sonka, M. & Yang, T. Large-scale robust deep AUC maximization: a new surrogate loss and empirical studies on medical image classification. In Proc. IEEE/CVF International Conference on Computer Vision 3040–3049 (IEEE, 2021).
-
Shulman, E. D. et al. AI-predicted spatial transcriptomics unlocks breast cancer biomarkers from pathology. Cell189, 4225–4240.e25 (2026).
-
Li, Z. et al. AI-enabled virtual spatial proteomics from histopathology for interpretable biomarker discovery in lung cancer. Nat. Med.32, 231–244 (2026).
-
Li, Y. et al. Cellular architecture and neighborhood-informed virtual spatial tumor profiling from histopathology. Cell189, 4241–4259.e9 (2026).
-
Cox, D. R. Regression models and life-tables. J. R. Stat. Soc. B34, 187–220 (1972).
-
Kvamme, H. & Borgan, Ø Continuous and discrete-time survival prediction with neural networks. Lifetime Data Anal.27, 710–736 (2021).
-
Loshchilov, I. & Hutter, F. Decoupled weight decay regularization. In Proc. 7th International Conference on Learning Representations (ICLR, 2019).
-
Loshchilov, I. & Hutter, F. SGDR: stochastic gradient descent with warm restarts. In International Conference on Learning Representations (ICLR, 2017).
-
Prechelt, L. in Neural Networks: Tricks of the Trade 55–69 (Springer, 2002).
-
Ling, X. et al. Agent aggregator with mask denoise mechanism for histopathology whole slide image analysis. In Proc. 32nd ACM International Conference on Multimedia 2795–2803 (ACM, 2024).
-
Luo, X. et al. nnMIL: a generalizable multiple instance learning framework for computational pathology. GitHubhttps://github.com/Luoxd1996/nnMIL (2026).
-
Neto, P. C. et al. iMIL4PATH: a semi-supervised interpretable approach for colorectal whole-slide images. Cancers14, 2489 (2022).
-
Goyal, P. et al. Accurate, large minibatch SGD: training imagenet in 1 hour. Preprint at https://doi.org/10.48550/arXiv.1706.02677 (2017).
-
Mandt, S., Hoffman, M. D. & Blei, D. M. Stochastic gradient descent as approximate Bayesian inference. J. Mach. Learn. Res.18, 4873–4907 (2017).
Acknowledgements
We acknowledge the BCCC36, BCNB41, BRACS37, EBRAINS38, IMP-CRC202439,70, PANDA40, MCO43,44, SURGEN42, NLST52, TCGA45, CAMELYON1647 and CAMELYON1748 consortia for making their datasets publicly available. We further thank the research team46 for providing aneuploidy scores and curating the whole-genome doubling and tumour mutational burden annotations.
Funding
This study was supported by the Himalaya Foundation Faculty Scholarship.
Authors and Affiliations
Contributions
X.L. conceived and designed the study, developed the model, curated the datasets, conducted all experiments, performed statistical analysis, created visualizations and drafted the manuscript. J.X. and Y.J. contributed to the conception of the study and provided advice on experimental design, visualization and manuscript writing. R.L. contributed to the conception and design of the study, interpreted the results, revised the manuscript, acquired funding and supervised the project. All authors reviewed and approved the final manuscript.
Ethics declarations
Competing interests
The authors declare no competing interests.
Peer review
Peer review information
Nature Biomedical Engineering thanks Junzhou Huang, Hari Subramoni and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data
Extended Data Fig. 1 Detailed ranking scores across disease diagnosis and subtyping, molecular biomarker detection, pan-cancer prognosis prediction and generalizability of prognostic models in four external cohorts.
a, Disease diagnosis and subtyping. b, Molecular biomarker detection. c, Pan-cancer prognosis prediction. d, Generalizability of prognostic models in external validation. All results are reported as mean ± standard deviation.
Extended Data Fig. 2 The relationship between prediction score and estimated slide-level uncertainty.
(a), shows results for disease subtyping tasks, where uncertainty distinctly separates correct from incorrect predictions across EBRAINS, BCCC, PANDA, and IMP-CRC2024. (b), presents molecular biomarker detection tasks, including ER, BRAF, IDH, and TMB, showing that the same relationship between uncertainty and prediction accuracy is preserved, underscoring the model’s consistent calibration properties. Purple open circles indicate correctly classified cases, whereas orange crosses denote misclassified samples.
Extended Data Fig. 3 Attention visualization of nnMIL aligned with pathological hallmarks.
Representative whole-slide images (WSIs from MCO cohort, features were extracted by Virchow2) with corresponding attention maps for low-risk and high-risk cases. High-attention regions predominantly localize to tumor cell–rich areas, and frequently co-localize with dense fibroblastic connective tissue, consistent with regions that are routinely emphasized by pathologists during diagnosis. This concordance suggests that nnMIL captures clinically relevant histopathological patterns rather than spurious background signals.
Extended Data Fig. 4 Correlation between predicted risk and uncertainty in external cohorts and KM survival analyses.
Correlation between predicted risk scores and uncertainty estimates in external cohorts, and Kaplan–Meier survival analysis stratified by uncertainty levels. (a), Correlation between predicted risk and estimated uncertainty across external cohorts, demonstrating that uncertainty estimates are well calibrated and convey meaningful information about prediction confidence. (b), Kaplan–Meier survival analyses based on estimated uncertainty scores in both low- and high-risk groups in each external cohort. Statistical significance of survival differences between high- and low-uncertainty groups in low- and high-risk groups was assessed using a two-sided log-rank test, respectively.
Extended Data Fig. 5 Ablation study of the nnMIL framework.
Ablation study of nnMIL. (a) Stepwise ablation of the nnMIL training strategy, progressively adding gradient accumulation (effective batch size = 32), patch sampling, and feature sampling starting from a baseline trained with a batch size of 1. Bars show mean values and error bars denote the standard error of the mean across all 35 tasks (40 cohorts). Statistical significance was determined using a two-sided Wilcoxon signed-rank test, where ns means non-significance, * indicates P < 0.05, ** indicates P < 0.01, and *** indicates P < 0.001. (b) Sensitivity analysis regarding the hyperparameters during the training strategies of nnMIL. Hyperparameters were chosen to represent practical training choices commonly adjusted in WSI-based MIL (e.g., model capacity, instance coverage, and computational efficiency), with default values following standard practice and preliminary stability checks. For the random seed, we used commonly adopted values in the machine learning literature, including 0, 1, 2, 42 (default setting in many machine learning packages), and 2026 (this year), to ensure fair and reproducible evaluation. (c) Ablation of two different data sampling strategies for mini-batch. In c, bars for EBRAINS and CRC BRAF represent mean values and error bars represent standard deviations, derived from 1,000 bootstrap replicates on each independent test set. And, bar of Survival and Mean show mean values and error bars denote the standard error of the mean across all cohorts/tasks, each dot represents one task.
Supplementary information
Supplementary Tables 1–28.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Luo, X., Xiang, J., Ji, Y. et al. nnMIL: a generalizable multiple instance learning framework for computational pathology.
Nat. Biomed. Eng (2026). https://doi.org/10.1038/s41551-026-01767-8
-
Version of record:25 August 2026
-
DOI
:https://doi.org/10.1038/s41551-026-01767-8
