| 研究生: |
蘇芳毅 Su, Fang-Yi |
|---|---|
| 論文名稱: |
邁向公平精準醫療之癌症病理人工智慧:資料生成、表徵去偏與智慧代理驅動之基礎模型整合 Toward Fair Precision Medicine with Cancer Pathology AI: Data Generation, Representation Debiasing, and Agentic Foundation-Model Integration |
| 指導教授: |
蔣榮先
Chiang, Jung-Hsien |
| 學位類別: |
博士 Doctor |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 106 |
| 中文關鍵詞: | 生成式人工智慧 、癌症病理 、演算法公平性 、基礎模型 、多體學預測 |
| 外文關鍵詞: | Generative AI, Cancer Pathology, Algorithmic Fairness, Foundation Models, Multi-Omics Prediction |
| 相關次數: | 點閱:2 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
精準醫療的核心目標,是為每一位病人提供個別化的癌症照護。然而,近年快速發展的人工智慧癌症病理診斷系統,已被觀察到在不同病人族群之間存在系統性的效能落差,進而挑戰了精準醫療「人人皆可受益」的基本承諾。本論文發展一套以生成式人工智慧為核心的整合框架,旨在使人工智慧輔助癌症病理診斷同時達到更高的公平性與更精準的診斷表現。研究從一個實證起點出發,將效能落差追溯至腫瘤生物學;接著沿著臨床人工智慧流程的三個層次(資料、表徵、決策)由最上游逐步推進至最下游進行介入;最後將這些方法對應到通往臨床落地的具體途徑。
本論文首先建立 Multi-Omics Multi-cohort Assessment(MOMA)平台,證明常規 H&E 病理切片中蘊含豐富的分子資訊,並可由人工智慧模型加以萃取與預測,包括微衛星不穩定性(MSI)、驅動基因突變、複本數變異,以及臨床預後等。此平台於三個獨立隊列、共 1,888 名病人中完成驗證。進一步延伸至 10 種癌別、共 9,217 名病人的分析則顯示,不同族群之間的驅動基因突變率差異,特別是 TP53 與 CDH1,與人工智慧模型在驗證過程中呈現的人口統計效能差距具有統計關聯。這些結果指出,病理人工智慧中的公平性問題並非單純的統計假象,而是與真實存在的腫瘤生物學差異密切相關。
在此基礎上,本論文於資料層先行介入:生成式人工智慧在最上游處(訓練資料分布的組成本身)校正偏差。本論文提出 Fairness DDPM,一種以人口統計屬性為條件的擴散生成模型,用以合成代表性不足族群的病理影像。實驗結果顯示,Fairness DDPM 在 33 種癌別與三個病理基礎模型上,可降低最高達 80% 的種族偏誤;在 UNI 模型的四項公平性指標中,更有三項完全消除了種族偏差。
於表徵層,本論文提出 FAIR-Path,一個公平性感知的對比學習架構,結合非歧視損失函數,改變模型從既有資料中學習表徵的方式。在涵蓋 28,732 張全玻片影像、14,456 名病人與 20 種癌別的泛癌資料中,FAIR-Path 於內部評估消除了 88.5% 的診斷不公平性,並於來自 7 個醫療中心、15 個獨立外部隊列的驗證中達到 91.1% 的效能差距縮減,且未犧牲整體診斷準確率。其後,一項機制性分析 UNVEIL 進一步打開黑箱,揭示這些被抑制的人口統計訊號其實編碼於基礎模型的特徵之中,並可追溯至中介了最高達 47.4% 效能差距的組織與核形態,使表徵層的偏差不僅可被校正,更可被看見。
於決策層,本論文將上述公平性導向的病理人工智慧流程整合至 METASIGHT,一個具有臨床部署潛力的泛癌轉移預測系統。METASIGHT 透過大型語言模型驅動的智慧代理集成整合多個病理基礎模型:由一個 agent 在 CHIEF、GigaPath、MUSK 與 KEEP 組成的面板上,搜尋依癌別而定的組合函數,而非仰賴人工調校的固定權重。此系統以 TCGA 開發,並於七個獨立外部隊列進行驗證,研究族群共 14,297 名病人、橫跨 23 種癌別,涵蓋 28,415 張全玻片影像與 2,895 個 TMA cores。結果顯示,METASIGHT 在轉移狀態預測達到 macro-AUROC 0.793、未來軌跡預測最高達 0.903,並一致優於最佳單一基礎模型、傳統集成與多變數臨床變數模型。進一步的核形態量化分析亦顯示,與轉移能力相關的核形態特徵在長期追蹤中具有穩定且保守的表現,為模型預測結果提供了由影像表徵連結至病理機制的可解釋基礎。本論文並進一步發展一種自我演化的延伸,參考 AlphaEvolve 的智慧代理程式演化範式,直接演化集成的組合函數,而非僅調整其權重。
綜合而言,本論文主張並以實證結果支持:公平性與診斷精準性並非彼此衝突的目標。相反地,真正可被稱為精準的醫療人工智慧系統,必須在不同病人族群之間皆能維持可靠表現;一個只對部分病人有效的系統,並不能稱為真正精準。本論文所發展的生成式人工智慧方法,在資料、表徵、決策三個層次上同時追求公平與精準。最後,本論文將上述貢獻對應到通往臨床落地的具體途徑(從作為臨床評估的數位 copilot,到跨多元族群的公平人工智慧)勾勒出一個內部既公平又精準的病理人工智慧,如何走入例行臨床照護。
Precision medicine promises individualized cancer care, yet AI-driven cancer pathology diagnostic systems exhibit performance disparities across patient subgroups that threaten this promise. This dissertation develops a unified framework that makes AI-based cancer pathology both fairer and more precise. Beginning from an empirical origin that traces these disparities to tumor biology, it intervenes at three successive levels of the clinical AI pipeline (data, representation, and decision), from the most upstream stage of the diagnostic workflow to the most downstream, and closes by mapping these methods onto concrete routes to clinical translation.
The work begins from an empirical origin. The Multi-Omics Multi-cohort Assessment (MOMA) platform establishes that routine hematoxylin and eosin (H&E)-stained histopathology images encode rich molecular information that AI can extract (including microsatellite instability (MSI), driver mutations, copy-number alterations, and clinical outcomes) validated across 1,888 patients from three independent cohorts. A follow-up analysis on 9,217 patients across ten cancer types shows that population-level differences in driver-gene mutation rates (notably TP53 and CDH1) are statistically associated with the demographic performance gaps observed during validation, grounding the fairness problem in concrete tumor biology rather than treating it as a statistical artifact.
At the data level, generative AI corrects bias at its most upstream point, the composition of the training distribution itself. Fairness DDPM, a demographic-attribute-conditioned denoising diffusion probabilistic model, synthesizes pathology images for under-represented minority groups and reduces racial bias by up to 80% across 33 cancer types and three foundation models, completely eliminating racial bias in three of four metrics for UNI.
At the representation level, FAIR-Path (Fairness-Aware Artificial Intelligence Review for Pathology), a fairness-aware contrastive learning framework with a non-discrimination loss, changes how the model learns from given data: it eliminates 88.5% of demographic diagnostic disparities across 28,732 whole-slide images from 14,456 patients spanning 20 cancer types, with a 91.1% gap reduction across 15 independent external cohorts from seven medical centers, without sacrificing diagnostic accuracy. A mechanistic analysis (UNVEIL) then opens the black box, showing that the suppressed demographic signals are encoded in foundation-model features and traceable to tissue and nuclear architecture that mediates up to 47.4% of the performance gap, making representation-level bias not only correctable but legible.
At the decision level, METASIGHT (Metastasis Evolution and Trajectory Alignment for Surveillance using Image-Guided Histological Topology) integrates multiple pathology foundation models into a clinically deployable system through a large language model (LLM)–driven agentic ensemble, in which an agent searches for a cancer-specific combination function over a panel of CHIEF, GigaPath, MUSK, and KEEP. Developed on The Cancer Genome Atlas (TCGA) and validated across seven independent external cohorts, on a study population of 14,297 patients (28,415 WSIs and 2,895 TMA cores) spanning 23 cancer types, METASIGHT achieves a macro-AUROC of 0.793 for metastasis status and up to 0.903 for future-trajectory prediction, consistently outperforming the best single foundation model, conventional ensembles, and multivariable clinical-variable models, while nuclear-morphometric analyses reveal conserved morphological signatures of metastatic competence that persist over extended follow-up. Following the AlphaEvolve paradigm of agentic program search, an evolutionary extension further evolves the ensemble combination function itself, rather than only its weights.
This dissertation argues, and provides empirical evidence, that fairness and diagnostic accuracy are not in opposition: the methods developed here achieve both, recognizing that a precision-medicine system which fails some patients is not truly precise for any. A closing analysis maps these contributions onto concrete pathways for clinical translation (from a digital diagnostic copilot to equitable AI across diverse populations), charting how internally fair and precise pathology AI can reach routine patient care.
[1] Ashley, E. A. Towards Precision Medicine. Nature Reviews Genetics 17, 507–522 (2016).
[2] Aygün, E. et al. An AI System to Help Scientists Write Expert-Level Empirical Software. Nature (2026).
[3] Benjamini, Y. & Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological) 57, 289–300 (1995).
[4] Benson, A. B., Venook, A. P., Al-Hawary, M. M., et al. Colon Cancer, Version 2.2021, NCCN Clinical Practice Guidelines in Oncology. Journal of the National Comprehensive Cancer Network 19, 329–359 (2021).
[5] Bommasani, R., Hudson, D. A., Adeli, E., et al. On the Opportunities and Risks of Foundation Models. arXiv preprint arXiv:2108.07258 (2021).
[6] Brown, T., Mann, B., Ryder, N., et al. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems (NeurIPS) 1877–1901 (2020).
[7] Cancer Genome Atlas Network. Comprehensive Molecular Characterization of Human Colon and Rectal Cancer. Nature 487, 330–337 (2012).
[8] Carbonneau, M.-A., Cheplygina, V., Granger, E., et al. Multiple Instance Learning: A Survey of Problem Characteristics and Applications. Pattern Recognition 77, 329–353 (2018).
[9] Caron, M., Touvron, H., Misra, I., et al. Emerging Properties in Self-Supervised Vision Transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 9650–9660 (2021).
[10] Carrot-Zhang, J., Chambwe, N., Damrauer, J. S., et al. Comprehensive Analysis of Genetic Ancestry and Its Molecular Correlates in Cancer. Cancer Cell 37, 639–654.e6 (2020).
[11] Caruana, R., Niculescu-Mizil, A., Crew, G., & Ksikes, A. Ensemble Selection from Libraries of Models. In Proceedings of the Twenty-First International Conference on Machine Learning (ICML) (2004).
[12] Chaffer, C. L. & Weinberg, R. A. A Perspective on Cancer Cell Metastasis. Science 331, 1559–1564 (2011).
[13] Chen, R. J., Chen, C., Li, Y., et al. Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 16144–16155 (2022).
[14] Chen, R. J., Wang, J. J., Williamson, D. F., et al. Algorithmic Fairness in Artificial Intelligence for Medicine and Healthcare. Nature Biomedical Engineering 7, 719–742 (2023).
[15] Chen, R. J., Ding, T., Lu, M. Y., et al. Towards a General-Purpose Foundation Model for Computational Pathology. Nature Medicine 30, 850–862 (2024).
[16] Chen, T., Kornblith, S., Norouzi, M., et al. A Simple Framework for Contrastive Learning of Visual Representations. In International Conference on Machine Learning (ICML) 1597–1607 (2020).
[17] Chiang, C.-J., Lo, W.-C., Yang, Y.-W., You, S.-L., Chen, C.-J., & Lai, M.-S. Incidence and Survival of Adult Cancer Patients in Taiwan, 2002–2012. Journal of the Formosan Medical Association 115, 1076–1088 (2016).
[18] Chouldechova, A. Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments. Big Data 5, 153–163 (2017).
[19] Ciga, O., Xu, T., & Martel, A. L. Self supervised contrastive learning for digital histopathology. Machine Learning with Applications 7, 100198 (2022).
[20] Cliff, N. Dominance Statistics: Ordinal Analyses to Answer Ordinal Questions. Psychological Bulletin 114, 494–509 (1993).
[21] Collins, F. S. & Varmus, H. A New Initiative on Precision Medicine. New England Journal of Medicine 372, 793–795 (2015).
[22] Coudray, N., Ocampo, P. S., Sakellaropoulos, T., et al. Classification and Mutation Prediction from Non-Small Cell Lung Cancer Histopathology Images Using Deep Learning. Nature Medicine 24, 1559–1567 (2018).
[23] Davis, M. B. & Martini, R. Precision oncology and genetic ancestry: The science behind population-based cancer disparities. Cancer Cell 43, 619–622 (2025).
[24] DeLong, E. R., DeLong, D. M., & Clarke-Pearson, D. L. Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. Biometrics 44, 837–845 (1988).
[25] Ding, J., Ma, S., Dong, L., et al. LongNet: Scaling Transformers to 1,000,000,000 Tokens. arXiv preprint arXiv:2307.02486 (2023).
[26] Ducreux, M., Chamseddine, A., Laurent-Puig, P., et al. Molecular Targeted Therapy of BRAF-Mutant Colorectal Cancer. Therapeutic Advances in Medical Oncology 11, 1758835919856494 (2019).
[27] Efron, B. Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics 7, 1–26 (1979).
[28] Esteva, A., Robicquet, A., Ramsundar, B., et al. A Guide to Deep Learning in Healthcare. Nature Medicine 25, 24–29 (2019).
[29] Fu, Y., Jung, A. W., Torne, R. V., et al. Pan-Cancer Computational Histopathology Reveals Mutations, Tumor Composition and Prognosis. Nature Cancer 1, 800–810 (2020).
[30] Galon, J., Costes, A., Sanchez-Cabo, F., et al. Type, Density, and Location of Immune Cells Within Human Colorectal Tumors Predict Clinical Outcome. Science 313, 1960– 1964 (2006).
[31] Ganesh, K., Stadler, Z. K., Cercek, A., et al. Immunotherapy in Colorectal Cancer: Rationale, Challenges and Potential. Nature Reviews Gastroenterology & Hepatology 16, 361–375 (2019).
[32] Garraway, L. A. Genomics-Driven Oncology: Framework for an Emerging Paradigm. Journal of Clinical Oncology 31, 1806–1814 (2013).
[33] Ghareeb, A. E. et al. A Multi-Agent System for Automating Scientific Discovery. Nature (2026).
[34] Gottweis, J. et al. Accelerating Scientific Discovery with Co-Scientist. Nature (2026).
[35] Guinney, J., Dienstmann, R., Wang, X., et al. The Consensus Molecular Subtypes of Colorectal Cancer. Nature Medicine 21, 1350–1356 (2015).
[36] Hardt, M., Price, E., & Srebro, N. Equality of Opportunity in Supervised Learning. In Advances in Neural Information Processing Systems (NeurIPS) 3323–3331 (2016).
[37] Harrell, F. E., Lee, K. L., & Mark, D. B. Multivariable Prognostic Models: Issues in Developing Models, Evaluating Assumptions and Adequacy, and Measuring and Reducing Errors. Statistics in Medicine 15, 361–387 (1996).
[38] Hasin, Y., Seldin, M., & Lusis, A. Multi-Omics Approaches to Disease. Genome Biology 18, 83 (2017).
[39] Haslam, A., Kim, M. S., & Prasad, V. Updated Estimates of Eligibility for and Response to Genome-Targeted Oncology Drugs among US Cancer Patients, 2006-2020. Annals of Oncology 32, 926–932 (2021).
[40] He, K., Zhang, X., Ren, S., et al. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 770–778 (2016).
[41] He, K., Fan, H., Wu, Y., et al. Momentum Contrast for Unsupervised Visual Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 9729–9738 (2020).
[42] Ho, J. & Salimans, T. Classifier-Free Diffusion Guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications (2022).
[43] Ho, J., Jain, A., & Abbeel, P. Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems (NeurIPS) 6840–6851 (2020).
[44] Hoadley, K. A., Yau, C., Hinoue, T., et al. Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer. Cell 173, 291–304 (2018).
[45] Ilse, M., Tomczak, J., & Welling, M. Attention-Based Deep Multiple Instance Learning. In International Conference on Machine Learning (ICML) 2127–2136 (2018).
[46] Kadota, K., Suzuki, K., Kachala, S. S., et al. A Grading System Combining Architectural Features and Mitotic Count Predicts Recurrence in Stage I Lung Adenocarcinoma. Modern Pathology 25, 1117–1127 (2012).
[47] Kaplan, E. L. & Meier, P. Nonparametric Estimation from Incomplete Observations. Journal of the American Statistical Association 53, 457–481 (1958).
[48] Kather, J. N., Pearson, A. T., Halama, N., et al. Deep Learning Can Predict Microsatellite Instability Directly from Histology in Gastrointestinal Cancer. Nature Medicine 25, 1054–1056 (2019).
[49] Khosla, P., Teterwak, P., Wang, C., et al. Supervised Contrastive Learning. In Advances in Neural Information Processing Systems (NeurIPS) 18661–18673 (2020).
[50] Klein, S. L. & Flanagan, K. L. Sex Differences in Immune Responses. Nature Reviews Immunology 16, 626–638 (2016).
[51] Komura, D. & Ishikawa, S. Machine Learning Methods for Histopathological Image Analysis. Computational and Structural Biotechnology Journal 16, 34–42 (2018).
[52] Lambert, A. W., Pattabiraman, D. R., & Weinberg, R. A. Emerging Biological Principles of Metastasis. Cell 168, 670–691 (2017).
[53] Le, D. T., Uram, J. N., Wang, H., et al. PD-1 Blockade in Tumors with Mismatch-Repair Deficiency. New England Journal of Medicine 372, 2509–2520 (2015).
[54] Lehman, J. & Stanley, K. O. Abandoning Objectives: Evolution Through the Search for Novelty Alone. Evolutionary Computation 19, 189–223 (2011).
[55] Li, B., Li, Y., & Eliceiri, K. W. Dual-Stream Multiple Instance Learning Network for Whole Slide Image Classification with Self-Supervised Contrastive Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 14318–14328 (2021).
[56] Lin, P.-J., Lin, S.-Y., Tsai, P.-C., et al. Mutation Rate Differences Across Populations and Association with Performance Disparities in Pathology AI Diagnostic Models. In Proceedings of the American Society of Clinical Oncology Annual Meeting (2025). Abstract 1603.
[57] Lin, S.-Y., Tsai, P.-C., Dong, Y., et al. Generative AI to Augment the Fairness of Foundation Models in Cancer Pathology Diagnosis. In Proceedings of the American Society of Clinical Oncology Annual Meeting (2025). Abstract e23230.
[58] Lin, S.-Y., Tsai, P.-C., Su, F.-Y., et al. Contrastive Learning Enhances Fairness in Pathology Artificial Intelligence Systems. Cell Reports Medicine 6, 102527 (2025).
[59] Lu, M. Y., Williamson, D. F., Chen, T. Y., et al. Data-Efficient and Weakly Supervised Computational Pathology on Whole-Slide Images. Nature Biomedical Engineering 5, 555–570 (2021).
[60] Macenko, M., Niethammer, M., Marron, J. S., et al. A Method for Normalizing Histology Slides for Quantitative Analysis. In 2009 IEEE International Symposium on Biomedical Imaging: From Nano to Macro 1107–1110 (2009).
[61] Marquart, J., Chen, E. Y., & Prasad, V. Estimation of the Percentage of US Patients With Cancer Who Benefit From Genome-Driven Oncology. JAMA Oncology 4, 1093–1098 (2018).
[62] Martinsson, E. WTTE-RNN: Weibull Time to Event Recurrent Neural Network. In Chalmers University of Technology and University of Gothenburg (2016). Master’s thesis.
[63] Mehrabi, N., Morstatter, F., Saxena, N., et al. A Survey on Bias and Fairness in Machine Learning. ACM Computing Surveys 54, 1–35 (2021).
[64] Moor, M., Banerjee, O., Abad, Z. S. H., et al. Foundation Models for Generalist Medical Artificial Intelligence. Nature 616, 259–265 (2023).
[65] Mouret, J.-B. & Clune, J. Illuminating Search Spaces by Mapping Elites. arXiv preprint arXiv:1504.04909 (2015).
[66] Nelder, J. A. & Wedderburn, R. W. M. Generalized Linear Models. Journal of the Royal Statistical Society: Series A (General) 135, 370–384 (1972).
[67] Niazi, M. K. K., Parwani, A. V., & Gurcan, M. N. Digital Pathology and Artificial Intelligence. The Lancet Oncology 20, e253–e261 (2019).
[68] Novikov, A. et al. AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery. arXiv preprint arXiv:2506.13131 (2025).
[69] Obermeyer, Z., Powers, B., Vogeli, C., et al. Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations. Science 366, 447–453 (2019).
[70] Oquab, M., Darcet, T., Moutakanni, T., et al. DINOv2: Learning Robust Visual Features Without Supervision. Transactions on Machine Learning Research (2024).
[71] Rajkomar, A., Dean, J., & Kohane, I. Machine Learning in Medicine. New England Journal of Medicine 380, 1347–1358 (2019).
[72] Robins, J. M., Rotnitzky, A., & Zhao, L. P. Estimation of Regression Coefficients When Some Regressors Are Not Always Observed. Journal of the American Statistical Association 89, 846–866 (1994).
[73] Romano, J., Kromrey, J. D., Coraggio, J., et al. Appropriate Statistics for Ordinal Level Data: Should We Really Be Using t-Test and Cohen’s d for Evaluating Group Differences on the NSSE and Other Surveys? In Annual Meeting of the Florida Association of Institutional Research 1–33 (2006).
[74] Rombach, R., Blattmann, A., Lorenz, D., et al. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 10684–10695 (2022).
[75] Seyyed-Kalantari, L., Zhang, H., McDermott, M. B., et al. Underdiagnosis Bias of Artificial Intelligence Algorithms Applied to Chest Radiographs in Under-Served Patient Populations. Nature Medicine 27, 2176–2182 (2021).
[76] Shao, Z., Bian, H., Chen, Y., et al. TransMIL: Transformer Based Correlated Multiple Instance Learning for Whole Slide Image Classification. In Advances in Neural Information Processing Systems (NeurIPS) 2136–2147 (2021).
[77] Siegel, R. L., Miller, K. D., Fuchs, H. E., et al. Cancer Statistics, 2021. CA: A Cancer Journal for Clinicians 71, 7–33 (2021).
[78] Singhal, K., Azizi, S., Tu, T., et al. Large Language Models Encode Clinical Knowledge. Nature 620, 172–180 (2023).
[79] Singhal, K., Tu, T., Gottweis, J., et al. Toward expert-level medical question answering with large language models. Nature Medicine 31, 943–950 (2025).
[80] Srinidhi, C. L., Ciga, O., & Martel, A. L. Deep Neural Network Models for Computational Histopathology: A Survey. Medical Image Analysis 67, 101813 (2021).
[81] Stögbauer, F., Beck, S., Ourailidis, I., et al. Tumour Budding-Based Grading as Independent Prognostic Biomarker in HPV-Positive and HPV-Negative Head and Neck Cancer. British Journal of Cancer 128, 2295–2306 (2023).
[82] Su, F.-Y., Marostica, E., Wang, X., et al. Bridging the Gap: Translating AI in Pathology into Clinical Impact. NEJM AI (2026). In press.
[83] Sung, H., Ferlay, J., Siegel, R. L., et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA: A Cancer Journal for Clinicians 71, 209–249 (2021).
[84] Topol, E. J. High-Performance Medicine: The Convergence of Human and Artificial Intelligence. Nature Medicine 25, 44–56 (2019).
[85] Tsai, P.-C., Lee, T.-H., Kuo, K.-C., et al. Histopathology Images Predict Multi-Omics Aberrations and Prognoses in Colorectal Cancer Patients. Nature Communications 14, 2102 (2023).
[86] Uno, H., Cai, T., Pencina, M. J., D’Agostino, R. B., & Wei, L.-J. On the C-Statistics for Evaluating Overall Adequacy of Risk Prediction Procedures with Censored Survival Data. Statistics in Medicine 30, 1105–1117 (2011).
[87] Vaswani, A., Shazeer, N., Parmar, N., et al. Attention Is All You Need. In Advances in Neural Information Processing Systems (NeurIPS) 5998–6008 (2017).
[88] Vorontsov, E., Bozkurt, A., Casson, A., et al. A Foundation Model for Clinical-Grade Computational Pathology and Rare Cancers Detection. Nature Medicine 30, 2924–2935 (2024).
[89] Wang, L., Ma, C., Feng, X., et al. A Survey on Large Language Model Based Autonomous Agents. Frontiers of Computer Science 18, 186345 (2024).
[90] Wang, X., Zhao, J., Marostica, E., Yuan, W., Jin, J., Zhang, J., Li, R., Tang, H., Wang, K., Li, Y., Wang, F., Peng, Y., Zhu, J., Zhang, J., Jackson, C. R., Zhang, J., Dillon, D., Lin, N. U., Sholl, L., Denize, T., Meredith, D., Ligon, K. L., Signoretti, S., Ogino, S., Golden, J. A., Nasrallah, M. P., Han, X., Yang, S., & Yu, K.-H. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature 634, 970–978 (2024).
[91] Wei, J., Tay, Y., Bommasani, R., et al. Emergent Abilities of Large Language Models. Transactions on Machine Learning Research (2022).
[92] Wen, J., Qiu, L., Benton, J., Kirchner, J. H., & Leike, J. Automated Weak-to-Strong Researcher. Anthropic Alignment Science Blog, 2026.
[93] Xiang, J., Wang, X., et al. A Vision–Language Foundation Model for Precision Oncology. Nature 638, 769–778 (2025).
[94] Xu, H., Usuyama, N., Bagga, J., et al. A Whole-Slide Foundation Model for Digital Pathology from Real-World Data. Nature 630, 181–188 (2024).
[95] Zafar, M. B., Valera, I., Gomez Rodriguez, M., et al. Fairness Constraints: Mechanisms for Fair Classification. In Artificial Intelligence and Statistics (AISTATS) 962– 970 (2017).
[96] Zavala, V. A., Bracci, P. M., Carethers, J. M., et al. Cancer Health Disparities in Racial/Ethnic Minorities in the United States. British Journal of Cancer 124, 315–332 (2021).
[97] Zhang, B. H., Lemoine, B., & Mitchell, M. Mitigating Unwanted Biases with Adversarial Learning. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society 335–340 (2018).
[98] Zhou, X., Sun, L., He, D., et al. Knowledge-Enhanced Pretraining for Vision–Language Pathology Foundation Model on Cancer Diagnosis. Cancer Cell 44, 777–791.e7 (2026).
[99] Zimmermann, E. et al. Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology. arXiv preprint arXiv:2408.00738 (2024).