Abstract Background The Montreal classification is fundamental for phenotyping inflammatory bowel disease (IBD), yet these descriptors are rarely structured in routine colonoscopy reports. Manual extraction is labor-intensive, and keyword-based natural language processing (NLP) often fails in complex narrative text. We developed and validated a hybrid NLP–large language model (LLM) framework to automatically extract Montreal categories from free-text reports and applied it to a large real-world cohort. Methods Colonoscopy narratives from IBD patients (2020–2025) at a tertiary center were analyzed. A 300-report test set (150 ulcerative colitis UC, 150 Crohn’s disease CD) was manually annotated as gold standard. A rule-based NLP pipeline identified segmental pathology descriptors and mapped them to Montreal extent (E1–E3) and location (L1–L3). Incomplete examinations were coded as E2* or L2* to denote left-sided colitis or colonic CD in which the proximal colon could not be fully intubated. An LLM was prompted to infer Montreal classification directly from the narrative. Accuracy, macro F1-score, and Cohen’s κ were calculated. Distributional agreement between NLP and LLM outputs was evaluated using Jensen–Shannon (JS) similarity. The validated NLP pipeline was applied to 2,150 real-world reports. Results In the test set, the LLM achieved 74% accuracy, macro F1 0.76, and κ = 0.71, with expected adjacent-class drift (E2↔E3; L1/L2↔L3). In the real-world cohort, ulcerative colitis (UC, n = 844) exhibited the following distribution of Montreal extent: E1 9.6%, E2 27.4%, E2* 5.8%, E3 33.5%, and unknown (UNK) 23.7%. Crohn’s disease (CD, n = 1,306) showed the following Montreal location distribution: L1 26.7%, L2 11.7%, L2* 7.4%, L3 16.4%, and UNK 37.8%. Overall, Montreal categories were classifiable in 67.7% of reports; remaining UNK cases reflected normal mucosa or insufficient segmental information. Cross-model comparison demonstrated strong distributional agreement between LLM and NLP outputs (JS similarity 0.82), indicating convergent Montreal phenotyping across independent computational methods. Conclusion A hybrid NLP–LLM approach enables accurate and scalable extraction of the Montreal classification from routine colonoscopy narratives. This framework supports automated real-world IBD phenotyping, accelerates digital registry development, and provides a foundation for large-scale phenotypic research without relying on structured reporting. Conflict of interest: Dr. Agargun, Besim Fazil: No conflict of interest Genc Ulucecen, Sezen: No conflict of interest Dağcı, Gizem: No conflict of interest Yağlı, Mehmet Akif: No conflict of interest Gurbanov, Asım: No conflict of interest Telli, Pelin: No conflict of interest Mammadov, Sabuhi: No conflict of interest Bilgin, Ersel: No conflict of interest Cavus, Bilger: No conflict of interest Ormeci Çifcibaşi, Asli: No conflict of interest Demir, Kadir: No conflict of interest Besisik, Fatih: No conflict of interest Kaymakoglu, Sabahattin: No conflict of interest Akyüz, Filiz: No conflict of interest
Agargun et al. (Thu,) studied this question.