RT Journal Article SR Electronic T1 Using large language models to identify prediagnostic clinical features of ovarian cancer from healthcare records: a population-based case–control study JF British Journal of General Practice JO Br J Gen Pract FD British Journal of General Practice SP e544 OP e551 DO 10.3399/BJGP.2025.0366 VO 76 IS 768 A1 Funston, Garth A1 Park, Namu A1 Thompson, Matthew A1 Yetisgen, Meliha A1 Goff, Barbara A A1 Kessler, Larry A1 Walter, Fiona M YR 2026 UL http://bjgp.org/content/76/768/e544.abstract AB Background Most women with ovarian cancer are diagnosed after developing symptoms. However, symptoms are often recorded as free text within electronic health records (EHRs), which is not readily accessible for research.Aim To use EHRs to examine associations between coded and large language model (LLM)-extracted free-text clinical features with ovarian cancer diagnosis.Design and setting Population-based case–control study using EHRs and cancer registry data from women attending primary care, outpatient, and emergency clinics associated with the University of Washington, US.Method In total, 136 women with ovarian cancer cases (diagnosed 2012–2019) were matched (age, clinic type) to 1360 control participants. Twelve months of prediagnosis coded and free-text data were extracted from EHRs. LLMs were tested on annotated notes, before extracting information on 17 prespecified clinical features. Univariate conditional logistic regression analyses were used to identify clinical features associated with ovarian cancer.Results There were 14 clinical features that were more commonly identified from free text using LLMs than from codes in both the case and control groups. There were 14 features that were significantly associated with ovarian cancer when using codes and LLM-extracted data, but only eight features were significant using codes alone. Using both coded and LLM-extracted data, 11 features had odds ratios >2. Thirteen features were significantly associated when restricting analysis to early-stage (I–II) diagnosis.Conclusion There is an identifiable ovarian cancer symptom signature within EHRs, with LLM-based natural language processing approaches enabling extraction of key non-coded symptom information. LLMs could support EHR-based research, while LLM-based clinical decision support tools may improve identification of patients with symptoms of possible cancer.