University of Edinburgh research group

Researching the language of healthcare.

We develop language technologies, health intelligence and privacy science to make information in healthcare narratives useful for rigorous, trustworthy research.

Why narratives?

Rich health information is written, not only coded.

Clinical notes, reports and letters capture symptoms, decisions, responses to care and the context around a person's healthcare journey. They can transform research, but they are sensitive, complex and created within a relationship of trust.

We ask two questions together: what can healthcare narratives tell us? And how can that information be used safely and responsibly?
Our research

Three connected themes.

We combine computational methods with clinical validation, privacy assessment, governance, and patient and public involvement.

01

Clinical text analytics

Developing and evaluating NLP, foundation models and information extraction methods for clinically meaningful information.

02

Health intelligence

Turning narratives into research-ready data for epidemiology, population health, healthcare analytics and translational research.

03

Privacy & responsible AI

Building evidence-based approaches to privacy risk, Trusted Research Environments and responsible clinical language models.

Featured work

Research with real-world purpose.

Our projects span discovery science, public health, cancer informatics and the infrastructure needed to work safely with sensitive narratives.

Wellcome · Clinical phenotyping

AMBER

Deriving robust measures of antidepressant exposure and response from GP and hospital narratives.

Programme PI: Cathryn Lewis · Clinical text informatics led by our group
CSO · Cancer informatics

Can AI tell the story of cancer?

Extracting disease, treatment and biomarker information to build longitudinal patient knowledge graphs.

Clinical Fellow: Sam McInerney · Supervisors: Peter Hall, Arlene Casey, David Lowe and Kathryn Cresswell
MRC · Privacy science

STAR-TRE

Understanding contextual privacy risk and developing safe access approaches for free text.

Principal Investigator: Arlene Casey
MRC · Responsible AI

TransPECT

Assessing when language models trained on sensitive healthcare data can be safely released.

PI: Arlene Casey · Co-Investigators: Pasquale Minervini and Richard Walls
Embedded collaboration
DataLoch

Turning research methods into sustainable capability

Our group works closely with DataLoch, connecting University of Edinburgh research with the secure data, governance and Trusted Research Environment expertise needed to make new approaches usable in practice.

The collaboration supports safe access to clinical free text, privacy-risk assessment, NLP-derived research variables and responsible model development.

Explore the DataLoch NLP programme ↗

Patient & public involvement

Public perspectives shape how we do research.

Technical performance alone cannot answer questions about sensitive healthcare narratives. Patients and members of the public help us understand what feels sensitive, which safeguards are expected and where human judgement should remain central.

SARA · DARE UK

Embedding public views into privacy-risk tools

Workshops and wider consultation explored privacy risk in clinical free text and data provenance. Public involvement informed the project methods and reinforced the importance of human oversight.

Read the project report ↗
STAR-TRE · UK-wide

Public expectations of AI-assisted de-identification

Our current work explores public perspectives on using AI and language models to identify privacy risks in sensitive free text and inform responsible access within secure environments.

Follow the work through DataLoch ↗
Papers & outputs

Sharing our research and ideas.

We publish peer-reviewed papers, preprints and accessible perspectives on our work in progress.

Our approach

Validation, context and trust are part of the science.

Healthcare text contains technical information alongside uncertainty, personal circumstances and highly sensitive details. Robust clinical language informatics needs interdisciplinary methods from the outset.

ListenStart with clinical, patient and public priorities.
DevelopBuild methods around meaningful research questions.
ValidateTest performance, generalisability and failure modes.
ProtectAssess risk and apply proportionate governance.
Who we are

An interdisciplinary group by design.

Our team brings together expertise in clinical NLP, health data science, machine learning, privacy, governance and public involvement.

Arlene Casey

Arlene Casey

Group Leader
Vivensa Senior Research Fellow
Strategic and Operational NLP Lead, DataLoch

STAR-TRE · TransPECT · AMBER · Senior Proleptic Fellowship
Franz Gruber

Franz Gruber

NLP Research Fellow

STAR-TRE · TransPECT
Judit Kuti

Judit Kuti

NLP Research Fellow

STAR-TRE · TransPECT
Fahrurrozi Rahman

Fahrurrozi Rahman

NLP Research Fellow

STAR-TRE · TransPECT
Matúš Falis

Matúš Falis

NLP Research Fellow

AMBER
Sam McInerney

Sam McInerney

Clinical Fellow · Cancer clinical language informatics

Can AI Tell the Story of Cancer?

Our projects are also shaped by clinical collaborators, data specialists, information governance experts, and patient and public contributors across universities, the NHS and Trusted Research Environments.

Collaborate

Let’s make healthcare narratives useful—and use them responsibly.

We welcome conversations with researchers, clinicians, public partners and organisations working on trustworthy health data research.

Get in touch