
Clinical NLP
The social story in clinical notes
How do conventional and transformer models compare for SBDH classification?

TIANYI YU / AI FOR HEALTH
Building AI for the people behind the data.
01 / MEET ME
I'm Tianyi Yu. I build research software to understand biomedical language, learn from medical images and examine predictions in health.
A clinical note holds someone's story. An image can carry a clue that is easy to miss. A risk score may shape what happens next. Those connections are what draw me to AI for health.
I'm especially interested in what happens after a promising result: does it survive a different hospital, a new population or the messiness of everyday care? I share the code so others can explore those questions with me.
02 / RESEARCH DIRECTIONS
Different data. Connected questions.
What can we learn—and what should we trust?
Clinical and biomedical language sit at the centre of this direction, alongside reading and rationale benchmarks that help examine how language models behave.
Medical images, reconstruction and elastography-derived measurements: useful signals, with explicit limits on what each pipeline actually analyses.
A strong headline metric is only the beginning. These projects examine calibration, sample reuse, source information and fragile evaluation choices.
Predictions live in time. I study how clinical models behave when the population, hospital, endpoint or analysis specification changes.
How we collect and record data is part of the story. These workflows connect clinical measurement, population surveys, surveillance and access to care.
From molecular variation to sleep and movement, I build tools to explore health questions across different scales.
03 / EXPLORE THE CODE
Scientific code, individual project guides,
and a clear path into the methods.
52 published projects

Clinical NLP
How do conventional and transformer models compare for SBDH classification?

Computer Vision & Biomedical Imaging
How do calibration and abstention change model evaluation?


Prediction Modelling
How does model performance change across time and ICU databases?

Health Informatics
How do eating-related symptoms, affect and maladaptive exercise connect across EMA surveys?

AI for Health
How does adaptive smartwatch task acquisition compare with fixed acquisition?

Clinical NLP · Trustworthy AI
Does better classification imply better calibration and perturbation consistency?


Clinical NLP · AI for Health
Can lexical and early-fixation features predict later processing load?

Health Informatics · AI for Health
Can aggregate surveillance signals anticipate next-week escalation?

Prediction Modelling
Does model complexity improve probability estimation under temporal shift?

Prediction Modelling
How does endpoint construction affect cross-database transport?

Prediction Modelling
How do temporal renal models behave under external transport?

Clinical NLP
How do lightweight sentence classifiers transfer between corpora?

AI for Health
How do local context, channel choice and smoothing affect sleep staging?

Clinical NLP · Trustworthy AI
Do supporting rationales remain faithful and robust?

Health Informatics
How do healthcare access and glucose-testing gaps vary in public-use survey data?

AI for Health
How sensitive is glomerular BEX3 prioritisation to preprocessing choices?

Computer Vision & Biomedical Imaging · Prediction Modelling
How do temporal evaluation and sensitivity targets affect liver elastography triage?

Clinical NLP
How do reference choices and answer mapping affect clinical-trial inference?

Clinical NLP
How does source availability affect clinical NLP evaluation?

Health Informatics · Prediction Modelling
How do measurement patterns affect retrospective ICU prediction?

Prediction Modelling
How do shock anchors and dynamic predictors affect cross-database risk modelling?

Prediction Modelling
How does temporal landmark modelling transport across ICU databases?

Health Informatics
What can routinely collected intensive care data tell us about laboratory measurements six hours ahead?

Health Informatics
How does measurement timing and source availability shape retrospective ICU evaluation?

AI for Health · Trustworthy AI
How does patient reuse affect apparent transcriptomic diagnostic performance?

Trustworthy AI
What can synthetic validation establish when source information is limited?

Trustworthy AI · Prediction Modelling
How do analysis specifications change retrospective ICU evaluation?

Health Informatics · Prediction Modelling
How stable and clinically interpretable are kidney risk models?

Computer Vision & Biomedical Imaging
How does excision selection affect melanoma assessment?

AI for Health · Trustworthy AI
How do false positives and sample reuse affect pulmonary-nodule assessment?

Computer Vision & Biomedical Imaging
What can MRI and clinical information tell us about pathological complete response?

Health Informatics
When care costs too much, who misses a test?

Prediction Modelling
What does a model add beyond routine sodium measurements?

Prediction Modelling
Do disease-specific treatment patterns add useful mortality predictions?

Trustworthy AI
How do endpoint recording and observation timing change ICU model evaluation?

Trustworthy AI
Which clinical prediction conclusions can the available evidence identify?

Trustworthy AI
When does a shared stochastic observation preserve decision order?

AI for Health
Which sleep epochs deserve review when estimating total sleep time?

Computer Vision & Biomedical Imaging
Can stain-perturbation responses help screen pathology classification errors?

Computer Vision & Biomedical Imaging
How does sensitivity-map calibration carry added noise into MRI reconstruction?

Trustworthy AI
How should a fixed-target forecast change when sensor messages arrive late?

AI for Health
Can fixed calibration recognize perturbed ECG inputs without mistaking simulated labels for clinical quality?

Health Informatics
What changes when a clinical model can use only the information actually available at prediction time?

Health Informatics
When does a laboratory documentation count reflect recording rules rather than a distinct clinical measurement?

Trustworthy AI
Do clinical timing features add information after measurement counts are matched?

Prediction Modelling
How does kidney-event prediction change across clinical landmarks when follow-up and cohort eligibility remain explicit?

AI for Health · Trustworthy AI
Does an augmentation gain remain meaningful when absolute predictive performance falls under observation loss?

Trustworthy AI
Do missingness stress tests change source-only model selection, or do competing rules choose the same candidate?

AI for Health
What can objective sleep architecture reveal without turning physiological disruption into a psychiatric label?

Prediction Modelling
What changes when an AKI progression model moves between clinical databases and requires local updating?
No projects match that search. Try a different term or research direction.
04 / HOW I WORK
I care about the assumptions behind a result as much as the result itself. Sharing a pipeline makes room for someone else to notice a fragile choice, ask a better question or take the work somewhere new.
Each repository documents its inputs, environment and analysis steps. Some projects require credentialed clinical data or separately acquired datasets; their guides explain those boundaries. The illustrations here are conceptual artwork.
05 / CONTINUE THE CONVERSATION
A different dataset. A difficult assumption. A question worth exploring.
I'd love to hear what you find.