The trial registry and the FDA's adverse-event reports had never been joined. I built the ETL that joins them, on Google Cloud and Snowflake. The joined data scores every running drug trial for its risk of stopping early, with a reason.
118 GB of FDA safety reports (20.7M of them) and 605k registry records, joined into one row per trial.
It ranks a terminated trial above a completed one 71% of the time, on trials it never saw.
Reviewing its top 10% finds terminations at 2.1× the rate of picking at random.
Two streams that had never met, joined as of the day each trial started.
Right 71% of the time, on trials it never saw.
Slide down the ranking: each step catches more terminated trials (up) at the cost of flagging more that completed (across). The more the curve bows up-left, the better.
Pick one trial that was terminated and one that completed. 71% of the time the model scores the terminated one higher. A coin flip gets 50%.
Trained on trials started 2008–14; tested on 10,242 from 2015–16. A model that scored 0.714 was rejected because four of its fields leaked the outcome.
Most running trials sit low and left. The ones to open first stand out.
Immunotherapy Using Tumor Infiltrating Lymphocytes for Patients With Metastatic Cancer
National Cancer Institute (NCI) · started 2010-08-26
Reasons explain the score, not why a trial would stop. Snapshot of model v6, 2026-10-07.
Stack
Google Cloud: Cloud Run Jobs, Cloud Storage, Dataproc Serverless (PySpark), Cloud Run. Snowflake: external stages, SQL layers, zero-copy clones. LightGBM · SHAP · FastAPI · GitHub Actions
A note on what's shown
The trials are a frozen snapshot of the v6 lookup published 2026-10-07. The live lookup ↗ searches all 111,118 trials; the code ↗ is public.