Loading
Fetching the latest research
Retrospective study
We scored the early publication of 166 drug assets whose Phase 3 result is now known, blind to that result, and tested every rigor criterion against later success or failure.
How to read this. Retrospective and hypothesis-generating. In-sample association, not a validated predictor.
Any model can grade a paper. The harder question is whether that grade means anything about what happens next. So we ran a clean test: score the early publication of 166 drug assets whose Phase 3 result is now known, blind to that result, and check every rigor criterion against later success or failure.
The short answer
We assembled a cohort of drug assets that each have two things: a pivotal early-phase publication, and a later Phase 3 read-out that is now public. The assets come from ClinicalTrials.gov Phase 1/2 to Phase 3 trajectories, built through the v2 API. Starting from 290 candidate trajectories, we de-duplicated correlated assets from the same program (so one drug is never counted twice) down to 224, then curated a benchmark cohort of 166.
Of the 166, 60 went on to succeed in Phase 3 and 106 went on to fail. The outcome was adjudicated independently of, and after, the rigor review: 83 are trial terminations read from the registry’s stated reason for stopping, and the remaining outcomes were adjudicated from the completed-trial results. The review engine itself saw only the early paper, frozen at the decision point, with no access to anything downstream.
Every early paper was scored by the real Adcurare rigor engine in full audit mode, nothing skipped, at a single pinned engine version. That is the same eight-dimension taxonomy and 1-to-5-star deduction scale the product ships, so nothing here is a special-purpose model built to win this benchmark.
Assets scored
60 later succeeded, 106 later failed
Rigor dimensions
The engine's predefined taxonomy, fixed before the cohort
To the outcome
The engine saw only the early paper, never the Phase 3 result
Read purely as a grader of evidence packages, the corpus is unremarkable and fairly typical of the clinical literature. The mean rating was 2.6 of 5 stars. Split into bands, 47 papers landed in the top half (Strong or Exemplary is uncommon), most sat in the middle, and 45 drew a critical rating that flags the paper for priority expert review. On the pass / concerns / fail verdict, the large majority carried at least one candidate concern.
When you line the grade up against what actually happened in Phase 3, the two barely separate. On the deterministic rigor score, the assets that later succeeded averaged 81 out of 100; the ones that later failed averaged 79. That gap is not distinguishable from noise. Expressed as discrimination, where 0.5 is a coin flip and 1.0 is perfect, the area under the ROC curve sits near 0.57 no matter which headline number you use.
AUC · rigor score
0.5 = chance (p = 0.15)
AUC · gated composite
Not significant (p = 0.34)
AUC · star rating
The product headline (p = 0.13)
The pass / concerns / fail verdict tells the same story. A paper the engine flagged was only 1.2 times as likely to belong to a later failure as a clean one, and that interval comfortably includes no effect (odds ratio 1.24, 95% CI 0.58 to 2.64). A single overall rigor grade, on its own, does not sort the winners from the losers.
The dominant driver of Phase 3 outcome in this cohort is not rigor at all. It is the therapeutic area. Oncology and neuropsychiatry assets failed the large majority of the time; cardiometabolic and renal assets succeeded far more often than they failed. Once you know the indication, you have already explained most of the apparent signal.
The proof is in the adjustment. The marginal discrimination of 0.57 falls to 0.54 once you stratify by indication, which means the small amount of separation the overall score appeared to have was largely indication mix, not rigor.
The overall score is a blunt instrument, but the criteria underneath it are not all equal. Testing each of the eight dimensions individually against later failure, one stands out and holds up after correcting for the fact that we tested many: biological variables, the dimension covering whether sex, age, strain, and other biological sources of variation are handled properly. A paper the engine faulted on that dimension was about three times as likely to belong to a later failure.
Odds of a later trial failure, per rigor criterion
Each row is one rigor criterion, plotted as its odds ratio for a later trial failure with a 95% interval, against the no-effect line at 1. We show the effect sizes but not the row-by-row labels: the mapping is the study’s finding, not a teaser. Four criteria clear 1; two are kept in deliberately because they show no association. After Benjamini-Hochberg correction, only the biological-variables dimension reaches q < 0.05.
Biological variables
Odds of later failure when flagged (95% CI 1.4-7.3)
Indication-adjusted
Mantel-Haenszel odds ratio, still elevated
FDR q-value
The only dimension clearing q < 0.05
A couple of finer-grained findings point the same way without clearing the bar: an overstated claim relative to the evidence (odds ratio ~4.4) and a partially-met biological-variables check (~4.9) each land at q ≈ 0.10. We report them honestly as leads, not confirmed effects.
We re-ran the discrimination inside five different slices of the cohort, to check the signal is not an artifact of one subgroup. Restricting to the highest-confidence marquee tier, to sponsor-matched early papers, to confirmed-full-text papers, and to failures that are outright trial terminations, the area under the curve stays in a narrow 0.53-to-0.62 band. The direction of effect is stable; the intervals stay wide. That is exactly the pattern of a real-but-small lead, not a decisive predictor.
The fastest way to judge the work is on an asset of your own. Request a demo and we’ll center it on a real BD or R&D decision you bring.