An LLM chatbot to facilitate primary-to-specialist care transitions: a randomized controlled trial.
Tao X, Zhou S, Ding K, Li S, Li Y, Wu B, Huang Q, Chen W, Shen M, Meng E, Chen X, Hu H, Zhang J, Zhou J, Zou L, Ma L, Han S
Loading
Fetching the latest research
Randomized clinical trials in AI- and software-driven care (AI imaging triage, LLM assistants, clinical decision support, assistive robotics), from NEJM · Lancet · Nature Medicine · BMJ (2024–2026)
Every paper starts at 5★ and loses stars for the concrete problems the review finds — never a vague average, just a running total:
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own.
Papers are reviewed on their original full text — any retraction notice is removed before analysis — so the findings reflect the science, not the label.
6 papers
Tao X, Zhou S, Ding K, Li S, Li Y, Wu B, Huang Q, Chen W, Shen M, Meng E, Chen X, Hu H, Zhang J, Zhou J, Zou L, Ma L, Han S
Elías-Cabot E, Romero-Martín S, Raya-Povedano JL, Rodríguez-Ruiz A, Álvarez-Benito M
Williams AJ, Rhodes CA, Cleare S, Borschmann R, Gross JJ, Petrova K, Posada L, Tench CR, Chapman-Nisar A, Martin L, Hollis C, Townsend E, Slovak P, Digital Youth research team
Zhang X, Ding L, Jing J, Wang C, Gu H, Jiang Y, Meng X, Liu T, Xie X, Xu M, Hu M, Zhang Y, Fu H, Liu P, Du C, Du K, Wang M, Li H, Gong X, Dong K, Xiong Y, Wang Y, Liu L, Zhang Z, Zang Y, Yang C, Xian Y, Peterson E, Fonarow GC, Schwamm LH, Zhao X, Wang Y, Li Z, GOLDEN BRIDGE II Investigators
Bean AM, Payne RE, Parsons G, Kirk HR, Ciro J, Mosquera-Gómez R, Hincapié M S, Ekanayaka AS, Tarassenko L, Rocher L, Mahdi A
Woznitza N, Smith L, Rawlinson J, Au-Yong I, George B, Djearaman MG, Nair A, Lee RW, Navani N, Ndwandwe S, Clarke CS, Creeden A, Newsome J, Das I, Abaokporo S, Tucker R, Hathorn J, Baldwin DR
How to read this. The engine surfaces specific, verifiable rigor problems. It does notpredict retractions, and a “pass” is not a guarantee of soundness. Papers are analyzed on their published full text with any retraction notice removed, so detection is content-based. Real-world outcome labels (retracted, landmark) are curator-assigned from the public record; the verdict and every surfaced problem are the engine’s own. These are selected examples.