An LLM chatbot to facilitate primary-to-specialist care transitions: a randomized controlled trial.
Tao X, Zhou S, Ding K, Li S, Li Y, Wu B, Huang Q, Chen W, Shen M, Meng E, Chen X, Hu H, Zhang J, Zhou J, Zou L, Ma L, Han S
Paper source
An LLM chatbot to facilitate primary-to-specialist care transitions: a randomized controlled trial.
This rigor review was examined and confirmed by Adcurare Editorial · July 7, 2026
How this rating was calculated▸
- IntegrityIntegrity concern ×3−1.5★
- ReportingData & code availability partially met−0.25★
- No reported statistical tests were found to recompute.
- Data/code availability incomplete
- Methods and results do not match
- Implausibly large reported effect
- Internal contradictions in the reported numbers
This Adcurare Rigor Review uses AI Rigor Reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The manuscript describes a methodologically rigorous pragmatic RCT of an LLM chatbot for care transitions, with strong reporting on most rigor dimensions. The main weaknesses are in data/code availability (vague data access statement, proprietary chatbot code) and minor statistical reporting gaps (normality test method not named, some p-values as inequalities).
This evaluation synthesizes three independent reviewer assessments and a copyedit pass. Of the eight rigor dimensions, three are fully applicable (scientific premise, study design, ethical approvals, biological variables, key resources, statistical analysis, data code availability, reporting transparency). The statistics verification component could not independently recompute any reported test statistic due to the format of reported results; the statistical analysis dimension's status reflects reporting adequacy, not mathematical verification against raw data.
12 major claims checked against the paper's own evidence: all adequately supported.
3 integrity concerns flagged (0 high).
5 copyedit issues flagged: mostly consistency, punctuation, typo.
Checked 57 references: 50 verified — 7 not checked.
2 data/code links checked; 2 live.
Registered (1 ID: Chinese Clinical Trial Registry). Reporting guideline cited: CONSORT.
Ready after minor pre-submission edits. The manuscript is in strong shape for submission, but should address the data availability statement (use a managed-access platform), specify the exact GPT model version, name the normality test used, and fix the F1 formula typo and minor formatting issues before submission.
- 1.HIGHdata codeReplace the email-based data access statement with a link to a managed-access platform (e.g., YODA, Vivli, or a study-specific data use committee) that provides a structured application process and approval timeline.Journal reviewers and editors increasingly expect data access to go through a formal platform rather than a simple email request, and the current statement is vague ('available on request').
- 2.HIGHstatisticsIn Methods, Statistical analysis, name the normality test used (e.g., Shapiro-Wilk or Kolmogorov-Smirnov) and state the threshold (e.g., P > 0.05) used to decide between parametric and non-parametric tests.Two reviewers flagged this as a reporting gap; naming the test strengthens reproducibility and statistical rigor.
- 3.HIGHreportingIn Methods, Model development, specify the exact OpenAI model version identifier (e.g., gpt-4o-mini-2024-07-18) rather than only the generic 'GPT-4.0 mini'.The investigational product is the key resource; knowing the precise model version enables reproducibility and version-specific safety/performance assessment.
- 4.MEDIUMstatisticsReplace all inequality p-values (e.g., P < 0.001) with exact values (e.g., P = 0.0003) where possible, or at least for the three primary comparisons.Exact p-values are preferred in clinical trial reporting for transparency and meta-analysis.
- 5.MEDIUMreportingIn Methods, add a sentence confirming that a completed CONSORT checklist was submitted as supplementary material (if true) or add the checklist fields to the Reporting Summary.One reviewer noted the lack of explicit mention of a CONSORT checklist submission in the paper text, though a Nature Portfolio Reporting Summary is linked.
- 6.MEDIUMreportingIn Methods, add a brief subsection on outlier handling and sensitivity analyses (e.g., winsorization, per-protocol vs ITT analysis, or a statement that no data points were excluded).All three reviewers noted outlier_handling as inadequate or not reported; this is a standard section in clinical trial methods.
- 7.MEDIUMreportingIn the Discussion, acknowledge the potential performance bias introduced by the single-blind design (patients unmasked) and discuss its possible impact on subjective outcomes like patient experience.The single-blind design is already described, but the Discussion does not explicitly address the risk of bias from unblinded patients on patient-reported outcomes.
- 8.LOWcopyeditFix the F1 formula in Methods, Statistical analysis: change 'F1 Score=TP/(2TP + FP + FN)' to 'F1 = 2TP/(2TP + FP + FN)' to match the standard harmonic mean definition.The copyedit pass identified a typo in the formula that could cause confusion; the standard F1 score is 2*TP/(2TP+FP+FN).
- 9.LOWcopyeditRemove the space before the final period in: 'identifier: ChiCTR2400094159 (https://...).' and remove spaces inside parentheses around the GitHub URL.Typographical inconsistencies flagged in the copyedit pass; minor but improves professional appearance.
- 10.LOWdata codeIf possible, deposit the anonymous non-dialogue dataset (e.g., aggregated data underlying Table 1 and main figures) in a repository like Zenodo with a DOI, and update the Data Availability section accordingly.Although individual-level patient data cannot be shared, aggregate data supporting the figures could be deposited to strengthen reproducibility.
- 11.LOWotherIn the Discussion, briefly clarify that the total patient time for PreA-only (3.51 min PreA interaction + 3.14 min physician consultation = 6.65 min) is longer than the No-PreA consultation (4.41 min), to prevent misinterpretation of 'consultation duration'.The integrity verification component noted a potential clarity issue where readers might conflate consultation duration with total patient time; a brief clarification would preempt confusion.
Adcurare assesses methodological rigor, not the importance of the findings. See how we evaluate →