Obafunsho John Abiola, Surgical Data Institute, Department of Applied Health Sciences, University of Birmingham, Birmingham, UK

DOI: 10.1038/s41746-026-02788-y


Most people who read a surgical trial never read the full paper. They read the abstract, a few hundred words that are supposed to explain what was done, who was included, how randomisation was handled, what happened to patients, and whether the trial was registered. The problem is that surgical trial abstracts routinely leave much of that out. We already knew this from previous research. What we wanted to know was whether it could be fixed systematically and at scale using a large language model.

Our approach was to build a constrained GPT-4o pipeline that took the full text of a surgical randomised controlled trial and rewrote its abstract under strict rules: no fabrication, no inference beyond the source text, and structured output only. We then scored every abstract, original and rewritten, against a validated 14-item CONSORT-derived rubric with a maximum score of 25. In total, we analysed 651 open-access surgical randomised controlled trials indexed in PubMed between 2005 and 2025.

Original abstracts scored a mean of 9.1 out of 25, remaining consistently incomplete across two decades of published trials. The rewritten versions scored 16.5 using a 250-word format and 17.1 using a 300-word format. Every one of the 14 CONSORT reporting items improved. The largest gains were in randomisation and allocation concealment, harms reporting, and trial registration, precisely the elements that are most important for interpreting a trial and among the most frequently omitted. The 300-word format preserved approximately the same length as the original abstracts while containing substantially more useful information. Readability improved slightly as well, although that was a secondary finding.

Figure 1. Mean total CONSORT scores (maximum = 25) for original, 250-word, and 300-word rewritten abstracts, demonstrating significantly higher completeness in both rewritten formats

We deliberately designed the pipeline to minimise opportunities for hallucination. The model was constrained to the source text at every step. Where information was absent from the full paper, the corresponding rubric item scored zero. The model did not fill gaps; it reported them. We used deterministic model settings, logged every API call together with model version identifiers, and validated scoring against two independent expert reviewers before applying the pipeline across the full dataset. Agreement between model and human scoring was good, and reproducibility across three independent model runs was high.

None of that eliminates the fundamental limitation of language models, which is that they can produce fluent, confident, and plausible text that is wrong. We did not perform a line-by-line numerical audit of every rewritten abstract. A high completeness score means that the required reporting elements are present; it does not guarantee that every reported value is correct. That distinction matters, and it is why human oversight remains essential whenever tools like this are used.

Having built and tested the pipeline, my view is that LLMs are most useful here in the same way that a rigorous colleague is useful. They can identify what is missing, propose a structured revision, and do so consistently across hundreds of papers in the time it would take a person to read only a handful. They are not a replacement for the trial investigators, the editor who understands the clinical context, or the reviewer who checks the numbers. The right model is the one our findings support: LLM as assistant, human as accountable.

The finding that matters most to me is not the improvement in the overall completeness score. It is what happened to reporting of randomisation and allocation concealment. These are the details that allow readers to judge whether the results of a trial are trustworthy, yet they appeared in fewer than 15% of the original abstracts. A constrained pipeline increased that to nearly 60%. In global surgery, where clinicians and researchers may rely heavily on abstracts when the full paper is inaccessible, that gap is more than a methodological footnote. It has the potential to influence how evidence is interpreted and, ultimately, how clinical decisions are made.

Better reporting does not make a trial better. It does, however, make it easier for clinicians, researchers, reviewers, and guideline developers to understand what was actually done and what the findings really mean. If language models can help achieve that consistently while remaining transparent, constrained, and accountable to human oversight, then I think they deserve a place in the research workflow.

The full paper is published in npj Digital Medicine and is open access. The code and data are available on GitHub.

Conflict of interest statement: None declared.

Corresponding author: Professor Aneel Bhangu, Director, Surgical Data Institute, University of Birmingham, UK. a.a.bhangu@bham.ac.uk

Full paper: Abiola OJ, Nepogodiev D, Glasbey J, Omar O, Linder C, Ledda V, Bhangu A. Feasibility and impact of a large language model pipeline for surgical trial abstracts. npj Digital Medicine. 2026. doi: 10.1038/s41746-026-02788-y