Fine-tuning and structured prompting strategies for question answering over full-text biomedical research articles

K Kaiming Tao R Rohit Satija J Jinru Zhou Z Zachary A. Osman V Vineet Ahluwalia C Chiara Sabatti R Robert W. Shafer

Abstract

Objectives The ability of large language models (LLMs) to answer targeted scientific questions by synthesizing information from research articles remains an open research challenge. Methods We evaluated the effects of fine-tuning and a question-specific prompting strategy to answer 16 pre-defined questions about HIV drug resistance studies, including whether viral genetic sequences were reported and the demographics and antiviral treatments of the individuals from whom sequences were obtained. For fine-tuning, we constructed an instruction set comprising 250 HIV drug resistance studies, with 16 questions per study and corresponding answers and explanations. For question-specific prompting, we developed a set of if-then rules tailored to each question. We compared the performance of three base models – GPT-4o-mini-2024-07-18 (GPT-4o), Meta Llama-3.1-70B-Instruct (Llama-3.1-70B), and Meta Llama-3.1-8B-Instruct (Llama-3.1-8B) – with their performance using fine-tuning, question specific prompting, and fine-tuning followed by question-specific prompting. Performance was assessed using accuracy, precision, recall, and F1 score, averaged over 150 held-out studies not used for fine-tuning. Comparisons were performed using Wilcoxon signed-rank tests. Results Fine-tuning increased precision by 5% for GPT-4o, 16% for Llama-3.1-70B, and 8% for Llama-3.1-8B, although this increase reached statistical significance only for Llama-3.1-70B. Fine-tuning also significantly increased recall for GPT-4o by 11%. Question specific prompting increased recall for all three models (6% for GPT-4o, 7% for Llama-3.1-70B, and 18% for Llama-3.1-8B), with statistically significant improvements observed only for Llama-3.1-8B. Applying question specific prompting to each of the fine-tuned models did not yield additional improvements beyond fine-tuning alone. When pooled across the three models, fine-tuning was associated with a greater effect on precision than recall (OR = 4.35; p = 0.001; Fisher’s exact test), whereas question-specific prompting led to a greater effect on recall than on precision (OR= 7.09; p = 0.0001; Fisher’s exact test). Conclusions In this domain-focused proof-of-concept study, fine-tuning and question-specific prompting each led to improvement in one or more metrics for each of the three models. Pooled analyses indicated that fine-tuning improved precision, whereas question specific prompting preferentially improved recall.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 6
Published June 24, 2026
Pages e0351631
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (7)

K

Kaiming Tao

R

Rohit Satija

J

Jinru Zhou

Z

Zachary A. Osman

V

Vineet Ahluwalia

C

Chiara Sabatti

R

Robert W. Shafer