Evaluating large language models in biomedical data science challenges through a classroom experiment

H Huifang Ma (College of Chemistry and Materials and National Engineering Research Centre for Carbohydrate Synthesis) Z Zhicheng Ji (Department of Biostatistics and Bioinformatics) T Tara Al-Hashimy A Austin Allen N Nan Cen O Orlando Chen Y Yongyin Chen Y Yutian Chen T Tong Cheng Y Yueqi Gu B Beijie Ji X Xiaohui Jiang F Fengnan Li P Peiyu Li Y Yueshan Liang B Bena Liu C Coco Liu E Elisa Ma Z Zhicheng Ma V Vicky Shao M Mengyao Shi J Jiang Shu L Leyi Sun R Rushi Tang H Hanyu Wang V Vivian Wang Y Yuxin Wang (Department of Chemistry) K Krissie Wilson R Ruobing Xue T Tianyi Yang A Alison Yu A Allison Yuan H Haiqi Zhang V Vera Zhang Y Yinuo Zhang

Abstract

Large language models (LLMs) have shown remarkable capabilities in algorithm design, but their effectiveness in solving data science challenges in real-world settings remains poorly understood. We conducted a classroom experiment in which graduate students used LLMs to solve biomedical data science challenges on Kaggle, focusing on tabular data prediction. While their submissions did not top the leaderboards, their prediction scores were often close to those of leading human participants. LLMs frequently recommended gradient boosting methods, which were associated with better performance. Among prompting strategies, self-refinement, where the LLM improves its own initial solution, was the most effective, a result validated using additional LLMs. While LLMs are capable of handling more complex data science tasks beyond tabular data prediction, their performance is substantially worse. These findings demonstrate that LLMs have the potential to design competitive machine learning solutions, even when used by nonexperts.

Article Details

Volume / Issue Vol. 122, Issue 50
Published December 16, 2025
ISSN 0027-8424
Publisher National Academy of Sciences

Authors (35)

H

Huifang Ma

College of Chemistry and Materials and National Engineering Research Centre for Carbohydrate Synthesis

Z

Zhicheng Ji

Department of Biostatistics and Bioinformatics

T

Tara Al-Hashimy

A

Austin Allen

N

Nan Cen

O

Orlando Chen

Y

Yongyin Chen

Y

Yutian Chen

T

Tong Cheng

Y

Yueqi Gu

B

Beijie Ji

X

Xiaohui Jiang

F

Fengnan Li

P

Peiyu Li

Y

Yueshan Liang

B

Bena Liu

C

Coco Liu

E

Elisa Ma

Z

Zhicheng Ma

V

Vicky Shao

M

Mengyao Shi

J

Jiang Shu

L

Leyi Sun

R

Rushi Tang

H

Hanyu Wang

V

Vivian Wang

Y

Yuxin Wang

Department of Chemistry

K

Krissie Wilson

R

Ruobing Xue

T

Tianyi Yang

A

Alison Yu

A

Allison Yuan

H

Haiqi Zhang

V

Vera Zhang

Y

Yinuo Zhang