Evaluating large language models in biomedical data science challenges through a classroom experiment
Abstract
Large language models (LLMs) have shown remarkable capabilities in algorithm design, but their effectiveness in solving data science challenges in real-world settings remains poorly understood. We conducted a classroom experiment in which graduate students used LLMs to solve biomedical data science challenges on Kaggle, focusing on tabular data prediction. While their submissions did not top the leaderboards, their prediction scores were often close to those of leading human participants. LLMs frequently recommended gradient boosting methods, which were associated with better performance. Among prompting strategies, self-refinement, where the LLM improves its own initial solution, was the most effective, a result validated using additional LLMs. While LLMs are capable of handling more complex data science tasks beyond tabular data prediction, their performance is substantially worse. These findings demonstrate that LLMs have the potential to design competitive machine learning solutions, even when used by nonexperts.
Article Details
Journal Info
Proceedings of the National Academy of Sciences
National Academy of Sciences
Authors (35)
Huifang Ma
College of Chemistry and Materials and National Engineering Research Centre for Carbohydrate Synthesis
Zhicheng Ji
Department of Biostatistics and Bioinformatics
Tara Al-Hashimy
Austin Allen
Nan Cen
Orlando Chen
Yongyin Chen
Yutian Chen
Tong Cheng
Yueqi Gu
Beijie Ji
Xiaohui Jiang
Fengnan Li
Peiyu Li
Yueshan Liang
Bena Liu
Coco Liu
Elisa Ma
Zhicheng Ma
Vicky Shao
Mengyao Shi
Jiang Shu
Leyi Sun
Rushi Tang
Hanyu Wang
Vivian Wang
Yuxin Wang
Department of Chemistry
Krissie Wilson
Ruobing Xue
Tianyi Yang
Alison Yu
Allison Yuan
Haiqi Zhang
Vera Zhang
Yinuo Zhang