Browse Articles

Discover research articles across all indexed journals

Dynamic time slot allocation method for deterministic communication in UAV formation

Scientific Reports Junfang Xiao, Jianming Huang, Qin Chen et al. Dec 03, 2025 DOI: 10.1038/s41598-025-30533-0

A structural framework for fire and explosion risk in EFRTs: empirical validation using EFA, CFA, and path analysis

Scientific Reports Parisa Moshashaei, Omid Akbarzadeh, Mohammad Asghari-Jafarabadi et al. Dec 03, 2025 DOI: 10.1038/s41598-025-30839-z

Effects of supplementation Senegalia macrostachya seed flour on the zootechnical performance of traditional chickens from Burkina Faso

Scientific Reports Hadidjatou Belem, Windmi Kagambega, Eliasse Zongo et al. Dec 03, 2025 DOI: 10.1038/s41598-025-29578-y

Improved multi-strategy secretary bird optimization for efficient IoT task scheduling in fog cloud computing

Scientific Reports K. Sangeetha, M. Kanthimathi Dec 03, 2025 DOI: 10.1038/s41598-025-30918-1

Study on variation law of rock mass dilatancy and stress threshold during loading and unloading

Scientific Reports Yanhong Du, Laigui Wang, Feng Chen et al. Dec 03, 2025 DOI: 10.1038/s41598-025-27012-x

Intelligent demand-side energy management via optimized ANFIS–gene expression programming in hybrid renewable–grid systems

Scientific Reports Noureddine Elboughdiri, Karim Kriaa, Mutiu Shola Bakare et al. Dec 03, 2025 DOI: 10.1038/s41598-025-26988-w

Benchmarking large language models on the United States medical licensing examination for clinical reasoning and medical licensing scenarios

Scientific Reports Md Kamrul Siam, Angel Varela, Md Jobair Hossain Faruk et al. Dec 03, 2025 DOI: 10.1038/s41598-025-31010-4

Abstract Artificial intelligence (AI) is transforming healthcare by assisting with intricate clinical reasoning and diagnosis. Recent research demonstrates that large language models (LLMs), such as ChatGPT and DeepSeek, possess considerable potential in medical comprehension. This study meticulously evaluates the clinical reasoning capabilities of four advanced LLMs, including ChatGPT, DeepSeek, Grok, and Qwen, utilizing the United States Medical Licensing Examination (USMLE) as a standard benchmark. We assess 376 publicly accessible USMLE sample exam questions (Step 1, Step 2 CK, Step 3) from the most recent booklet released in July 2023. We analyze model performance across four question categories: text-only, text with image, text with mathematical reasoning, and integrated text-image-mathematical reasoning and measure model accuracy at three USMLE steps. Our findings show that DeepSeek and ChatGPT consistently outperform Grok and Qwen, with DeepSeek reaching 93% on Step 2 CK. Error analysis revealed that universal failures were rare ( $$\le$$ 1.60%) and concentrated in multimodal and quantitative reasoning tasks, suggesting both ensemble potential and shared blind spots. Compared to the baseline ChatGPT-3.5 Turbo, newer models demonstrate substantial gains, though possible training-data exposure to USMLE content limits generalizability. Despite encouraging accuracy, models exhibited overconfidence and hallucinations, underscoring the need for human oversight. Limitations include reliance on sample questions, the small number of multimodal items, and lack of real-world datasets. Future work should expand benchmarks, integrate physician feedback, and improve reproducibility through shared prompts and configurations. Overall, these results highlight both the promise and the limitations of LLMs in medical testing: strong accuracy and complementarity, but persistent risks requiring innovation, benchmarking, and clinical oversight.

Correction: Predicting current and future habitat of Indian pangolin (Manis crassicaudata) under climate change

Scientific Reports Siddiqa Qasim, Tariq Mahmood, Bushra Allah Rakha et al. Dec 03, 2025 DOI: 10.1038/s41598-025-27761-9

Satellite swarms set to photobomb more than 95% of some telescopes’ images

Nature Jenna Ahart Dec 03, 2025 DOI: 10.1038/d41586-025-03953-1

Antimicrobial use, prescribing quality, and therapy-related problems among hospitalized patients in Northeast Ethiopia

Scientific Reports Mengistie Yirsaw Gobezie, Seble Zewdu, Nuhamin Alemayehu Tesfaye et al. Dec 03, 2025 DOI: 10.1038/s41598-025-30634-w

Evaluating intraoperative C2 slope as a radiographic guide for cervical deformity correction

Scientific Reports Namhoo Kim, Kyung-Soo Suk, Ji-Won Kwon et al. Dec 03, 2025 DOI: 10.1038/s41598-025-26914-0

Acceptance of disease as a mediator between social support and quality of life in women with Hashimoto’s disease—a pilot study

Scientific Reports Kamilla Bargiel-Matusiewicz, Magdalena Wnuk-Grzybowska, Natalia Ziółkowska et al. Dec 03, 2025 DOI: 10.1038/s41598-025-27479-8

Latent topic-driven cyber intelligence model for tactics, techniques, and procedures (TTPs) detection using hybrid framework and Birch-inspired optimisation

Scientific Reports Musaed Mutared Alanazi, Ainuddin Wahid Abdul Wahab, Mohd Yamani Idna Idris Dec 03, 2025 DOI: 10.1038/s41598-025-27451-6

Examining associations among caregiver stress, social support, and the infant gut microbiota

Scientific Reports Sarah C. Vogel, Francesca R. Querdasi, Bridget L. Callaghan et al. Dec 03, 2025 DOI: 10.1038/s41598-025-30553-w

Correction: A comparison among employees in Germany and Denmark of associations between quality of leadership and subsequent 5-year development of mental distress

Scientific Reports Hermann Burr, Norbert Kersten, Kathrine Sørensen et al. Dec 03, 2025 DOI: 10.1038/s41598-025-30371-0

Theoretical investigation of electrodialysis-driven salt ion transport in pillared graphene membranes

Scientific Reports Amin Hamed Mashhadzadeh, Maryam Zarghami Dehaghani, Salah A. Faroughi et al. Dec 03, 2025 DOI: 10.1038/s41598-025-26125-7

HyperGraph-based capsule temporal memory network for efficient and explainable diabetic retinopathy detection in retinal imaging

Scientific Reports Mishmala Sushith, N. Malligeswari, M. Anlin Sahaya Infant Tinu et al. Dec 03, 2025 DOI: 10.1038/s41598-025-30128-9

Application of the principle of collapsed distributions to the detection of faulty elements in square grid antenna arrays

Scientific Reports David Michael Parkinson-Oreiro, María Elena López-Martín, Juan Antonio Rodríguez-González et al. Dec 03, 2025 DOI: 10.1038/s41598-025-27379-x

China accounts for more than half of leading output in the applied sciences

Nature Benjamin Plackett Dec 03, 2025 DOI: 10.1038/d41586-025-03715-z

SMENN-hybrid: an efficient technique combining the synthetic minority oversampling technique with ensemble learning for diabetes prediction

Scientific Reports Essam H. Houssein, Ibrahim A. Ibrahim, Amir Mostafa et al. Dec 03, 2025 DOI: 10.1038/s41598-025-26583-z