An information-matching approach to optimal experimental design and active learning

Y Yonatan Kurniawan (Brigham Young University 1 , Provo, Utah 84602,) T Tracianne B. Neilsen (Brigham Young University 1 , Provo, Utah 84602,) B Benjamin L. Francis (Achilles Heel Technologies 2 , Orem, Utah 84097,) A Alex M. Stankovic (SLAC National Accelerator Laboratory 3 , Menlo Park, California 94025,) M Mingjian Wen (Institute of Fundamental and Frontier Sciences University of Electronic Science and Technology of China Chengdu P. R. China) I Ilia Nikiforov (University of Minnesota 5 , Minneapolis, Minnesota 55455,) E Ellad B. Tadmor (University of Minnesota 5 , Minneapolis, Minnesota 55455,) V Vasily V. Bulatov (Lawrence Livermore National Laboratory 6 , Livermore, California 94550,) V Vincenzo Lordi M Mark K. Transtrum (Brigham Young University 1 , Provo, Utah 84602,)

Abstract

The efficacy of mathematical models heavily depends on the quality of the training data, yet collecting sufficient data is often expensive and challenging. Many modeling applications require inferring parameters only as a means to predict other quantities of interest (QoI). Because models often contain many unidentifiable (sloppy) parameters, QoIs often depend on a relatively small number of parameter combinations. Therefore, we introduce an information-matching criterion based on the Fisher information matrix to select the most informative training data from a candidate pool. This method ensures that the selected data contain sufficient information to learn only those parameters that are needed to constrain downstream QoIs. It is formulated as a convex optimization problem, making it scalable to large models and datasets. We demonstrate the effectiveness of this approach across various modeling problems in diverse scientific fields, including power systems and underwater acoustics. Finally, we use information-matching as a query function within an active learning (AL) loop for materials science applications. In all these applications, we find that a relatively small set of optimal training data can provide the necessary information for achieving precise predictions. These results are encouraging for diverse future applications, particularly AL in large machine-learning models.

Article Details

Volume / Issue Vol. 128, Issue 6
Published February 09, 2026
ISSN 0003-6951
Publisher American Institute of Physics

Journal Info

Applied Physics Letters

American Institute of Physics

ISSN: 0003-6951 Physical Sciences

Authors (10)

Y

Yonatan Kurniawan

Brigham Young University 1 , Provo, Utah 84602,

T

Tracianne B. Neilsen

Brigham Young University 1 , Provo, Utah 84602,

B

Benjamin L. Francis

Achilles Heel Technologies 2 , Orem, Utah 84097,

A

Alex M. Stankovic

SLAC National Accelerator Laboratory 3 , Menlo Park, California 94025,

M

Mingjian Wen

Institute of Fundamental and Frontier Sciences University of Electronic Science and Technology of China Chengdu P. R. China

I

Ilia Nikiforov

University of Minnesota 5 , Minneapolis, Minnesota 55455,

E

Ellad B. Tadmor

University of Minnesota 5 , Minneapolis, Minnesota 55455,

V

Vasily V. Bulatov

Lawrence Livermore National Laboratory 6 , Livermore, California 94550,

V

Vincenzo Lordi

M

Mark K. Transtrum

Brigham Young University 1 , Provo, Utah 84602,