InterFeat: a pipeline for finding interesting scientific features

D Dan Ofer M Michal Linial D Dafna Shahaf

Abstract

Abstract Finding interesting phenomena is the core of scientific discovery, but the notion of interestingness is vaguely defined and heavily reliant on manual judgment. We present InterFeat, an integrative pipeline for automating the discovery and ranking of inter esting feat ures ( InterFeat ) in structured biomedical data. The pipeline combines machine learning, knowledge graphs, literature search and large language models. We formalize “interestingness” as a combination of novelty, utility and plausibility. In a time-split evaluation, InterFeat was trained only on historical data, and managed to surface risk factors years ahead of their eventual discovery. Across eight major diseases, up to 21% of suggested factors appeared in the literature after the time cut-off. In a human evaluation, four senior physicians annotated InterFeat’s suggestions, deeming 28% of them interesting. Out of highly-ranked candidates, 40–53% were interesting, vs. 0–20% for SHAP and L1 baselines. InterFeat addresses the challenge of operationalizing “interestingness” scalably for any target with existing literature. Code and data: https://github.com/LinialLab/InterFeat

Article Details

Volume / Issue Vol. 16, Issue 1
Published March 18, 2026
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (3)

D

Dan Ofer

M

Michal Linial

D

Dafna Shahaf