Interpreting machine learning models trained on molecular data remains challenging, particularly at the level of individual atoms. Although many attribution methods can indicate which atoms influence a prediction, the field still lacks standardized benchmarks for determining whether these explanations are chemically and structurally meaningful.

To address this gap, this work introduces a controlled benchmark based on rule-derived pharmacophore annotations. Predefined geometric constraints describe the spatial arrangement of functional features within a molecule and provide reference atom sets against which model explanations can be evaluated.

The benchmark is accompanied by PharmacoScore, a metric that quantifies how closely atom-level attributions align with these pharmacophore-derived references. The framework was used to compare explanation methods across several model families, ranging from fragment-based approaches to distance-aware transformers that explicitly encode interatomic geometry.

Distance-aware transformers consistently achieved higher PharmacoScore values, indicating that the atoms emphasized by their explanations more closely reflected the expected spatial relationships. Together, the benchmark and PharmacoScore provide a unified framework for systematically evaluating and comparing atom-level explainability methods in molecular machine learning.

Autors: Adam Sułek, Jakub Klimczak, Jakub Jończyk, Tomasz Kosciolek, Tomasz Danel, Barbara Pucelik

Read the article

This publication is one of the outcomes of the FIRST TEAM FENG project FNP Foundation for Polish Science, under which Łukasiewicz – Krakow Institute of Technology is developing new approaches to the identification and evaluation of therapeutic candidates in hormone-dependent breast cancer.