CHIMIYA-1: An Autoselection Foundation Model for ADMET Property Prediction, Rigorously Benchmarked Against the Therapeutics Data Commons ADMET Group

Avatar
Poster
Voice is AI-generated
Connected to paperThis paper is a preprint and has not been certified by peer review

CHIMIYA-1: An Autoselection Foundation Model for ADMET Property Prediction, Rigorously Benchmarked Against the Therapeutics Data Commons ADMET Group

Authors

Varghese, R.; Tiwary, P.; Oswal, K.

Abstract

Accurate, generalizable prediction of absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties remains one of the highest-leverage unsolved problems in computational drug discovery, and late-stage attrition driven by ADMET liabilities continues to be a dominant cost driver in pharmaceutical research and development. The Therapeutics Data Commons (TDC) ADMET Group has emerged as the fields most widely adopted public benchmark, comprising 22 endpoints under standardized scaffold-split evaluation. In this work we report a comprehensive evaluation of CHIMIYA-1, a proprietary autoselection foundation model developed by Covenant Biosciences, against the full TDC ADMET Group. Departing from common practice in the field, every reported score is the mean and standard deviation of five independently seeded end-to-end evaluation runs (TDCs own minimum submission standard, which we find is not met by all public leaderboard entries), and all 22 endpoints were additionally subjected to an explicit train/test structural-overlap audit prior to reporting, finding zero overlaps on any endpoint. Despite this deliberately conservative evaluation standard, CHIMIYA-1 ranks first among all publicly listed methods on four endpoints, places within the top decile of the field on twenty of twenty-two endpoints (91%), and attains a mean percentile standing near the 74th percentile across the full benchmark, with particular strength on toxicity and physicochemical-property endpoints. We further show that several top-ranked public comparators on this benchmark have been independently found to exhibit confirmed data leakage, a finding that, if anything, understates CHIMIYA-1s relative standing. All results were obtained on commodity single-GPU workstation hardware without recourse to distributed or cloud-scale training infrastructure. We discuss these results in the context of benchmark reporting norms in molecular machine learning and outline ongoing extensions, including continuous prospective-data retraining and CUDA-level throughput optimization of the underlying selection pipeline.

Follow Us on

0 comments

Add comment