Deep Origin claims 100-fold hit rate gain with DODock screening tool
Deep Origin has published a bioRxiv preprint describing DODock and DOScore, a virtual screening architecture that pairs physics-based energy calculations with machine learning to predict how small molecules bind to protein targets. In prospective wet-lab validation across four targets, the South San Francisco company reported hit rates substantially above those produced by conventional AI docking tools.
The headline result came from a screen against CD73, a nucleotidase target of broad oncology interest, where 56 of 183 synthesised compounds were confirmed as active inhibitors below 500 µM, giving a 30.6% hit rate. A prior large-scale machine-learning screen against the same target, conducted by AtomNet across 318 targets, confirmed zero inhibitors below 100 µM among 335 compounds, yielding a rate of approximately 0.3%. Deep Origin characterises the comparison as a roughly 100-fold improvement in screening efficiency. The company also reported a 15.0% hit rate on the kinase IRAK4, a 4.3% rate on Factor XIa, and a 3.1% rate on the protein-protein interface target IL-17A.
The "memorisation trap"
A central argument in the preprint is that current AI docking models are systematically overestimated on published benchmarks because training and test data sets share structural similarity. The company describes this as a "memorisation trap": models that have seen near-identical protein-ligand pairs in training score well on standard benchmarks but collapse when confronted with genuinely novel biology. On the independent Runs N' Poses benchmark, leading co-folding models including AlphaFold 3, Boltz-1, Chai-1 and Protenix achieve pose accuracy of 75% to 88% on familiar targets but fall below 25% on novel ones. Deep Origin reports DODock maintains above 50% accuracy in the same novel-target regime.
DODock addresses this by inserting a physics engine, DOFast, between the machine-learning pose generation step and a final AI ranking step. Rather than relying purely on patterns learned from historical data, the intermediate engine calculates molecular forces explicitly across 80 parameters, anchoring predictions in first principles. The company says this hybrid approach achieves 80% pose accuracy on OpenBind, a benchmark constructed specifically from targets structurally distinct from common training sets, compared with 4% to 28% for co-folding comparators.
In a blind prospective test on PCSK9, DODock predicted the binding pose of AstraZeneca's oral Phase 3 candidate laroprovstat at an atypical C-terminal domain site. Deep Origin subsequently solved the crystal structure experimentally, confirming the prediction to 1.2 Å heavy-atom RMSD, a result the company says demonstrates the method's generalisability to genuinely unknown binding geometries.
Market context and competitive landscape
AI-driven virtual screening has attracted substantial capital across the sector, with established players such as Schrödinger, Recursion Pharmaceuticals and Insilico Medicine, as well as a growing cohort of university spinouts, all pursuing improved hit-finding from large chemical libraries. The persistent criticism of the field, shared by Deep Origin itself, is that computational predictions frequently fail to translate into active compounds in the laboratory. The company's decision to accompany its preprint with wet-lab confirmation data and more than 80 pages of supplementary methodology is a deliberate attempt to address that credibility gap.
Deep Origin is backed by over $50 million in private capital and more than $32 million in non-dilutive funding, including an ARPA-H CATALYST award. Chief executive Michael Antonov, who co-founded Oculus before moving into computational biology, said the company tested its model specifically on cases "built to make it fail." The preprint has not yet been peer-reviewed, and independent replication of the prospective hit rates will be the next credibility test for the platform.
Chief scientific officer Garegin Papoian said: "For science to advance, model performance reporting has to adhere to strict data rigor and transparency. My hope is that this represents a genuine step forward for docking and scoring, helping move the field forward alongside our own drug discovery programs."