When Novo Nordisk and OpenAI announced their partnership earlier this year, it was one of several signals that pharma is committing to AI at enterprise scale. For example, Merck signed a deal with Google Cloud worth up to $1 billion, while also expanding its partnership with Tempus AI on precision medicine biomarkers. And Roche deployed what it described as the pharmaceutical industry’s largest hybrid-cloud AI factory in collaboration with NVIDIA.
The deals are getting bigger, the timelines more ambitious, and the projected returns more confident. But when I speak with leaders at pharma companies about the impact of their AI work, the response is often the same: they have yet to see operational returns or effective real-world adoption from their teams.
In my experience leading AI teams at CAS, the organizations seeing real results from AI in drug discovery are those with the prerequisite data foundation. What matters most is the quality of the data AI draws from, and whether its outputs can be verified.
Reasoning like a scientist is not enough
Reasoning and knowledge are two different capabilities. Frontier AI models now surpass PhD-level human performance on GPQA Diamond, a benchmark of expert scientific reasoning in biology, chemistry, and physics. That reasoning capability isn’t the same thing as doing reliable science, though. The question now becomes: what scientific information is the model reasoning over?
In drug discovery, the same molecule can appear in hundreds of different forms across databases and published literature, often with terse and ambiguous coded names. Protein targets are annotated at different levels of specificity across sources, often with unclear species or variant mapping, missing assay conditions, or unspecified truncation and tagging. What may seem like minor inconsistencies to a non-expert often represent critically meaningful differences from a scientific perspective.
A model reasoning on messy data may provide a plausible answer but cannot reliably provide a correct one. For example, ask a general AI agent whether there are known benzimidazole-based inhibitors of a particular transporter target, and it might answer yes, citing real compounds and pointing to real paper references. But the cited compounds are piperidine-ureas, not benzimidazoles. The answer is directionally defensible, but the evidence is structurally false. A medicinal chemist would catch it immediately, but a biologist selecting tool compounds, a project manager scoping competitive space, or an investor doing due diligence isn’t likely to. As AI scales beyond domain experts, the verification burden increasingly migrates to those least equipped to carry it.
The details that matter most to a chemist or biologist – stereochemistry, reaction conditions, atom mapping, target annotations, bioactivity data, and assay conditions – can also be easily lost. When any of these are stripped out or misrepresented, the model's output may look right, but acting on it can ultimately lead to failed experiments, wasted cycles, and delayed pipeline decisions.
What data investment can achieve
In practice, clean data means AI can discern whether two compounds are the same chemical represented differently or are truly different substances. It has the context to understand whether assays leveraged conditions that allow them to be compared. It can also recognize when two differently named proteins are the same target, or when two similarly named ones are not.
Those at the cutting edge are increasingly raising the importance of data for successful AI. Isomorphic Labs, the Google DeepMind spinout that recently raised $2.1 billion in Series B financing, attributes the performance of its drug design engine in part to its “dataverse” – a curated collection of life science data that powers its research environment. On the most difficult protein-ligand structure prediction benchmarks, the company’s IsoDDE model more than doubles the accuracy of AlphaFold 3, the system whose underlying methodology won the 2024 Nobel Prize in Chemistry. While we so often hear the story of the powerful model, the innovation is not possible without the curated data infrastructure underneath.
Our own experimentation at CAS bears this out: a model trained on a curated dataset for drug-protein binding achieved twice the prediction accuracy of a model trained on a large public dataset, using 50 percent fewer records. Better accuracy translates to better pipeline decisions, helping teams root out false leads before they enter the pipeline and reduce time spent validating unpromising results.
Curated data with verified structures, complete experimental conditions, and standardized annotations give models a chance to recognize meaningful scientific relationships rather than statistical coincidences. Achieving this structure requires ongoing review and revision from people who understand the science, and this investment is essential for reliable AI outputs.
The impact of improved predictions in discovery
When the data foundation is right, efficiency gains follow: fewer false leads enter the pipeline; less time is spent validating results that won’t hold up; go/no-go decisions happen earlier, with more confidence behind them. These are the gains that AI investment promises, but they only materialize with the right foundation.
The stakes are also higher than ever. The industry is moving from AI as a discrete tool to AI as infrastructure, with agentic systems and coordinated agent “swarms” that act with increasing autonomy to reason across data sources, take sequences of actions, and draw on external knowledge in real time to support discovery workflows. In that architecture, the quality of the knowledge source determines the quality of everything the agent produces. An agent connected to a curated, consistently structured scientific knowledge base produces fundamentally different outputs than one drawing from low-quality sources. The agent amplifies whatever it connects to.
The foundation determines the outcome
The industry's enthusiasm for AI in drug discovery is warranted, but model power and data volume alone will not deliver results. Discovery teams that want AI to meaningfully accelerate their work should evaluate the fundamentals first: is the data curated for scientific accuracy, and is it structured for the specific problem being solved? Those questions matter more than which foundation model a team chooses, because they determine whether the model has anything useful to work with in the first place.
The organizations that get this right will be the ones where AI delivers the efficiency and ROI the industry is projecting. The ones that do not will spend their time re-verifying what their models produced, slowing the very pipeline decisions AI was supposed to accelerate.
