Skip to content

AI in Drug Discovery: What a Model Can Move and What It Cannot

5 min read · updated August 3, 2026

This page is about method, not about any particular medicine, and nothing here is medical advice. It is written to answer one question: when a company says a drug was discovered with AI, which part of a decade-long process is that sentence about?

The pipeline, and where the years go

Roughly, and with enormous variation: pick a target, find molecules that do something to it, optimise those molecules into something drug-like, test in animals and in safety assays, then run the clinical stages — first for safety in a small number of people, then for efficacy in patients, then in a large confirmatory trial — and then apply to a regulator. Start to finish is usually over a decade.

Two facts about that pipeline determine everything else on this page. The first is that the calendar and the money are dominated by the clinical stages, not the discovery ones. The second is that failure is the normal outcome, and it is concentrated where the drug first meets human biology: a candidate can be a beautiful molecule, hit its target exactly as designed, and still not help anyone, because the target was the wrong thing to hit.

Where models are genuinely used

ApplicationDescription
Virtual screeningScore enormous make-on-demand chemical libraries against a target site far faster than physics-based docking can. The output is a shortlist to synthesise and assay, and it replaces a search, not an experiment.
Generative chemistryPropose molecules conditioned on a target, a scaffold or a set of property constraints, rather than picking from a catalogue. Whether the molecule can be made at all is a separate model.
Property predictionSolubility, permeability, metabolic stability, cardiac ion channel liability. These filter a list early and cheaply. They are trained on assay data and inherit its coverage: they are most reliable on chemistry that resembles what has been tested.
RetrosynthesisPlan a route from purchasable starting materials. This is the application closest to a solved problem, because the verification step is a chemist reading the route, and then the route either works or does not.
Target identificationUse omics, genetic association and literature to nominate what to go after. Highest leverage in principle, weakest evidence in practice, because the feedback loop is the whole pipeline.
Trial designPatient stratification, site selection, enrichment for likely responders. Unglamorous, and plausibly where the largest real effect on timelines sits.

Two structural points about that table. Virtual screening changed because the accessible chemical space changed: commercial make-on-demand libraries grew to a size where physics-based docking of every compound is no longer affordable, so a learned model is used to rank the library and docking is reserved for the survivors. The model did not replace an experiment; it replaced an earlier, slower filter, and the final decision is still a synthesis and an assay. Property prediction has the opposite shape: it is trained on accumulated assay data, which covers the chemistry pharmaceutical companies have historically been interested in, and it is least reliable exactly on the unusual scaffolds a generative model is most likely to propose. Those two facts interact badly, and noticing that they do is most of reading a discovery platform’s architecture diagram.

Why speeding up the front does little

Suppose a model takes hit finding and lead optimisation from four years to one. That is a genuine, valuable saving, and it is a saving on the cheapest and shortest part of the process. The clinical stages are unchanged. Total time falls by the three years you saved and by nothing else, and the probability the programme ultimately succeeds is unchanged unless the molecules coming out are better, not merely faster.

This is the general form of the argument in the case for a compressed research cycle: accelerating one stage of a serial process buys you that stage and no more, and the fraction of total duration it represents caps the benefit. It is why “discovered in months rather than years” is a true and much smaller claim than it sounds.

The version of the claim that would matter is different: that model-designed candidates fail less often in the clinic. That is an empirical question about outcomes, it requires a matched comparator and a decent number of programmes to answer, and it is the question worth waiting for. Individual candidates originating from computational pipelines have entered clinical trials, and some have been discontinued after early readouts — which is the expected pattern for any candidate and tells you very little either way.

The translation problem underneath

It is worth being precise about what fails and why, because it explains why a better molecule model does not fix it. Broadly, candidates fail for safety reasons — the compound does something harmful — or for efficacy reasons — the compound does what it was designed to do and the patients do not improve.

Efficacy failure is usually a failure of the biological hypothesis, not of chemistry. The target was implicated by association rather than causation; the pathway is redundant so blocking one node changes nothing; the animal model does not represent the human disease; the disease is heterogeneous and the drug helps a subgroup too small to show up in the trial. None of these are problems a better binder solves. A model that designs an exquisite inhibitor of the wrong protein has produced an exquisite inhibitor of the wrong protein.

This is why target identification is the application with the highest theoretical leverage and the least demonstrated success. To learn which targets are worth pursuing you need labelled examples of targets that worked and did not, and each label costs a decade and a programme. The training signal is exactly what the field is short of.

Reading an announcement

  • Which stage is the milestone? Entering first-in-human testing is a statement about safety plans and manufacturing, not about whether the drug works. The efficacy readout comes years later.
  • What was the model’s contribution? Nominating the target, generating the molecule, ranking a library somebody else built, and planning the synthesis are different claims. Announcements frequently leave this unspecified.
  • What is the endpoint being reported? A binding affinity, a cell assay, an animal model and a patient outcome are four different kinds of evidence, and the distance between the first and the last is where the whole difficulty lives.
  • From when is the timeline measured? “Eighteen months to a candidate” usually starts after target selection, which is often the part that took years.
  • Is the chemistry novel? A known scaffold rediscovered quickly is a real efficiency result and a weak novelty one. Both are worth reporting; they are not the same claim.

None of this makes the applications unserious. Screening spaces too large to dock, filtering candidates before spending bench time, and planning syntheses are all genuinely useful, and the retrosynthesis case in particular has the property that makes automation work — a cheap, fast, unambiguous verification step. The mistake is only ever in which part of the pipeline the reader thinks has moved.

AI in Drug Discovery: What a Model Can Move and What It Cannot · Multigrid