Ryan Cawood and Raphaël Ognar on What Happens When AI Gets Drug Discovery Wrong
What you’ll learn
- Where AI actually fails in drug discovery, and why the failure arrives fluent rather than obvious
- Who carries the liability when a model produces data that reaches an FDA submission
- Why the expensive errors are temporal, surfacing months later at the door of a GMP facility
- How failure is priced differently at big pharma and at a small biotech
- What the AI business is being optimized for, and why neither guest advertises AI in a pitch deck
The anecdote that opens this episode is small enough to be funny and specific enough to be alarming. Raphaël Ognar was doing market analysis on radioligand therapeutics, and the AI he was using handed him a set of references. He does what a scientist does with a reference: he went to find it. One of them did not exist. The names looked legitimate. The title looked legitimate. It was not a real paper.
When he pushed, the model admitted it had invented the citation. Then it explained itself.
I didn’t want to fall short by giving you an unsatisfactory answer. And I literally prompted back. I said I don’t know is more than satisfactory and it’s definitely better than lying.
Ognar is President, CEO, Co-Founder and Chairman of NKILT Therapeutics, and before that spent years inside big pharma running drug development programs, carrying two small molecules through to approval. Ryan Cawood, Co-Founder and CEO of Lab Thread, has hit the same wall from a different direction. Their conversation with host Lori Ellis is not a debate about whether AI belongs in drug discovery. Both of them use it. It is about the specific places where a confident wrong answer becomes expensive, and about what neither of them is willing to claim in public.
A confident answer about a piece of DNA
Cawood ran his own version of the test. He pasted in a DNA sequence he already knew the origin of and asked the model to identify it. It did not say it was unsure. It gave him an answer, and the answer was wrong.
That is the pattern both of them keep returning to. The failure mode is not that the model refuses. It is that it produces something fluent and plausible in the exact register of a correct answer. Cawood’s framing is that the tool is genuinely useful and should be treated with the skepticism you would apply to a person.
it really does need to be treated like a colleague. You wouldn’t trust everything that a colleague told you as gospel all the time.
The FDA does not care where the data came from
In most industries a wrong answer is a wasted afternoon. In drug development it has a regulatory owner. Cawood is direct about where the liability lands:
the FDA has been very clear that a company that submits data to the FDA is 100% legally responsible for that data regardless of where it comes from
There is no version of that sentence where the algorithm absorbs the consequence. Ognar extends it to the people rather than the entity, pointing out that chief medical officers and physicians carry personal responsibility for what reaches a patient. His argument is that this is the reason the human review layer cannot be removed, not a reason to distrust the technology. The liability question sits on top of an ownership question most contracts never settle: who actually owns the data in a biotech collaboration agreement.
The mistake happens at the desk, not at the bench
Asked what an AI failure actually looks like inside a working lab, Cawood corrects the premise. The robots are not the problem, because the robots are not there yet: “AI I wouldn’t say is actually in the physical lab yet.” The errors are upstream of the bench. Someone uploads a file and asks the model to interpret it without saying what the data is. The model assumes a standard curve that is not a standard curve, and returns a number. Nothing looks broken. The experiment that follows is built on it. The institutional version of the same problem has its own failure pattern, which the FDA’s first CIO argues is budgetary rather than technical.
Why the real risk is nine months long
This is the part of the conversation that should worry anyone running a preclinical program, and it is a point about time rather than about accuracy. “The real danger though is temporal,” Cawood says, and he walks through it concretely. A piece of DNA goes into a cell line. The cell line is grown and passaged over many months toward a GMP handoff. If the sequence was never validated and turns out to contain something it should not, the error does not surface when it is made. It surfaces at the door of the manufacturing facility.
the cell line you’ve spent 9 months making is now no longer viable in a GMP environment
The same shape applies to the paperwork. Chain of custody for a cell line is still frequently recorded by hand, batch by batch, and a single missing record can invalidate the package. Cawood’s view is that this is exactly where AI earns its place, checking the completeness of records weekly rather than generating conclusions. But the value is entirely dependent on the underlying data being real, because “if you give them flawed data, you will get a flawed answer.”
Ognar’s practical answer is to stop treating a prompt as a single transaction. Build the checkpoints into the request itself so the work can be inspected on the way rather than accepted at the end.
Failure is not priced the same at both ends of the industry
Ognar has run programs at both scales, and he is blunt about the asymmetry.
big pharma doesn’t characterize failure the same way a small biotech will just for a very simple reason. Big pharma has cash and we don’t as small biotechs
A delay inside a company generating tens of billions in annual revenue is absorbed. The same delay at a small biotech is one fewer chance of reaching the finish line. He notes that despite three or four decades of genuine improvement in how drugs are developed, roughly one program in ten still makes it, and argues that some of that gap sits in the distance between a trial and the world the drug enters. He is unusually candid about what a trial is designed to do.
I don’t design a clinical trial with my team to just design answer the best scientific response. I designed a clinical trial to ensure that we can get approved because that’s the goal is to make the drug approved
He would like to see AI applied to that gap, modeling population level tradeoffs that are currently argued rather than calculated. Cawood’s answer to that is where the two of them divide the labor: he would rather people made the ethical decision and left the model to run the numbers.
What the AI business is being built for
The last third of the conversation turns to the economics, and both guests arrive at a similar concern from opposite sides. Cawood, as someone deciding how much AI to put inside his own product, points at a cost structure most buyers never see.
these companies are actually underwriting the cost of service serving you the AI result to the degree of 90%
His concern is not this quarter’s bill. It is building a product on a service whose real cost is being deferred, and watching the tokens per dollar shrink with each new model release. Ognar’s objection is broader, and it is about what the system is being optimized for. His formulation is that the industry needs to design “not just based on greed but based on outcome.”
He also raises the energy and water footprint of the data center buildout as a cost nobody is currently pricing, which prompts the driest line in the episode from Cawood: “It’ll be quite ironic if we stop burning oil and then start burning water, which I think is where we’re heading.”
Neither of them puts AI in the pitch deck
Cawood’s objection to the current moment is not the technology but the compulsion. He describes software that now rewrites his email whether or not he asked, and points out that a colleague behaving that way “would be the most annoying colleague on the planet.”
Ognar takes it to the place it costs him something. He is raising money for a biotech in a market where investors expect the acronym, and he has left it out.
my pitch deck for example doesn’t have one mention of AI and everybody’s telling me you need to put AI in your pitch deck
Ognar’s last word on AI is not about AI at all. Asked what drew him to the work, he answers that he wants one patient to benefit from something he built, then closes on an arithmetic he clearly expects to land.
Remember that the largest pharma makes approximately 50 or 60 billion dollar a year. The largest health insurance company makes five time that. Okay. So let’s see who is the bad guy here. It’s not the people that are trying to develop drugs and make money with it.
He is answering a charge nobody in the room had made, which is its own comment on how often he has heard it.
His reasoning is that a company using AI well does not need to announce it, and that the claim will not survive contact with people who ask what it actually does. It is the same standard both of them applied to the technology all the way through: the question is never whether AI is impressive, it is whether it earns its place. Lori Ellis closes with the two questions she always asks, and the answers, on creative independence and on the cost of leaving a comfortable career, are worth the last ten minutes on their own.
I didn't want to fall short by giving you an unsatisfactory answer.
Key takeaways
- The failure mode is confidence, not refusal. Both guests describe models returning fluent, plausible answers in the exact register of a correct one.
- Verification is a habit, not a feature. Ognar found the invented citation because he checks references by reflex.
- The sponsor owns the data. Provenance does not transfer liability, and no company can point at an algorithm in a submission.
- Mistakes happen at the desk. AI is not yet doing physical lab work, so the errors enter upstream through interpretation.
- The real risk is time. An unvalidated input can sit inside a cell line for months before it fails at the GMP handoff.
- Prompt in checkpoints. Ognar builds milestones into the request so work can be inspected on the way rather than accepted at the end.
- Small biotechs have no buffer. The same delay big pharma absorbs is one fewer chance at the finish line.
- Neither puts AI in the pitch deck. Their shared test is whether it earns its place, not whether it impresses.
Key Questions, Answered
What happens when an AI invents a source?
there was that one publication that I couldn't find ... by scratching and then digging further I had the AI admitting that they invented that publication
Raphaël Ognar asked for references, went to verify one, and found the paper did not exist.
Who is legally responsible for AI-generated data submitted to the FDA?
the FDA has been very clear that a company that submits data to the FDA is 100% legally responsible for that data regardless of where it comes from
Ryan Cawood on why the sponsor owns the submission no matter what produced the data.
Who is personally liable when a model gets it wrong?
the chief medical officers the physicians they are liable for everything that comes out and that affects patients
Ognar on the individuals, not the entity, who answer for what reaches a patient.
Where do AI mistakes actually happen in a lab?
it's in the office where the mistakes are made ... you're uploading a file and you're saying interpret this data without guardrails in place
Cawood corrects the premise: the robots are not in the lab yet, so the errors are at the desk.
Why is an early data error so expensive in cell line development?
the cell line you've spent 9 months making is now no longer viable in a GMP environment
A validation gap does not surface when it is made. It surfaces at the GMP handoff.
How should a scientist structure an AI prompt?
Build in your prompt some milestones so you can check and verify that it's still addressing the question that you need to answer
Ognar on treating a prompt as a process with checkpoints rather than a single transaction.
Why does failure cost a small biotech more than big pharma?
big pharma doesn't characterize failure the same way a small biotech will just for a very simple reason. Big pharma has cash and we don't as small biotechs
Ognar has run programs at both scales and is blunt about the asymmetry in buffer.
Are clinical trials designed for the real world?
Let's not kid ourself. ... I designed a clinical trial to ensure that we can get approved because that's the goal is to make the drug approved
Ognar is unusually candid about what a trial is actually optimized to achieve.
Why is a confident AI answer dangerous?
there is a temptation to trust it beyond its real capabilities
Cawood on fluency being mistaken for accuracy, and models now training on their own output.
What should AI development be optimized for?
we need to try to think about designing systems and leveraging processes and AI technology not just based on greed but based on outcome
Ognar argues the bubble is being designed around extraction rather than results.
Who is really paying for AI services?
these companies are actually underwriting the cost of service serving you the AI result to the degree of 90%
Cawood on the deferred cost sitting underneath every AI subscription.
Should a biotech put AI in its pitch deck?
my pitch deck for example doesn't have one mention of AI and everybody's telling me you need to put AI in your pitch deck
Ognar is raising money in a market that expects the acronym, and has left it out.
Resources
- FDA draft guidance on AI in regulatory decision-making Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products.
- 21 CFR 312.50, general responsibilities of sponsors The regulation behind the point that the sponsor answers for the data it submits.
- Ryan Cawood on LinkedIn Co-Founder and CEO, Lab Thread.
- Raphaël Ognar on LinkedIn President, CEO, Co-Founder and Chairman, NKILT Therapeutics.
Need the life-sciences signal but short on time?
Get the free quarterly briefing: every guest from the quarter, in one sitting. What decides whether a therapy reaches a patient, gets funded, and can be trusted.




