How to Prompt AI for Scientific Research, From Two CEOs

The single most useful change two biotech CEOs made to how they use AI was to stop treating a prompt as a question and start treating it as a process. Raphaël Ognar, President, CEO, Co-Founder and Chairman of NKILT Therapeutics, states the underlying principle first, and it is deliberately unglamorous.
if you don't ask the right question, you cannot expect the right answer
That sounds obvious and is routinely ignored. His argument is that people underestimate how completely the phrasing of a request determines the output, and that a model given a vague question will produce a confident answer to a different question than the one you meant to ask. In scientific work, where the answer looks the same either way, that is not a small problem.
What does a good scientific prompt actually contain?
Ognar's practical answer is the most transferable thing in the conversation. Do not ask for the destination. Ask for the route, with stops.
Build in your prompt some milestones so you can check and verify that it's still addressing the question that you need to answer
The point of the milestones is inspectability. If a request is structured so the work is visible in stages, a wrong turn is caught at the stage it happens rather than inferred backwards from an answer that looks wrong at the end. If the request is a single transaction, the only thing you can audit is the conclusion, and by then you have lost the information you would need to audit it.
Ognar's framing is that this is management, not prompt engineering. At this stage of the technology he treats the model like a capable teammate who needs coaching and checkpoints, because it is learning from human beings and will inherit both their competence and their errors.
Is prompting a skill you have to learn?
Ryan Cawood, Co-Founder and CEO of Lab Thread, thinks so, and is direct that everyone is still on the curve.
As we learn how to use AI, we have to learn how to prompt it correctly and it becomes an iterative process
His objection is to the mental model where a prompt is a vending machine.
it's not just a plug in a prompt and get the answer and take it away and everything's going to be fine
Where he lands is slightly counterintuitive, and useful precisely because it is uncomfortable. Getting reliable output means asking in a way you would never speak to a qualified colleague, spelling out what you would normally leave implicit between experts.
you need to really put the guard rails around what you're questioning
Two experienced scientists talking to each other rely on enormous shared context. A model has no shared context, only a very large amount of general context, so anything you leave unsaid it will fill in from somewhere. Stating the obvious is not talking down to the tool; it is removing the space in which it improvises.
Why does the input data matter more than the prompt?
Because a perfect prompt over bad inputs still returns a confident answer, and that is the failure mode with the longest tail.
if you give them flawed data, you will get a flawed answer
Cawood's view is that AI models are extraordinarily capable and are only as good as the material handed to them. If the data going in is uncurated, incomplete or simply wrong, the model will still detect a pattern and still deliver a conclusion, and the conclusion will arrive in exactly the same register as a correct one. That is why the prompting discipline and the data discipline are the same discipline, and why an error introduced this way can sit undetected for months, which we cover in why AI errors in biotech surface months too late.
Does the regulator care how you prompted it?
Not directly, and that is the point people miss. The FDA's January 2025 draft guidance Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products sets out a risk-based framework for establishing the credibility of an AI model for a particular context of use. It is concerned with whether you can demonstrate the output is trustworthy, not with your technique for obtaining it.
Which is exactly why the checkpoint approach is more than a productivity habit. A prompt structured in inspectable stages produces something you can show. A single-transaction prompt produces an answer and nothing else, and if you cannot demonstrate where it came from, the responsibility does not move anywhere, as we set out in who is liable for AI mistakes in drug development.
What does this look like as a working habit?
Reduced to practice, both men are describing the same four moves. State what the data is rather than assuming the model will infer it. Ask for the work in stages you can inspect. Check each stage against the question you actually needed answered. And treat a fluent answer as a draft rather than a result, particularly when it is one you were hoping for.
None of that requires new tooling, and none of it is specific to life sciences. It is ordinary scientific method applied to a source that happens to be very fast and very confident. Both men work this way every day, and neither puts it in a pitch deck, which is a separate call we look at in should startups put AI in their pitch deck.
The convener's read
Neither guest reached for a technique or a formula. They described a posture instead: assume the answer is wrong until the work behind it is visible. That is harder to package than a template, and considerably more durable as the models keep changing underneath everyone.
The full conversation is on the episode page, along with more from Ryan Cawood and Raphaël Ognar. If your organization wants to reach the people making these calls, partner with Open Door Salon.
Methodology: every quotation above is drawn verbatim from the recorded, on-the-record conversation between Ryan Cawood, Raphaël Ognar and host Lori Ellis on Open Door Salon, and was checked against the episode transcript before publication.
