Who Owns the Data in a Biotech Collaboration Agreement?

In most biotech collaboration agreements, the answer is that nobody clearly owns the data, because the contract barely addresses it. Anne Herold Li, Shareholder and New York Managing Partner at Brownstein Hyatt Farber Schreck and a first-chair intellectual property trial lawyer, says that unless an agreement was written in roughly the last two or three years, it likely contains very little language about who owns the data a collaboration generates or what each party may do with it. That silence was tolerable when data was a byproduct of the work. It is not tolerable now that data is the input to model training, drug discovery, and commercial strategy. On Open Door Salon, Herold Li calls it the most pressing question in the industry today, and locates the exposure in a specific and easily missed place: a short provision buried deep in a joint development agreement.
Why has data ownership become the pressing contractual question?
Because the asset changed value while the contracts stayed still. Herold Li marks the shift plainly.
"Data has always been interesting, but now data is king."
What changed is what data can be turned into.
"Data is what people are using to train AI models on, discover new drugs, to figure out how best to market and do the business. And all these collaborations now generate new data."
A collaboration between two companies produces a third thing that neither party existed with beforehand. When the agreement was drafted, that output was understood as research results. It is now understood as training material with independent commercial value, and the contract language did not move with it.
"And when all of these collaboration contracts were written, unless it's been the last two or three years, they have very little contractual language about who owns the data and who owns what rights to do what with the different data sets. And that is becoming the most pressing question in the industry today."
Where does the exposure actually sit in the document?
In a provision nobody had time to read. Herold Li's description of how these clauses get approved will be familiar to anyone who has closed a deal under time pressure.
"CEOs, I think, are just totally unaware that this weird little provision, like in paragraph 75 of your joint development agreement, that some IP lawyer looked at for seven minutes and nobody gave them time until the deal was getting signed the next day to read it. And they weren't allowed to read the whole contract."
The failure is procedural rather than legal. The clause was reviewed by someone competent, working without the context or the hours to see what it did. Then it sat unexamined.
"No one realizes that that little four sentence vulnerability exists and they're just not paying attention to it until it's too late."
Four sentences is the scale. This is not a systemic drafting problem across a hundred-page agreement. It is a specific, findable, short passage that a company can go look at this week.
Why can this not be fixed after the fact?
Because a data-rights problem does not behave like a breach of contract that damages can repair. Herold Li is direct about the asymmetry.
"Prevention is worth more than the cure really here because there is no cure."
She then explains why, and the reason is technical rather than legal.
"There's no cure once that data is leaked. It's been put into another AI model. It's out there in the world. There's no getting it back."
Once a data set has been used as training material, it is no longer a discrete object that can be returned, deleted, or enjoined. It has been absorbed into the statistical behavior of a model. A court can award damages. It cannot un-train a system. That is why she treats this as a pre-signature question rather than a dispute-resolution question.
Does putting information into an AI system waive privilege?
Herold Li raises a consequence that sits outside the ownership question and catches teams by surprise.
"It definitely has issues for privilege. You've waived it if you put it in AI to be clear."
Her broader point is that the exposure comes from the system learning, not merely from the system storing.
"So it's because the AI is learning from the people that is an added level of threat that most users of AI don't consider."
For a company whose legal team is routinely working through sensitive licensing, litigation, and diligence questions, that is an operational rule with immediate consequences, and it is one most AI usage policies were not written to address.
Is everything that needs an AI actually getting one?
Herold Li thinks the reflex is running ahead of the need, and the comparison she reaches for is one most executives lived through.
"It's sort of like the internet bubble of the early 2000s where everyone had a web page or you're on the web. Now it's like, do you have an AI?"
"Even when an AI is not the best solution for your problem right now, an old fashioned computer, technology works, algorithms work just as well, software works just as well. But now it has to be an AI because that's the product du jour."
This matters for data rights because each AI initiative creates a new demand for data, and each new demand creates pressure to move data across a boundary that a contract may not have contemplated. A project that did not need a model has still generated a rights question by the time it ships.
What should a company do about this now?
Go find the clause. The specificity in Herold Li's account is what makes it actionable: pull the joint development and collaboration agreements signed more than three years ago, locate the data provisions, and determine what each party is permitted to do with the data the collaboration produced, including using it to train a model. Where the answer is unclear, that ambiguity is the finding.
The same discipline applies going forward, and it belongs in the same conversation as the vendor obligations covered in what a biotech vendor contract should require. The question of which data sets are most sensitive has its own counterintuitive answer, addressed in why anonymized clinical data is a national security risk. And for companies with cross-border partners, the audit obligations arriving through the COINS Act, enacted through the FY2026 NDAA, give the exercise a compliance deadline as well as a commercial one.
The same delay problem shows up in the lab, where a model’s wrong answer surfaces months later at the manufacturing door: why AI errors in biotech take months to find.
Ownership is one question; legal responsibility is another, and the FDA answers it bluntly. Ryan Cawood and Raphaël Ognar on who is liable when a model produces the data.
Ownership settles who holds the data. It does not settle who answers for it: who is actually liable when AI produces it.
This post draws on the recorded, on-the-record conversation with Raja Mikkili and Anne Herold Li on Open Door Salon. Watch or listen to the full episode, Nation States Aren't Targeting Big Pharma, or read more about Anne Herold Li and Raja Mikkili.
Open Door Salon brings the people who build, fund, and fight for global healthcare into one room. If your organization wants to reach that audience, see our sponsorship options.
