A business development view from the other side of the table - where the AI conversation in regulated clinical development actually is, versus where the headlines say it is.
My colleague Sofia Chalkiadaki, our Head of Engineering, wrote earlier this month about the architectural question behind our work at TriloDocs: whether the reasoning in a regulatory document should ever be probabilistic. Her perspective comes from an unusual combination of drug discovery and software engineering - she felt the cost of "approximately right" firsthand, at the bench and in the code.
Mine is a different vantage point.
This year I've had more than a hundred conversations with senior IT, Digital, Clinical Development, Regulatory and AI leaders across more than sixty pharma and biotech companies. Some are already building their own AI capabilities. Some are evaluating multiple vendors. Others are only beginning to work out where AI fits into regulated workflows.
I don't see the code. I see what happens when people have to decide whether they'd actually trust the technology.
Over the last year, I've noticed that conversation is changing.
Not long ago, simply generating a credible regulatory document with AI was impressive. It still makes for a good demo - upload source documents, wait a few minutes, and suddenly there are pages of polished scientific prose where previously there were none.
But most conversations now move past that fairly quickly.
The questions have shifted to:
That last question comes up a lot.
There's no shortage of impressive AI productivity claims in our industry - a document that took weeks, drafted in hours. That's valuable.
But experienced medical writing teams tend to get interested in what happens after the generation.
If a system produces eighty pages in ten minutes, but a medical writer then has to verify every number, statement and interpretation against source material, the drafting has been accelerated - but a large share of the workload is still there.
And that work is expensive precisely because it requires the same people whose time you were supposedly trying to save.
This is one of the most consistent themes I hear when companies talk about previous AI evaluations.
They don't have a generation problem anymore. They have a verification problem.
Which is probably why I've grown sceptical of generation speed as the headline measure of AI productivity in regulated work. The meaningful measure isn't time to first draft. It's time for an output the organisation can confidently stand behind.
Ten minutes to generate and three days to verify can easily be worse than two hours to generate and two hours to verify. The first number looks better in a demo. The second is what actually determines whether a document ships on time.
Nearly every regulated AI conversation eventually reaches those four words.
Of course there should be a human in the loop. But I've come to think it's not a particularly useful measure on its own.
The more important question is what humans are actually being asked to do.
There's a big difference between a medical writer reviewing a system's reasoning, applying scientific judgement, and refining the final communication - and a medical writer independently re-checking hundreds of AI-generated claims because the underlying system can't give them enough confidence in the output.
Both technically have a human in the loop.
They are completely different operating models, and that distinction seems to matter more every time I hear it come up.
One of the more interesting things about selling into pharma is how many functions end up in the room.
A Medical Writing leader asks how much QC their team will still have to do. Someone in Quality asks how the system can be validated. An IT or AI leader asks whether the output is reproducible. Regulatory asks whether they can trace where a conclusion came from.
They sound like different questions.
I've come to think they're versions of the same one:
What can we safely stand behind?
That's a much harder question than whether an LLM can write a paragraph - and it's a question that increasingly gets asked before a system is built, not just before it's signed off.
The comparison point isn't always the old manual process anymore.
More companies we speak to have already tried another AI solution, run an internal build, or are evaluating several approaches side by side.
That changes the conversation from "should we use AI" to "which approach actually reduces the burden without introducing a new one."
It's a healthier comparison, because two systems can both produce an impressive-looking CSR and leave the team responsible for reviewing it in very different positions.
The output can look similar.
The operating model underneath it often isn't.
Something else has changed this year that's worth naming directly: for a growing number of the people I'm talking to, the question is no longer whether to bring in an outside AI platform.
It's already there.
Bristol Myers Squibb is deploying Claude across research, development, manufacturing and commercial operations, reaching more than 30,000 employees. It's not alone: Novo Nordisk and AstraZeneca have struck similarly broad partnerships with frontier model providers this year. Even CROs are moving the same way, with ICON announcing a multi-year collaboration with Anthropic to bring Claude across the clinical trial lifecycle.
This isn't about whether pharma is interested in AI anymore. The bigger organisations are already making decisions about how these models fit into the enterprise.
That changes what a specialised vendor conversation is actually about.
It used to be: "Should we adopt an outside AI tool?"
Increasingly, it's: "What value gets added around a model we already have?"
And that's a more interesting question.
A frontier model is extraordinary at reasoning over uncertainty and expressing complex information clearly. It was never built to be the thing that decides on its own whether a specific clinical conclusion is true.
That's a different, narrower job.
And it can sit alongside a Claude or ChatGPT deployment rather than necessarily competing with it.
If I oversimplify the shift I've seen over the last three years:
And increasingly: "Does my Quality function need to be in this conversation before I build, not after?"
That last one is the most interesting shift I've seen this year.
Quality and Regulatory used to be the functions you looped in to sign off on a pilot. Increasingly, they're shaping the architecture decision before anything gets built.
That tells you something about how far "governed AI" has moved from marketing language toward an actual procurement requirement.
Sofia's argument is an engineering one: the shape of the problem should determine the shape of the tool. Discovery rewards probability and exploration. Regulated documentation demands a different level of reproducibility and control.
What I find interesting is that I hear pharma and biotech teams arriving at a similar place from the opposite direction - not through an architecture debate, but through the very practical experience of finding out what "approximately right" costs once someone has to trace it back to source, claim by claim.
The novelty of generation is wearing off.
What's replacing it - verification, reproducibility, traceability, trust - is a harder, and I think a much more interesting, conversation. And it leaves one more question sitting on top of it:
If the frontier model is already in the building, what should we build around it?
If your organisation is working through the same question, I'd be glad to compare notes: will.ewart@trilodocs.com.