ABOUT THIS EPISODE
AI is now good enough to change how computational biology teams actually work, but most companies are still adopting it like it’s 2023. Sonia Timberlake, R&D strategy consultant for Timberlake & Maclsaac Biopharma Consulting, speaks with host Eleanor Howe about what’s real: agentic coding, high-throughput data workflows, and the practical limits that still slow teams down. Timberlake digs into benchmarks for capabilities and end-to-end tasks, as well as multimodal chart understanding, source verification, and where human review remains non-negotiable. The conversation also explores beyond AI to what biopharma risks missing, why novel targets still matter, and where investment interest is clustering right now. Plus, tune in to get a preview of Timberlake’s workshop at Bio-IT World Conference & Expo in Boston next month!
If this helped you think more clearly about AI in drug development and computational biology, subscribe, share the show with a colleague, and leave a review so more builders can find the conversation.
Links from this episode:
Workshop: AI Upskilling for Computational Biology Teams
Bio-IT World
BioTeam
Diamond Age Data Science
Bio-IT World’s Trends from the Trenches podcast delivers your insider’s look at the science, technology, and executive trends driving the life sciences through conversations with industry leaders.
IN THIS EPISODE
SHOW NOTES 🔗
TRANSCRIPT 🔗
00:00:01.199 --> 00:00:02.399
Hello there, everyone.
00:00:02.399 --> 00:00:04.639
Welcome to Trends from the Trenches.
00:00:04.639 --> 00:00:13.759
Joining us today is Sonia Timberlake, a biotech executive who advises companies on research and analytics strategies for developing novel therapeutics.
00:00:13.759 --> 00:00:16.800
Sonia has spent 13 years building biotechs.
00:00:16.800 --> 00:00:23.120
She was formerly head of research at Finch Therapeutics, where she built the new discovery platform and lead pipeline strategy.
00:00:23.120 --> 00:00:26.239
Before that, she was head of data science at Juno Therapeutics.
00:00:26.239 --> 00:00:28.079
Sonia, welcome to the trenches.
00:00:28.079 --> 00:00:31.280
Do you want to tell us a little bit about what you do now?
00:00:31.600 --> 00:00:32.079
Yeah.
00:00:32.079 --> 00:00:41.359
So about three and a half years ago, I started a consulting company and it's focused on how to leverage high-throughput data for RD, which is what I did in biotech.
00:00:41.359 --> 00:00:52.399
And during that time, I've continued to work as an operator at a cell therapy company, but I've also advised in SaaS, worked in venture creation, worked across multiple biotechs and VCs.
00:00:52.399 --> 00:01:06.079
And that's been just a really fun, fantastic learning experience to see how different teams approach problems and sort of the diversity of approaches and different industries.
00:01:06.079 --> 00:01:16.879
So I spent a lot of my time in the last year thinking about the frontier of AI generative models and specifically how we can bring them into drug dev.
00:01:16.879 --> 00:01:22.319
And I am, you know, full disclosure, I'm totally AI pilled.
00:01:22.319 --> 00:01:27.120
So everybody's gonna get all the propaganda.
00:01:28.239 --> 00:01:28.959
That sounds great.
00:01:28.959 --> 00:01:35.120
Well, and it's all your experience in all of these spaces that made me say, oh, I really want to interview Sonia.
00:01:35.120 --> 00:01:36.560
So thank you for coming.
00:01:36.560 --> 00:01:50.159
So then, you know, my first question for you is, and you've already, I think, told us the answer, but what forces are you seeing impacting these companies that you're working with now, especially since this is trends, compared to a year ago?
00:01:50.159 --> 00:01:50.719
Yeah.
00:01:51.040 --> 00:01:59.120
So of course, we have to mention the backdrop here that, you know, we're still, we still have a lot of headwinds in biopharma from a lot of different angles.
00:01:59.120 --> 00:02:03.120
But, you know, that's like that's cyclical and we're gonna come out of that.
00:02:03.120 --> 00:02:10.319
But compared to a year ago, the, you know, the really hopeful and exciting thing I see is how impactful AI is in our sphere.
00:02:10.319 --> 00:02:19.680
And I think a year ago you could say with a straight face, like, hey, I don't really think AI is impacting, you know, drug development or computational biology.
00:02:19.680 --> 00:02:22.879
I just don't think like the the capabilities are are there yet.
00:02:22.879 --> 00:02:30.080
And I really don't think you can say that with a straight face today or in the last, you know, four to six months.
00:02:30.479 --> 00:02:32.400
So then let's dig into that a bit.
00:02:32.400 --> 00:02:33.520
I'd like to hear more.
00:02:33.520 --> 00:02:36.800
You know, the first thing that leaps to mind is agenc coding models.
00:02:36.800 --> 00:02:40.240
So how are you seeing them impact computational biology?
00:02:40.240 --> 00:02:40.960
Yeah.
00:02:41.199 --> 00:02:50.719
So I see like the capability is enormous, but the mean impact is still quite low in my experience.
00:02:50.719 --> 00:03:00.800
And so I'd say, you know, I think a lot about like in software development, sort of mainstream software, not scientific software, like 95% of code.
00:03:00.800 --> 00:03:11.759
It is like very common for a company to say that today or in the last six months, is written by AI agentic agents, where I think it's it's the opposite for scientific code bases in in biotech and pharma.
00:03:11.759 --> 00:03:13.120
It's probably more like 5%.
00:03:13.120 --> 00:03:20.000
And I've talked to a lot of my friends, you know, not just the my clients, but a lot of my friends in in pharma.
00:03:20.000 --> 00:03:28.240
And so, so okay, so like I think there's this huge discrepancy between the capabilities today and and the adoption.
00:03:28.240 --> 00:03:33.759
And and that's like common with any new technology, and that's really common and it has to diffuse through society.
00:03:33.759 --> 00:03:45.840
Okay, so like specifically, I think that like pipelines, if you're doing, you know, RNA-seq or something that's well documented, it's sort of in the canon.
00:03:45.840 --> 00:03:48.960
Like, definitely no human should be writing that.
00:03:48.960 --> 00:03:51.759
There's tons of training examples.
00:03:51.759 --> 00:03:54.080
There's really good verification.
00:03:54.080 --> 00:04:00.560
Like, there's no real the domain expertise doesn't come from writing the pipeline, right?
00:04:00.560 --> 00:04:04.479
It it comes from maybe curating the inputs and sort of interpreting the outputs.
00:04:04.479 --> 00:04:10.960
Like there's still a big role there for humans, and there's a big gap in terms of the AI capabilities.
00:04:10.960 --> 00:04:20.079
But for coding in general, even for exploratory data analysis, for producing figures, for understanding figures, I think that the capabilities are there.
00:04:20.079 --> 00:04:27.279
And I think that you, I think that's reflected in the benchmarks, if not everybody's usage today.
00:04:27.839 --> 00:04:36.480
And and I want to ask you about benchmarks, but before I do, I mean, I we've seen very similar things as far as the utility of coding agents.
00:04:36.480 --> 00:04:43.360
My team has found their work is sped up tremendously because they don't have to write these pipelines by hand.
00:04:43.360 --> 00:04:47.040
They get a big boost from their AI assistant.
00:04:47.040 --> 00:04:49.120
And so the productivity goes way up.
00:04:49.120 --> 00:04:54.959
It's really impressive the changes that we're seeing in just day-to-day work for a bioinformatician.
00:04:54.959 --> 00:04:57.360
Is that kind of what you're seeing as well?
00:04:57.680 --> 00:04:58.160
Yeah, yeah.
00:04:58.160 --> 00:05:08.639
And I think about it both in terms of like the efficiency of things that you were already doing, but also your ability to do things that you wouldn't have already done, right?
00:05:08.639 --> 00:05:23.600
Like if I'm a single person single member computational team at a biotech, well, you know, two years ago, I probably was either a specialist in single cell or maybe bulk or, you know, maybe spatial, but like could I do protein modeling?
00:05:23.600 --> 00:05:24.000
No.
00:05:24.000 --> 00:05:32.879
And I think, you know, today the ability to do that, at, you know, like at a at a at least a basic level is is it's just like that.
00:05:32.879 --> 00:05:35.680
So so I think there are those two axes.
00:05:35.680 --> 00:05:39.920
And I have to say, like, I was impressed that that Diamond Age was an early adopter.
00:05:39.920 --> 00:05:53.199
I remember, I still remember John Hutchinson's talk was just probably, you know, over a year ago where he was using agentic coding for you know really impressive sort of zero shot or one-shot applications.
00:05:53.199 --> 00:05:59.839
And I think in in one case it actually coded up the application in a language he didn't know because like that was the best application for the task, right?
00:05:59.839 --> 00:06:05.839
So there's this there's this access of efficiency and there's this access of like, you know, new new capabilities.
00:06:05.839 --> 00:06:13.439
So so I'm excited to to see Diamond Age sort of like picking the right problems and and applying AI in the right way.
00:06:13.439 --> 00:06:14.000
Thanks.
00:06:14.079 --> 00:06:16.959
Yeah, I'm I'm excited to see it too, as you might imagine.
00:06:16.959 --> 00:06:21.199
So can can you tell me a bit about the benchmarks that you were talking about?
00:06:21.199 --> 00:06:24.240
Like how do you use those to get meaningful information?
00:06:24.560 --> 00:06:25.199
Yeah, yeah.
00:06:25.199 --> 00:06:30.079
And yeah, so I'm a I'm a big fan of benchmarks for AI.
00:06:30.079 --> 00:06:38.879
And I want to distinguish there's like benchmarks for capabilities and for tasks and then for processes.
00:06:38.879 --> 00:06:40.720
And I think those exist at different levels.
00:06:40.720 --> 00:06:47.279
But I think fundamentally that like this is the way, you know, we always ask, like, how do I trust the output or like how do I trust this process?
00:06:47.279 --> 00:06:50.959
And I think you have to, you have to have these objective measures.
00:06:50.959 --> 00:07:01.120
And they're not cheap to make, and they they're a public good, and so they suffer from all of the issues that public goods tend to suffer from.
00:07:01.120 --> 00:07:09.279
But I think they're they're so, so critical for for us all to you know continue hill climbing in the capabilities, but also build trust in what we have today and and encourage adoption.
00:07:09.279 --> 00:07:12.240
So, what are the the specific types of benchmarks?
00:07:12.240 --> 00:07:17.519
So for for bioinformatics, there are three that I know of that are really relevant.
00:07:17.519 --> 00:07:20.800
There's Bix Bench, which was published by Future House about a year ago.
00:07:20.800 --> 00:07:22.560
That was one of the first.
00:07:22.560 --> 00:07:24.800
And those are all human-curated.
00:07:24.800 --> 00:07:30.560
Like, here is a paper, and we've verified that you can go get the the data.
00:07:30.560 --> 00:07:45.120
And then here is not just like the tab output, but like here's a figure from the paper, and here is the human curated correct peer-reviewed responses that you should be able to reproduce.
00:07:45.120 --> 00:07:51.360
Comp Biobench, actually, just like three days ago, this is out of Genentech, published.
00:07:51.360 --> 00:07:57.199
They put up a preprint and a whole GitHub repo on a new computational biology benchmark.
00:07:57.439 --> 00:08:05.600
And they they benchmark then a bunch of you know of the frontier models, sort of with tools, without tools, are they allowed to think indefinitely?
00:08:05.600 --> 00:08:06.560
Are they not, right?
00:08:06.560 --> 00:08:10.879
So there's all sorts of different constraints and they measure the performance.
00:08:10.879 --> 00:08:12.000
And this is all open source.
00:08:12.000 --> 00:08:33.200
So you can go, like if you built an agent with your own domain expertise from Diamond Age that, you know, obviously has a different set of capabilities than, you know, Claude, like vanilla Claude, then you can go and you can benchmark it and you can show, like, in a like third-party verified objective way, like this is where we sit.
00:08:33.200 --> 00:08:39.440
There are also benchmarks for specific capabilities, for example, chart X or Chart QA.
00:08:39.440 --> 00:08:51.440
So your ability to for the LLM model to understand visual, visually presented information and like what kind of trend did this chart show, and what can I conclude from it?
00:08:51.440 --> 00:08:58.720
Those are nowhere near as good as the text understanding, as like for obvious reasons.
00:08:58.720 --> 00:09:03.679
Like that hasn't been the focus of training, but they're getting much better, like those multimodal models.
00:09:03.679 --> 00:09:15.120
And if you think about like the language that scientists use to communicate, like so often it's through chart or a figure, not even like a p-value or an effect side.
00:09:15.120 --> 00:09:20.639
Like so many of the decisions that we make as humans looking at at charts and and and figures and interpreting them.
00:09:20.639 --> 00:09:23.200
So I think that's like a key capability.
00:09:23.200 --> 00:09:27.759
But so I said, like there's there's capabilities, there's tasks, and then there's processes.
00:09:27.759 --> 00:09:29.120
Tasks are much harder.
00:09:29.120 --> 00:09:33.120
So tasks like actually Comp Biobench is a great example of tasks.
00:09:33.120 --> 00:09:44.639
It's end-to-end, like a you know, a pretty hard question that you would give to a PhD scientist, and they would go away and they would work for, you know, a month or so to code this up.
00:09:44.799 --> 00:09:49.120
Like, okay, here's a spatial transcriptomics data set.
00:09:49.120 --> 00:09:57.679
Can you like analyze all of the single cell profiles and like look at, you know, different gene differences between treatment and control, for example, right?
00:09:57.679 --> 00:10:02.000
Like that's you know, that's not that's not trivial, it's not totally protocolized.
00:10:02.000 --> 00:10:04.240
There are like many steps and gotchas.
00:10:04.240 --> 00:10:07.519
And so it end to end, like, how does it do on that test?
00:10:07.519 --> 00:10:15.600
Now you can imagine that that curating like a set of ground truth positive control responses to that is is very expensive.
00:10:15.600 --> 00:10:22.080
And then there's the third category, processes, which I think we just don't have yet today, but that's where the field is going.
00:10:22.080 --> 00:10:26.639
So a process would be like, oh, I generated this pre-IND document.
00:10:26.639 --> 00:10:30.320
It, you know, it has sources data from 10 different teams.
00:10:30.320 --> 00:10:37.440
They've all curated it subjectively through their different processes, and everybody's got a different thing.
00:10:37.440 --> 00:10:43.039
But like I somehow I trust that the team has put together something good and I'm gonna submit it to the FDA.
00:10:43.039 --> 00:10:45.039
Like that's a whole different level.
00:10:45.039 --> 00:10:48.399
And AI for knowledge workers in general just isn't on that level.
00:10:48.399 --> 00:10:51.279
So that was that was a long answer.
00:10:51.279 --> 00:10:55.600
But like I think, I think benchmarks and our understanding them is so important.
00:10:55.600 --> 00:11:00.399
And also like the field is moving so fast that they change every few months.
00:11:00.399 --> 00:11:06.399
And but I think it's just it's so critical for us to understand the capabilities of the tools that we're using.
00:11:06.799 --> 00:11:15.519
And it's great to know that there are some people out there professionally tracking those things because it is overwhelming to try to keep up with the movement in the field right now.
00:11:15.519 --> 00:11:16.320
It really is.
00:11:16.320 --> 00:11:17.519
I agree with you completely.
00:11:17.519 --> 00:11:18.320
It's tremendous.
00:11:18.320 --> 00:11:18.799
Yeah.
00:11:18.960 --> 00:11:19.200
Yeah.
00:11:19.200 --> 00:11:22.639
And I really applaud like Genentech in in particular, right?
00:11:22.639 --> 00:11:27.200
Like you think of benchmarks, public goods, that's sort of the realm of academia usually.
00:11:27.200 --> 00:11:38.240
And that sometimes can make it really hard to apply them in industry because we have like different sets of things that we weigh than academic research.
00:11:38.240 --> 00:11:44.480
So, so to have an industry group put some thought into it and put it out there for the community, I think is really laudable.
00:11:45.120 --> 00:11:46.000
Yeah, agreed.
00:11:46.000 --> 00:11:48.480
Then what about other knowledge work?
00:11:48.480 --> 00:11:50.559
You mentioned regulatory a bit.
00:11:50.559 --> 00:11:54.080
What about the non-bioinformatics side of things?
00:11:54.080 --> 00:11:55.120
What are you seeing there?
00:11:55.360 --> 00:12:04.000
Yeah, I'm seeing like there was this wave of adoption of like, I'll get it to write my emails and search my outlook and stuff like that.
00:12:04.000 --> 00:12:10.720
And then there's, I think there's really like two camps that I see now downstream.
00:12:10.720 --> 00:12:20.960
There's a set of people who maybe their company said you need you can't use LLMs until we spend nine months coming up with a company policy.
00:12:20.960 --> 00:12:27.919
And like you can't use Chat GPT as if everybody isn't using it in their private browser anyway to do their work.
00:12:27.919 --> 00:12:31.120
And then, but like, you know, they're they're they're handcuffed, right?
00:12:31.120 --> 00:12:35.679
You're copying and pasting and like, you know, trying to avoid putting company information.
00:12:35.679 --> 00:12:40.559
And then there's like super adopters, and it's remarkable to see them.
00:12:40.559 --> 00:12:53.200
Like the CSO that I work with, he is not a coder, but I showed him some stuff about like, you know, deterministic literature review and like how to sandbox his stuff on Claude Code.
00:12:53.200 --> 00:13:08.480
And now he has these like interactive HTML with drop-down menus and like it's this huge clinical trial review, and everything is source-verified, and it's like this, you know, beautiful work product.
00:13:08.480 --> 00:13:12.159
And I know he doesn't have the time to like put in a lot of time to this.
00:13:12.159 --> 00:13:21.279
I know he's just like gotten really good at leveraging AI systems for like really true, like dependable sort of knowledge work products.
00:13:21.519 --> 00:13:23.200
And this is all custom work that he did.
00:13:23.200 --> 00:13:24.960
It's not something that he bought from someone.
00:13:25.200 --> 00:13:30.240
No, no, it's just like it's Claude Code, and I showed him like, oh, here's a bunch of skills.
00:13:30.240 --> 00:13:41.919
So like Claude can learn use these to call tools to make, you know, go to the Clint Trials API or the PubMed API and not just use sort of the internet training corpus.
00:13:41.919 --> 00:13:54.320
And here's how, like, instead of using the chat box, like use one of the coding environments, and so that like unlocks it to write a interactive like HTML document for you.
00:13:54.720 --> 00:14:00.320
So we can benchmark the tools that we're using, but what about tools that someone is trying to sell to us?
00:14:00.320 --> 00:14:10.159
How do you how do you cut through the hike and and evaluate these marketed products and services and determine which of them have any value?
00:14:10.399 --> 00:14:14.960
Yeah, and there's so much out there, it can be exhausting, right?
00:14:14.960 --> 00:14:19.440
And everybody's got a new app and it it all looks good too, right?
00:14:19.440 --> 00:14:22.240
Because AI is really good at making things look good.
00:14:22.240 --> 00:14:36.000
And so, yeah, so I do some of this for VCs that I work with and I diligence, you know, new biotechs, or you know, sometimes it's like a software like SaaS platform or an AI discovery platform.
00:14:36.000 --> 00:14:50.080
So there's a set of questions that we go to, and you know, like it's doing it outside of like doing it in a non-con, like a non-confidential format, I think is can be difficult.
00:14:50.080 --> 00:14:59.200
But I think what I'll do in a non-con environment is is ask people, well, can you critique some other past models?
00:14:59.200 --> 00:15:11.600
And I think just looking at the founding team and their ability to think through and explain similar problems, like that's sort of a baseline level, right?
00:15:11.600 --> 00:15:23.919
And can they explain why, hey, like that was not technically sound or why that was technically sound, but it didn't have a good moat and for competitors.
00:15:23.919 --> 00:15:26.159
So, like going through that process.
00:15:26.159 --> 00:15:35.919
And like that's really common diligence, like half of the time you're diligencing the product, and half the time you're just diligencing the founding team because they're gonna have to pivot, right?
00:15:35.919 --> 00:15:40.960
Like you're gonna before your like between your seed and the time you get to the clinic, you're gonna pivot six times.
00:15:41.039 --> 00:15:44.720
So you're really like you are investing in the team more than anything.
00:15:44.720 --> 00:15:56.320
Once we get into like more of a under NDA, you know, I'll ask them to walk me through, okay, like if it's a foundation model, show me that the logistic regression didn't work.
00:15:56.320 --> 00:15:57.759
And like, did you did you try that?
00:15:57.759 --> 00:15:59.039
Did you try that?
00:15:59.039 --> 00:16:06.000
Like, and that sounds like simple, but like that's a sort of a place to start.
00:16:06.000 --> 00:16:10.399
Or, you know, what are other people's tools that you could have built on?
00:16:10.399 --> 00:16:16.320
Or like if you go back six months ago with what you know now, like how would you do things differently?
00:16:16.320 --> 00:16:23.600
So I think like a lot of this is evaluating the way they the way they think about things and the way they attack problems.
00:16:23.600 --> 00:16:29.200
But I think at the end of the day, there's no like free lunch here.
00:16:29.200 --> 00:16:39.919
Like AI and cutting edge AI and foundational models does rely on a lot of specialized training and you have to dig into the technical bits with the founder sometimes.
00:16:39.919 --> 00:16:43.039
And not all VCs want to do that today.
00:16:43.039 --> 00:16:53.679
And I think it makes it really harder for them to distinguish between like what's sort of the veneer of AI Shazam and what's going to be transformational.
00:16:53.919 --> 00:17:02.080
And for people who are thinking about not necessarily a startup, but rather a product that's on the market and who maybe doesn't have access to the founding team.
00:17:02.080 --> 00:17:03.759
Do you have advice for those folks?
00:17:04.079 --> 00:17:08.240
Yeah, like what kind of what kind of products are you thinking of is that?
00:17:08.720 --> 00:17:15.440
Many products claiming all kinds of things, like we will discover all of your new drug targets for you, right?
00:17:15.440 --> 00:17:27.359
There are companies that say things like that, and maybe, maybe they are doing it, but but let's say that I work for a pharma company and I want to know how do you how do you evaluate this?
00:17:27.599 --> 00:17:27.920
Yeah.
00:17:27.920 --> 00:17:30.160
My first question is show me your benchmarks.
00:17:30.160 --> 00:17:35.119
So not to be rep, not to be repetitive, but like a lot of people don't have an answer to that, right?
00:17:35.119 --> 00:17:40.480
And they'll say, like, oh, that you know, this is a great new process we've developed.
00:17:40.480 --> 00:17:48.319
And I'm like, okay, like if you just use clawed code without your proprietary harness, can you show me what that would look like?
00:17:48.319 --> 00:17:49.759
And do you have some metrics?
00:17:49.759 --> 00:17:57.680
Because like I think it's very reasonable to say you should have internal metrics for whether your product is getting better or not.
00:17:57.680 --> 00:17:59.519
I'm sure you're building right now.
00:17:59.519 --> 00:18:02.720
So can you like show me those internal metrics?
00:18:02.720 --> 00:18:09.920
Like, maybe there isn't some like unbiased third-party benchmark that's out there that's perfect for this.
00:18:09.920 --> 00:18:11.119
Maybe it's not a perfect fit.
00:18:11.119 --> 00:18:16.880
Maybe you can do something adjacent, but at least show me internally that you've thought about measuring this in an objective way.
00:18:17.200 --> 00:18:19.200
Are you enjoying the conversation?
00:18:19.200 --> 00:18:20.960
We'd love to hear from you.
00:18:20.960 --> 00:18:24.240
Please subscribe to the podcast and give us a rating.
00:18:24.240 --> 00:18:27.680
It helps other people find and join the conversation.
00:18:27.680 --> 00:18:31.599
If you've got speaker or topic ideas, we'd love to hear those too.
00:18:31.599 --> 00:18:33.839
You can send them in a podcast review.
00:18:34.160 --> 00:18:36.319
What about stuff out there that's not AI?
00:18:36.319 --> 00:18:44.160
I know you're I know we're in the AI till world, but like, what are we what are we missing because we're real busy with all of this AI stuff?
00:18:44.640 --> 00:18:45.039
Yeah.
00:18:45.039 --> 00:18:49.839
Well, we aren't investing enough, I think, in new targets, right?
00:18:49.839 --> 00:18:53.119
Biopharma is really in a Me Too kind of phase.
00:18:53.119 --> 00:18:59.039
We got a little bit risk averse with the bubble popping.
00:18:59.759 --> 00:19:03.359
I think there's a lot of biology left on the table.
00:19:03.359 --> 00:19:14.240
And it's I don't think it's because we're doing AI stuff, though I do think like tech bio and tech investing is sucking up like an enormous amount of capital.
00:19:14.240 --> 00:19:28.720
And like any idea you have to sell in drug dev, I think, you know, for some investor pools, you are like their opportunity cost is investing in the next like AI SaaS thing.
00:19:28.720 --> 00:19:33.039
And the growth trajectory of that field right now is is really enormous.
00:19:33.039 --> 00:19:35.599
So, like, yeah, there like there is some competition.
00:19:35.599 --> 00:19:39.279
But yeah, so I think we're leaving biology on the table.
00:19:39.279 --> 00:19:44.000
I think that AI can help there though, too.
00:19:44.000 --> 00:20:01.039
Like, I think that we can de-risk new targets and you know, sh like show that, okay, like, yeah, you're taking more target risk, but maybe you'll take less commercial risk because like you'll have a first in class, right?
00:20:01.039 --> 00:20:05.519
And I think we are, I'm hoping that we are just on the verge of that.
00:20:05.519 --> 00:20:14.559
Like I see, you know, so like if I think of computational biology and like generative AI, like the protein modeling is like very, very mature, right?
00:20:14.559 --> 00:20:17.200
And it was immediately obvious like how we're gonna use that.
00:20:17.200 --> 00:20:20.960
And no, it's not gonna like automatically just produce a drug for you, right?
00:20:20.960 --> 00:20:23.279
But it like it is still transformational.
00:20:23.279 --> 00:20:47.680
I think a lot of the target discovery stuff, for example, like the target discovery or target evaluation, like the DNA language models, predicting transcription, predicting single cell perturbation or even bulk RNA seq perturbation, those aren't like super exciting, but not as mature as the protein, the protein language modeling.
00:20:47.680 --> 00:20:57.599
And so I'm very hopeful that those can really de-risk some of those new targets for us and we don't have to keep developing the same old ones.
00:20:57.920 --> 00:20:59.279
Yeah, that would be amazing.
00:20:59.279 --> 00:21:14.079
Do you think the the the thing I'm skeptical about, the transcriptional profiling, modeling, and prediction is that the data that they used from you know the PDB database to build Alpha Fold and the structural you know models is was massive.
00:21:14.079 --> 00:21:21.519
And and the data that we have for transcription, I'm not convinced that the data we have now is good enough to build those models in the way that we need them.
00:21:21.519 --> 00:21:23.279
Do you have a do you have a ballpark?
00:21:23.279 --> 00:21:31.359
I mean, I'm totally asking you to speculate how much data do we actually need to build the the alpha fold for transcription?
00:21:31.680 --> 00:21:32.079
Yes.
00:21:32.079 --> 00:21:35.839
So I saw this quote and I'm trying to find it real time.
00:21:35.839 --> 00:21:39.200
I found it.
00:21:39.200 --> 00:21:41.440
Thank you, Google AI overview.
00:21:41.440 --> 00:21:44.240
So I like that is such a hard question, right?
00:21:44.240 --> 00:21:57.680
But there was a Viv Regev Genentech paper that took a stab at this, and I forget if she was publishing with Genentech or part of the virtual cell like Chan Zuckerberg, but they it was just like back of the envelope.
00:21:57.680 --> 00:22:06.640
What was what was the amount of Data for Chat GPT versus what is on the SRA.
00:22:06.640 --> 00:22:09.680
And actually, the SRA has 14 petabytes.
00:22:09.680 --> 00:22:11.200
And this was at the time of writing.
00:22:11.200 --> 00:22:12.319
So like whatever.
00:22:12.319 --> 00:22:14.400
It's right order of magnitude.
00:22:14.400 --> 00:22:21.200
And this is that's a thousand times bigger than the data set used to train chat GPT for.
00:22:21.200 --> 00:22:24.240
So like a bunch of asterisks is on that, right?
00:22:24.240 --> 00:22:26.640
Like, how do you measure real information content?
00:22:26.640 --> 00:22:29.920
It's not like, you know, ACGT and stuff like that.
00:22:29.920 --> 00:22:34.880
But like order of magnitude were a thousand times bigger.
00:22:34.880 --> 00:22:38.400
So like maybe there's some there of there.
00:22:38.400 --> 00:22:49.200
And that was for a virtual cell, which I think, you know, it's it's only looking at the SRA and and and it's like measuring against this target of a virtual cell, which I think is like a harder target.
00:22:49.200 --> 00:22:58.559
So yeah, the from some people who've thought about this more and much smarter than I am, there's like an order of magnitude.
00:22:58.880 --> 00:23:06.720
Yeah, I'm really curious about the the basically the information content difference given that SRA is full of human data and humans are very similar to each other, right?
00:23:06.720 --> 00:23:10.240
Like how much differential data is there among all of those sequences?
00:23:10.240 --> 00:23:18.480
I mean I'm not asking you to answer me because I don't I don't think either of us knows, but yeah, it's a really great question about like what what will it really take to build these?
00:23:18.720 --> 00:23:18.960
Yeah.
00:23:18.960 --> 00:23:28.480
I mean, I guess on the other side, if you wanted to critique like the internet training there, a whole bunch of like Reddit garbage and Quora garbage.
00:23:28.480 --> 00:23:33.359
Like I don't know how that was all filtered, but it it cuts both ways.
00:23:33.759 --> 00:23:34.480
It does indeed.
00:23:34.480 --> 00:23:39.920
Okay, then to move again away from AI, although I'm sure we'll be right back there.
00:23:39.920 --> 00:23:41.119
Sorry.
00:23:41.119 --> 00:23:42.160
That's fine.
00:23:42.160 --> 00:23:43.200
This is what it's all about.
00:23:43.200 --> 00:23:44.559
Biology areas.
00:23:44.559 --> 00:23:48.400
What are you finding that's popular or interesting for investment right now?
00:23:48.400 --> 00:23:57.519
And I'm thinking along the lines of disease modalities, sorry, therapeutic modalities, diseases, technology platforms, what are people investing in?
00:23:57.519 --> 00:24:00.559
What do they want to invest in that people should start building?
00:24:00.799 --> 00:24:01.279
Yeah.
00:24:01.279 --> 00:24:06.960
I think so obviously cardiometabolic is huge, right?
00:24:06.960 --> 00:24:17.200
Anything that you can combine with a glip or modify with a glip, or how can we look at patient subpopulations, right?
00:24:17.200 --> 00:24:20.640
So these are like miracle drugs, but they're first generation.
00:24:20.640 --> 00:24:37.359
And like the idea of painting that patient population with a single brushstroke measured by like percent body weight loss is like it's such such a yeah, it's such a broad brushstroke and such a sort of gross measure.
00:24:37.359 --> 00:24:45.359
So like obviously there's a lot of funding there, and it's like you know, very impactful for a huge patient population, huge TAN too.
00:24:45.359 --> 00:24:58.400
Miro is is still hot, and I can't really put it point a finger at the source other than like there, you know, there was the Alzheimer's approval, and that made it a little bit exciting.
00:24:58.400 --> 00:25:03.920
Again, there's another company that I just that had the schizophrenia readout.
00:25:03.920 --> 00:25:09.440
I can't remember their name, but I think there's there's been a little bit of of momentum in Neuro.
00:25:09.440 --> 00:25:26.960
In terms of tech platforms, I think spatial is getting pretty mature, but like high throughput perturbations and having proprietary data on that is seen as really valuable, right?
00:25:26.960 --> 00:25:33.599
Because we we have a pretty good measure of like what is, you know, steady state single cell data.
00:25:33.599 --> 00:25:36.079
And that's been transformative for drug development.
00:25:36.079 --> 00:25:54.559
Actually, like the the effect size on your probability of your drug getting approved if it was based on a single cell or if there's like a clear single cell data signal is about 2x, which is like the same figure that people quote for, you know, if it if your drug was has a genetic signal, a human genetic signal.
00:25:54.559 --> 00:26:04.000
So it's like it's a pretty big, it's a it's a powerful technology, but we just I think we can we're we're trying to build on it by looking at the the perturbations, and they're they're hard to predict, right?
00:26:04.000 --> 00:26:05.680
It's out of distribution.
00:26:05.680 --> 00:26:09.119
So yeah, so those are some of the themes that I see.
00:26:09.279 --> 00:26:14.720
And of course, like I'm more interested in the ones that intersect with AI capabilities, right?
00:26:14.720 --> 00:26:18.000
So, so technology platforms and high throughput data.
00:26:18.000 --> 00:26:25.920
And I've seen firsthand, so I was working closely with a founding team with founded around a spatial tech.
00:26:25.920 --> 00:26:28.240
So I'm one of the advisors there.
00:26:28.240 --> 00:26:52.160
And just like through the back half of last year, like I made a an estimate for like, okay, what would it take for us to pull in, you know, all the spatial data in our top three indications and like reprocess it in a consistent way and like combine it all and put like some of the founders' proprietary models on top.
00:26:52.160 --> 00:26:57.119
And my estimate was like, you know, two FTEs for six months.
00:26:57.119 --> 00:27:04.240
And then I looked at that same exercise in November when we were, you know, doing some DD and I was refreshing it.
00:27:04.240 --> 00:27:09.279
And I was like, wow, I, you know, I could do this myself in like two weeks.
00:27:09.279 --> 00:27:11.920
And so just like that was that was very concrete for me.
00:27:11.920 --> 00:27:19.359
Yeah, it's kind of anecdotal, but like I did the same exercise at two points in time, just like six months apart, and and it was amazing to see.
00:27:19.359 --> 00:27:33.119
So that is to say that like my personal investment thesis is that if if you have something where AI can really accelerate it, like drug development is is hard.
00:27:33.119 --> 00:27:45.359
And if so, if it's a meaningful, if it's a meaningful step and and AI can accelerate it, you know, 100x, then that's that's part of your whole investment thesis.
00:27:45.680 --> 00:27:46.480
That makes perfect sense.
00:27:46.480 --> 00:27:56.880
And yeah, that that working with those data sets, that's another thing that we see is just it's tremendously faster to collect data and harmonize it together than it used to be.
00:27:56.880 --> 00:27:58.160
It's fantastic.
00:27:58.160 --> 00:28:03.359
Well then, you are going to be teaching a workshop by OIT World this year.
00:28:03.359 --> 00:28:06.319
Do you wanna do you wanna talk a little bit about what you're teaching?
00:28:06.319 --> 00:28:07.839
What are you covering in that workshop?
00:28:08.079 --> 00:28:08.559
Yeah, yeah.
00:28:08.559 --> 00:28:13.519
The workshop is focused on agentic coding for computational biologists.
00:28:13.519 --> 00:28:27.599
So you're already a, you know, a great scientist, great computational biologist, but you haven't, you know, you have a day job, you haven't had time to keep up with the like frontier of AI tools that are available to you.
00:28:27.599 --> 00:28:41.359
And I think a lot about like, you know, this 95.5 sort of divide that I mentioned, where like I think 5% of our code is is written by AI and it's like 95% in software dev.
00:28:41.359 --> 00:28:44.720
So, what can we learn from classic software dev?
00:28:44.720 --> 00:28:47.680
What practices do we need to learn?
00:28:47.680 --> 00:28:50.079
And what do we need to adapt?
00:28:50.079 --> 00:28:55.759
Because their models don't have the same capabilities in the scientific domain.
00:28:55.759 --> 00:29:09.759
How can like, what are the most efficient ways for us to inject domain expertise, domain-specific tools to sort of rectify the fact that their models weren't trained as much on our problems as they were on mainstream coding?
00:29:09.759 --> 00:29:20.400
So what those are the two pieces and you know, adapting your software dev practices, adopting, adopting their software dev practices and and and adapting their their models.
00:29:20.400 --> 00:29:27.839
And I'm doing that with a friend and former colleague of mine, Ryan Belmore, and we're doing it like hands-on workshop.
00:29:27.839 --> 00:29:32.559
You will do this and you will leave with like an awesome AI coded product.
00:29:33.039 --> 00:29:36.640
And you and Ryan together, I think that's gonna be a fantastic workshop.
00:29:36.640 --> 00:29:38.960
That's I didn't realize we were working with Ryan.
00:29:38.960 --> 00:29:39.759
That's amazing.
00:29:39.759 --> 00:29:40.240
Okay.
00:29:40.240 --> 00:29:41.440
That's gonna be great.
00:29:41.440 --> 00:29:43.200
And so that's at BioIT Worlds.
00:29:43.200 --> 00:29:45.920
And which day is that on BioIT World?
00:29:46.079 --> 00:29:50.640
I think that is the it's in the afternoon on the 19th.
00:29:50.640 --> 00:29:52.960
So yeah, I'd love to see people there.
00:29:52.960 --> 00:29:58.480
And if you have questions about the class in in the lead up, I'm I'm happy to talk more.
00:29:58.480 --> 00:30:05.200
And Eleanor, I know you're doing a bunch of stuff and Diamond Age is too, bio IT.
00:30:05.200 --> 00:30:05.519
Right.
00:30:05.519 --> 00:30:06.960
Do you want to tell me about that?
00:30:07.279 --> 00:30:09.440
Oh, oh, well, now you're interviewing me.
00:30:09.440 --> 00:30:10.559
Oh my gosh.
00:30:10.559 --> 00:30:14.079
Yeah, I'm giving the Trends from the Trenches talk.
00:30:14.079 --> 00:30:19.119
So I am gonna be, I see, I am not entirely AI pilled quite as much as you are.
00:30:19.119 --> 00:30:30.480
I see a lot of a lot of benefit from it, but there are also some things that I am absolutely going to try to pop the bubble about because I think there's some stuff that is not necessarily very helpful.
00:30:30.480 --> 00:30:34.000
But there's still also a lot that's that's that is helpful.
00:30:34.000 --> 00:30:37.039
And so that's what my talk is gonna be about is things like this.
00:30:37.039 --> 00:30:37.920
What are the trends?
00:30:37.920 --> 00:30:45.279
What am I seeing from you know my customer base, my my colleagues, my network, what's changing in our work?
00:30:45.279 --> 00:30:47.039
That's what I'm gonna be talking about there.
00:30:47.039 --> 00:30:48.480
So I'm very excited about it.
00:30:48.480 --> 00:30:50.640
I was really, really happy to be asked to.
00:30:51.680 --> 00:30:52.960
Yeah, I'm I'm looking forward to that.
00:30:52.960 --> 00:30:57.359
And I love it like a feisty debate, and I love when people are bubble popping.
00:30:57.359 --> 00:31:06.880
And I think like, I think it can be simultaneously 100% true that like AI is amazing and transformative for computational biology, and AI is like overhyped and over.
00:31:07.519 --> 00:31:12.960
Yes, okay, yes, and that is absolutely the case because some parts of it are actually that transformative.
00:31:12.960 --> 00:31:23.680
Coding the AI, the agentic coding agents and the structural biology prediction are the day-to-day life of the people who use those things are completely different than they were a year, two years ago.
00:31:23.680 --> 00:31:28.880
Like every day is different now because of the impact of those tools, and you can't really understate that.
00:31:28.880 --> 00:31:31.279
And then there are some other things that are just not sense.
00:31:31.279 --> 00:31:34.319
So yeah, yeah, absolutely.
00:31:34.319 --> 00:31:35.680
And we need to benchmark them.
00:31:35.680 --> 00:31:46.400
So then to wrap up, you're giving a workshop and you're gonna be teaching, but maybe just as a quick preview, like what would you recommend just as a short thing, like for people with getting started?
00:31:46.400 --> 00:31:46.799
Yeah.
00:31:46.799 --> 00:31:49.920
What what kind of recommendations do you have for those folks?
00:31:50.400 --> 00:31:50.799
Yeah.
00:31:50.799 --> 00:32:05.839
So I think first you have to see a reason to believe because we're all busy and it's hard to, you know, devote time to learning a new technology unless you have the confidence that it's gonna pay off.
00:32:05.839 --> 00:32:08.720
And so how do you see reason to believe?
00:32:08.720 --> 00:32:11.920
So, like go to this Genentech paper.
00:32:11.920 --> 00:32:26.559
You can see all of the prompts that it gave and all of the results, and they have it all up in GitHub, and you can see that, you know, this is what a vanilla agentic coding off-the-shelf model can do with no help.
00:32:26.559 --> 00:32:30.480
So imagine what it could do with my help and and my domain expertise.
00:32:30.480 --> 00:32:33.039
So see, see that reason to believe.
00:32:33.039 --> 00:32:35.920
And then, and then I think you can.
00:32:35.920 --> 00:32:47.519
I've I've had experience, I've taught like many workshops with my clients and like one-on-ones, and in the span of, you know, 30 to 60 minutes, have transformed someone.
00:32:47.519 --> 00:32:57.920
Like they'll come back two days later and they're like, oh my gosh, it can just like it wrote all my code for me, it wrote my tests, I wrote my docs, I can just like, I can write the validations, I can like review the code.
00:32:57.920 --> 00:33:09.440
So it, it really like if you can invest the 60 minutes with somebody who's proficient in these tools, but also understands your your work, right?
00:33:09.440 --> 00:33:14.640
It can't be, you know, your neighbor who does it, right?
00:33:14.640 --> 00:33:17.920
Like they have to under, they have to understand your work.
00:33:17.920 --> 00:33:20.799
They have to be able to speak your language and have the domain expertise.
00:33:20.799 --> 00:33:28.319
So I really think like, yeah, if you can put in, you know, that's that's two hours, and if you can commit that, like this isn't going away.
00:33:28.319 --> 00:33:30.720
This is it's gonna change the way we work.
00:33:30.880 --> 00:33:33.359
Yeah, yeah, it already has, but not for everybody.
00:33:33.359 --> 00:33:36.799
So, and that's what you're doing is is making it more available.
00:33:37.039 --> 00:33:39.759
And I guess I should plug myself that I do these workshops.
00:33:39.759 --> 00:33:48.240
I'm sure they could also call up Diamond Age, and Diamond Age has a a bunch of people who with expert expertise in agenda.
00:33:48.720 --> 00:33:51.200
Okay, so thank you for taking the time.
00:33:51.200 --> 00:33:52.880
It's been a really fun conversation.
00:33:52.880 --> 00:33:56.880
It's been fun, and let's do it again in a in a few months, and everything will be different.
00:33:56.880 --> 00:33:58.400
Everything will be different.
00:33:58.400 --> 00:34:00.960
Yes, actually, yes, let's do that.