00:00:00.160 --> 00:00:09.759
Hardware and software verifications, so you can actually verify things about chip design because they are actually quite clearly specified.
00:00:10.080 --> 00:00:13.279
Most podcasts talk to founders after the story's already written.
00:00:13.439 --> 00:00:17.039
When the latest trade is announced, the product has shipped and the narrative is clear.
00:00:17.359 --> 00:00:20.800
We wanted to find out why they decided to take the founder leap in the first place.
00:00:20.960 --> 00:00:24.800
What idea captured their imagination and how they handle the pressure to build?
00:00:24.960 --> 00:00:26.079
Welcome to First Commit.
00:00:26.160 --> 00:00:27.039
I'm Laura Hamilton.
00:00:27.280 --> 00:00:30.079
I'm Adea Lon, and we're investors at Federal Capital.
00:00:30.480 --> 00:00:33.600
We spend a lot of time with technical founders who are on the cutting edge.
00:00:33.759 --> 00:00:37.600
If you're thinking about what it's like to start a company in the AI era, this one's for you.
00:00:37.840 --> 00:00:38.960
This is First Commit.
00:00:39.119 --> 00:00:40.320
Let's get into it.
00:00:42.000 --> 00:00:48.079
Today we are welcoming Karina Hong, who's the CEO and founder of Axiom, building the AI mathematician.
00:00:48.240 --> 00:00:49.039
Welcome, Karina.
00:00:49.439 --> 00:00:50.560
Hi, great to be here.
00:00:50.880 --> 00:00:52.159
Excited to hop in.
00:00:52.399 --> 00:00:53.520
So let's go back.
00:00:53.679 --> 00:00:59.920
You're at Stanford, you're young, you started a company that has already raised over$200 million 15 months ago.
00:01:00.079 --> 00:01:05.439
What did you know in March 2025 that made you sure it was the right moment to start Axiom?
00:01:05.760 --> 00:01:11.040
Yeah, I think March 2025 was an interesting time, just in general in the islandscape.
00:01:11.200 --> 00:01:15.920
We are about six months after the first reasoning model launch, OpenAS01.
00:01:16.400 --> 00:01:26.719
We are more than, I think, well, actually, we're more than 2.5 years after LIM4, which is the programming language that we depend on to do math proofs.
00:01:26.799 --> 00:01:30.879
The latest version, LIM4, has been rolled out in September 2023.
00:01:31.200 --> 00:01:32.560
Am I calculating right?
00:01:32.799 --> 00:01:35.200
I think it's it's roughly rough, oh, it's 1.5 years.
00:01:35.599 --> 00:01:36.480
We'll plug in an axiom.
00:01:36.560 --> 00:01:36.959
Yeah, yeah, yeah.
00:01:38.239 --> 00:01:54.719
It's kind of an interesting sort of intersection of time between kind of the reasoning model's capability, or not just sort of informal reasoning, but also coding, as well as formal math development, making lean a lot less finicky and much better suited for industry-scale engineering.
00:01:54.799 --> 00:01:57.920
So those three things I think are happening at the same time.
00:01:58.159 --> 00:02:05.120
It's also a time where people are feeling pressure, I think, from various big tech places to play catch up.
00:02:05.599 --> 00:02:14.639
So there are a lot of, I think, efforts and initiatives to try to train close-to-frontier models, but didn't actually necessarily pay off.
00:02:14.800 --> 00:02:16.800
So a lot of churns at big places.
00:02:16.879 --> 00:02:18.639
So talents are starting to float.
00:02:18.800 --> 00:02:20.400
There was no new lab at the time.
00:02:32.080 --> 00:02:32.719
A lot has changed.
00:02:32.960 --> 00:02:40.800
So I'm sure our re our listeners are familiar with reasoning models and chain of thought one and the whole evolution, but maybe much less familiar with lean.
00:02:41.360 --> 00:02:45.919
Maybe you want to share a bit what lean is and how it came to be.
00:02:46.319 --> 00:02:48.639
Lean builds something called type theory.
00:02:48.719 --> 00:02:52.560
It has a lot of objects, each have its type, and there are like logical dependencies.
00:02:52.960 --> 00:02:54.159
So not type script.
00:02:54.400 --> 00:02:54.639
No.
00:02:54.960 --> 00:02:55.520
Type theory.
00:02:55.759 --> 00:02:56.400
No, type theory.
00:02:56.639 --> 00:02:56.719
Okay.
00:02:56.960 --> 00:02:59.520
TypeScript is actually kind of on the other end of the spectrum.
00:02:59.599 --> 00:03:01.199
It's a dynamically typed language.
00:03:01.280 --> 00:03:02.479
So it's quite hard to deal with.
00:03:02.639 --> 00:03:06.800
Well, actually, no, TypeScript is statically typed versus JavaScript is dynamically typed.
00:03:07.039 --> 00:03:10.479
So Lynn is this sort of formal language where you can write mass proofs in.
00:03:10.560 --> 00:03:13.360
You can also do other things in Lynn like you can build an autograd in Lin.
00:03:13.520 --> 00:03:14.639
Nothing's solving you.
00:03:14.800 --> 00:03:21.680
So it's a functional programming language where you can do various things in, but you can also write mass proofs that become computer programs where you can execute.
00:03:21.840 --> 00:03:28.080
So once you execute it and you see a check mark, you can be pretty confident that the statement that you asked the system to prove has been proven.
00:03:28.240 --> 00:03:30.879
Now, obviously, if your statement is wrong, then there's no guarantee.
00:03:31.039 --> 00:03:40.159
So people still need to check the statements compared to say three-line, five-line statements versus like thousands of lines of proof, you're definitely making the work a lot, a lot, a lot better.
00:03:40.479 --> 00:03:43.439
So a way to think about it is a proof compiler of sorts.
00:03:43.759 --> 00:03:44.879
In a way, yes, that's right.
00:03:45.039 --> 00:03:50.400
When it's developed in, I think, 20, 2013, 2014 by Leo DeMora and folks.
00:03:50.560 --> 00:03:51.680
He was at Microsoft.
00:03:51.919 --> 00:03:54.719
It was a great, great nonprofit, the Lynn FRO.
00:03:54.960 --> 00:03:55.280
Very cool.
00:03:55.360 --> 00:03:55.759
Thank you for that.
00:03:55.919 --> 00:03:59.280
Can you just touch on like the broader vision of Axiom and where you guys are today?
00:03:59.599 --> 00:04:03.039
Back to March 25, it's like these three things are coming together.
00:04:03.280 --> 00:04:06.960
There has not been an industry-scale effort to make AI for mass a reality.
00:04:07.199 --> 00:04:10.479
I was seeing Google DeepMind doing work on Mass Olympiads.
00:04:10.719 --> 00:04:13.840
It was also at a time where I feel like that effort has kind of ramped down.
00:04:14.000 --> 00:04:21.680
There are kind of efforts starting in China, Deep Seek and Biden's, but it wasn't very clear that geographically that's gonna work for me personally.
00:04:21.920 --> 00:04:27.839
And so yeah, I think I wanted to start this effort to make AI for mass discoveries a reality.
00:04:28.000 --> 00:04:33.199
One thing I was very much insisting on is research math, research mass, research math.
00:04:33.360 --> 00:04:38.160
That has to be the priority and goal beyond Olympiad mass benchmark problems.
00:04:38.319 --> 00:04:47.519
Because the thing is, if you think of Olympiad mass benchmark problems and the role they play in a lot of models, Avel by, you know, Frontier Large Language Model.
00:04:47.600 --> 00:04:48.720
There's so many benchmarks there.
00:04:48.800 --> 00:04:50.879
There are like even like a healthcare benchmark somewhere.
00:04:51.120 --> 00:04:54.399
There's RKGI that shows reasoning and superintelligence.
00:04:54.639 --> 00:04:57.199
And there's nothing sort of special about math.
00:04:57.279 --> 00:04:58.800
And for me, math is really special.
00:04:58.879 --> 00:05:01.920
And for the people that later joined Axiom, math is very special.
00:05:02.000 --> 00:05:04.560
It is the fundamental of a lot of scientific domains.
00:05:04.800 --> 00:05:08.879
It can be expressed in code and by the Carver Howard correspondence.
00:05:08.959 --> 00:05:15.600
And then for for the existence of Lynn and other formal languages, like other versions like Isabel or Coch.
00:05:15.759 --> 00:05:28.160
And math can sort of, math and code can control a lot of scientific domains, and you can run experiments about the physical world, and you already have rewards, you can control reward, it's a reward kind of flywheel compounding.
00:05:28.319 --> 00:05:33.600
So math is very special for us for our personal passion reason and for these reasons.
00:05:33.759 --> 00:05:40.000
And just like I believe, for example, Encerampic believes in coding is going to transfer, we also believe that math is going to transfer.
00:05:40.160 --> 00:05:48.000
So that was that was, I think, the initial spark of a shin in March 2025, which is, and I was very sure it's gonna happen.
00:05:48.079 --> 00:05:55.680
I was like, by 2026, late half, we're going to have AI approved at least a research math result.
00:05:56.160 --> 00:05:57.680
Turns out that's way earlier.
00:05:57.759 --> 00:06:01.199
I mean, beginning of the year, we wrote out seven, eight papers now on archive.
00:06:01.519 --> 00:06:09.759
Two of them recently got accepted to professional math journals, Archive de Mathematique, and Induction Nice TK.
00:06:10.160 --> 00:06:21.279
So we have seen instances from Axiom Prover and also from Google DeepMinds Alethea agent and also MTMS, amateur mathematicians using frontier models in all three cases, prove a lot of research results in math.
00:06:21.600 --> 00:06:31.839
Karina, maybe you can walk us through a bit what do you think is wrong with current AI systems and why they're not good enough to do formal math.
00:06:32.319 --> 00:06:33.279
There is this belief.
00:06:38.800 --> 00:06:41.199
And they also I think don't like Excel now.
00:06:41.600 --> 00:06:45.600
So it's like they're at the interesting, they're at an interesting point, right?
00:06:45.759 --> 00:06:48.639
Where you see all these like AI for Excel companies.
00:06:48.959 --> 00:06:50.079
They don't want to do Excel.
00:06:50.160 --> 00:06:53.199
They don't want to, like, you know, reason in that abstraction.
00:06:53.680 --> 00:06:56.000
They don't trust large language models output.
00:06:56.240 --> 00:06:59.920
Um, even the coding models, which I think have a lower chance of hallucinating.
00:07:00.000 --> 00:07:03.839
I think it transfers that way, are still hallucinating when it comes to math.
00:07:04.160 --> 00:07:19.040
And we're at a point where we feel like there's so much need for something that can be a thought partner, but can also give like actual correct like conclusions, quantitative ones.
00:07:19.439 --> 00:07:22.480
I think that's kind of the future we're kind of starting to see.
00:07:23.040 --> 00:07:30.560
And I know that there's a history toward, okay, well, if you want to kind of try to formalize language, you're going to hit this roadblock and the other one.
00:07:30.800 --> 00:07:38.879
Chomsky has all this work and and attempts, and other people were also trying to do like full-stack for language, I think, since like 1980s.
00:07:39.040 --> 00:07:43.360
But we're finally at the point where we feel like the the intersection of them is the reality.
00:07:43.439 --> 00:07:49.360
I mean, it is not purely formally verify everything because we'll all kind of wait until the end of the universe.
00:07:49.519 --> 00:07:55.439
It's also not like wipe coding to the end of the universe because then you're going to see blow-up at some point.
00:07:55.680 --> 00:07:57.519
The reality is somewhere in the middle.
00:07:57.759 --> 00:08:05.519
I think current AI lacks the ability to consistently, reliably give you the correct output.
00:08:05.680 --> 00:08:11.040
And there are value in getting that or you know, not making the mistakes or catching the edge cases.
00:08:11.199 --> 00:08:17.839
There's another additional layer of value of trust, trust that you can actually deploy that, you know, output.
00:08:18.000 --> 00:08:32.320
And I think people who are working, for example, AI for science, talking about like for them, like the decision to go into like as actual physical word, like experiments, very expensive and costly, and a lot of theoretical level stuff needs to be like 100% like soundproof.
00:08:32.480 --> 00:08:35.200
And I think those are the realities we're seeing.
00:08:35.279 --> 00:08:43.759
And if there is a technical, technically plausible solution, actually with not much technical risk because it's proven, that is where we should go.
00:08:43.840 --> 00:08:45.360
And that's kind of where we want to go next.
00:08:45.679 --> 00:08:47.679
To help conceptualize this a bit, walk us through.
00:08:47.759 --> 00:08:55.519
So lean is this project you're using, but help us understand how Axiom's own model fits in.
00:08:55.679 --> 00:08:56.960
What's the interaction model?
00:08:57.360 --> 00:09:04.960
Lean is the substrate, and I'm sure reasoning comes a lot in the feedback from lean serves your model as a reasoning layer.
00:09:05.200 --> 00:09:06.879
Yeah, so Axiom Prover is a system.
00:09:06.960 --> 00:09:09.120
It's an ensemble of multiple models.
00:09:09.279 --> 00:09:10.720
Some are informal, some are formal.
00:09:10.960 --> 00:09:15.600
There are a set of tools that we have open released called AXO, Axiom Lean Engine.
00:09:15.919 --> 00:09:18.720
We wrote it out, I think, like about eight weeks ago in March.
00:09:18.960 --> 00:09:25.039
And people around the world have been using different kinds of ATR improving models on top of Axel infrastructure.
00:09:25.279 --> 00:09:33.200
And this whole system, the models underlying it, trains on like lean data, which are mass proofs written in lean, the programming language.
00:09:33.440 --> 00:09:42.080
And so therefore the system can take any natural language questions and produce a lean proof, which to me at the beginning felt like magic.
00:09:42.240 --> 00:09:44.799
I mean, this is this is magic because of a few reasons.
00:09:44.960 --> 00:09:48.080
One, everyone was telling me that data is really scarce, right?
00:09:48.159 --> 00:09:56.080
You're functioning in a data scarce domain, which is could could be a reason why I say like the frontier labs are not doing it, because they have a lot of compute.
00:09:56.399 --> 00:10:03.440
They will have much more of a strategic advantage to play in the data rich field, to utilize those compute.
00:10:03.600 --> 00:10:05.039
And one is data scarce.
00:10:05.120 --> 00:10:10.480
The other one is that just it hasn't been shining in the Mass Olympiads before.
00:10:10.639 --> 00:10:14.879
You always see the informal models are like really crushing the formal ones.
00:10:15.120 --> 00:10:16.960
I think that was a lesson people learned.
00:10:17.279 --> 00:10:20.559
Maybe quickly explain the difference between a formal model and an informal model.
00:10:20.960 --> 00:10:26.240
So informal model is a model that produces output mostly in natural language.
00:10:26.480 --> 00:10:41.840
And when asked to do math in the specific context, and the formal model is one that when I give you math question in natural language or lean, it gives me the output in in lean, which you can be sure that is a hundred percent correct by simply running the program.
00:10:42.240 --> 00:10:46.799
Interestingly, you don't need to understand once it's translated into a formal format.
00:10:47.120 --> 00:10:47.200
Yeah.
00:10:47.360 --> 00:10:51.919
And you can't do this today with the current models, like foundation models, because they're not the capabilities.
00:10:52.320 --> 00:10:57.279
Yeah, people find general purpose models to be lousier link capability.
00:10:57.679 --> 00:11:01.039
So an example would be like not in math, just to simplify for our listeners.
00:11:01.200 --> 00:11:07.679
So let's say I'm designing a system, and in simple language, I say as a payment system, I want to say uh never pay twice.
00:11:07.759 --> 00:11:09.679
It's a banking bank transfer system.
00:11:09.919 --> 00:11:14.240
That sounds very intuitive for us as humans, but as a system, it might fail at the database layer.
00:11:14.399 --> 00:11:19.120
It might have network, a timeout, a user might be clicking too many times.
00:11:19.360 --> 00:11:26.960
So lean would translate that natural language statement into a set of clear requirements that then can be proven.
00:11:27.200 --> 00:11:28.159
Proven in a formal way.
00:11:28.559 --> 00:11:30.080
So there are obviously nuances to that.
00:11:30.159 --> 00:11:31.279
So for example, of course.
00:11:31.440 --> 00:11:32.320
Whereas simplifying it.
00:11:32.720 --> 00:11:33.279
Yeah, yeah, yeah.
00:11:33.440 --> 00:11:48.000
But in a lot of the cases that we observe is that if you have the specification in mind, which is the things you want to verify, in this case, never pay twice, it's proving is, you know, which is Axiom Prover's job and becomes a lot easier.
00:11:48.399 --> 00:11:48.559
Yeah.
00:11:48.879 --> 00:11:53.039
For the set of these statements, yeah, the following claim is true or not true.
00:11:53.360 --> 00:12:03.600
I I think that Axiom Prover has surpassed a level of reasoning ability, formal reasoning ability, that if you give me clear specification, I have a high chance of proving it.
00:12:03.919 --> 00:12:09.679
If I don't prove it, you will not be misled into thinking that it is correct or satisfied.
00:12:09.919 --> 00:12:15.120
But if you don't give me the specification, you say I want this to be safe or good.
00:12:15.360 --> 00:12:15.759
Yeah.
00:12:16.000 --> 00:12:17.120
It becomes a lot trickier.
00:12:17.440 --> 00:12:19.840
How do you envision people using Axiom?
00:12:19.919 --> 00:12:23.120
Or is it you're always on math co-pilot?
00:12:23.519 --> 00:12:28.159
I think that will be a very interesting sort of like, you know, first kind of research, you know, product.
00:12:28.320 --> 00:12:33.759
But the the general sort of domains we're trying to go in are hardware and software verifications.
00:12:33.919 --> 00:12:41.759
So you can actually verify things about chip design because they are actually quite clearly specified.
00:12:41.919 --> 00:12:42.080
Right.
00:12:42.240 --> 00:12:46.399
So for example, I will never want a race condition to happen.
00:12:46.639 --> 00:12:52.720
Or I want to make sure that, you know, not certain kinds of components are full at the same time, because otherwise that's a shin.
00:12:52.960 --> 00:12:56.080
All these things are already presented as clear specs.
00:12:56.320 --> 00:13:00.080
And with those specs, we could try to have action prover generate.
00:13:00.480 --> 00:13:11.840
Now there's a little bit of advantage of that versus existing EDA tools, in that existing EDA tools, a lot of them rely on something called SMT, SMT software, which is Boolean logic.
00:13:12.080 --> 00:13:20.720
And it can get very high complexity very quickly because your exponentially, you know, this exact space becomes ridiculous.
00:13:23.519 --> 00:13:23.759
Right.
00:13:23.919 --> 00:13:26.320
And that becomes quickly unmanageable.
00:13:26.480 --> 00:13:28.399
And this will not happen with lean.
00:13:28.720 --> 00:13:34.000
And in the case of software verification, we're seeing obviously being a more messy domain than chip verification.
00:13:34.320 --> 00:13:35.279
Just because of the abstractness.
00:13:35.519 --> 00:13:38.480
Yeah, people like safe software, and like you don't know what that means.
00:13:38.720 --> 00:13:47.679
But in a lot of places, actually, a large amount of high complexity industry projects in the B2B world, they do have specifications, right?
00:13:47.759 --> 00:13:50.320
They want to make sure that there's no security loophole.
00:13:50.480 --> 00:13:57.440
They want to be able to formalize, say, a hypervisor as a certain component instead of humans, right?
00:13:57.519 --> 00:14:02.879
Like some sort of cryptographic assurance about some what whatever workload or transaction.
00:14:03.840 --> 00:14:04.080
Yeah.
00:14:04.240 --> 00:14:14.639
So I think those are the starting cases, which is that everything you can define, you can execute, and just like that, everything you can specify, you can prove it.
00:14:15.039 --> 00:14:17.679
It sounds like there's a lot to do with this.
00:14:17.840 --> 00:14:18.080
Yeah.
00:14:18.320 --> 00:14:24.960
It also, my guess will be, you know, we're a bunch of nerds in this room, but most people, this is a bit too abstracted for them.
00:14:25.120 --> 00:14:31.279
You had this intuition in in March 25, and one was just coming out, and lean was already there.
00:14:31.440 --> 00:14:38.159
I wonder a lot of the applications, was that something, the use cases, is that something you had in mind prior?
00:14:38.240 --> 00:14:38.320
Yeah.
00:14:38.559 --> 00:14:39.440
Or it evolved over time.
00:14:39.759 --> 00:14:41.200
You see, I I personally don't.
00:14:41.360 --> 00:14:42.720
I have the ambition for math.
00:14:42.799 --> 00:14:44.240
I want like scientific breakthroughs.
00:14:44.320 --> 00:14:45.279
I was I was so excited.
00:14:45.360 --> 00:14:50.320
I was like, I remember it was honestly like I had to start this company because I already couldn't do anything else.
00:14:50.480 --> 00:14:53.440
Like it was a state of almost like total consumption.
00:14:53.519 --> 00:14:55.519
Like, I mean, like you're high.
00:14:55.600 --> 00:14:56.480
You're high on your own supply.
00:14:56.799 --> 00:14:57.519
I am high, yes.
00:14:57.600 --> 00:15:02.320
Like I and I was just like so confused about I've never seen this state of me like before.
00:15:02.559 --> 00:15:08.240
I mean, like, maybe if I'm crazy about a boy, like, you know, like like that, that sort of in love with the idea.
00:15:08.480 --> 00:15:11.600
Like being totally like just obsessed with that.
00:15:11.679 --> 00:15:12.639
And it's addictive.
00:15:12.879 --> 00:15:15.919
I was like, well, I mean, I'm not gonna be productive any any other way.
00:15:16.159 --> 00:15:17.519
Like I was supposed to do that.
00:15:17.679 --> 00:15:18.080
Yeah.
00:15:18.240 --> 00:15:20.879
But but the other people who join my team, they have that vision.
00:15:21.039 --> 00:15:22.879
I think a lot of them are industry veterans.
00:15:22.960 --> 00:15:25.519
They some of them have been doing software testing.
00:15:25.759 --> 00:15:31.279
My really good friend and CTO, he did software testing work all across the work of standard across Meta.
00:15:31.519 --> 00:15:34.799
And there are other people who were doing chips for for many years.
00:15:34.879 --> 00:15:42.559
And there are people who did chips and then go into formal methods and then going to reinforcement learning and thinking about the combination of formal methods and reinforcement learning, right?
00:15:42.639 --> 00:15:47.120
Like these people, they have these other sort of application areas in mind.
00:15:47.360 --> 00:15:56.399
I was thinking a lot just about math, and I was, I also did, I think, a bit of you know, experience at XTX, which is a really, really great quant trading firm.
00:15:56.480 --> 00:15:57.679
So I did some math there.
00:15:57.840 --> 00:16:00.399
So I thought math has to be making money, it has to be useful.
00:16:00.480 --> 00:16:02.480
Like it just first principle-wise.
00:16:03.360 --> 00:16:06.480
You've brought in these individuals who have a lot of experience.
00:16:07.039 --> 00:16:10.399
How is it for you leading this organization?
00:16:10.559 --> 00:16:13.519
Where do you find that you need to set the tone?
00:16:13.600 --> 00:16:15.840
And where do you take a step back and allow others?
00:16:16.639 --> 00:16:23.200
I think we have very good research taste as an organization as a whole.
00:16:23.679 --> 00:16:29.200
When it comes to research taste, that's where like it's something you just have to insist and you cannot get it wrong.
00:16:29.679 --> 00:16:32.000
And research taste can't be formalized.
00:16:32.320 --> 00:16:33.039
Yeah, it's hard.
00:16:33.200 --> 00:16:33.840
It's an art.
00:16:34.000 --> 00:16:41.039
I think research taste, there are consensus by the community, especially great ideas coming from academia communities.
00:16:41.279 --> 00:16:49.120
And I actually attend, I think, a lot of conferences and discussion seminars, and those are things I love to attend.
00:16:49.279 --> 00:16:51.600
And I think I'll bring that taste back.
00:16:51.679 --> 00:16:53.200
And also we share papers.
00:16:53.279 --> 00:16:55.840
I mean, it's a very, very research organization.
00:16:56.000 --> 00:17:01.840
So that's where I would have a very strong voice if I feel like the research taste is diverging from the ground truth.
00:17:02.080 --> 00:17:05.759
Where I take a step back is, you know, more on the, I think, execution.
00:17:05.920 --> 00:17:10.720
Like we have really great work-class engineers and they they have sensible plans about where to go.
00:17:10.799 --> 00:17:12.960
And I I will just, you know, let them cook.
00:17:13.039 --> 00:17:24.799
And it's also, I think, an interesting mix of you mentioned they're experienced folks, but there are also some other very young, high-agency folks with very good discipline and training in a way.
00:17:24.880 --> 00:17:32.160
Like, as in like they're very good coding machines, and they are very good system designers and they know the info, like in and out.
00:17:32.400 --> 00:17:38.559
Those folks are also very important, I think, as a complementary to the more research scientist vibe of folks.
00:17:38.720 --> 00:17:44.480
And we try to have full-stack folks, but sometimes I feel like that is is also a very important, important piece.
00:17:44.880 --> 00:17:49.279
You touched on taste a lot, and I feel like that's your unique term, probably, that you bring.
00:17:49.680 --> 00:17:54.400
Do you feel like that taste is the unifying substrate between these different types of individuals?
00:17:54.480 --> 00:17:59.119
And maybe it's the first time for those people working with different personas.
00:17:59.599 --> 00:18:01.680
I think there's taste and there's optimistic taste.
00:18:02.000 --> 00:18:03.680
So it's a very, very interesting concept.
00:18:03.839 --> 00:18:09.839
I think as as founder or CEO, considered chief energy excitement officer.
00:18:10.000 --> 00:18:11.759
Um I want that title.
00:18:12.079 --> 00:18:12.319
Yeah.
00:18:12.480 --> 00:18:17.119
You want to bring that taste away from anything that is not constructivism.
00:18:17.440 --> 00:18:22.240
So, you know, there are a lot of nuances, I think, especially by experts, about what is not possible, what is not possible.
00:18:22.319 --> 00:18:25.680
That is hard, that is obscure, that is difficult to define.
00:18:25.920 --> 00:18:27.039
So what is the can do?
00:18:27.119 --> 00:18:29.359
Like what can we do, like incrementally, right?
00:18:29.519 --> 00:18:37.119
Even and where do we know, like critically, you know, where are we aware of the boundaries and within that space, how much we can push?
00:18:37.279 --> 00:18:42.799
And I think that is something I feel like having higher agency folks really help.
00:18:43.279 --> 00:19:06.079
I also think that people who come from, you know, AI versus formal verification, the CS people versus automated theorem proving, the non-AI, AI for math, well, the NI, I guess, like machines for math, like, you know, automated theory improving people, and then the lean people and different, you know, substreams within that, like math slip people, metaprogramming people, they all have their sort of like know-house.
00:19:06.640 --> 00:19:08.880
And sometimes they don't speak on the same language.
00:19:08.960 --> 00:19:16.480
I I don't think it's the taste diverging being the challenge, but more like the terminology and language diverging being a very clear inertier.
00:19:16.640 --> 00:19:25.440
But then once that is solved by good kind of coordination, communication, good culture, once that translation is done, then the taste becomes actually a lot.
00:19:25.680 --> 00:19:35.359
I actually find it beautiful that the underlying taste, as in the part, the optimistic part, are converging.
00:19:35.680 --> 00:19:37.839
This is the nice thing I think about doing research.
00:19:37.920 --> 00:19:39.359
This is a fun part of my life.
00:19:39.519 --> 00:19:44.960
It's that you see the same thing you can borrow from other literatures that can be applied here.
00:19:45.200 --> 00:19:51.359
For a long time, I will go into James Stowe, who is a Stanford professor, pushing AI for science, not necessarily improving.
00:19:51.599 --> 00:19:55.839
I would go to his lab and just look at the papers and take ideas to math.
00:19:56.000 --> 00:19:58.880
I mean, this is like a source, uh a stream of ideas.
00:19:58.960 --> 00:20:08.720
And the other stream of ideas, I look at the hypothesis generation, you know, how people generate hypothesis and then AI experiment and try to have AI submit to ICML and those like machine learning papers.
00:20:08.799 --> 00:20:10.079
I also borrow ideas from there.
00:20:10.160 --> 00:20:17.440
And Schubert was a visionary as well, and so on, and Ken, a lot of our research scientists and folks.
00:20:17.680 --> 00:20:22.960
I think that we are very clear on AI format is just a starting point.
00:20:23.440 --> 00:20:42.880
And I say that just in the verification being the best first market sense, but in the there's so much to be done on self-improving reasoner, where we need conjecturing to fit with proving, where we need auto-formalization as a way of hierarchical planning, honestly.
00:20:43.039 --> 00:20:50.720
Autoformalization refers to converting math in natural language, like English, German, French, Chinese, to lean.
00:20:51.119 --> 00:20:54.799
And auto-informalization refers to converting it back.
00:20:54.960 --> 00:21:00.640
And we believe auto-formalization, auto-informalization forms a bi-directional circle between different abstractions.
00:21:00.720 --> 00:21:02.720
And it is a way of hierarchical planning.
00:21:02.799 --> 00:21:13.680
It's a fundamental technology, not just to generate data for proving model, but but uh as crucial, important, and maybe even at times harder than proving, because when you're auto-formalizing the statement, you don't have the reward.
00:21:13.920 --> 00:21:24.160
So these are, I think, things that draw also talents to axiom because they they see these sort of hot takes and they're like, well, that is actually a grounded way to think about things.
00:21:27.519 --> 00:21:27.920
Hot takes.
00:21:28.240 --> 00:21:28.720
Yeah, yeah.
00:21:29.039 --> 00:21:34.400
How do you think as you're recruiting and building a research-oriented company, how do you balance product and research?
00:21:34.799 --> 00:21:40.480
See, you have to make money at some point, but the research is the core of what makes the product incredible.
00:21:40.880 --> 00:21:49.039
We currently have our core research engineering folks working on some very important future customers' cases, like just directly.
00:21:49.119 --> 00:21:51.599
So there's no sort of separation so far.
00:21:51.759 --> 00:21:59.680
At one point, we might need to consider, you know, having product engineers and it will be kind of a little bit different, sort of composition of the team.
00:22:00.160 --> 00:22:07.839
But currently what we are really seeing is there's a lot of capability overhang of our Action Improver technology.
00:22:08.000 --> 00:22:17.920
And we try to go into these domains where we want to pursue and learn from the experts there and try to see how we collaborate.
00:22:18.079 --> 00:22:30.240
And it's actually surprisingly working well when you have people who are nerds to try to sell to nerds, and we find that to be working quite well for our early stage.
00:22:30.640 --> 00:22:32.000
It's not B2B, it's end to end.
00:22:32.240 --> 00:22:33.039
Yeah, it's funny.
00:22:33.440 --> 00:22:35.359
We ran into math a lot.
00:22:35.599 --> 00:22:36.640
You talk a lot about taste.
00:22:36.880 --> 00:22:41.519
Something I've always been explained to me is beautiful in math is how can you prove something with a minimal set?
00:22:41.839 --> 00:22:42.160
That's right.
00:22:42.240 --> 00:22:42.559
That's right.
00:22:42.720 --> 00:22:43.839
That is what you're describing right now.
00:22:44.400 --> 00:22:48.640
You can prove from the fundamental axiom that zero multiplied by any number is zero.
00:22:48.799 --> 00:22:52.160
I was asked to prove that when I was like, I think 15.
00:22:52.400 --> 00:22:53.759
Like I was like, well, of course.
00:22:53.839 --> 00:22:55.839
Like zero multiplied by everything is zero.
00:22:56.160 --> 00:22:59.359
And like to prove that from those axioms was was fascinating.
00:22:59.440 --> 00:23:04.960
Like, and it really re-revires my brain, I think, back then about how I think about something.
00:23:05.039 --> 00:23:08.799
And then the same way, I go all the way to quadratic reciprocity.
00:23:09.039 --> 00:23:14.240
So I go from there all the way to quadratic reciprocity where I don't take anything for granted.
00:23:14.400 --> 00:23:15.599
Just one by another.
00:23:15.680 --> 00:23:17.119
And it's it's fascinating.
00:23:17.359 --> 00:23:20.720
That was fascinating to me, the deductive nature of things, deductive logic.
00:23:20.799 --> 00:23:23.359
And I think there's this other thing that was fascinating to me.
00:23:23.440 --> 00:23:25.599
It's how things magically connect.
00:23:26.079 --> 00:23:31.839
When you have mirror symmetry or things that are coming from theoretical physics, you can try to do it in math.
00:23:32.079 --> 00:23:36.079
Even in now, AI for math math is being applied on AI for theoretical physics.
00:23:36.160 --> 00:23:41.920
I was at Stanford for Human AI Institute yesterday on a panel with Professor Kyle Kramer.
00:23:42.160 --> 00:23:52.000
He's from the in the same sort of I know him from France Law and Conservative collaboration on having how a transformer would predict call-out sequences, which is a pure math problem.
00:23:52.160 --> 00:23:54.000
The underlying technique, the same technique.
00:23:54.240 --> 00:24:01.359
It can predict the coefficient of the amplitudes of the young males theory, or something of deep interest to theoretical physicists.
00:24:01.440 --> 00:24:03.839
And I think they're submitting a Genesis project together.
00:24:03.920 --> 00:24:14.720
And those are fascinating to me, how things connect, have both real world implications, but also it's a bi-directional discovery.
00:24:14.960 --> 00:24:17.359
So sometimes physicists bring their thing to you.
00:24:17.599 --> 00:24:21.279
Sometimes you do something and think that might be related to, say, particles.
00:24:21.440 --> 00:24:22.799
It's bi-directional as well.
00:24:22.880 --> 00:24:30.880
And ideas fuse more, I think, now with AI between the different scientific subjects, but it doesn't flow in random direction, right?
00:24:30.960 --> 00:24:35.759
You could argue that you need a theoretical layer for experimental subjects.
00:24:36.000 --> 00:24:41.440
And I did one year, this weird year in neuroscience, where at least half the year was in AI.
00:24:41.599 --> 00:24:43.920
So a very, very brief time there.
00:24:44.160 --> 00:24:50.640
I remember when theory of neuroscience and by various methods, some computational, some just mathematics, honestly.
00:24:50.720 --> 00:24:54.559
It's like yours, you have actual equations and stuff, programming it.
00:24:54.880 --> 00:24:59.839
And how that kind of converges with the experimental side of things.
00:25:00.079 --> 00:25:04.160
And that is, I think, what neuroscientists consider beautiful.
00:25:04.480 --> 00:25:08.559
I I think that is another property why I think math is beautiful.
00:25:08.720 --> 00:25:10.079
Yeah, it solves these problems.
00:25:10.240 --> 00:25:14.079
But the the problem is most mathematicians don't talk to applied scientists enough.
00:25:14.400 --> 00:25:19.599
Yeah, they it's it's easy to get lost in the theory and not understand where it can be applied.
00:25:20.079 --> 00:25:22.079
I mean, for for a long time didn't know the definition.
00:25:22.240 --> 00:25:24.160
I didn't know the definition of DNA and RNA.
00:25:24.240 --> 00:25:29.039
I mean, like, and then to to go deep into a field of neurobiology, right?
00:25:29.119 --> 00:25:32.160
Like it's just it takes a lot of the translation here.
00:25:33.039 --> 00:25:37.119
But if they have an AI mathematician, I think that's gonna be interesting.
00:25:37.440 --> 00:25:51.920
The shocking result I saw of MIT Ela Feet's lab is Chinese remainder theorem, which is a rudimentary elementary number theory result, was used to calculate how neurons space out, like the neural capacity.
00:25:52.240 --> 00:26:05.279
So you just think that if you give all the applied scientists axiom prover, like how many more things can be discovered on the theoretical level, and how many of them will find experimental grounding.
00:26:06.000 --> 00:26:10.480
And that's I think when things get really, really like weird.
00:26:10.960 --> 00:26:14.079
You think we're gonna see the equivalent of vibe coding with a vibe science?
00:26:14.319 --> 00:26:14.640
I don't know.
00:26:14.799 --> 00:26:25.279
The 2005 Nobel Prize finding, grid sales, play sales, they file in like honeycomb, hexagonal, you know, six six polygon patterns.
00:26:25.680 --> 00:26:26.799
You have a math proof.
00:26:27.279 --> 00:26:28.799
People who give a math proof for that.
00:26:29.200 --> 00:26:31.920
It's something about minimizing the energy spent.
00:26:32.079 --> 00:26:33.279
It's fascinating to me.
00:26:33.440 --> 00:26:35.440
In that case, it's a proof, right?
00:26:35.599 --> 00:26:38.559
We we want AI where it can conjecture.
00:26:39.039 --> 00:26:40.880
That's gonna be freaky.
00:26:41.039 --> 00:26:53.440
I mean, it's like the the math intuitions which is captured in both the act of conjecturing and in a lot of the cases the pre-conjecturing, just like you're starting to see patterns and you're not sure how they fit.
00:26:53.599 --> 00:26:56.799
And you see something that don't fit the pattern, but you believe they are not important.
00:26:56.960 --> 00:27:00.880
You have that optimistic taste because you know you're going to have your conjecture.
00:27:00.960 --> 00:27:05.920
You have to you have to ignore some of those weird, you know, initial rebellious patterns.
00:27:06.240 --> 00:27:11.200
And and and that process, if scientists can have that, I think that's interesting.
00:27:11.359 --> 00:27:15.759
That's why we're, you know, focused on not just proving, but also constructions, right?
00:27:15.920 --> 00:27:18.400
AI for finding outlier constructions.
00:27:18.559 --> 00:27:27.599
And this is called discovery, the discovery line of work, which are not related to formal theorem proving, but they they should integrate and they should be one workflow, I think.
00:27:28.160 --> 00:27:29.680
This is all super fascinating.
00:27:29.759 --> 00:27:33.759
And Axiom is clearly on a beautiful path of success right now.
00:27:33.839 --> 00:27:37.119
But I mean, and it will fundamentally change what it means to be a mathematician.
00:27:37.200 --> 00:27:41.039
But when you think about it evolving, what is the potential resistance you foresee happening?
00:27:41.359 --> 00:27:57.680
Resistance happening, I think one we are in a little bit of a crazy market right now, in terms of both, I think, the labor market and if you have a lot of heat in the market, then fragmentation of a field is a very natural probable consequence.
00:27:57.920 --> 00:27:59.920
So why would someone join a company?
00:28:00.079 --> 00:28:04.240
We'll all start companies, and then this field will have 1,000 companies and nothing gets done.
00:28:04.480 --> 00:28:09.920
All the time is sort of spent on the the initial co-start activation energy.
00:28:10.160 --> 00:28:13.759
That's one thing I worry about, I think, in terms of resistance.
00:28:13.839 --> 00:28:15.680
And it also I guess comes to that, right?
00:28:15.759 --> 00:28:19.839
Which is that you need to execute very, very fast and and really compound from there.
00:28:20.079 --> 00:28:24.720
Otherwise, I think that's an important point, and I think it might be relevant in other parts.
00:28:24.799 --> 00:28:28.160
I just want to make sure we understood you.
00:28:28.319 --> 00:28:30.880
So your point is capital is so available.
00:28:31.039 --> 00:28:33.200
Yeah, anyone can start a company.
00:28:33.440 --> 00:28:42.240
But in order to get to solve a really hard problem in the world, you actually need people to commit to a cause, not fragment it into thousands of causes.
00:28:42.640 --> 00:28:42.880
That's right.
00:28:43.039 --> 00:28:45.680
And I think that's how you tell categorical.
00:28:46.319 --> 00:28:55.839
So if you see a field with way too many attempts, and and like if you look at the probability of like what is the expected number of good researchers in the field.
00:28:56.480 --> 00:28:58.319
Maybe we'll formalize here a property of above.
00:28:58.640 --> 00:29:00.640
Yeah, I would I would have a theory on that, right?
00:29:00.799 --> 00:29:04.960
And and that is, you know, potentially helpful to Francis VCs, right?
00:29:05.039 --> 00:29:10.480
Like you you look, you can you can call that fragmentation index of a category, right?
00:29:10.559 --> 00:29:16.640
You you you you you estimate the number of good researchers from there based on a couple of factors, like how how new the field is it?
00:29:16.720 --> 00:29:18.000
Does it have an academia history?
00:29:18.480 --> 00:29:21.279
Strong assumptions on the scarcity of the category, yes.
00:29:21.519 --> 00:29:21.599
Yes.
00:29:21.759 --> 00:29:31.039
And it's not just, I think, starting companies, when I mean capital is available, we also see sort of the pay packages of big tech being quite interesting.
00:29:31.200 --> 00:29:40.160
And sometimes you buy projects to kill it, which I find that to be not in favor of constructivism, right?
00:29:40.319 --> 00:29:49.680
And and this the whole concept of constructivism, as Richard Posner's famous judge who loves math, put it, there are certain economic activities like locks.
00:29:49.920 --> 00:29:53.039
If you build a lock, you're preventing your neighbor from stealing from you.
00:29:53.119 --> 00:29:54.720
It doesn't create economic value.
00:29:54.799 --> 00:29:55.839
It doesn't grow the pie.
00:29:56.000 --> 00:30:03.359
And there are certain other economic activities that are constructive, which means growing the pie and you create value and then you capture value.
00:30:03.599 --> 00:30:11.920
So I think there are certain sort of um phenomena we see in the market dynamics that are not necessarily for constructivism.
00:30:12.400 --> 00:30:15.440
Do you think there's limits to what can be formalized?
00:30:15.759 --> 00:30:17.039
I think so, yeah, of course.
00:30:17.200 --> 00:30:20.400
I think that a lot of the world is still fuzzy.
00:30:20.880 --> 00:30:26.240
And I would note that a lot of humans have the best intuition about data.
00:30:27.039 --> 00:30:32.000
This is really, really freaky, how good we are to see outliers in data.
00:30:32.079 --> 00:30:37.359
And we're trying to solve a little bit of that by pushing on the discovery work stream.
00:30:37.519 --> 00:30:41.279
But that is something where I think human intuition like really trumps right now.
00:30:41.519 --> 00:30:45.119
There are cases where there are obviously a lot of formats of data, right?
00:30:45.200 --> 00:30:52.319
I think like, you know, because time series and its implication in trading and the whole art trading literature that's understood a lot more by AI.
00:30:52.400 --> 00:30:56.400
But um the other, I think, fuzzy things just like there are large part of, you know, language.
00:30:56.480 --> 00:31:05.440
I think of language as a manifold, it's a high-dimensional manifold, and it has like distributions around words and math or or code, actually, each programming language.
00:31:05.519 --> 00:31:15.359
It's weird that I call math a programming language, and there's this sort of very provocative line of math is one of the last programming languages we haven't concurred by AI.
00:31:15.599 --> 00:31:20.000
And it's a it's a sub-manifold of that where the distribution changes rapidly.
00:31:20.240 --> 00:31:31.440
And a lot of these submanifolds, say in the family of formal languages, say in the family of broader programming languages, they have similar distributions where you can expect transfers.
00:31:31.599 --> 00:31:39.119
But there are a lot of other parts of this high-dimensional surface that are not related to this family of submanifolds.
00:31:39.279 --> 00:31:40.559
That's how I think about it.
00:31:40.880 --> 00:31:41.519
Interesting.
00:31:41.839 --> 00:31:42.640
Awesome.
00:31:42.880 --> 00:31:44.079
Well, thank you, Karina.
00:31:44.240 --> 00:31:44.480
Thank you.
00:31:44.640 --> 00:31:45.359
This was so much fun.