Host: Bogdan Cristei Guest: Aurora Feng, Founder and CEO, Neural Motion Runtime: 21:16 Video: https://youtu.be/fncAqaeVA2w Status: Cleaned transcript. Disfluencies removed, ASR term errors corrected, speaker labels added. Timestamps preserved from the Riverside export.
[00:00:00] BOGDAN: This is the OPTIM Update. I'm Bogdan Cristei, and today's guest is Aurora Feng, founder and CEO of Neural Motion. Neural Motion is building a generative world model for robot manipulation. It takes a task recorded on one robot and produces that same task on a robot it has never seen, video and action together, with no retraining and nothing collected on the new body. Aurora founded Saturday Robotics, the largest robotics and world model research forum in Silicon Valley, took it from zero to 2,600 researchers in three months, and recruited most of her team out of it. Before that she was founding head of North America at LimX Dynamics, and invested at Pear VC and ZhenFund. We get into why every robot foundation model lab is capped by the data it physically collected, and why co-training on every robot dataset in the world doesn't remove that cap. Without further ado, here's my conversation with Aurora.
[00:00:57] BOGDAN: Aurora, welcome in. How are you? How are you feeling today?
[00:01:00] AURORA: Good. How are you?
[00:01:01] BOGDAN: I'm very good. I'm excited to talk to you about robot data, world models, everything in between. Are you ready to start?
[00:01:09] AURORA: Thanks for inviting me.
[00:01:10] BOGDAN: Awesome. So before we get into what you're building, I'd love to hear the backstory, because the team you put together is impressive. It's not the team a first-time founder usually gets. I see a bunch of top-program PhDs leaving what they were doing to join a small company. I'm curious how you pulled that off.
[00:01:30] AURORA: Whenever I talk about my company, I always talk about my community first. So I'm going to start with my community, Saturday Robotics. I built the community before I built the team.
I became interested in robotics when I was a senior in college back in 2024, and post-grad I spent a year at LimX Dynamics, a Chinese robotics startup that's about to go public this year. I was their head of North America, so we jumped between the China world and the US world in terms of robotics, world models, and everything in between. I got to know a lot of researchers, and a lot of the people working on robotics are Chinese people in the Bay Area and Chinese people in China.
[00:02:16] AURORA: There are geographical and cultural gaps among those researchers, and there's also a lack of community that brings the researchers together. If you look at events in SF, we have so many networking events, founder-VC events, happy hours. But people are always pitching their companies. People aren't talking about what's happening and what strikes them intellectually. I want to be intellectually stimulated, and I'm sure a lot of people do too.
That's why I founded Saturday Robotics, so people can have a platform to talk about research, to have a VC-free, startup-free environment to debate the problems that matter. Surprisingly, a lot of the people who came to the Saturday Robotics researcher group are VCs and founders themselves. But on top of the researchers, everyone else was also talking about research. Even founders weren't pitching their startups, and VCs weren't necessarily investing or looking for investment. They're just here to learn. I think that's awesome.
[00:03:14] AURORA: So as I was talking to people training world models, policies, collecting data, I kept hearing different versions of the same frustration: is data the bottleneck of physical AI? If data collection is the only way to solve the data bottleneck, then the debate I was hearing was, is teleoperation the best data collection method? The problem with teleop is that it scales slowly and it's costly. So people said, maybe we should move to egocentric data collection. Everyone said, let's scale ego data, that's the future.
But if we take a step back, is data collection the only way to scale? Probably not. You're still relying on human labor, and no matter how cheap it is, human labor always has limitations. You realistically cannot have the entire world collect the entirety of the data in the world.
[00:04:12] AURORA: So I took a step back. I thought, what if we find all the data in the market that's already been collected and have robots learn from each other, learn from the robots that have already collected the task, so you don't have to replicate the same task in different robot environments over and over again.
Then I came back to my reading club with that thought in mind. I thought, this is going to be great. This is going to be the most groundbreaking thing in robotics. If we can solve that, we pretty much solve the entire physical AI data bottleneck. But I need a world-class team to help me build it, because I'm not a researcher myself.
When I came back to my community, I had this good friend I'd met through the reading club. He was one of the first guest lecturers there, and he was talking about his work. I thought, wait, I've read your paper before.
[00:05:11] AURORA: When I became interested in cross-embodiment, I was already going through thousands of papers myself. So I said, your paper describes exactly what my company is trying to solve. But my company was basically myself. I needed someone to build it with me.
When I talked to him about this idea, he got excited, because he said, I wrote this paper, I've been working on it since two years ago, but people underestimate how important it is. People think it's a data editing problem. And data editing is such tiring work that nobody else wants to do it. Everyone in academia is trying to build the greatest policy, and no one's focused on the dirty work of dealing with data. He said, it's my first time having somebody tell me they read my paper because they're interested in the problem. So we clicked, and that's how I was able to form the rest of my team, because we're all interested in the same problems.
[00:06:05] BOGDAN: Beautiful. Now let's get into the technical talk a bit. First of all, I'd love if you could define some of the terms you're going to be using. Maybe start with embodiment.
[00:06:18] AURORA: That just means the body of the robot, like an arm, a gripper. Anything you see physically doing a task. That's an embodiment.
[00:06:29] BOGDAN: And on the model side, you say that before a model can generalize in the physical world, it has to know what changes when the robot body changes and what stays invariant.
[00:06:41] AURORA: What changes is the embodiment. Say I've collected data on a Franka arm in the past, but now I want data collected on a UR arm. The embodiment changes from Franka to UR, but what the arm has been doing stays the same. Say the Franka arm is picking up a pencil from a table. When transferred, the new arm is still doing the same thing. It's still doing the same task. So the task stays invariant, and the embodiment is what changes.
The reason people can't simply learn it is because there are embodiment gaps. There's an observation gap, because robots look different, and there are also action space gaps.
[00:07:24] AURORA: For all the reasons above, you simply cannot replicate the same experience of one task from one robot to another robot.
[00:07:32] BOGDAN: Got you. So the field's answer to cross-embodiment has been co-training. Think Open X-Embodiment, Octo, OpenVLA, Pi-Zero. One model trained on every robot dataset you can possibly find. Why isn't that enough?
[00:07:46] AURORA: We're talking about two different approaches to the problem. One is from the data side, the other is from the model side. The field's common approach always comes from the model. With co-training, you throw in a bunch of random datasets, all from different embodiments, and you don't know what's going to happen to them, because you let the model figure it out. And boom, magic. It somehow works. The model learns everything.
But the issue is you never know: if you do cross-embodiment before pre-training, can the model perform better? That's a black box, because we don't know.
[00:08:22] AURORA: And because the foundation model teams never had the time and energy to figure it out themselves, that becomes a question left for us. If we were to figure out the data recipe before we put all those datasets into the pre-training pile, what's going to happen? Our hypothesis is that it's going to be much better. That's why we're focused on solving the hardest pre-training data problem.
[00:08:43] BOGDAN: So where's the ceiling for the labs? What does a policy do when the task sits outside what its body collected?
[00:08:51] AURORA: Most models only learn within the distribution of the task. Say they've collected millions of hours of data on five arms, and each arm has 200 tasks in total. They're only going to learn those 200 tasks, because that's what the model has been fed. But if you have a 201st task, and that task is something the model hasn't seen before, it doesn't know how to react to it.
That's the issue, because realistically you can never collect the entirety of tasks in the world. There are always going to be out-of-distribution tasks. So the best thing the model can do is try to generalize within the task distribution. Maybe 200 tasks is already good enough. It can do most things, but it can never be the best, because it's still bottlenecked by...
[00:09:46] AURORA: ...the amount of data and the amount of tasks you're able to collect. That's why we think true in-context learning, allowing robots to learn out-of-distribution tasks and expanding the distribution of tasks during pre-training, is so important.
[00:10:03] BOGDAN: At least one large lab told you they can't reuse demonstration data across its own hardware generations. How common is that?
[00:10:16] AURORA: That's pretty common with a lot of the labs we've spoken to, because they always change their embodiments. I asked them, what do you do with the data collected on old embodiments? They said, we try to reuse some if we can, but most of it we can't reuse because the embodiment gap is too big. We'll just dump the data or leave it there. We don't use it.
I think that's sad, because it takes time to collect such high quality teleoperation data. That's why I said, instead of collecting new data on every single new embodiment, why don't we make use of the old embodiments' data that's already been collected and the tasks that have already been learned?
[00:10:58] BOGDAN: Some folks have told me that cross-embodiment is a kinematics problem. You have the poses, the joint positions, you retarget, and you're done. In your view, why is that wrong, or at least not enough?
[00:11:15] AURORA: It's definitely not enough, because kinematics solves correspondence in the configuration space, but we need correspondence in both observation space and action space. That goes back to what I said earlier about why you can't simply let one embodiment learn from the task on another embodiment. It's because of both the observation gap and the action space gap.
If all I cared about was moving an end effector from embodiment A to embodiment B, then sure, classical retargeting could be useful. But that's not the goal we want. I don't care about just moving an end effector from this embodiment to that embodiment. I want the data to contain tasks from different embodiments, so it can feed a better policy. That's why we need more than retargeting, and we need a learning-based approach...
[00:12:08] AURORA: ...to generate data. That's the world action model we're building.
[00:12:12] BOGDAN: And does it work? Arm to humanoid, single arm to bimanual. Where does it stop?
[00:12:20] AURORA: We mostly do single-arm transfer, because we're focused on cross-embodiment learning for manipulation. We're not talking about transferring humanoids to a quadruped, or two legs to four legs. That would be a stretch. I also don't think locomotion is as complicated as manipulation. It's much simpler.
So we're not too focused on humanoids right now, but the core we're solving is manipulation. If you want a humanoid to do a task on top of all the lower body motion it can already do, we can apply our approach to the manipulation part of the humanoid.
[00:12:55] BOGDAN: And how is what you do different from real-to-sim-to-real?
[00:13:01] AURORA: That's a great question, and I get that a lot. The way we do it is we don't oversimplify anything. We try to make the data remain the same, except that the embodiment is different. Say we have teleoperation data as our source data. When transferred to the target, the target output is still teleoperation data. It's just no longer collected on the source embodiment. It's as if it had been collected on the target embodiment.
Whereas for real-to-sim-to-real...
[00:13:36] AURORA: ...to jump to sim, you have to oversimplify a lot of things, because you're creating rules in simulation for the simplified simulation environment to understand how complicated things are in the real environment. By the time the sim environment is created, there are already a lot of losses and biases introduced.
And when you go from sim back to real, closing the real-to-sim-to-real loop, that automatically introduces new biases, because it has to make that simplified environment complicated again to adapt to the real environment. By the time we close the whole loop, often the data itself is already too different from the original data, and you can't even use it. That's why the real-to-sim gap is a forever challenge, and nobody knows how to solve it.
[00:14:25] BOGDAN: Got it. So when will I be able to walk into a room and watch a Franka arm do a task that was only ever recorded on another body?
[00:14:36] AURORA: We're launching Q1 next year, so after our technical report launch we'll have a setup of robots and people can play around with that. I'll be excited for you and for all our listeners to come play with the robot and watch magic happen.
[00:14:57] BOGDAN: Amazing. Sign me up. So you're connected to folks in China, and you spent time there meeting dozens of teams training their own models. What did you learn, and where do you think they are relative to the US labs?
[00:15:15] AURORA: What's interesting is the way things got started. Most Chinese companies started as hardware companies, and it's the opposite for the US, because most people here are building brains first. It's almost impossible to build hardware in the US, and people nowadays are still debating whether manufacturing in America even exists.
The reason Chinese companies are so focused on building brains right now is that pretty much all the hardware side is solved. But you can't do anything with just the hardware. So they said, what if we add some LLMs into it? What if we add some world models into it? They're playing around with what type of model would make sense for the hardware they built.
[00:15:57] AURORA: Versus in the US, the people who started with models are now looking for the hardware deployment that can contain the model. From an ideal perspective, hardware and models should go hand in hand. It's hard to customize hardware to a model, and there's so much more complicated work you have to do from the beginning. So it's interesting to see how things converge from both sides, but from exactly opposite directions.
[00:16:24] BOGDAN: Fascinating. Maybe a couple of questions, more rapid fire style, but feel free to go long on these as well. My only ask is to not overthink it. So the first one: if conversion works, what happens to the teleop data vendors? Are you going to be their best customer or their worst problem?
[00:16:48] AURORA: Neither. We're going to be their best friend. First of all, we are not a vendor, and we will never be a vendor. We don't ever collect data, so we still need teleoperation data, and the teleoperation companies and the data they've collected are almost our fuel for all the model training. We definitely need them.
On the other hand, I'm also not selling to any of the teleoperation companies. I'm helping them turn their datasets, whether off-the-shelf already collected datasets or future exclusive datasets that can only be sold once, to now be augmented and transitioned into much bigger batches of data to be sold. That's exciting for them and for us.
[00:17:35] BOGDAN: This is one I like to ask: what will look obvious in five years that sounds non-obvious today?
[00:17:42] AURORA: In five years, I think we'll pretty much be done with scaling data. Hopefully, if Neural Motion works out, none of the companies in the market training models should ever worry about data. The only thing left to worry about is compute, because that's something that has to physically exist, and we can't magically generate compute.
The bottleneck of training models today is both data and compute. Hopefully in five years we remove the data bottleneck, and the only real bottleneck that exists is compute.
[00:18:14] BOGDAN: People in general say ego data is cheaper than teleop. I wonder why you don't remove teleop altogether, so there are no cross-embodiment issues at all.
[00:18:26] AURORA: A lot of people have asked me that. What if we purely use human data to train a foundation model? Human data gives you enormous task semantic coverage, because the human hand is pretty much able to do anything, and it doesn't require slow teleoperation of the robot.
But the issue is that it leaves the robot embodiment undetermined. You never know what the teleoperation embodiment should look like, and therefore it doesn't have a good benchmark to compare the human data to. So our hypothesis is that once the model has explicitly learned the mapping between the task and the robot embodiment...
[00:19:08] AURORA: ...then human experience becomes much more efficiently usable as just another source embodiment. But if we don't have any robot embodiment as the reference, then even if you scale millions and millions of hours of human data, the marginal benefit gets lower and lower, because of that same issue: you don't have a good reference of the teleop robot.
[00:19:33] BOGDAN: Maybe one more. You have PhD students who joined a startup, leaving their academic roots. Do you have a piece of advice for researchers weighing that jump? Should I leave academia and go work at a company?
[00:19:51] AURORA: From my perspective, the greatest innovation always happens in industry, not in academia. And within industry, the greatest innovation happens in startups rather than big tech. So I would 100% encourage PhD students who want a more innovative environment to achieve whatever they want, whether it's academic goals or commercial goals, to try startups and see what you can do within a short period of time with a small group of people, and how you can move fast and iterate fast. At the same time, be well budgeted for time and compute consuming research.
[00:20:33] BOGDAN: Before we wrap, is there anything you need?
[00:20:37] AURORA: We're always looking for top talent to join us. If you're interested in in-context learning, cross-embodiment, or pre-training robot foundation models, definitely reach out to us. We'll also be preparing our technical report and our big launch for next year. Stay tuned for that.
[00:20:55] BOGDAN: And with that, Aurora, where can people learn more about Neural Motion?
[00:20:59] AURORA: You can go to our website, our LinkedIn page, our X page. Everything is secretive right now, but you'll figure it out once we launch. It'll be everywhere.
[00:21:09] BOGDAN: Really awesome to have you on. Thank you so much for your time today. It was great talking to you.
[00:21:14] AURORA: Thank you. Thanks for inviting me.
[00:21:16] BOGDAN: Of course.
End of transcript.