**Episode title:** TBD | Sze Yuan Cheong & Elijah Ong, Devol Robots
**Guests:** Sze Yuan Cheong, Co-Founder & CEO, and Elijah Ong, Co-Founder & CTO, Devol Robots
**Host:** Bogdan Cristei
*Cleaned transcript of the live conversation. The solo host intro is recorded separately and is not included. Timestamps are carried over from the raw recording.*
---
**Bogdan Cristei (00:00)**
Hey guys, so awesome to have you on. We've been talking about doing one of these for three or four months now. Super excited to see you here.
**Sze Yuan Cheong (00:07)**
Great to see you, Bogdan.
**Bogdan Cristei (00:10)**
Super excited to have you on. So before we start the discussion, let's start with the name.
**Sze Yuan Cheong (00:17)**
It's Devol. The name comes from George Devol. He invented the first industrial robot arm, the Unimate. The idea is that he revolutionized the robotics field by building something programmable. We were thinking, the field hasn't changed in the last 60 years, and we wanted to be the other Devol, the one that changes the status quo.
**Bogdan Cristei (00:50)**
Interesting, very cool. Let's start with a few questions to get going. Sze, let's start with you. When we were chatting, I remember that you'd been running factories for quite some time, maybe even a decade, before you ever touched a robot. So I wanted to start there. Help me see the problem. Walk me through a plant that already owns robots. What works? What is still being done by hand, and why?
**Sze Yuan Cheong (01:18)**
Imagine walking into a factory, any kind of factory, from consumer electronics, where everything is really small, to the big ones like metal fabrication. If you start from the end of the line, the packaging line, the palletizing line, that will most likely be automated if the factory is big. The robot can palletize everything that's already packaged.
But everything in between - plugging in different parts, assembling your phone, handling metal fabrication parts from station to station - all the things that have a ton of variation, humans are doing all of that. Because variation kills automation.
**Bogdan Cristei (02:11)**
I was just on a call with another founder today, and he's been hearing from factories that the number one reason they can't automate is that too many things change all the time. That's the problem with automation. Interesting.
So if you're trying to automate with robotics, but the robot can't feel, the factory tends to compensate with fixtures, jigs, integrator hours, things of that nature. What does that workaround actually cost, and who pays for it?
**Sze Yuan Cheong (02:44)**
The funny thing is, like I said, I started and have been running factories for a decade, and that's basically what we do. The entire field of industrial engineering is about figuring out how to be more efficient and break tasks down into their smallest units. Say I pick something up and put it somewhere. That particular task can be done either by a machine or by a human. The way we think about it is this: if we can have a fixture that gives an absolute position every single time, then we can automate it. And even in that case, you need to tune the machine or the robot every single day to make sure there is no drift. Those are the kinds of things we tend to automate, either with a robot or with a machine.
The human comes in when that variation cannot be fixed, when there are a lot of SKU changes. When you're integrating machines, one of the biggest problems is that those jigs and fixtures require a lot of design work: hardware integration, and ultimately programming and fine-tuning for the hardware to work. That is a huge industry by itself, making sure the factory actually runs and can automate.
I think that is the biggest challenge we have. And what I notice especially is that the younger generation, people like us, tend to go into computer science and now AI. No one really wants to do that work anymore.
**Bogdan Cristei (04:54)**
Interesting. So let's talk about one of your deployments. I know you're working with a customer in optics. If you can talk a little bit about that without giving away too many secrets - I would assume there are thousands of lens variants, there's probably a jig for every single process, and every pocket has a slightly different fit. From some of the videos I've seen you post, I'd imagine that's a nightmare for a conventional robot. I'm curious, what does a failed attempt cost? How do you automate something like that?
**Sze Yuan Cheong (05:27)**
As we usually say, what is automatable is, for now, already automated. That means all the major processes - washing, ultrasonic cleaning, coating, things like that - are already automated. But think about it: there are 20 process steps, and 20 process steps require 20 different types of jigs for every variation of the lens. So how do you transfer the lenses between those jigs? And for certain processes, such as washing, you don't need such a high-precision jig, so you tend to build a manually tuned jig that is basically a loose fit, to try to fit everything in.
Imagine the nightmare. You have a jig that is itself non-deterministic, and you're moving lenses that are really small and easily scratched from that jig into another jig that has only 10 to 20 microns of tolerance. The lens gets stuck, because it's glass, and if it gets stuck and you try to push it in, you will almost definitely scratch the lens, and you get a reject. The nightmare is handling all of this uncertainty throughout the entire process.
The funny thing is, even a human trying to do this needs at least six to nine months of training to be able to do it. It's really hard to do. I couldn't do it, actually. It is a nightmare. That's why our theory is that you train a model and let the robot do it.
**Bogdan Cristei (07:22)**
Okay. So now that we're talking about theory, let's move over to Elijah, because I have a feeling he's got some thoughts on this.
Elijah, I want to paint a picture for you. Imagine a robot arm resting on a table, and then the same arm pressing down with 50 newtons. To a camera, those two are the same picture, right? And that really matters. Could you define stiffness and damping for us? Maybe think about what my arm is doing when I plug in a charger in the dark. Could you talk a little bit about that?
**Elijah Ong (08:02)**
In classical robotics, stiffness is what lets us control and regulate force and position at the same time. Damping is the way you regulate velocity, tracking a target velocity through forces.
If you look at a robot pressing down on a table with 50 newtons, from a pixel perspective you only get spatial understanding. The robot is sort of touching the table, but the image doesn't tell you what it is actually doing. Information like forces, stiffness, and damping becomes really important to tell us, the way a human would feel it, what the actual interaction is. You have to use forces, stiffness, and damping to really understand these interactions, so the robot can learn in a much more efficient way instead of just through pixel space.
**Sze Yuan Cheong (09:18)**
I'll ground this with a quick example. Imagine you come home and you're trying to grab your key from your pocket. You have no vision at all, but you can do it, every time, and really fast. But if you try to imagine the trajectory your hand and your fingers are following inside the pocket, you can't. It is really, really intricate. That entire feeling is based on stiffness and damping. They help us capture that kind of interaction. It's a really good medium to capture it.
**Bogdan Cristei (10:03)**
That makes a lot of sense. Talking about force control - Elijah, you turned down a Stanford PhD to join a three-person startup in Austin building a force-controlled arm from scratch. I'm curious, what did you learn there that you would not have learned in the PhD?
**Elijah Ong (10:20)**
I was really fascinated by teaching robots to grasp things, to manipulate objects. At the time, about ten years ago, robotics researchers were still pondering grasp metrics and kinematic analysis - what kinds of constraints and what kinds of metrics a robot could learn, whether from human heuristics, deep learning, or reinforcement learning. All of these methods were just trying to teach robots to grasp things better. And I kept running into the same situation: a robot holding a cup, with no other information, purely from the joint angles and the pixels - that doesn't tell you the actual interaction.
After I finished grad school, I got an offer from Stanford for a PhD program, and the thesis and research direction was mostly robotic grasping and manipulation, much the same direction. I couldn't really figure out what was going on with it. I felt that direction wasn't leading anywhere at the time, and that I had to get out there and try to find the answer myself.
Then I came across this startup in Texas doing force control, specifically impedance control, and it really fascinated me. I joined them and helped them build an entire robotic arm from scratch. They had this very interesting technology called the series elastic actuator. It isn't common, and they were building their own actuators and sensors.
That's where I found out that impedance - forces, stiffness, damping - is the complementary element, the right element, to serve as a physical representation for the robot to really understand what's going on behind the scenes. Because of that, I called Sze at something like 3 a.m.: hey, I found this, it's so interesting, we can finally solve robotic manipulation. For me, impedance is the most critical piece of the puzzle. It can serve as a physics prior for robot learning.
But there was a fundamental piece missing. When we both started the company, we tried to train the model ourselves on forces, stiffness, damping, the impedance parameters and so on, and the model didn't really learn well. That became another research question I've been thinking about for quite a long time.
Robot data has been flattened all along. One very simple example is poses. When a pose turns from zero degrees to 180 degrees, then suddenly, from 180 degrees, in terms of Euler angles it becomes zero degrees again. That phenomenon in robotics is what we call gimbal lock. So when we treat this data in a flattened way, so we can use it to train a neural net, there's a huge change in the numbers, but in the physical interaction there isn't. There's an information gap there, and what it tells us is that robot data actually lives in curved space. We should really treat robot data as curved. It shouldn't just live in a flat, one-dimensional vector. Every piece of robot data should be projected onto its right curve, its right manifold, so that it can be learned much more efficiently. You can see this reflected in something quite common in VLAs today: a lot of VLAs need tons of pretraining data and fine-tuning data to teach the robot even one simple pick-and-place task. The majority of them treat the data as one-dimensional, whether it's robot state or robot action. They flatten it and then they train on it.
So, with all that said, what we believe is that we need the right physics prior for the robot to understand physical interaction, and underneath that, we need to project every piece of robot data onto the right curve, so the model can learn way more efficiently than any other model in the world, and the robot can really understand the actual interaction behind every manipulation.
**Bogdan Cristei (15:28)**
So let's talk through one particular task, phase by phase. Take an Ethernet plug. There's the alignment, the clip compressing, the snap, the seating. What is the model doing with stiffness at each moment? And what happens to a position-controlled policy at, let's say, the snap?
**Elijah Ong (15:47)**
When you pick up the cable from the table and move it closer to the port, think of that motion as moving in free space. The robot is moving in free space. It's compliant, it's soft. When you get close to the port, that's when you really need to maintain stiffness. That's when you need stiff control guiding the robot toward the port. At the same time, you need to regulate the forces and the damping well, because you don't want it to be too fast, and you want it to be a bit more reactive when it touches the edges of the port.
When it's sliding in, you need to be more precise. You need to maintain your position along that axis, along that trajectory. That's where stiff control becomes really important. By learning through impedance, we have this kind of compliance schedule in our model. It's what we call a visuo-impedance map: at this moment, at this point in time, what stiffness and damping you need to exert. And it's directional - stiffness is definitely directional. In each direction, it's what you need to exert so the motion is guided well and the insertion finishes successfully.
**Bogdan Cristei (17:09)**
When we were chatting earlier, you described a minimum of a thousand hours to get a task to work well. Talk a little bit about how that compares - your approach versus the big VLA approach.
**Elijah Ong (17:24)**
On the thousand hours of pretraining: we want to show readers and the audience that a thousand hours is the minimum we've found that enables a model to really understand physical interaction. It doesn't mean we stop there. We're trying to show that a thousand hours is enough to give the robot real physical understanding, so we can teach it to do physical tasks, and beyond that, contact tasks. In our first paper, IWM 1.0, the benchmark tasks slowly increase in contact complexity. We show that with these thousand hours, our model can learn simple assembly, plug insertion, Ethernet cable insertion - these kinds of contact-rich manipulation tasks.
The typical VLA, by contrast, needs a lot of video data, and even robot data, to work out, based on what the robot sees at the moment, what the best trajectory is and what the best motion is, so that it ends up being successful. But most of the time, a typical VLA is open loop. Based on what the robot sees, it just keeps rolling out actions, and it isn't really physics-grounded. It doesn't really have closed-loop control in that sense.
**Sze Yuan Cheong (18:59)**
That shows up in the benchmark we did. With that thousand hours, we're trying to show that with only a thousand hours of training data, the robot learns the physics right. How do we show that? We compare it to other state-of-the-art models. The funny thing is, they are already in the million-hour range. Then we gave them some post-training data. For ours, we gave 20 trajectories for each task, and for the comparators, we gave them a hundred. The end result: our success rate is over 90 percent, and the two comparators are at 30 and 40.
**Bogdan Cristei (19:52)**
Gotcha. So let's say a manufacturer calls you tomorrow with a new insertion task. What do they buy? What happens on their floor? How long until it's running in production?
**Sze Yuan Cheong (20:06)**
Say they call us with a particular insertion task. First, let's say they have a robot, and the robot is in our ecosystem. They use our model and show the robot what to do - basically, a demonstration from their operators. The model identifies what the objective is and what the object is, and from there, there are two paths.
The first path is that the robot plays with the object, remembers the objective, and learns about the interaction. It learns for about an hour, basically trying different ways to do the task. It fails a bunch of times, and it succeeds. Most importantly, it collects that data for post-training, to come out with a policy that actually works, because we're using the actual interaction data to ground the policy in the physics of that particular task.
The other path is for when they can't do that - say they don't have a robot yet, and they want to train a policy and see it succeed before they deploy. Then we can use other training methods, such as UMI grippers. With about 20 or 30 minutes of collected data, we can post-train the model and come up with a successful policy.
**Bogdan Cristei (21:59)**
Okay, let's talk a little bit about the bitter lesson. Over to Elijah. The bitter lesson says that scale beats hand-built structure every time. The VLA labs have billions of dollars. If they add force data and keep scaling, do you think they arrive at the same place?
**Elijah Ong (22:18)**
I'm not saying scaling is bad. It's that we need to scale in the right paradigm. I think current VLAs and models are scaling in the wrong direction. What I mean is that we shouldn't look at this as just an AI problem. We should look at it as a robotics control problem. What does a robot actually need? What kind of information, what kind of training data, does a robot need to really understand the physical world? That's the belief we've held since we started this company.
A physical prior is the first thing. Without a physical prior, a robot won't be able to really understand physical interaction. Second, we shouldn't treat robot data as just numbers. We should let the data live in curved space, so the robot can understand what is going on inside the data, inside the interaction, and learn it far more efficiently. With that kind of paradigm, the scaling law actually works, and it will be far more efficient. We need to scale more efficiently.
**Sze Yuan Cheong (23:34)**
I'd summarize it as two things we believe to be true. The first is the foundation, the model architecture itself. We don't believe the current paradigm is the ultimate paradigm. We have not yet found a way to embody all the data and train an efficient model. So, the model architecture.
The second is a lesson from classical robotics. In classical robotics, we always try to abstract. We find the correct abstraction, and then we work with the abstraction. That lesson has been forgotten. What we're trying to do is push the boundary on both the model architecture and the abstraction. Stiffness, impedance, and so on are the abstraction framework. Then comes a model architecture that can embody that abstraction framework, to train and then actually execute in the real world.
**Bogdan Cristei (24:40)**
Gotcha. Correct me if I'm wrong, but when I think of the installed base of robot arms, most of them are position controlled, and your approach commands stiffness. I'm wondering how much of the world's existing robots you can run on, and what you lose if an arm can't do impedance control.
**Elijah Ong (25:01)**
Our model can accommodate all kinds of embodiments, regardless of whether it's a force-controlled robot, a position-controlled robot, or even a velocity-controlled robot. In terms of impedance, there are two kinds. The first is explicit impedance, which you can collect directly from a force-controlled robot, or from a position-controlled industrial robot with a force-torque sensor at its end effector. The second is what we call derived impedance.
Imagine a factory. Across multiple trajectories, there's a part of the trajectory where the motion is compliant, and a part where the motion is stiff. Back to the Ethernet insertion example: when you pick up the cable and move it in front of the port, the compliance schedule is soft. When you try to insert it, the compliance schedule is stiff. Our model is able to extract these impedance parameters, and they serve as a supervising signal for the model to really understand the interaction behind the motion.
So even today, if we're using industrial robots or cobots, our model can be deployed, and the robot can use those derived signals to understand what acceleration and deceleration it needs so that it ends up with a successful motion.
**Bogdan Cristei (26:32)**
Gotcha. And where does the model break today? You have deformable objects, long-horizon assembly, cable routing. Where would it break?
**Elijah Ong (26:42)**
A very extreme example would be threading a needle.
**Bogdan Cristei (26:49)**
Wow, a thread through a needle.
**Elijah Ong (26:51)**
Because the thread is very soft, you can barely get a force signal from it. Most of the time you rely on position control and the image itself, so impedance data becomes less critical at that point. But if the sensor is good enough for us to collect the forces that matter for the motion, our model can still do it. Other than that, for most tasks, we're able to learn the interaction.
**Bogdan Cristei (27:38)**
In closing, I want to ask maybe two or three questions. I don't want to call them hot takes, but I'd love to hear your thoughts. Either of you, or both of you, can answer whichever question. The first one: what is the most technically wrong thing robot learning teams are still doing in 2026?
**Elijah Ong (28:08)**
We really shouldn't scale from pixels anymore. We shouldn't just scale from videos anymore. We should look underneath those videos, that egocentric data. How do we derive the right physical interaction? How do we derive the correct physical representation for the robot to really understand? It shouldn't just be video in, video out, and then predict the action. I truly think that's not the right way to interpret things.
**Sze Yuan Cheong (28:41)**
He covered one part. I'd add the other part, which I touched on just now: the model architecture. We're actually releasing a paper at the end of this month. It will be submitted to ICLR, and it's about our architecture, called Devol One. It's a mixture-of-transformers that blends in basically everything. Currently there is a divide between VLAs and world action models, and then there is the JEPA route. What we're doing in this paper is blending everything together.
Imagine a world action model: two distinct routes within the architecture. One part figures out the latent space, and the other part is the action decoder. The fundamental question we asked ourselves is, why couldn't this be sequential? Why couldn't you combine both? And why couldn't you use JEPA? What we did is take the output of the latent and feed it into the action decoder. We used that to create a latent-conditioned action model, and that action model is, in a way, a world model itself.
With this architecture, which basically blends VLA and JEPA - and without even putting in our framework of geometry and impedance - we tested it on various benchmarks, and it ranked number one on every single one. That gave us a lot to think about, because what everyone assumes to be right at the moment might be wrong. We might be spending the money on data and scaling on the wrong side. There may be a lot of fundamental breakthroughs that can still happen, and if we find them, we might find a significantly more efficient way to scale.
**Bogdan Cristei (31:20)**
I see. That leads me to my next question, and I have a feeling I know what you're going to say. Five years from now, what will be obvious that sounds non-obvious today?
**Sze Yuan Cheong (31:31)**
Wow, your questions. Okay, I'll answer first. I think there's a huge divide between classical robotics, with what's been learned from decades of development there, and where modern AI is right now. One of the very first things - and it's slowly happening now - is that people from the AI world are rediscovering knowledge that was already gained in classical robotics. It's a bit weird if you think about it. Do you want to continue?
**Elijah Ong (32:11)**
I've noticed that researchers and physical AI people have started to realize that low-level control matters quite a lot. There's research on a kind of slow brain and fast brain. The VLM serves as the slow brain, reasoning about what is going on in the general situation, and the fast brain runs closed-loop control based on the robot's feedback and the image, deciding what action is needed. That kind of research is going on. People are starting to appreciate that we need to think about how to make robots precise again, instead of just letting the actions roll out and leaving it at that.
Those issues are why VLAs and action models still can't do precise tasks, force-relevant tasks: they ignore this low-level, fundamental control in robotics.
**Sze Yuan Cheong (33:18)**
What the entire field is currently trying to solve is merely the planning end. The execution end, on the robot side, is not solved. And dare I say, no one is actually trying to solve it, or seeing it as an actual problem. Which it is.
**Elijah Ong (33:41)**
A very typical example is the action head itself. Everyone is using diffusion, flow matching, even autoregressive heads. It's purely open loop. It isn't closed loop. When it comes to robot control, they just say, okay, I ask the robot to go there, and it just goes.
**Bogdan Cristei (34:05)**
Very good. Maybe the last question. Do you have one piece of advice for robotics founders starting today?
**Sze Yuan Cheong (34:11)**
It comes from our experience - our own bitter lesson. When we started, we had a very ignorant idea of building everything in house. We had this idea of the representation we wanted to build. We work in the AI space, but as a classical roboticist, you tend to romanticize things. You think: if I'm going to build the best model, I need to build the hardware, the entire hardware stack - our own actuators, maybe our own motors, definitely our own drivers and our own custom firmware - so we know how to control everything and optimize everything.
As a startup, you can't really do that. We learned that the hard way. It's good that we got a lot of lessons from it, but we almost died.
So don't try to do everything yourself. Get into the ecosystem, work with your peers, work with other companies, because robotics is such a huge problem to solve. Find the thing you are good at and solve that. Don't try to solve everything.
**Bogdan Cristei (35:40)**
Wisdom from pain and suffering. Very good. So with that, where can people learn more about Devol, and who do you want to hear from? Who do you want to reach out to you?
**Sze Yuan Cheong (35:54)**
Our website is devolrobots.ai - that's D-E-V-O-L, robots with an s, dot ai. We'd like to work with researchers in this area, roboticists who are interested in control, and obviously customers who have problems they want solved.
And the last thing: anyone who's interested in robotics, we'd like to hear from you, even if you have no experience in robotics at all. We've always believed in building a team with variety, bringing in people with very different backgrounds and experiences, and mainly trying to find outliers, so that as a team we can think outside the box and build a solution that is fundamentally impactful in this industry.
**Bogdan Cristei (37:00)**
Well said. Diversity. Awesome. Really good to have you both on. Thank you so much for your time, and excited to keep the conversation going.
**Sze Yuan Cheong and Elijah Ong (37:10)**
Thank you, Bogdan. Thank you so much. Always good to talk to you.