1
00:00:00,000 -->
00:00:10,640
Hello there, everyone! Welcome to episode number 676 of this here electronic engineering podcast
2
00:00:10,640 -->
00:00:18,879
called Amelia's Weekly Fish Fry, brought to you by eejournal.com and written, produced,
3
00:00:18,879 -->
00:00:26,160
and hosted by yours truly, Amelia Dalton. Folks, I have been excited to share this interview for
4
00:00:26,160 -->
00:00:34,720
some time. My guest is an absolute rock star in the world of electronic engineering. Sandra Rivera
5
00:00:34,720 -->
00:00:42,639
joins me this week. Sandra and I chat about what drew her to Visora, an emerging AI accelerator
6
00:00:42,639 -->
00:00:50,160
startup from France that is taking the world of AI inference by storm, the importance of inference
7
00:00:50,160 -->
00:00:57,520
for AI proliferation, and what sets Visora apart from the rest of the pack. Oh yeah, and a little
8
00:00:57,520 -->
00:01:05,199
bit about llamas, too. So, without further ado, please welcome Sandra to Fish Fry. Hi, Sandra,
9
00:01:05,199 -->
00:01:10,800
thank you so much for joining me. Thank you for having me on, Amelia. Absolutely. Okay, so let's
10
00:01:10,800 -->
00:01:18,559
talk about Visora. So, what attracted you to Visora, an emerging AI accelerator startup from France
11
00:01:18,559 -->
00:01:24,959
that wants to take on the world of AI inference? Well, I think you just stated it best, an emerging
12
00:01:24,959 -->
00:01:30,559
AI inference chip company based in France. Like, those things don't typically go together,
13
00:01:30,559 -->
00:01:38,239
and I have to say that when I first was introduced to the company and the CEO, it was really more
14
00:01:38,239 -->
00:01:43,519
out of curiosity that I wanted to meet them because I built these chips in my previous lives
15
00:01:43,519 -->
00:01:47,519
working for a very large semiconductor manufacturing company, and I had never
16
00:01:47,519 -->
00:01:54,959
heard of Visora. So, when I met them and met the CEO, I was just so fascinated by the fact that
17
00:01:54,959 -->
00:02:00,160
here's a company that is on leading-edge process technology, leading-edge packaging capability,
18
00:02:00,160 -->
00:02:07,599
building what is arguably a chip for the most exciting growth part of AI going forward, which
19
00:02:07,599 -->
00:02:13,279
is not so much training because that's relegated to a few that can afford training those frontier
20
00:02:13,360 -->
00:02:20,479
models, but more the deployment in the inference space. And I was, again, curious, and then when I
21
00:02:20,479 -->
00:02:26,800
learned that they were indeed on leading-edge technology with such a small, nimble team,
22
00:02:26,800 -->
00:02:32,240
and that that team had been together for many years building and delivering to the market
23
00:02:32,240 -->
00:02:36,559
successful chips, I felt it was highly differentiated because many of the startups,
24
00:02:36,559 -->
00:02:41,279
Amelia, as you know, in this space, they're very bright, very capable architects and engineers,
25
00:02:41,279 -->
00:02:44,800
but they don't have a history of working together and delivering products to the market. And here's
26
00:02:44,800 -->
00:02:50,320
a company that had something like 14 successful chips delivered in the past and now working on
27
00:02:50,320 -->
00:02:57,039
the most complex one that they'd ever done around AI inference. So, Sandra, why has inference become
28
00:02:57,039 -->
00:03:05,199
so important to AI proliferation, and what makes Visora's architecture especially valuable for
29
00:03:05,279 -->
00:03:13,360
inference in particular? Great question. So, the reason inference is the highest growth part of
30
00:03:13,360 -->
00:03:20,080
that, call it AI continuum, is because it is actually what happens when you go and deploy
31
00:03:20,720 -->
00:03:27,279
the AI models, whether that's in a data center, whether that's in an enterprise infrastructure
32
00:03:27,279 -->
00:03:33,440
environment, or whether that's all the way out to edge devices and physical AI, robots,
33
00:03:33,440 -->
00:03:40,960
autonomous vehicles, drones, those types of technologies and devices. And this is where the
34
00:03:41,839 -->
00:03:48,960
implementations and deployments vary broadly in terms of the actual use cases, but where things
35
00:03:48,960 -->
00:03:56,559
like cost per token, power constraints, area constraints, weight, physical weight and area
36
00:03:56,559 -->
00:04:04,000
constraints are really driving a lot of the obstacles that need to be overcome for these
37
00:04:04,000 -->
00:04:10,559
types of applications to be deployed at scale. And so, because it is so wide open in terms of
38
00:04:10,559 -->
00:04:16,480
the number of applications, the use cases, the physical environment, those constraints, the type
39
00:04:16,480 -->
00:04:23,440
of almost unlimited compute power and energy and money that goes into training frontier
40
00:04:23,440 -->
00:04:28,320
foundational models is not applicable to inference when you're talking about much,
41
00:04:28,320 -->
00:04:32,959
much smaller devices or much more power and cost constraints. And more frankly,
42
00:04:33,600 -->
00:04:40,239
latency and determinism really matter. So, that latency, that time to that first token in terms
43
00:04:40,239 -->
00:04:47,519
of responsiveness of the platform, and also in terms of just the deterministic outcome, meaning
44
00:04:47,519 -->
00:04:54,239
that if you're doing some type of call it robotic surgery, like you really need for it to be the
45
00:04:54,239 -->
00:04:58,559
same every time in terms of the responsiveness when you're actually working with the different
46
00:04:58,559 -->
00:05:04,880
tools. So, this is what makes it so different from training and why we believe, and it's not just us,
47
00:05:04,880 -->
00:05:09,839
one of the things that we've seen is that there's actually a lot of startups and a lot of competitors
48
00:05:09,839 -->
00:05:15,440
in this AI inference space that are coming to market. But this is where there seems to be a
49
00:05:15,440 -->
00:05:24,079
much broader appetite from customers for solutions that are more bespoke to their problem statement,
50
00:05:24,079 -->
00:05:31,679
which is different from training with, again, many more constraints. But it's also where we see that
51
00:05:31,679 -->
00:05:37,600
it is not a one size fits all. So, if you have a general purpose GPU that does very well in the
52
00:05:37,600 -->
00:05:43,839
training domain, it will do the job in inference, Amelia, but you'll pay more, it'll be less
53
00:05:43,839 -->
00:05:49,760
deterministic, it'll be more power hungry, and it'll be more inefficient than something that was
54
00:05:49,760 -->
00:05:55,760
more specifically designed for an inference deployment. Sure. So, this inference challenge
55
00:05:55,760 -->
00:06:03,600
in data center chips is closely watched and a very competitive market. So, how are you guys different?
56
00:06:03,600 -->
00:06:11,119
Yes. So, the focus that this company has had to sort has really been around addressing the memory
57
00:06:11,200 -->
00:06:19,920
wall problem. And just very simply, what happens in any AI deployment is that you have the compute
58
00:06:20,480 -->
00:06:26,480
and you have the memory. And the data that moves between the compute and the memory
59
00:06:27,040 -->
00:06:35,040
is actually something that most of the developers are trying to tackle because that data movement
60
00:06:35,040 -->
00:06:42,000
creates latencies, a lot of power inefficiencies, just the power that is required to move between,
61
00:06:42,000 -->
00:06:48,480
again, the compute and the memory. And it also introduces a lot of costs if you think about
62
00:06:48,480 -->
00:06:55,279
the compute sitting idle, waiting to be fed data. So, what Visora has done, and they've done it with
63
00:06:55,279 -->
00:07:01,279
a patented approach in terms of how they built their software architecture, is that they have
64
00:07:01,279 -->
00:07:09,359
addressed this memory wall by creating a software architecture that really collapses down many of
65
00:07:09,359 -->
00:07:17,119
the layers of memory between the processing unit and that external memory. And by collapsing those
66
00:07:17,119 -->
00:07:25,200
layers, fusing the operators, combining it into one set of instructions and really through their compiler,
67
00:07:25,279 -->
00:07:35,359
it just allows that memory to look much more like registers and near memory. And from that perspective,
68
00:07:35,359 -->
00:07:42,320
it allows the actual operation to happen faster and for that movement to be a lot less distant,
69
00:07:42,320 -->
00:07:46,480
if you will, from the processing unit to the external memory. So, just collapsing those layers
70
00:07:47,040 -->
00:07:54,399
allows more efficient data movement, more efficient software compilation, and less idle
71
00:07:54,399 -->
00:07:58,399
time of that processing unit sitting there waiting to be fed data.
72
00:07:58,399 -->
00:08:03,519
Can you give my listeners an update about Visora's recent successful tape out?
73
00:08:04,160 -->
00:08:09,760
Yes. So, this is pretty exciting. This is a big year for us, Amelia, because I mentioned earlier
74
00:08:09,760 -->
00:08:15,920
that this is a team that's delivered many chips to market in the past, both in a previous instantiation
75
00:08:15,920 -->
00:08:21,760
of a company that had a successful exit, as well as Visora when they were formed in 2015,
76
00:08:21,760 -->
00:08:28,720
which was around an automotive and an autonomous vehicle type of focus. They pivoted to AI inference
77
00:08:28,720 -->
00:08:37,679
in 2022, and they really had focused on this unique software architecture, addressing the memory wall
78
00:08:37,679 -->
00:08:43,760
problem, and then designing a very complex chip on latest process and packaging technology.
79
00:08:43,760 -->
00:08:51,919
We taped out that chip at the end of last year. We are expecting it back from TSMC in early May,
80
00:08:51,919 -->
00:08:56,719
and we're very excited because all of our simulated measurement that we've done and all
81
00:08:56,719 -->
00:09:02,320
the analysis that we've done shows that we will have a 3x performance improvement over
82
00:09:02,320 -->
00:09:09,359
the leading GPU and alternative solutions in the market at half of the power. So, much more power
83
00:09:09,359 -->
00:09:16,080
efficient, much greater performance throughput, and with this greater level of determinism, lower
84
00:09:16,080 -->
00:09:23,359
latency, and, of course, lower cost. So, this year, we get our chip back. We will be putting that in
85
00:09:23,919 -->
00:09:30,000
an OEM module. We're working with server providers, with rack providers, so that we can actually
86
00:09:30,000 -->
00:09:36,000
deliver to customers that are waiting to take the product at the end of the year. They can take it
87
00:09:36,000 -->
00:09:42,320
in a way that's more consumable and easy for them to deploy by providing the modules, the
88
00:09:42,320 -->
00:09:48,080
server, and the rack architecture, which, again, we're doing with our ecosystem partners. So, between
89
00:09:48,080 -->
00:09:54,479
getting the chip back, publishing our real measured results when we have our silicon back, publishing
90
00:09:54,479 -->
00:10:01,679
the results of the MLPerf benchmarks, which we plan to do in the late third quarter, and then
91
00:10:01,679 -->
00:10:07,840
having real customer deployments. It's just a critical, huge year for Missouri in 2026.
92
00:10:07,840 -->
00:10:14,400
Absolutely. I love it. All right. So, before I let you go, Sandra, let's talk about your family
93
00:10:14,400 -->
00:10:22,479
llama farm. So, how did that come about? Talk to me about that. So, my husband is a big outdoorsman,
94
00:10:22,479 -->
00:10:30,880
and we actually, as a family, love to camp and hike. And when he was in the mountains of Colorado
95
00:10:30,880 -->
00:10:36,479
on his two-week trip, a hunting trip with his friends, they realized that carrying up all of
96
00:10:36,479 -->
00:10:41,679
that weight, and certainly if they were successful in hunting, everything that they would need to
97
00:10:41,679 -->
00:10:48,719
carry down after a successful hunt, they actually wanted to have some help with that. And llamas are,
98
00:10:48,719 -->
00:10:55,679
as you know, pack animals that carry weight. So, he started with the fab four, as we call them,
99
00:10:55,679 -->
00:11:01,760
the first four llamas. And as a family, we took the llamas with us. They carried all of our tents
100
00:11:01,760 -->
00:11:08,080
and camping gear, our food, our water. I mean, it was wonderful. We were hiking big, long, steep
101
00:11:08,080 -->
00:11:14,400
trails here in the wonderful Sierras, Nevada, the mountains of California. And it was a beautiful
102
00:11:14,400 -->
00:11:23,200
thing. And our first four llamas, the kids helped name them, Dolly Llama, Llama Bean. We had Drama
103
00:11:23,200 -->
00:11:30,320
Llama and Michello Llama. So, that was super fun. But then my husband's hobby turned into an
104
00:11:30,320 -->
00:11:38,479
obsession. And today on our llama farm here in the East Bay of Silicon Valley, we have 32 llamas.
105
00:11:38,479 -->
00:11:44,960
All names, personalities, my husband knows them all. I can't confess to knowing all about them,
106
00:11:44,960 -->
00:11:50,000
but certainly he's the llama guy. And the neighbors love the llamas. They help feed the llamas.
107
00:11:50,400 -->
00:11:59,440
We take them to the local little school on their garden day. I've taken the llamas to our company's
108
00:11:59,440 -->
00:12:06,000
family picnic. I mean, llamas are beautiful pack animals that are just wonderful to have and
109
00:12:06,000 -->
00:12:13,440
very easy to care for. I love that so much. That is wonderful. Well, Sandra, it was a pleasure
110
00:12:13,440 -->
00:12:18,000
having you on. Thank you so much for joining me. Thank you for having me, Amelia. It was really
111
00:12:18,000 -->
00:12:24,159
delightful to talk with you as well. Well, folks, I must say that this interview is one of my
112
00:12:24,159 -->
00:12:31,520
favorite episodes of all time. Did you know that I have an Amelia's Favorite Episodes playlist
113
00:12:31,520 -->
00:12:41,119
on YouTube? This playlist includes my 500th episode with Mayman Aerospace CEO, David Mayman.
114
00:12:41,679 -->
00:12:49,200
In this episode, David and I chat about his motivation to design jetpacks and air utility
115
00:12:49,200 -->
00:12:58,640
vehicles and the evolution of Mayman Aerospace's Speeder VTOL air utility vehicle. Also included
116
00:12:58,640 -->
00:13:05,200
in this playlist is my discussion with Evan Coopersmith about the newest advancements in
117
00:13:05,200 -->
00:13:12,559
brain-computer interface technology. Evan and I also chat about the details of the Neural Latence
118
00:13:12,559 -->
00:13:19,760
Benchmark Challenge and how AI Studio is leveraging their expertise in machine learning
119
00:13:19,760 -->
00:13:27,359
to encourage innovation in BCI technology. And of course, my all-time favorite playlist
120
00:13:27,359 -->
00:13:35,520
would also include the Great Shark Cafe. In this episode, Rich Stump from Fathom and I
121
00:13:35,520 -->
00:13:42,640
chat about how Fathom is working with the Monterey Bay Aquarium Research Institute to create
122
00:13:42,640 -->
00:13:51,440
3D-printed video tracking devices for great white sharks in the California coastal region.
123
00:13:51,440 -->
00:13:57,599
And keeping with that theme, another absolute favorite episode of mine is about the Elephant
124
00:13:57,599 -->
00:14:05,440
Edge Challenge. Adam Benzion and I chat about this super cool challenge, what the Open Caller
125
00:14:05,440 -->
00:14:11,760
Initiative is all about, and how you can get involved to help save this vulnerable species
126
00:14:11,760 -->
00:14:18,640
from extinction. And amongst the rest of my favorite episodes is an oldie but a goodie
127
00:14:18,640 -->
00:14:27,679
called VisionTech's Dirty Dealings in ICs. In this episode, I dig into the dirty dealings
128
00:14:27,679 -->
00:14:35,359
of component house VisionTech. I investigate how one company from Clearwater, Florida
129
00:14:35,359 -->
00:14:44,400
managed to dupe over a thousand unsuspecting customers over $16 million and how those
130
00:14:44,479 -->
00:14:51,840
counterfeit ICs have affected our industry. And there are even more amazing episodes for your
131
00:14:51,840 -->
00:14:58,400
listening pleasure in this playlist. And of course, I've included a link below the player
132
00:14:58,400 -->
00:15:05,039
on this week's Fish Fry-In page that will take you right there. And you can also find this playlist
133
00:15:05,039 -->
00:15:11,760
under the playlist heading on our eeJournal YouTube channel as well. And if you'd like
134
00:15:11,760 -->
00:15:18,559
more information about Vasora, I've also included a link below the player on this week's Fish Fry-In
135
00:15:18,559 -->
00:15:25,919
page on eeJournal.com and in the description for this week's YouTube episode as well.
136
00:15:26,719 -->
00:15:34,080
Hey, have you checked out eeJournal on social media yet? Well, you should! You can find us
137
00:15:34,080 -->
00:15:42,000
at Facebook.com slash eeJournal. If LinkedIn is more your thing, I dig it. You can follow me or
138
00:15:42,000 -->
00:15:50,239
us on LinkedIn. And we are also on Blue Sky Social and Mastodon, too. And we have that YouTube
139
00:15:50,239 -->
00:15:57,919
channel I just mentioned, YouTube.com slash eeJournal. Folks, it is chock full of all kinds
140
00:15:57,919 -->
00:16:07,280
of techie videos, including our very popular Chalk Talk webcast series hosted by me. And of course,
141
00:16:07,280 -->
00:16:14,799
you can subscribe to our eeJournal YouTube channel as well. Thank you, everyone, for tuning in. If
142
00:16:14,799 -->
00:16:21,280
you know of any cool new technology or heck, you just want to chat, shoot me a line at Amelia,
143
00:16:21,280 -->
00:16:30,719
that's A-M-E-L-I-A at eeJournal.com. Or post a comment on our forums on eeJournal.
144
00:16:30,719 -->
00:16:40,479
For the week of April 10, 2026, I'm Amelia Dalton, and you've been fried.