WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 05/06/2026
23:39:21Duration: 3514.250
Channels: 1
1
00:00:11.360 -->
00:00:15.440Hello, and welcome to the Data Engineering Podcast, the show about modern data management.
2
00:00:16.395 -->
00:00:26.235Your host is Tobias Maci, and today I'm interviewing Robert Nishihara about the challenges of maximizing the utility of your available hardware for AI and data intensive applications.
3
00:00:26.395 -->
00:00:32.720Robert, can you start by introducing yourself? Thanks for having me on. I'm Robert. I'm one of the co founders of AnyScale,
4
00:00:32.800 -->
00:00:37.680and we are commercializing Ray, which is an open source project distributed
5
00:00:37.680 -->
00:00:39.040system that
6
00:00:39.360 -->
00:00:41.360a lot of companies use to scale
7
00:00:41.520 -->
00:00:44.879and run compute intensive AI workloads
8
00:00:44.454 -->
00:00:48.774ranging from model training to training data preparation,
9
00:00:49.254 -->
00:00:56.454inference, reinforcement learning. We can go into a lot more detail on all of those, but this is an open source project that we started
10
00:00:57.000 -->
00:01:00.920as grad students at Berkeley, and then started any scale to commercialize.
11
00:01:01.480 -->
00:01:05.240And do you remember how you first got started working in the data and AI space?
12
00:01:05.479 -->
00:01:14.335Yeah, well, I actually got my start more with AI research, more on the theoretical side of things. This was just around the time that deep
13
00:01:15.375 -->
00:01:19.535learning was taking off. If you remember in 2012,
14
00:01:19.535 -->
00:01:20.4152013,
15
00:01:20.735 -->
00:01:23.775deep learning was delivering amazing results, groundbreaking
16
00:01:24.175 -->
00:01:25.775results
17
00:01:24.980 -->
00:01:30.740in computer vision, and the whole world was It felt like AI was just exploding at the time,
18
00:01:31.860 -->
00:01:33.780so that's when I got into AI research.
19
00:01:35.380 -->
00:01:43.215I had no background or interest in distributed systems at the time, on the infrastructure side of things. My interest was really more on the algorithmic
20
00:01:43.215 -->
00:01:46.415side. Can we come up with better algorithms for learning
21
00:01:46.495 -->
00:01:52.255from data, for training these models? Can we develop better general purpose techniques for learning?
22
00:01:52.735 -->
00:01:58.140So at the time, we were doing research on deep learning training methods, optimization algorithms,
23
00:01:58.140 -->
00:01:59.900reinforcement learning algorithms.
24
00:02:00.300 -->
00:02:08.745And the bottleneck that we faced, or one of the bottlenecks, was that in order to AI is very empirical, so if you come up with a new algorithm,
25
00:02:08.825 -->
00:02:14.905in order to validate whether it's good or not, you have to try it out and see how it performs in practice.
26
00:02:15.705 -->
00:02:24.180Often need to really convince yourself you need to not just try it out on a small toy problem, you need to try it out at some meaningful scale with a largish model,
27
00:02:24.739 -->
00:02:28.020a large amount of data, and see if it can really solve the problem.
28
00:02:29.060 -->
00:02:32.660In order to run those experiments, you need to scale your
29
00:02:32.900 -->
00:02:35.495algorithm across a bunch of machines, across
30
00:02:35.895 -->
00:02:37.735a bunch of GPUs,
31
00:02:37.895 -->
00:02:44.295or you have to And so we found ourselves spending all of our time I'm talking about my fellow grad students and myself.
32
00:02:44.535 -->
00:02:52.230We were spending all of our time building tools for managing clusters, for handling machine failures, like moving data across machines,
33
00:02:52.470 -->
00:02:53.750getting stuff to run on
34
00:02:54.470 -->
00:02:55.190cheaper
35
00:02:55.270 -->
00:02:57.430spot instances or on GPUs,
36
00:02:57.510 -->
00:03:00.150and we thought, Wow, surely there's
37
00:03:00.630 -->
00:03:10.275some reusable tooling that can be built here so that not everyone has to build their own tools and redo all of this all the time. And that led us to start RAID. Basically,
38
00:03:10.515 -->
00:03:10.915we
39
00:03:11.394 -->
00:03:26.420thought that AI was taking off and that the need for scale was only going to grow, right? That the era of doing machine learning research on a single machine was going to be a thing of the past, and so the need for scale and the degree of scale was only going to grow.
40
00:03:26.980 -->
00:03:35.675And that was just going to introduce a lot of hard systems and infrastructure challenges, and so there 's a big opening to try to build something useful there.
41
00:03:36.395 -->
00:03:43.195Initially, we were just building it for ourselves, but we thought, Hey, distributed computing is really hard. We'd love to build tools that are useful for a lot of people.
42
00:03:43.515 -->
00:03:47.290That's kind of how we got started. And it ended up being very prescient
43
00:03:47.290 -->
00:03:49.530and also coincided
44
00:03:49.530 -->
00:03:53.530fairly closely with the introduction and rise of Kubernetes.
45
00:03:53.610 -->
00:03:56.970I'm wondering what are some of the evolutions
46
00:03:56.970 -->
00:04:02.325from when you first started Ray and introduced it as an available project
47
00:04:02.645 -->
00:04:06.085through to where we are now where we have moved beyond
48
00:04:06.165 -->
00:04:20.370the niche aspect of these deep learning. Despite the popularity that it grew to, it was still fairly bespoke in terms of who was using it to LLM training and inference, which has gained much broader adoption,
49
00:04:20.530 -->
00:04:24.370and the distributed compute ecosystem has also evolved substantially.
50
00:04:24.370 -->
00:04:29.665Just wondering if you can talk through some of that growth and how Ray has taken advantage of it.
51
00:04:30.225 -->
00:04:32.945Yeah, and you mentioned far more people are building
52
00:04:33.265 -->
00:04:35.505models today than before,
53
00:04:35.825 -->
00:04:44.330and that is really going to accelerate with the rise of coding agents, which it's still, even today, it's hard to build models and make use of your data.
54
00:04:45.450 -->
00:05:01.905As coding agents like Cursor and Cloud Code and others reduce that barrier and the amount of expertise required to really make use of your data and train models, I think far more people are going to do it. And you're right, the landscape has changed a tremendous amount.
55
00:05:02.305 -->
00:05:03.264Kubernetes
56
00:05:03.264 -->
00:05:07.824was not anything like what it is today, the standard it is today.
57
00:05:09.000 -->
00:05:20.840Today, the vast majority of Ray users are using Ray on top of Kubernetes, and Kubernetes has really just emerged as the dominant standard for container orchestration and managing
58
00:05:21.160 -->
00:05:22.040provisioning
59
00:05:22.345 -->
00:05:24.185compute and managing container
60
00:05:24.585 -->
00:05:27.225life cycles across different clouds.
61
00:05:27.544 -->
00:05:30.745And that was not the case when we started Ray.
62
00:05:32.665 -->
00:05:36.025We were in a distributed systems and AI lab at Berkeley,
63
00:05:36.569 -->
00:05:38.490and so there was a lot of The
64
00:05:38.889 -->
00:05:53.455people who had created Apache Spark were in that lab, and so we had a lot of experience with other distributed systems, and we launched a bunch of Spark clusters, and the standard way to launch Spark clusters was not on Kubernetes, it was to
65
00:05:53.615 -->
00:06:04.095run this script that another grad student had written to directly talk to EC2 and spin up virtual machines and SSH to them and install all the relevant stuff, and
66
00:06:04.630 -->
00:06:10.230that has changed quite a bit. But I would say the change you've seen over those years is
67
00:06:10.390 -->
00:06:11.430the shift
68
00:06:11.670 -->
00:06:14.310from a high degree of fragmentation to
69
00:06:14.870 -->
00:06:15.670consolidation
70
00:06:15.670 -->
00:06:18.245of the infrastructure tech stack.
71
00:06:18.405 -->
00:06:23.125And Kubernetes is one example. There used to be a number of container orchestrators and
72
00:06:24.405 -->
00:06:26.324ways of doing container orchestration.
73
00:06:26.724 -->
00:06:28.965At another layer, think about deep learning frameworks.
74
00:06:29.890 -->
00:06:31.890Of course, you're familiar with
75
00:06:32.370 -->
00:06:33.970PyTorch and TensorFlow.
76
00:06:34.210 -->
00:06:36.050There used to be Theano
77
00:06:36.050 -->
00:06:38.530was a popular one, Torch.
78
00:06:38.770 -->
00:06:39.330Actually,
79
00:06:40.850 -->
00:06:41.410folks at Berkeley
80
00:06:42.765 -->
00:06:46.205created Caffe, which was one of the early popular
81
00:06:46.525 -->
00:06:51.725deep learning frameworks. There were a dozen of these. There are a ton of different deep learning frameworks,
82
00:06:51.885 -->
00:06:52.365and
83
00:06:52.765 -->
00:06:54.685now it is primarily PyTorch.
84
00:06:54.685 -->
00:06:55.565So what
85
00:06:55.840 -->
00:07:04.240often happens when you have a new use case or new workload, like the emergence of AI, there's a proliferation of different frameworks
86
00:07:04.639 -->
00:07:06.160and then consolidation,
87
00:07:06.160 -->
00:07:17.455because you tend to have a standard emerge over time that gains momentum that many people are using and contributing to, especially with open source. Open source really lends itself to the
88
00:07:18.014 -->
00:07:19.855emergence of a standard.
89
00:07:22.255 -->
00:07:24.815And so the PyTorch layer
90
00:07:24.815 -->
00:07:28.600that we see, the Kubernetes layer, those are really
91
00:07:28.680 -->
00:07:36.520some of the most important layers that we see, pieces of software that we see people using together with Ray. And so typically,
92
00:07:36.840 -->
00:07:42.275our users are not just using Ray on its own, they're using Ray plus PyTorch plus Kubernetes,
93
00:07:42.514 -->
00:07:44.755often plus a VLLM or SGLANG
94
00:07:44.755 -->
00:07:47.235to implement their overall workloads.
95
00:07:48.035 -->
00:07:48.595And
96
00:07:48.914 -->
00:07:51.555I know too that one of the
97
00:07:52.210 -->
00:07:56.530early standout features of Ray was the RayTune
98
00:07:56.530 -->
00:07:57.409library,
99
00:07:57.409 -->
00:07:59.090which was focused on hyperparameter
100
00:07:59.090 -->
00:08:02.770tuning, which was a very frequent topic of conversation
101
00:08:02.930 -->
00:08:08.305and one that I hear a lot less now, particularly because there are so many more of them.
102
00:08:08.865 -->
00:08:30.320And I'm just wondering if you can talk to some of the ways that the focus of Ray has shifted from when you first started it and in those early days of deep learning growth and adoption to where we are now, where deep learning is still valuable and there are areas of utility for it, but the majority of the focus is on these transformer based language models or vision models?
103
00:08:31.040 -->
00:08:31.360Yeah,
104
00:08:32.105 -->
00:08:38.345hyperparameter tuning itself is a fun area to talk about, and then I'll come back to
105
00:08:38.745 -->
00:08:44.265how Ray relates to all of that. But was always hard to difficult to optimize
106
00:08:44.750 -->
00:08:49.550neural networks, because there are many choices you have to make. Choices around the architecture,
107
00:08:50.029 -->
00:08:57.790choices of the parameters of the optimization algorithm, learning rates, things like dropout and so forth. And
108
00:08:58.385 -->
00:09:00.865if you got the parameters a little bit wrong,
109
00:09:01.265 -->
00:09:02.225it was easy
110
00:09:02.945 -->
00:09:04.305for the optimization
111
00:09:04.305 -->
00:09:06.225to not work. And so
112
00:09:08.065 -->
00:09:10.785there were a lot of papers written and a lot of experiments done to
113
00:09:11.350 -->
00:09:15.190figure out the best way to just search over these different parameters.
114
00:09:16.870 -->
00:09:18.870Hyperparameter search is basically
115
00:09:19.110 -->
00:09:29.635train the model a bunch of times with different settings of the parameters and see what works the best. And you can be more clever about that, you can be less clever about that and search randomly.
116
00:09:30.035 -->
00:09:33.155In fact, random search is a very strong baseline.
117
00:09:33.955 -->
00:09:34.515And
118
00:09:35.155 -->
00:09:37.875of course that lends itself well to Ray because
119
00:09:39.340 -->
00:09:40.940if you're training the
120
00:09:40.940 -->
00:09:45.980same model a bunch of times, it's a compute intensive workload, you're using a bunch of things in parallel,
121
00:09:46.380 -->
00:09:51.420and so you need a good way to express that. And it can get more complex than just
122
00:09:52.325 -->
00:09:54.405run a bunch of experiments in parallel,
123
00:09:54.645 -->
00:09:57.365because you're often looking at the results of some of the experiments,
124
00:09:57.525 -->
00:09:58.725stopping them early,
125
00:09:58.965 -->
00:10:05.045investing more resources in the promising ones, spinning up new experiments based on what worked well or what worked poorly,
126
00:10:05.285 -->
00:10:06.965but it's not that complicated.
127
00:10:07.480 -->
00:10:08.120Now,
128
00:10:08.600 -->
00:10:10.760what has changed in hyperparameter
129
00:10:10.760 -->
00:10:11.960search? First,
130
00:10:12.440 -->
00:10:13.640a few things have changed.
131
00:10:14.600 -->
00:10:15.560One is
132
00:10:16.120 -->
00:10:18.280we're starting to train really big models,
133
00:10:18.520 -->
00:10:27.985and you just, for your big run, you can't actually run 100 copies of the training run, because you just don't have the compute resources to do that. So that
134
00:10:28.225 -->
00:10:29.265is no longer
135
00:10:29.745 -->
00:10:31.425viable in a lot of cases.
136
00:10:31.905 -->
00:10:32.625Second,
137
00:10:33.185 -->
00:10:36.865we came up with a lot better initialization techniques for
138
00:10:37.230 -->
00:10:38.269neural networks.
139
00:10:39.470 -->
00:10:45.709Lot of the problems that hyperparameter search was solving was just not knowing how to initialize
140
00:10:45.790 -->
00:10:48.029the weights of the neural network or
141
00:10:48.589 -->
00:10:56.205how to set learning rates and things like that, and that has been a lot more best practices have emerged around that, so you can get it right more frequently.
142
00:10:56.925 -->
00:10:58.525And the third thing is
143
00:10:59.085 -->
00:10:59.725where
144
00:11:01.404 -->
00:11:07.029the emphasis is now is not about run a bunch of experiments and take the best result.
145
00:11:07.270 -->
00:11:08.470It's really more,
146
00:11:08.709 -->
00:11:11.110I'm only going to do one big training run.
147
00:11:11.510 -->
00:11:12.150So
148
00:11:12.630 -->
00:11:14.230I need to get that one right.
149
00:11:14.790 -->
00:11:20.834I can do a bunch of small scale experiments first. And so how do I do small scale experiments
150
00:11:20.915 -->
00:11:28.355and learn how to set the parameters at that small scale and configure things properly so that when I do my big run,
151
00:11:30.274 -->
00:11:32.130will work? And so there's a
152
00:11:32.610 -->
00:11:33.810progression
153
00:11:33.810 -->
00:11:42.370of run more experiments at a small scale, use that to set a smaller number of experiments at the next scale, and gradually go up in scale,
154
00:11:42.690 -->
00:11:44.050but reduce the number of experiments,
155
00:11:44.574 -->
00:11:53.295and use what you learned at the smaller scale to ensure that the larger scale runs go well. So there's a lot of shift in perspective
156
00:11:53.454 -->
00:11:56.495of how to do what it means to do hyperparameter search.
157
00:11:57.100 -->
00:11:59.340Now, just to say hyperparameter
158
00:11:59.340 -->
00:12:03.660search is one use case that people used Ray for. Another very early
159
00:12:03.980 -->
00:12:08.620use case for Ray was reinforcement learning. This was actually kind of the motivating use case because
160
00:12:09.195 -->
00:12:15.755the whole world was excited about this previous wave of reinforcement learning with Atari and MuJoCo and AlphaGo,
161
00:12:16.395 -->
00:12:20.395and we wanted to do research on those types of algorithms.
162
00:12:20.955 -->
00:12:21.115And
163
00:12:22.090 -->
00:12:26.170now, of course, we can talk about how the use cases we see
164
00:12:26.970 -->
00:12:29.450people running with Ray have shifted, but
165
00:12:29.770 -->
00:12:36.410reinforcement learning, of course, has made a comeback. And now we see a ton of people using Ray for reinforcement learning
166
00:12:36.915 -->
00:12:38.755for post training LLMs.
167
00:12:38.995 -->
00:12:39.555And
168
00:12:39.875 -->
00:12:40.435one of the
169
00:12:41.475 -->
00:12:45.635actually, just the other day, Cursor released their Composer two model,
170
00:12:45.875 -->
00:12:46.435which
171
00:12:47.075 -->
00:12:49.235uses a lot of reinforcement learning to
172
00:12:49.714 -->
00:12:51.714build a great coding model. And
173
00:12:52.050 -->
00:12:53.970they use Ray for all of that for
174
00:12:54.370 -->
00:12:56.210basically the RL infrastructure.
175
00:12:57.730 -->
00:12:59.490Then in order to be able to
176
00:12:59.810 -->
00:13:01.090build these models,
177
00:13:01.330 -->
00:13:03.570whether you're doing a from scratch
178
00:13:03.650 -->
00:13:04.529training
179
00:13:04.529 -->
00:13:18.255run or you're doing fine tuning of the model, there's also a lot of data preparation and data manipulation involved, which Ray is also very well situated for, particularly given the fact that it's easily integrated into the broader Python ecosystem,
180
00:13:18.255 -->
00:13:22.095which has become the de facto language for a lot of these workloads.
181
00:13:22.690 -->
00:13:24.690And for people who are
182
00:13:25.170 -->
00:13:28.290building these data workflows, data pipelines,
183
00:13:28.370 -->
00:13:53.360there are numerous tools to choose from. I'm just wondering if you can give some of the ways that you think about the juxtaposition of Ray versus something like an orchestrator like Airflow or Dagster or a distributed compute system like Spark and some of the selection criteria for when somebody would use Ray for a particular use case versus reaching to some of the other tools that are more of these, I'm going to say,
184
00:13:53.920 -->
00:13:55.920pipeline native workflows?
185
00:13:56.240 -->
00:13:57.920Yeah, that's a great question.
186
00:13:58.320 -->
00:14:05.775So data preparation, training data preparation and preprocessing is something that has changed so much since
187
00:14:06.095 -->
00:14:08.895the early days of deep learning. If you remember
188
00:14:09.855 -->
00:14:12.735early on when people were training deep learning models,
189
00:14:12.895 -->
00:14:15.135you would preprocess your data, but
190
00:14:15.455 -->
00:14:23.330the preprocessing you would do was very limited. You might in the case of ImageNet, a computer vision data set, you might try to
191
00:14:24.530 -->
00:14:29.810take each data point and generate multiple data points to augment your data set. And you could do that by
192
00:14:30.130 -->
00:14:34.445taking each image and randomly cropping or scaling it to get some
193
00:14:35.485 -->
00:14:37.404more versions of that same image.
194
00:14:43.324 -->
00:14:48.045Going back to that point in time, all of the research was on model architecture.
195
00:14:48.459 -->
00:14:49.420The
196
00:14:49.420 -->
00:14:52.700ImageNet dataset was a fixed dataset
197
00:14:52.779 -->
00:14:57.980and split into your training set and your test set, so it's this fixed benchmark.
198
00:14:58.220 -->
00:14:59.580All of the research was,
199
00:14:59.820 -->
00:15:05.025how do I choose the best optimization algorithm and the best neural network architecture
200
00:15:05.105 -->
00:15:11.425so that I can when I train that on my training set and then test it on my test set, I get the best results, get the best score.
201
00:15:12.145 -->
00:15:16.625That has almost entirely flipped. Of course, there's still interesting research happening
202
00:15:17.000 -->
00:15:17.640on
203
00:15:18.040 -->
00:15:23.160model architectures, but people have largely converged on the transformer architecture.
204
00:15:24.920 -->
00:15:26.920And optimization algorithms are largely
205
00:15:27.160 -->
00:15:29.160variations of stochastic gradient descent,
206
00:15:29.895 -->
00:15:31.175and despite
207
00:15:31.335 -->
00:15:33.655tons and tons of research in that area.
208
00:15:34.535 -->
00:15:35.095And
209
00:15:35.895 -->
00:15:37.415people have realized that
210
00:15:37.655 -->
00:15:38.855really getting
211
00:15:39.015 -->
00:15:41.095great results comes down to getting the data right.
212
00:15:43.230 -->
00:16:00.375And so there was this maybe previous mental block where the dataset was treated as static. It was treated as kind of a given that you don't optimize, and now that's the thing you really optimize over, and you invest money in collecting data, and you spend a lot of your experimentation
213
00:16:00.375 -->
00:16:01.815and compute budget
214
00:16:02.295 -->
00:16:04.535on curating the data. And so
215
00:16:05.015 -->
00:16:06.375now we see
216
00:16:06.615 -->
00:16:17.640it's not simple data cleaning where you just strip trailing white space from your sentences and you crop your images so they're all the same size and things like that, and normalize your data.
217
00:16:17.959 -->
00:16:18.600You are
218
00:16:19.240 -->
00:16:19.959really
219
00:16:20.199 -->
00:16:22.519doing a ton of experimentation
220
00:16:22.215 -->
00:16:25.975and actually running a ton of models to filter out low quality data,
221
00:16:26.295 -->
00:16:29.575to augment your data with high quality synthetic data,
222
00:16:29.895 -->
00:16:31.495and annotate your data.
223
00:16:31.735 -->
00:16:32.535For example,
224
00:16:32.775 -->
00:16:38.370often using models to generate a lot of structure in your data. For example, if
225
00:16:38.770 -->
00:16:46.130you have a data set of images or videos, you might use a vision language model to generate captions for those videos and images,
226
00:16:46.290 -->
00:16:54.175and then to later use in training. You may compute embeddings. It's very common to run dozens of classifiers
227
00:16:54.255 -->
00:16:56.415in your data preparation
228
00:16:56.415 -->
00:16:57.215process
229
00:16:57.295 -->
00:16:58.975so that you can use
230
00:16:59.775 -->
00:17:04.415those tags, you can filter out subsets of your data to train on specific subsets,
231
00:17:05.149 -->
00:17:05.710and
232
00:17:06.190 -->
00:17:08.589you can filter out low quality data.
233
00:17:08.750 -->
00:17:16.190So it's very common to see data curation pipelines that involve dozens of stages of filtering and annotation
234
00:17:16.190 -->
00:17:16.989and
235
00:17:17.230 -->
00:17:17.950are really,
236
00:17:19.065 -->
00:17:20.344really complex.
237
00:17:20.505 -->
00:17:30.664It's just enormously different from how people did prepared training data in the past. So a lot of things have changed about this data preparation stage.
238
00:17:30.825 -->
00:17:32.505Some of the big differences
239
00:17:32.760 -->
00:17:38.360are that it's now model driven and GPU driven instead of CPU driven.
240
00:17:38.520 -->
00:17:39.240Basically,
241
00:17:39.800 -->
00:17:40.040are
242
00:17:40.680 -->
00:17:43.640More and more data processing is shifting to GPUs
243
00:17:43.880 -->
00:17:44.600because
244
00:17:44.680 -->
00:17:46.440more and more data processing
245
00:17:47.414 -->
00:17:49.014is being done with inference,
246
00:17:49.174 -->
00:17:49.734and
247
00:17:49.975 -->
00:17:55.974that is a massive shift. So you ask about how does Ray relate to the data world,
248
00:17:56.615 -->
00:17:59.815and when would you use Ray versus Apache Spark,
249
00:18:00.135 -->
00:18:01.255or when would you use
250
00:18:01.920 -->
00:18:07.440Airflow and workflow orchestrators like that? Let me start with the workflow orchestrators,
251
00:18:07.520 -->
00:18:12.880because those are actually quite complementary with Ray. They operate at different levels of granularity.
252
00:18:13.680 -->
00:18:15.520Of course, there's always overlap,
253
00:18:15.680 -->
00:18:18.075but I think of Ray as
254
00:18:18.555 -->
00:18:19.915running an individual
255
00:18:19.915 -->
00:18:20.794workload.
256
00:18:20.955 -->
00:18:21.915I'm
257
00:18:21.915 -->
00:18:24.475running my data processing pipeline, and it has
258
00:18:24.955 -->
00:18:28.875this op stage where I download the data and
259
00:18:29.435 -->
00:18:33.049decompress the videos, and then I stream that into
260
00:18:33.290 -->
00:18:34.810some CPU
261
00:18:34.810 -->
00:18:38.010based heuristic filtering, and then I stream that into some
262
00:18:38.650 -->
00:18:40.410model based filtering, which
263
00:18:40.970 -->
00:18:43.530judges aesthetic score and
264
00:18:43.684 -->
00:18:51.445filters out explicit content or these kinds of things. And then this streams into some vision language model stage where I'm running using transformers,
265
00:18:51.605 -->
00:18:56.165and then maybe I'm writing out to Parquet or writing to some blob storage.
266
00:18:56.770 -->
00:18:57.409And
267
00:18:57.809 -->
00:19:01.169Ray would be used to express this workload,
268
00:19:01.170 -->
00:19:01.809assign
269
00:19:03.010 -->
00:19:05.970different compute resources to each stage of computation,
270
00:19:06.690 -->
00:19:07.809manage the processes
271
00:19:09.645 -->
00:19:16.684sort of that are executing each operation, stream data between them, handle back pressure if one stage is too slow,
272
00:19:17.005 -->
00:19:20.605handle failures if one process dies and needs to be recreated,
273
00:19:20.685 -->
00:19:22.925handle auto scaling of
274
00:19:23.460 -->
00:19:26.979the compute resources to match the throughput of all the different operations
275
00:19:27.299 -->
00:19:30.259and things like that, really solving these core distributed
276
00:19:30.419 -->
00:19:31.700computing challenges.
277
00:19:32.100 -->
00:19:34.259Now, that data processing
278
00:19:34.740 -->
00:19:35.539workload
279
00:19:35.905 -->
00:19:37.264might just be one
280
00:19:37.425 -->
00:19:39.585part of some overall broader
281
00:19:39.665 -->
00:19:47.025workload. For example, maybe you want to set up something like you're collecting new data every day. You're a robotics company. You have
282
00:19:48.230 -->
00:19:57.110robots out in the real world that are streaming video back, so you've got new data coming in every day. Every day, you want to run that data processing job to curate your data.
283
00:19:57.510 -->
00:19:58.950Once that finishes,
284
00:19:59.030 -->
00:20:00.790you want to take the curated data,
285
00:20:01.265 -->
00:20:03.505maybe copy that over to a different cloud,
286
00:20:03.745 -->
00:20:09.425and then your Neo Cloud, where you run training, and then kick off a training job. That kind of coarse grained
287
00:20:09.425 -->
00:20:18.370workflow orchestration is where you would use something like Airflow or a different one, and that's very compatible with Ray. You might use Ray to run the individual components of each
288
00:20:18.770 -->
00:20:19.570stage.
289
00:20:19.570 -->
00:20:21.970So we see Ray being used together with
290
00:20:22.370 -->
00:20:25.409Airflow, Flight, and all of these different engines.
291
00:20:25.810 -->
00:20:29.345The Spark comparison is interesting. So there are
292
00:20:30.304 -->
00:20:31.264First,
293
00:20:31.904 -->
00:20:33.825the high order thing, of course, is that
294
00:20:34.304 -->
00:20:45.230Ray is used for a huge range of workloads, and Spark is specific for big data processing. So Ray is also used for big data processing, but it's also used for reinforcement learning, inference, training, and
295
00:20:45.630 -->
00:20:47.869stuff like that. But you
296
00:20:48.030 -->
00:20:54.830can ask the question specifically about data processing. If I have a bunch of data and I need to transform and manipulate my data,
297
00:20:55.405 -->
00:20:57.884when would I use Ray and when would I use Spark?
298
00:20:58.285 -->
00:21:02.764They are actually designed for very different use cases. So to oversimplify,
299
00:21:02.765 -->
00:21:07.325Spark was built at a time when, with the emergence of big data,
300
00:21:08.000 -->
00:21:08.720when
301
00:21:09.120 -->
00:21:11.440people were not really thinking about GPUs.
302
00:21:11.600 -->
00:21:17.920So if I to, if I have a bunch of tabular data, if I have a bunch of data that is nicely structured in tables,
303
00:21:18.320 -->
00:21:19.200and I want to
304
00:21:19.855 -->
00:21:23.774do analytics, I want to join a bunch of tables together
305
00:21:24.014 -->
00:21:25.934and run SQL queries
306
00:21:26.095 -->
00:21:33.615and do analytics like that, Spark is a fantastic choice. Where Spark starts to run into limitations
307
00:21:33.890 -->
00:21:38.850is when I have more multimodal data. Perhaps I'm working with images,
308
00:21:38.930 -->
00:21:39.570video,
309
00:21:40.210 -->
00:21:41.090robotic
310
00:21:41.090 -->
00:21:42.210sensor data,
311
00:21:42.450 -->
00:21:44.130all of these types of things.
312
00:21:44.450 -->
00:21:44.930And
313
00:21:45.330 -->
00:21:50.225the way that I manipulate this type of data, the way I get value out of this really
314
00:21:50.225 -->
00:21:51.105unstructured,
315
00:21:51.105 -->
00:21:55.184really multimodal data is not by running SQL queries on it.
316
00:21:55.665 -->
00:21:59.985Instead, it is by running inference on it. So what am I gonna do with
317
00:22:00.305 -->
00:22:02.385just a PDF or a video?
318
00:22:02.870 -->
00:22:11.429I'm going to feed it into a model. And that because that model is is able to understand it and and and reason about it. And so
319
00:22:11.590 -->
00:22:16.390we're entering this regime where data processing is becoming
320
00:22:16.835 -->
00:22:25.154an inference heavy workload and is running on GPUs. You've still got CPUs, of course. You've got it's a mixture of CPUs and GPUs. And
321
00:22:25.715 -->
00:22:27.075that heterogeneity,
322
00:22:27.075 -->
00:22:30.355this scenario where you have tons of multimodal data
323
00:22:30.790 -->
00:22:35.190and it is being processed both with inference and with regular processing
324
00:22:35.350 -->
00:22:37.669on a mixture of CPUs and GPUs,
325
00:22:38.070 -->
00:22:45.385that is a scenario where Ray is the best system for is really designed for using that for that kind of scenario.
326
00:22:45.625 -->
00:22:47.705And so to oversimplify,
327
00:22:47.785 -->
00:22:50.985Spark is amazing for running SQL queries
328
00:22:51.145 -->
00:22:52.505on tabular data,
329
00:22:52.665 -->
00:22:56.905and Ray is amazing for running inference on multimodal data.
330
00:22:58.169 -->
00:22:58.970You
331
00:22:58.970 -->
00:22:59.690mentioned
332
00:23:00.169 -->
00:23:00.970GPUs
333
00:23:00.970 -->
00:23:02.809being a critical
334
00:23:03.130 -->
00:23:06.809hardware need for a lot of these data processing workloads.
335
00:23:06.809 -->
00:23:07.769You mentioned
336
00:23:08.250 -->
00:23:16.535the focus on inference that Ray has been investing in. And we also briefly touched on some of the distributed systems
337
00:23:16.615 -->
00:23:17.495capabilities
338
00:23:17.495 -->
00:23:19.655around Ray and Kubernetes
339
00:23:19.655 -->
00:23:21.414and some of the overlap there.
340
00:23:21.655 -->
00:23:43.534And I think that all of that focuses in on an interesting question as well that a lot of people are dealing with right now is how do I actually get the most out of my hardware because these GPUs are super expensive. I wanna make sure that they're not just sitting idle when I could be putting them to some good use or deallocating them from my cloud environment if you're in a cloud and using some form of auto scaling.
341
00:23:43.695 -->
00:23:46.575I also know that, in particular, Kubernetes,
342
00:23:46.575 -->
00:23:48.014their orchestration
343
00:23:48.335 -->
00:23:50.495system is optimized for
344
00:23:51.010 -->
00:23:55.489a different style of use case than what a lot of these high data throughput
345
00:23:55.490 -->
00:24:16.254workloads need. And so you might sometimes find yourself fighting with the Kubernetes orchestrator where it's saying, hey. You're done over there, and you're saying, no. I'm really not. And I'm just wondering if you could talk to some of those aspects of how teams are dealing with some of this question of resource optimization at the hardware level to be able to get the most out of their expensive compute.
346
00:24:16.495 -->
00:24:18.095Yeah, that's a great question.
347
00:24:18.255 -->
00:24:19.295I'll talk about
348
00:24:20.600 -->
00:24:25.559the Ray Kubernetes relationship, and also about just really getting compute optimization,
349
00:24:25.559 -->
00:24:26.919getting the best utilization.
350
00:24:27.080 -->
00:24:29.399I want to say one more thing on the data topic,
351
00:24:29.559 -->
00:24:34.805which is that I mentioned Ray is amazing for this world of GPU data processing,
352
00:24:34.965 -->
00:24:39.685processing data with inference, and multimodal data. I think it's worth pointing out that
353
00:24:39.925 -->
00:24:43.685all of this multimodal data was previously useless.
354
00:24:43.845 -->
00:24:51.459So if you think about And I'm exaggerating a little bit, but think about all the random PDFs and documents sitting in your organization,
355
00:24:51.940 -->
00:24:54.339or video recordings of meetings,
356
00:24:54.740 -->
00:24:56.820or of sales calls, or
357
00:24:57.220 -->
00:24:59.459audio recordings of things.
358
00:24:59.620 -->
00:25:06.355You would store that data and then not do anything with it. Think about how many meetings people have been in where
359
00:25:06.515 -->
00:25:10.595you might record the meeting, but no one's going to go back and actually listen to it.
360
00:25:10.915 -->
00:25:14.195And so that data was There's tremendous
361
00:25:14.595 -->
00:25:23.590value and information in all of this data. It's just really hard to manipulate and get insights out of it. So the thing that has changed is that
362
00:25:24.070 -->
00:25:27.110AI is making it possible to programmatically
363
00:25:27.510 -->
00:25:28.550analyze
364
00:25:28.550 -->
00:25:29.350and manipulate
365
00:25:29.735 -->
00:25:33.495all different types of data because you have powerful multimodal models.
366
00:25:33.895 -->
00:25:37.335And so of course, reason Spark
367
00:25:37.415 -->
00:25:46.059and other systems have been primarily working with tabular data is not that that's the only valuable data, but rather it's just the easiest to work with and to
368
00:25:46.460 -->
00:25:47.740ask questions about.
369
00:25:49.020 -->
00:25:51.500Tabular
370
00:25:51.500 -->
00:26:00.334data is really a tiny, it's a miniscule fraction of the world's data, And now that we can unlock value in all the rest of the data,
371
00:26:00.815 -->
00:26:12.040we're going to start storing way more of it. We're going to start using it all the time, and this is going to be tremendously valuable for her. So I wanted to share that. Now, on your point about
372
00:26:12.840 -->
00:26:17.400cost of GPUs, yes, GPUs are tremendously expensive. And so
373
00:26:19.160 -->
00:26:21.960getting good utilization of your expensive resource
374
00:26:22.435 -->
00:26:26.034is a hard problem, and it has to be solved.
375
00:26:27.635 -->
00:26:34.434And this is why a lot of companies have infrastructure teams that are responsible for managing all of the compute
376
00:26:34.755 -->
00:26:36.595and making that compute
377
00:26:36.940 -->
00:26:39.740available to their AI researchers
378
00:26:39.740 -->
00:26:40.299and
379
00:26:40.540 -->
00:26:41.340practitioners.
380
00:26:41.340 -->
00:26:48.620So think about the challenge. It's also much harder at different scales. There are many different dimensions of
381
00:26:48.995 -->
00:26:49.875complexity.
382
00:26:50.195 -->
00:26:53.475So it's one thing if you have one AI person
383
00:26:53.715 -->
00:26:54.595running
384
00:26:54.595 -->
00:27:00.195one training job on one cluster. But if I have five different Kubernetes clusters
385
00:27:00.195 -->
00:27:01.635across different clouds,
386
00:27:02.035 -->
00:27:03.554and I've got a team of
387
00:27:03.795 -->
00:27:05.820dozens of researchers
388
00:27:05.820 -->
00:27:28.134that need to share that compute. All of a sudden, I need to solve for a few things. I need my researchers to be productive. I need them to be able to really move quickly, debug easily, run things at scale, search for capacity across different clouds, and take advantage of compute resources wherever they are so that they can be productive. I also need to get great utilization, right? So I need
389
00:27:28.375 -->
00:27:29.254to be able to
390
00:27:29.575 -->
00:27:31.654prioritize different workloads against each other.
391
00:27:33.100 -->
00:27:38.619I may have a big training run that needs all the GPUs and needs to run all at once. But when that finishes,
392
00:27:38.940 -->
00:27:42.940I don't want the GPUs to just sit idle. So I might need some background
393
00:27:43.260 -->
00:27:50.554elastic job that can just soak up all the unused compute and just expand elastically to fill the unused compute.
394
00:27:51.274 -->
00:27:54.554And how do I set that up? How do I enable
395
00:27:54.875 -->
00:27:58.475teams to run different workloads with different priorities and appropriately
396
00:27:58.620 -->
00:28:09.500share their compute resources among them. The kinds of challenges that we see these infrastructure teams needing to solve are, one, making their developers and their end users, researchers
397
00:28:09.660 -->
00:28:10.539productive.
398
00:28:10.540 -->
00:28:23.855Two is having a standardized interface to all of their compute so that you can easily plug in new sources of compute. That's important for two reasons. One is it lets you shop around for the cheapest GPUs and plug them in. And second,
399
00:28:24.430 -->
00:28:33.549it lets researchers and the users run their workloads everywhere, so that factors into the productivity point. So that's a standardized interface to your compute. Third is
400
00:28:33.790 -->
00:28:35.150being able to
401
00:28:35.390 -->
00:28:40.424really do a good job of assigning the right resources to the right workloads. And that will mean
402
00:28:40.825 -->
00:28:42.345workload prioritization.
403
00:28:42.345 -->
00:28:44.984It will mean elasticity of low
404
00:28:45.225 -->
00:28:47.624priority workloads so that they can expand and
405
00:28:48.025 -->
00:28:49.544make sure you always have
406
00:28:49.865 -->
00:28:58.229stuff running so that your GPUs are not sitting idle. And then fourth is being able to search for capacity across different regions and different clouds,
407
00:28:58.470 -->
00:29:00.710which if you have bursty jobs,
408
00:29:00.870 -->
00:29:05.365you're going back and reprocessing all of your data to compute some new feature,
409
00:29:05.605 -->
00:29:13.045being able to quickly acquire a lot of capacity across a bunch of different regions is very helpful for that. So those are some of the
410
00:29:14.005 -->
00:29:17.125in order to do a good job with GPU utilization,
411
00:29:17.880 -->
00:29:25.000those are all of some of the challenges you have to solve. That is in addition to Those are challenges that sit around
412
00:29:25.480 -->
00:29:26.679what Ray does.
413
00:29:27.160 -->
00:29:40.054Ray is responsible for running your training workload, your RL workload, your data workload, and running that individual workload and making sure it is performance and reliable and fast and cost efficient. And then everything else I just described
414
00:29:40.215 -->
00:29:52.869sits outside of the workload and has to be solved at a different layer. So you have to think about many different layers of the stack. Now, Ray and Kubernetes, you asked about the relationship between Ray and Kubernetes. They're highly complementary,
415
00:29:53.110 -->
00:30:07.554and they sit at different layers of the stack. So for the AI infrastructure stack, I like to think about the PyTorch layer, the Ray layer, and the Kubernetes layer. All of this is the software stack that sits on top of your GPUs and cloud providers.
416
00:30:07.635 -->
00:30:18.350So each layer on its own is not sufficient. Each layer solves some fraction of the infrastructure problems that you need to solve. What PyTorch is responsible for is
417
00:30:18.670 -->
00:30:25.150squeezing the most performance out of the model on the GPU, just running the model both for training and inference
418
00:30:25.310 -->
00:30:27.550in the most performant way possible.
419
00:30:27.870 -->
00:30:28.430And
420
00:30:29.215 -->
00:30:30.174that is
421
00:30:30.414 -->
00:30:37.695not just PyTorch. There's a rich ecosystem around PyTorch. So think about frameworks like VLLM and SGLANG for optimizing
422
00:30:37.774 -->
00:30:40.815inference with transformers, or frameworks like Megatron for
423
00:30:41.660 -->
00:30:42.299training.
424
00:30:45.020 -->
00:30:57.625But fundamentally, the responsibility of that layer is about squeezing the most performance out of the chips by running the model. Ray, we talked about a bunch. The responsibility at that layer is about solving
425
00:30:57.625 -->
00:30:59.705the distributed computing challenges.
426
00:30:59.785 -->
00:31:05.785So this means process management, process lifecycle management, process coordination
427
00:31:05.785 -->
00:31:06.905and communication,
428
00:31:06.905 -->
00:31:07.784data movement,
429
00:31:08.105 -->
00:31:13.279data ingest, failure handling, because a lot of the hardware is unreliable,
430
00:31:13.520 -->
00:31:17.919resource allocation to different portions or components of your workload,
431
00:31:18.240 -->
00:31:23.520stuff like that. And Kubernetes is responsible for container orchestration,
432
00:31:23.520 -->
00:31:27.325so provisioning of the compute, managing container life cycles,
433
00:31:27.805 -->
00:31:30.125and things like that. And those are all
434
00:31:30.285 -->
00:31:31.565very complementary.
435
00:31:31.645 -->
00:31:35.164And we've also seen them co evolve with each other. So
436
00:31:35.645 -->
00:31:39.245the Ray open source community and the Kubernetes open source community
437
00:31:39.880 -->
00:31:42.119collaborate deeply to really
438
00:31:43.160 -->
00:31:44.279share information
439
00:31:44.520 -->
00:31:49.239and have the right interfaces and be able to optimize in a way that they
440
00:31:49.880 -->
00:31:54.325couldn't without each other. For example, Kubernetes has the ability to
441
00:31:54.885 -->
00:31:57.284resize containers, resize
442
00:31:57.525 -->
00:32:00.245containers to add more memory or things like that.
443
00:32:01.285 -->
00:32:02.325Kubernetes
444
00:32:02.325 -->
00:32:09.940on its own doesn't know when it's appropriate to resize the container. On the other hand, Ray is running the workload inside of those containers.
445
00:32:10.100 -->
00:32:14.340And so Ray has knowledge of what the workload is and what its resource requirements are.
446
00:32:14.900 -->
00:32:17.780And so Ray is actually in the perfect position to
447
00:32:17.940 -->
00:32:25.335say, hey, we need more resources over here. But Ray is not managing the containers, and so is not able to execute that without Kubernetes'
448
00:32:25.335 -->
00:32:28.455help. And so there are lots of things like this where,
449
00:32:28.775 -->
00:32:30.855by working together and co evolving,
450
00:32:30.935 -->
00:32:33.655the different layers of the infrastructure stack are
451
00:32:33.895 -->
00:32:35.655developing to work better together.
452
00:32:36.340 -->
00:32:39.940And this is also especially true with Ray and VLLM,
453
00:32:39.940 -->
00:32:41.459where the
454
00:32:43.059 -->
00:32:46.820VLM community and the Ray community have worked very deeply together
455
00:32:47.059 -->
00:32:47.699to
456
00:32:47.860 -->
00:32:50.340make it possible to do performant
457
00:32:50.340 -->
00:32:55.595cross node inference. So LLM inference gets far more complex
458
00:32:55.675 -->
00:32:58.395when you have large models that span many machines.
459
00:32:58.715 -->
00:33:00.875If you have a single model on a single machine,
460
00:33:01.115 -->
00:33:02.475that's a simpler scenario.
461
00:33:02.875 -->
00:33:11.740But a large model that spans multiple machines is its own little distributed system. A single query to the model may touch many different experts
462
00:33:12.220 -->
00:33:18.539your expert layers, and those experts may be sharded across different machines. And so the query is getting routed around in
463
00:33:18.860 -->
00:33:19.820complex ways.
464
00:33:20.220 -->
00:33:23.794Different parts of the computation may get disaggregated
465
00:33:23.794 -->
00:33:27.874into separate pools of compute. There are many things like this where
466
00:33:28.195 -->
00:33:33.235Ray and VLM work together to enable complex forms of cross node parallelism.
467
00:33:34.130 -->
00:33:42.289The whole idea of sharding the models is also interesting, particularly for people who aren't deep in the weeds of the actual model architectures.
468
00:33:42.450 -->
00:34:01.865And I know that the models themselves, they present as being this monolithic thing when in reality, they're just a hierarchy of the actual neural layers. I'm wondering if you can maybe talk a little bit too to some of those challenges of being able to distribute the model effectively across different GPUs for the case where you can't fit it all on a single chip.
469
00:34:02.690 -->
00:34:06.049This is an area that's become far more complex recently,
470
00:34:06.370 -->
00:34:10.290because in the past, when you were scaling inference,
471
00:34:10.450 -->
00:34:12.850you would just stick your model in a container,
472
00:34:12.930 -->
00:34:14.770and then you would replicate the container
473
00:34:15.170 -->
00:34:18.915however many times you need it. And it really just wasn't that complicated.
474
00:34:19.235 -->
00:34:23.475If you need to scale more, you replicate the container more. Now
475
00:34:23.715 -->
00:34:26.115you may have many containers and many
476
00:34:26.435 -->
00:34:27.315machines
477
00:34:27.315 -->
00:34:30.730that are running a single replica of the model.
478
00:34:31.210 -->
00:34:32.730And that is
479
00:34:33.210 -->
00:34:35.690so questions like elasticity
480
00:34:35.690 -->
00:34:42.250become a lot more complex, failure handling become a lot more complex, because you can't just think of failure handling at the container level.
481
00:34:42.570 -->
00:34:43.530Locality matters.
482
00:34:46.005 -->
00:34:49.045I mentioned when you are running a big model,
483
00:34:49.444 -->
00:34:52.245you may separate So with transformers especially,
484
00:34:52.645 -->
00:34:56.244there's this pre fill stage where you process the input tokens,
485
00:34:56.660 -->
00:35:15.915which might be more compute bound, GPU compute bound. Then you've got your decode stage where you're generating the output tokens one at a time, which might be more GPU memory bandwidth bound. And it might make sense to separate out those two stages into different pools of compute and assign different GPUs to each one. But there may be
486
00:35:16.155 -->
00:35:28.060certain shards of your prefill workers and certain shards of your decode workers that correspond to each other, and actually you want them to be co located. And so now you need to think about co locating containers in
487
00:35:28.300 -->
00:35:29.260the same nodes.
488
00:35:29.580 -->
00:35:40.940And that never happened before when you were thinking about just, I have a single model and a single container and replicating it. This has especially gotten more complex with large mixture of expert models,
489
00:35:41.260 -->
00:35:41.660where
490
00:35:42.174 -->
00:35:44.734the expert stages are often
491
00:35:45.135 -->
00:35:47.535sharded across a bunch of different GPUs
492
00:35:47.694 -->
00:35:48.415and
493
00:35:49.375 -->
00:35:53.055a single query needs to be routed around to different experts.
494
00:35:53.214 -->
00:35:54.654So that is something that is,
495
00:35:55.350 -->
00:35:56.630and there are many different
496
00:35:56.950 -->
00:36:02.870ways of sharding and partitioning your model. It is not like, oh, there's just one strategy that
497
00:36:03.350 -->
00:36:07.670always works the best. So that is an area that's grown far more complex.
498
00:36:08.815 -->
00:36:11.135In order to be able to
499
00:36:11.535 -->
00:36:14.175use and manage these systems effectively,
500
00:36:14.175 -->
00:36:16.575obviously, there's also the question of
501
00:36:16.655 -->
00:36:17.775observability
502
00:36:17.775 -->
00:36:21.615and being able to understand what is being executed,
503
00:36:21.615 -->
00:36:27.830how efficient it's being executed. I'm wondering if you can just talk to some of the leading reasons for
504
00:36:28.070 -->
00:36:30.230wasted or inefficient compute.
505
00:36:30.230 -->
00:36:35.350We obviously talked a lot about some of the ways that Ray helps to alleviate that situation,
506
00:36:35.430 -->
00:36:38.015but also just some of the organizational
507
00:36:38.015 -->
00:36:46.975and team education that's required to make sure that they understand how best to effectively apply the capabilities of a framework like Ray to
508
00:36:47.055 -->
00:36:49.775such a complex and multivariate
509
00:36:49.775 -->
00:36:50.415problem space.
510
00:36:51.300 -->
00:36:52.100Yeah.
511
00:36:52.100 -->
00:36:54.980One of the biggest reasons, I would say, is not
512
00:36:55.780 -->
00:37:02.020having a good way to share GPUs between training and inference. Fundamentally, and this is at the fleet organizational
513
00:37:02.020 -->
00:37:02.660level,
514
00:37:02.820 -->
00:37:07.935not at the individual workload level. If I have training workloads and inference workloads,
515
00:37:08.415 -->
00:37:08.895and
516
00:37:09.375 -->
00:37:14.575I am partitioning my GPUs between them and trying to provision
517
00:37:15.055 -->
00:37:18.350each one or inference for peak capacity,
518
00:37:18.430 -->
00:37:22.910then there are going to be a lot of times when my inference GPUs are idle and not
519
00:37:23.390 -->
00:37:24.990being used because
520
00:37:24.990 -->
00:37:26.670we're not at the peak capacity.
521
00:37:26.910 -->
00:37:28.590And so being able to
522
00:37:29.305 -->
00:37:34.025efficiently share GPUs between training and inference in a way that doesn't
523
00:37:34.505 -->
00:37:35.225risk
524
00:37:35.385 -->
00:37:37.625your most important production workloads,
525
00:37:37.865 -->
00:37:38.985but allows
526
00:37:39.385 -->
00:37:42.825excess capacity to be used by lower priority workloads
527
00:37:43.080 -->
00:37:43.960is
528
00:37:44.120 -->
00:37:45.720probably the number one thing.
529
00:37:46.280 -->
00:37:52.440And that is complex to get right. That's the kind of problems that are solved by the MECL platform
530
00:37:52.600 -->
00:37:55.240and the layers that sit around Ray.
531
00:37:55.640 -->
00:37:56.200And
532
00:37:56.440 -->
00:37:58.200I would say that's
533
00:37:58.775 -->
00:38:00.215outside of the workload.
534
00:38:00.535 -->
00:38:02.855Within the workload itself,
535
00:38:03.415 -->
00:38:05.335there can be many different bottlenecks.
536
00:38:06.695 -->
00:38:09.895It's important to be able to get
537
00:38:09.895 -->
00:38:11.575good performance within a single workload.
538
00:38:12.180 -->
00:38:15.300It's very important to have the tools to tackle
539
00:38:15.619 -->
00:38:19.860whatever bottleneck is you're running into at the moment. So
540
00:38:20.099 -->
00:38:22.740what do I mean by that? I mentioned with
541
00:38:22.900 -->
00:38:23.700in inference,
542
00:38:24.494 -->
00:38:29.295might separate out the prefill and decode stages into separate pools of compute, because
543
00:38:29.535 -->
00:38:34.255they are different shaped workloads. They have different types of bottlenecks, and so
544
00:38:35.055 -->
00:38:41.420it's natural that you might want different compute resources for each one into potentially different accelerators
545
00:38:41.579 -->
00:38:42.380and to
546
00:38:42.619 -->
00:38:44.940right size the compute for each stage.
547
00:38:45.180 -->
00:38:51.340The same thing can be true of a data processing pipeline or a training pipeline. You might have stages that are GPU bound.
548
00:38:51.660 -->
00:38:56.755You might have stages that are IO bound. You might have components that are memory bound.
549
00:38:57.235 -->
00:39:03.875And if you don't have a system that can support a large degree of heterogeneity
550
00:39:04.115 -->
00:39:04.915and
551
00:39:05.075 -->
00:39:05.875assigning,
552
00:39:05.875 -->
00:39:08.435breaking down a workload into different pieces,
553
00:39:08.950 -->
00:39:15.430assigning different compute resources to each piece, and rightsizing and scaling those resources appropriately,
554
00:39:15.910 -->
00:39:16.470then
555
00:39:16.950 -->
00:39:19.910you are going to likely have inefficiencies.
556
00:39:20.390 -->
00:39:25.675And this is just to give an example with training. As you scale training on more GPUs,
557
00:39:25.835 -->
00:39:34.955it's very easy to become bottlenecked by the data ingest and preprocessing side of things. I might be using some It might be very CPU heavy processing.
558
00:39:35.115 -->
00:39:39.650I might be using some models in my preprocessing right before it gets fed into training.
559
00:39:39.810 -->
00:39:40.450Now,
560
00:39:41.170 -->
00:39:42.370if I can't
561
00:39:42.690 -->
00:39:44.610separate out that preprocessing
562
00:39:44.610 -->
00:39:46.850into a different pool of compute
563
00:39:46.930 -->
00:39:49.890and then scale that pool of compute independently from training,
564
00:39:50.625 -->
00:39:51.265then
565
00:39:51.425 -->
00:39:55.025I'm not going have a way to eliminate the data in just bottleneck.
566
00:39:55.105 -->
00:40:00.145And so this is really where Ray shines. It's giving you the control
567
00:40:00.465 -->
00:40:04.865to separate out different pieces of your workload into different pools of compute
568
00:40:05.570 -->
00:40:09.730of any different type, connect them all together, scale them independently,
569
00:40:10.050 -->
00:40:13.410manage them independently, handle failures independently.
570
00:40:14.530 -->
00:40:17.890And so when there is a bottleneck in one component of your workload,
571
00:40:18.675 -->
00:40:22.595you can address that without everything being so tightly coupled.
572
00:40:22.835 -->
00:40:23.715And that is,
573
00:40:24.035 -->
00:40:25.075I would say,
574
00:40:25.715 -->
00:40:27.315the key for eliminating
575
00:40:27.715 -->
00:40:32.730bottlenecks within a workload. So that's probably the biggest thing inside the workload. Then
576
00:40:32.970 -->
00:40:35.210the biggest thing outside of the workload is
577
00:40:35.530 -->
00:40:39.450sharing resources between workloads, especially training and inference.
578
00:40:41.290 -->
00:40:43.450Yeah, the sharing resources
579
00:40:43.450 -->
00:40:45.610is one of the pieces that I think is
580
00:40:46.455 -->
00:40:47.575most complex
581
00:40:47.575 -->
00:40:48.775and probably
582
00:40:49.015 -->
00:40:55.655least effectively executed by most teams because it requires so much of that coordination of understanding
583
00:40:55.895 -->
00:40:57.495what the workloads are,
584
00:40:57.815 -->
00:41:04.930which also brings up the question of, well, what if I have two different Ray clusters running or I have one system?
585
00:41:04.930 -->
00:41:11.330Maybe I just have an independently deployed VLLM that's using up a portion of my GPU, and then I have Ray
586
00:41:11.490 -->
00:41:14.050using the other portion of that GPU for
587
00:41:14.575 -->
00:41:15.935data preprocessing,
588
00:41:15.935 -->
00:41:27.375and I'm just curious how you can give some level of visibility to Ray to be aware of the other workloads that are co located within that same pool of hardware.
589
00:41:28.730 -->
00:41:33.450Yeah. This is a great example of how you can do a better job by
590
00:41:33.690 -->
00:41:36.090co designing different layers of the stack,
591
00:41:36.410 -->
00:41:36.970because
592
00:41:37.210 -->
00:41:40.090the layer outside of Ray, the
593
00:41:40.170 -->
00:41:40.890overall
594
00:41:41.130 -->
00:41:46.255compute provisioning and orchestration layer, is the layer that's responsible for
595
00:41:46.495 -->
00:41:47.295deciding
596
00:41:47.295 -->
00:41:50.415which resources to allocate to which workloads,
597
00:41:50.575 -->
00:41:53.375which nodes to preempt, things like that.
598
00:41:55.135 -->
00:42:01.510But to do a great job of that, you really want to know what's running inside the workload, and that's information that Ray has,
599
00:42:01.670 -->
00:42:03.670information that exists at the
600
00:42:04.950 -->
00:42:07.350distributed compute, like the workload layer.
601
00:42:07.750 -->
00:42:08.310And
602
00:42:08.550 -->
00:42:08.950so
603
00:42:09.735 -->
00:42:15.255if you can combine those pieces of information, then you can do a much better job of preempting
604
00:42:15.255 -->
00:42:16.295non critical
605
00:42:16.935 -->
00:42:21.255portions of your workload that can be more easily scaled down or restarted later.
606
00:42:21.990 -->
00:42:22.630And
607
00:42:22.870 -->
00:42:27.830that is something that the kind of thing that we think about when co designing
608
00:42:28.230 -->
00:42:30.870Ray with these other layers of the stack. So
609
00:42:31.110 -->
00:42:33.190you're calling out a very important problem.
610
00:42:34.704 -->
00:42:35.585As
611
00:42:35.585 -->
00:42:36.865you have been
612
00:42:37.425 -->
00:42:38.065working
613
00:42:38.385 -->
00:42:39.025on
614
00:42:39.185 -->
00:42:47.025and building the AnyScale company around it and just evolving along with this fast moving ecosystem,
615
00:42:47.025 -->
00:42:59.500what are some of the most interesting or innovative or unexpected ways that you've seen Ray being applied and maybe some of the lessons that you've learned from these unexpected uses that have helped guide the future trajectory of the project?
616
00:43:00.220 -->
00:43:01.020Yeah,
617
00:43:01.180 -->
00:43:02.540it's been really
618
00:43:03.625 -->
00:43:05.225interesting to see how
619
00:43:05.465 -->
00:43:07.945the broad categories of workloads
620
00:43:07.945 -->
00:43:11.465have largely remained the same. If you think about training,
621
00:43:12.185 -->
00:43:12.825inference,
622
00:43:13.625 -->
00:43:14.905data processing,
623
00:43:15.465 -->
00:43:17.305but each one has
624
00:43:18.039 -->
00:43:22.280evolved a ton. We talked about how inference has grown in complexity with
625
00:43:22.440 -->
00:43:24.920large models and multi node models.
626
00:43:25.960 -->
00:43:29.160We talked a bit about how the data processing side has evolved,
627
00:43:29.480 -->
00:43:30.280to
628
00:43:30.839 -->
00:43:31.640accommodate
629
00:43:31.855 -->
00:43:35.135all this multimodal data and really these heterogeneous
630
00:43:35.135 -->
00:43:36.895inference heavy pipelines.
631
00:43:37.215 -->
00:43:41.535On the training side, we've seen the evolution of, or the
632
00:43:42.255 -->
00:43:47.970resurgence of reinforcement learning, which is far more complex than regular training because,
633
00:43:48.450 -->
00:43:52.849well, it has all the complexities of regular training. You're still doing regular training,
634
00:43:53.170 -->
00:43:55.650but you're also doing inference.
635
00:43:55.730 -->
00:44:00.755Are also running To generate new data, you're running simulations or environments
636
00:44:01.315 -->
00:44:03.075that are tightly coordinating
637
00:44:03.075 -->
00:44:06.355with the inference, going back and forth to generate data.
638
00:44:06.595 -->
00:44:13.315You are shuffling data around. You're moving the data back to training. You're moving the new model weights over to the inference portion.
639
00:44:14.070 -->
00:44:14.710You
640
00:44:14.870 -->
00:44:17.510can have failures in each of these different stages.
641
00:44:17.830 -->
00:44:21.110Training can fail, of course, but as
642
00:44:21.670 -->
00:44:25.030models, we have more powerful agentic models that can
643
00:44:25.350 -->
00:44:28.230reason and take actions over long periods of time,
644
00:44:28.905 -->
00:44:30.745you need to think about failures
645
00:44:30.905 -->
00:44:36.025on the rollout side of things, the data generation side, because if rollout
646
00:44:36.905 -->
00:44:38.505is happening over the course of
647
00:44:39.385 -->
00:44:41.945ten hours or a day or multiple days,
648
00:44:42.425 -->
00:44:48.310you really don't want to lose that progress. And so you have to handle failures on that side of things as well.
649
00:44:49.910 -->
00:44:53.910So this introduces a ton of both algorithmic
650
00:44:53.910 -->
00:44:57.270complexities as well as infrastructure complexities.
651
00:44:57.670 -->
00:44:58.230And
652
00:44:59.095 -->
00:45:01.255Ray is, the more
653
00:45:02.055 -->
00:45:02.935challenging
654
00:45:02.935 -->
00:45:04.615AI workloads become,
655
00:45:05.015 -->
00:45:07.255the more important Ray becomes.
656
00:45:07.415 -->
00:45:08.055Because
657
00:45:10.055 -->
00:45:10.615if
658
00:45:11.255 -->
00:45:12.375you're just taking
659
00:45:12.740 -->
00:45:17.859a single process and replicating that a bunch of times and running the identical thing
660
00:45:18.660 -->
00:45:19.540everywhere,
661
00:45:20.020 -->
00:45:20.580then
662
00:45:21.380 -->
00:45:23.460that's fairly simple to do and you
663
00:45:23.859 -->
00:45:24.660don't need Ray.
664
00:45:25.595 -->
00:45:26.155When
665
00:45:26.395 -->
00:45:27.994there is heterogeneity,
666
00:45:27.994 -->
00:45:28.795when your
667
00:45:28.875 -->
00:45:38.714workload is broken down into different components that have different responsibilities and need to coordinate with each other, like in reinforcement learning, like in multi node inference, like in these
668
00:45:39.450 -->
00:45:41.050AI data pipelines,
669
00:45:41.530 -->
00:45:43.610then you really need Ray.
670
00:45:44.890 -->
00:45:47.770And to that point of statefulness
671
00:45:47.770 -->
00:45:49.530and failure recovery,
672
00:45:49.530 -->
00:45:50.570it also brings
673
00:45:51.135 -->
00:45:55.295question of systems such as Temporal that are focused on that.
674
00:45:55.935 -->
00:46:03.535They're termed crash proof, which is a little bit of a misnomer, but I'm wondering what you're seeing as some of the ways that people are integrating with some of those more
675
00:46:03.960 -->
00:46:05.640stateful workflow management
676
00:46:06.359 -->
00:46:16.280systems like Temporal to be able to do some of that failure recovery, for instance, if you're in the middle of a training run and you want to be able to make sure that you can resume from the appropriate checkpoint or whatever.
677
00:46:16.895 -->
00:46:19.135Yeah, failure recovery is essential,
678
00:46:19.135 -->
00:46:25.695right? Because failures happen all the time. And not only is failure recovery important, but fast
679
00:46:25.775 -->
00:46:28.095failure recovery is important because
680
00:46:29.055 -->
00:46:31.454the larger your cluster, the more frequent the failures.
681
00:46:32.090 -->
00:46:35.770And if you have a failure in extreme cases happening
682
00:46:35.850 -->
00:46:37.050every few hours,
683
00:46:37.610 -->
00:46:40.250and you spend thirty minutes recovering,
684
00:46:40.250 -->
00:46:44.890reloading from a checkpoint, that is a significant fraction of your overall time.
685
00:46:45.290 -->
00:46:45.530So
686
00:46:46.375 -->
00:46:47.735systems that
687
00:46:48.295 -->
00:46:48.695Great
688
00:46:49.495 -->
00:46:52.935recovery from failure is really a non negotiable.
689
00:46:53.415 -->
00:46:53.895And
690
00:46:54.935 -->
00:46:57.655one interesting trend we've seen there is that
691
00:47:00.270 -->
00:47:06.910the application needs to have quite a bit of control over how to handle failures, because the appropriate way to handle failures
692
00:47:07.070 -->
00:47:08.830may be quite custom.
693
00:47:09.230 -->
00:47:13.550And we see this more and more with the GB200s
694
00:47:13.550 -->
00:47:17.595and 300s, some of the latest generation of NVIDIA hardware,
695
00:47:17.915 -->
00:47:18.475which
696
00:47:19.115 -->
00:47:21.035previously you would reason about
697
00:47:21.355 -->
00:47:23.915a node of maybe eight GPUs.
698
00:47:24.315 -->
00:47:25.195And now
699
00:47:26.635 -->
00:47:27.435the
700
00:47:27.515 -->
00:47:33.410right level to reason about is not a node, but a rack. You might have 72
701
00:47:33.970 -->
00:47:36.130GPUs in a rack where
702
00:47:36.290 -->
00:47:37.010each
703
00:47:37.090 -->
00:47:40.850rack has a number of nodes and communication within the rack is far more
704
00:47:41.445 -->
00:47:42.725faster
705
00:47:42.725 -->
00:47:44.645than communication between racks.
706
00:47:46.725 -->
00:47:49.125And there are multiple levels of topology here.
707
00:47:49.605 -->
00:47:52.085And because of the performance difference,
708
00:47:52.245 -->
00:47:53.205the application
709
00:47:53.445 -->
00:47:56.450needs to precisely
710
00:47:56.450 -->
00:47:57.250control
711
00:47:57.410 -->
00:48:02.690which parts of the application run-in the same rack, which ones run-in different rack. And so if
712
00:48:03.250 -->
00:48:04.450a GPU fails,
713
00:48:05.010 -->
00:48:07.170let's say there's 72 GPUs in the rack,
714
00:48:08.865 -->
00:48:10.865I know these things are going to fail, and they're
715
00:48:11.185 -->
00:48:12.305going to fail frequently.
716
00:48:12.465 -->
00:48:15.905And so I may start my application and run without
717
00:48:16.865 -->
00:48:21.985using all of the GPUs. Maybe I'll use 64 and keep some of them as backups so that I can
718
00:48:22.900 -->
00:48:27.700swap them in when something goes wrong. And now I need to know, okay, GPU dies.
719
00:48:27.700 -->
00:48:28.180My
720
00:48:28.980 -->
00:48:32.260application needs to have control over which precise
721
00:48:32.340 -->
00:48:40.025processes to tear down and making sure we can recreate them in the same rack, and it needs to track how many spares it has,
722
00:48:40.345 -->
00:48:40.825and
723
00:48:41.224 -->
00:48:44.185then wire everything up with specific other processes.
724
00:48:44.505 -->
00:48:47.785But then it may need to do something different when I've used up the spares.
725
00:48:48.025 -->
00:48:49.065It may need to
726
00:48:49.545 -->
00:48:50.425recreate,
727
00:48:50.744 -->
00:48:53.385tear down a bunch of stuff and acquire a whole new rack.
728
00:48:54.640 -->
00:48:56.080Or maybe it will decide
729
00:48:56.560 -->
00:48:57.840at that point, Actually,
730
00:48:58.160 -->
00:48:59.920we can't acquire a new rack.
731
00:49:00.320 -->
00:49:01.280We need to
732
00:49:01.600 -->
00:49:11.355proceed with training on just a smaller number of workers, a smaller number of racks, and that might mean adapting batch sizes and things like that. And so there is a ton of
733
00:49:11.915 -->
00:49:13.915logic and
734
00:49:14.234 -->
00:49:16.555control that the application needs to have
735
00:49:16.795 -->
00:49:17.595over
736
00:49:18.395 -->
00:49:22.120what to do in the presence of a failure, And that is something that
737
00:49:23.800 -->
00:49:29.080Ray uniquely provides, is really this level of control over process failures
738
00:49:29.320 -->
00:49:30.840and what to do about them.
739
00:49:31.480 -->
00:49:32.040And
740
00:49:32.280 -->
00:49:33.960in your work of
741
00:49:34.135 -->
00:49:39.175building these technologies and understanding the use cases and applications,
742
00:49:39.175 -->
00:49:47.975what are some of the most interesting or unexpected or challenging lessons that you've learned personally in your work of this very data and AI heavy workloads?
743
00:49:49.130 -->
00:49:49.850Yeah.
744
00:49:51.290 -->
00:49:57.450First of all, there are many potential bottlenecks all over the place. And when you're trying
745
00:49:57.450 -->
00:50:00.730to build a flexible tool like Ray that supports a variety of use cases,
746
00:50:01.835 -->
00:50:03.515it's not enough to
747
00:50:04.154 -->
00:50:06.635do one thing well, you have to really
748
00:50:06.795 -->
00:50:10.795have the tools to eliminate any bottleneck that could come up.
749
00:50:11.275 -->
00:50:11.915Now,
750
00:50:12.474 -->
00:50:13.835I would say another lesson,
751
00:50:14.690 -->
00:50:18.370there is a or there can be a tension between
752
00:50:18.450 -->
00:50:19.250building
753
00:50:19.570 -->
00:50:21.090general purpose tools
754
00:50:21.250 -->
00:50:24.130and really optimizing for a specific use case.
755
00:50:24.690 -->
00:50:25.490And
756
00:50:26.690 -->
00:50:28.290this is actually a common
757
00:50:29.335 -->
00:50:34.135question people will ask about Ray. They'll say, Ray is so general purpose, it's so flexible.
758
00:50:34.295 -->
00:50:35.415Doesn't that mean
759
00:50:35.815 -->
00:50:39.335it's less optimized for each individual use case?
760
00:50:40.215 -->
00:50:40.455And
761
00:50:41.330 -->
00:50:44.130there is some truth to that, but the way that Ray,
762
00:50:44.850 -->
00:50:48.610the way we approach that with Ray is to really
763
00:50:49.250 -->
00:50:50.690build two layers.
764
00:50:51.090 -->
00:50:53.010There's the lower level primitives,
765
00:50:53.010 -->
00:50:54.130which are highly flexible.
766
00:50:54.945 -->
00:51:02.225And really, this is the array as an actor framework. So the ability to The low level primitives is basically the ability to
767
00:51:02.545 -->
00:51:06.065spin up a Python class as an actor and to manage processes.
768
00:51:07.500 -->
00:51:08.140And
769
00:51:08.620 -->
00:51:11.740that is very flexible. You can build anything with functions
770
00:51:11.740 -->
00:51:16.300and classes, and so you can build any kind of distributed system with actors.
771
00:51:18.540 -->
00:51:24.335And at that level, that core part of Ray, we focus on keeping it simple,
772
00:51:24.575 -->
00:51:28.655having a very small number of highly flexible primitives,
773
00:51:29.214 -->
00:51:30.975and making sure they are
774
00:51:31.615 -->
00:51:33.615battle tested, they're hardened,
775
00:51:33.694 -->
00:51:34.655their performance.
776
00:51:34.974 -->
00:51:39.760And we are very conservative about adding new functionality at that layer.
777
00:51:40.400 -->
00:51:41.040Then
778
00:51:41.520 -->
00:51:49.760there is the library ecosystem on top of Ray, in the same way that you have Python and then the core primitives, then you've got the library ecosystem around Python.
779
00:51:50.605 -->
00:51:53.405And there are libraries built on top of Ray for
780
00:51:53.485 -->
00:51:57.405data processing, for training, for reinforcement learning, for inference.
781
00:51:57.565 -->
00:51:58.125And
782
00:51:58.605 -->
00:52:02.445that's the layer at which you can specialize and really
783
00:52:03.485 -->
00:52:08.590deeply optimize for a specific use case, building on top of the lower level primitives.
784
00:52:08.670 -->
00:52:09.390And so
785
00:52:09.710 -->
00:52:11.790that's the approach we've taken to
786
00:52:12.350 -->
00:52:14.270trying to get the best of both worlds,
787
00:52:14.670 -->
00:52:17.310combining optimization with generality.
788
00:52:17.950 -->
00:52:18.510And
789
00:52:19.075 -->
00:52:20.755I think that's worked quite well.
790
00:52:21.155 -->
00:52:26.275We've actually seen a rich ecosystem emerge on top of Ray. So nearly every
791
00:52:26.595 -->
00:52:29.075open source RL framework
792
00:52:29.315 -->
00:52:30.915is built on top of Ray.
793
00:52:31.155 -->
00:52:40.670Companies like NVIDIA are building their data curation frameworks on top of Ray. We see inference frameworks and agent frameworks built on top of Ray. So that is
794
00:52:42.190 -->
00:52:43.470one of the lessons.
795
00:52:44.029 -->
00:52:50.734And as you continue to build and invest in this technology and ecosystem,
796
00:52:51.214 -->
00:53:00.255what are some of the predictions that you have for some of the next set of challenges or the next evolution of use cases as we
797
00:53:00.579 -->
00:53:03.780maybe move into newer model architectures
798
00:53:03.780 -->
00:53:09.700or invest more in optimizing the existing set of transformer applications?
799
00:53:10.339 -->
00:53:13.619Yeah. Complexity is going to continue to grow
800
00:53:14.095 -->
00:53:15.855along every axis.
801
00:53:16.095 -->
00:53:16.655So
802
00:53:17.055 -->
00:53:18.495on the hardware side, there
803
00:53:19.295 -->
00:53:20.655will be more and more
804
00:53:21.295 -->
00:53:25.775levels of topology that we need to reason about, more and more topology aware
805
00:53:28.030 -->
00:53:29.070scheduling.
806
00:53:30.990 -->
00:53:32.190Hardware failure
807
00:53:32.270 -->
00:53:34.109rates will continue
808
00:53:34.430 -->
00:53:37.710to be very prominent. There will be more heterogeneity,
809
00:53:37.710 -->
00:53:40.190just way more types of hardware accelerators
810
00:53:40.190 -->
00:53:43.385and different cloud providers that people will use. So
811
00:53:43.865 -->
00:53:46.265the complexity and heterogeneity
812
00:53:46.585 -->
00:53:47.305will
813
00:53:47.465 -->
00:53:49.545continue to grow on the hardware side.
814
00:53:49.865 -->
00:53:51.945The same will be true on the application side.
815
00:53:52.265 -->
00:53:53.705Data scale will grow,
816
00:53:53.945 -->
00:53:55.145model scale will grow,
817
00:53:55.839 -->
00:53:56.560the
818
00:53:56.720 -->
00:54:00.240types and heterogeneity of data will grow. So for example,
819
00:54:00.800 -->
00:54:06.320one thing you don't really have with tabular data, with tabular data, your data roughly all looks the same.
820
00:54:06.960 -->
00:54:09.839With multimodal data, you might have a video that's
821
00:54:10.845 -->
00:54:13.725ten seconds long, and you might have one that's three hours long.
822
00:54:14.045 -->
00:54:14.605And
823
00:54:15.245 -->
00:54:16.205you're
824
00:54:16.205 -->
00:54:18.285going to have to deal with both.
825
00:54:18.525 -->
00:54:19.085And
826
00:54:19.405 -->
00:54:20.125really,
827
00:54:20.845 -->
00:54:25.245every different type of data you work with might introduce different systems challenges.
828
00:54:25.610 -->
00:54:28.410Working with PDFs is different from working with textbooks,
829
00:54:28.810 -->
00:54:29.450and
830
00:54:29.930 -->
00:54:32.410there's going to be more heterogeneity there.
831
00:54:32.650 -->
00:54:38.490You're going to see much longer running and much larger scale workloads. Like with reinforcement learning, you're going to see
832
00:54:39.395 -->
00:54:40.915rollouts that are happening
833
00:54:41.234 -->
00:54:43.235at much larger timescales.
834
00:54:45.155 -->
00:54:47.715You're going to see models continue to grow in size.
835
00:54:48.115 -->
00:54:48.755So
836
00:54:49.555 -->
00:54:50.915all of these are
837
00:54:51.540 -->
00:54:53.780things that will continue to happen.
838
00:54:54.340 -->
00:54:54.900Are
839
00:54:55.140 -->
00:55:00.180there any other aspects of the Ray project, the ecosystem,
840
00:55:00.180 -->
00:55:05.300the use cases that it supports that we didn't discuss yet that you'd like to cover before we close out the show?
841
00:55:07.145 -->
00:55:08.345I would say that
842
00:55:09.225 -->
00:55:10.025the
843
00:55:10.744 -->
00:55:12.505shift to the world becoming
844
00:55:12.585 -->
00:55:14.825GPU centric really
845
00:55:14.825 -->
00:55:16.744plays to Ray's strengths.
846
00:55:18.990 -->
00:55:21.470A lot of our investment areas are in
847
00:55:23.230 -->
00:55:23.950really
848
00:55:24.110 -->
00:55:25.550natively supporting
849
00:55:26.590 -->
00:55:27.710high
850
00:55:28.350 -->
00:55:29.550performance communication
851
00:55:29.550 -->
00:55:30.670between GPUs,
852
00:55:30.945 -->
00:55:33.665natively supporting RDMA and Ray,
853
00:55:35.025 -->
00:55:35.825different
854
00:55:35.905 -->
00:55:38.385transport back ends. This is
855
00:55:38.705 -->
00:55:41.825growing in importance and is really essential for performance
856
00:55:42.385 -->
00:55:43.985across a variety of use cases.
857
00:55:45.210 -->
00:55:59.415Well, thank you for taking the time. For anybody who wants to get in touch with you, I'll have you add your preferred contact information to the show notes. And as the final question, I'm interested in getting your perspective on what you see as being the biggest gap in the tooling technology
858
00:55:59.415 -->
00:56:03.175or training that's available for data and AI management today.
859
00:56:04.055 -->
00:56:07.015Yeah, it is still not a solved problem.
860
00:56:08.295 -->
00:56:10.375Would say one of the biggest gaps is
861
00:56:11.140 -->
00:56:14.260the tooling around multi cloud workloads,
862
00:56:14.500 -->
00:56:15.140and
863
00:56:15.300 -->
00:56:17.300it's very common now for
864
00:56:17.460 -->
00:56:18.900companies to work with
865
00:56:20.500 -->
00:56:21.860one hyperscaler
866
00:56:22.740 -->
00:56:24.820and one or more Neo Clouds,
867
00:56:25.295 -->
00:56:27.615and it's quite complex to
868
00:56:27.775 -->
00:56:35.935find capacity in the right location at the right time to make your data available in a performance and cost efficient way in the right location
869
00:56:36.175 -->
00:56:41.290at the right time to make your workloads portable so they can run-in different locations
870
00:56:41.690 -->
00:56:42.410and
871
00:56:42.570 -->
00:56:43.530elastic,
872
00:56:43.530 -->
00:56:45.130fault tolerance so they can
873
00:56:45.450 -->
00:56:54.315take advantage of whatever compute resources are available. So at the same time, it's incredibly valuable when you can get it right and you can have a setup so that
874
00:56:54.635 -->
00:56:56.395your infra team can just
875
00:56:56.635 -->
00:56:57.675continuously
876
00:56:58.315 -->
00:57:01.515shop around for the cheapest compute, the cheapest GPUs,
877
00:57:01.755 -->
00:57:06.875the latest hardware accelerator, plug that in into a uniform interface that
878
00:57:07.115 -->
00:57:07.515your
879
00:57:08.130 -->
00:57:11.490researchers and AI teams can take advantage of. So that
880
00:57:11.809 -->
00:57:15.570is a lot of the stuff that we're excited about solving at any scale,
881
00:57:15.970 -->
00:57:18.369and one of the biggest challenges we see.
882
00:57:18.849 -->
00:57:27.275All right. Well, I appreciate you taking the time today to join me and share all the great work that you and your team are doing on the Ray project and helping to
883
00:57:27.595 -->
00:57:29.755address the myriad complexities
884
00:57:29.755 -->
00:57:31.995of all of the technological
885
00:57:31.995 -->
00:57:38.430utility that we get from these large language models and the whole ecosystem of ML and data processing.
886
00:57:39.150 -->
00:57:46.510It's definitely a very complex problem as we discussed, so I appreciate all of the energy that you're putting into it. And I hope you enjoy the rest of your day. Thank you so much.
887
00:57:54.365 -->
00:57:57.485Thank you for listening, and don't forget to check out our other shows.
888
00:57:57.645 -->
00:57:58.605Podcast.net
889
00:57:58.605 -->
00:58:07.800covers the Python language, its community, and the innovative ways it is being used. And the AI engineering podcast is your guide to the fast moving world of building AI systems.
890
00:58:08.360 -->
00:58:18.385Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@dataengineeringpodcast.com
891
00:58:18.385 -->
00:58:24.545with your story. Just to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.