درباره این اپیزود
This is a cross-over episode from our new show The Machine Learning Podcast, the show about going from idea to production with machine learning.
SummaryThe majority of machine learning projects that you read about or work on are built around batch processes. The model is trained, and then validated, and then deployed, with each step being a discrete and isolated task. Unfortunately, the real world is rarely static, leading to concept drift and model failures. River is a framework for building streaming machine learning projects that can constantly adapt to new information. In this episode Max Halford explains how the project works, why you might (or might not) want to consider streaming ML, and how to get started building with River.
Announcements- Hello and welcome to the Machine Learning Podcast, the podcast about machine learning and how to bring it from idea to delivery.
- Building good ML models is hard, but testing them properly is even harder. At Deepchecks, they built an open-source testing framework that follows best practices, ensuring that your models behave as expected. Get started quickly using their built-in library of checks for testing and validating your model’s behavior and performance, and extend it to meet your specific needs as your model evolves. Accelerate your machine learning projects by building trust in your models and automating the testing that you used to do manually. Go to themachinelearningpodcast.com/deepchecks today to get started!
- Your host is Tobias Macey and today I’m interviewing Max Halford about River, a Python toolkit for streaming and online machine learning
- Introduction
- How did you get involved in machine learning?
- Can you describe what River is and the story behind it?
- What is "online" machine learning?
- What are the practical differences with batch ML?
- Why is batch learning so predominant?
- What are the cases where someone would want/need to use online or streaming ML?
- The prevailing pattern for batch ML model lifecycles is to train, deploy, monitor, repeat. What does the ongoing maintenance for a streaming ML model look like?
- Concept drift is typically due to a discrepancy between the data used to train a model and the actual data being observed. How does the use of online learning affect the incidence of drift?
- Can you describe how the River framework is implemented?
- How have the design and goals of the project changed since you started working on it?
- How do the internal representations of the model differ from batch learning to allow for incremental updates to the model state?
- In the documentation you note the use of Python dictionaries for state management and the flexibility offered by that choice. What are the benefits and potential pitfalls of that decision?
- Can you describe the process of using River to design, implement, and validate a streaming ML model?
- What are the operational requirements for deploying and serving the model once it has been developed?
- What are some of the challenges that users of River might run into if they are coming from a batch learning background?
- What are the most interesting, innovative, or unexpected ways that you have seen River used?
- What are the most interesting, unexpected, or challenging lessons that you have learned while working on River?
- When is River the wrong choice?
- What do you have planned for the future of River?
- @halford_max on Twitter
- MaxHalford on GitHub
- From your perspective, what is the biggest barrier to adoption of machine learning today?
- Thank you for listening! Don’t forget to check out our other shows. The Data Engineering Podcast covers the latest on modern data management. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used.
- Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
- If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@themachinelearningpodcast.com) with your story.
- To help other people find the show please leave a review on iTunes and tell your friends and co-workers
- River
- scikit-multiflow
- Federated Machine Learning
- Hogwild! Google Paper
- Chip Huyen concept drift blog post
- Dan Crenshaw Berkeley Clipper MLOps
- Robustness Principle
- NY Taxi Dataset
- RiverTorch
- River Public Roadmap
- Beaver tool for deploying online models
- Prodigy ML human in the loop labeling
The intro and outro music is from Hitman’s Lovesong feat. Paola Graziano by The Freak Fandango Orchestra/CC BY-SA 3.0
Sponsored By:
- Linode: Do you want to try out some of the tools and applications that you heard about on Podcast.\_\_init\_\_? Do you have a side project that you want to share with the world? With Linode's managed Kubernetes platform it's now even easier to get started with the latest in cloud technologies. With the combined power of the leading container orchestrator and the speed and reliability of Linode's object storage, node balancers, block storage, and dedicated CPU or GPU instances, you've got everything you need to scale up. Go to [pythonpodcast.com/linode](https://www.pythonpodcast.com/linode) today and get a $100 credit to launch a new cluster, run a server, upload some data, or... And don't forget to thank them for being a long time supporter of Podcast.\_\_init\_\_!
یادداشت ها را نشان دهید 🔗
رونوشت 🔗
NOTE
Transcription provided by Podhome.fm
Created: 7/6/2024 5:24:58 PM
Duration: 4582.926
Channels: 1
1
00:00:13.505 --> 00:00:24.880
Hello, and welcome to podcast dot in it, the podcast about Python and the people who make it great. When you're ready to launch your next app or want to try a project you hear about on the show, you'll need somewhere to deploy it. So check out our friends over at Linode.
2
00:00:25.475 --> 00:00:32.775
With their managed Kubernetes platform, it's easy to get started with the next generation of deployment and scaling powered by the battle tested Linode platform,
3
00:00:33.170 --> 00:00:39.190
including simple pricing, node balancers, 40 gigabit networking, and dedicated CPU and GPU instances.
4
00:00:39.969 --> 00:00:49.580
And now you can launch a managed MySQL, Postgres, or Mongo database cluster in minutes to keep your critical data safe with automated backups and failover. Go to python podcast dotcom/linode
5
00:00:50.440 --> 00:01:03.595
today to get a $100 credit to try out their new database service, and don't forget to thank them for their continued support of this show. The biggest challenge with modern data systems is understanding what data you have, where it is located, and who is using it.
6
00:01:03.975 --> 00:01:04.475
SelectStar's
7
00:01:04.870 --> 00:01:15.034
Data Discovery platform solves that out of the box with a fully automated catalog that includes lineage from where the data originated all the way to which dashboards rely on it and who is viewing them every day.
8
00:01:15.415 --> 00:01:23.034
Just connect it to your DBT, Snowflake, Tableau, Looker, or whatever you're using, and select star will set everything up in just a few hours.
9
00:01:23.470 --> 00:01:24.770
Go to python podcast.com/selectstar
10
00:01:26.430 --> 00:01:31.090
today to double the length of your free trial and get a swag package when you convert to a paid plan.
11
00:01:31.645 --> 00:01:36.945
Your host as usual is Tobias Macy, and this month, I'm running a series about Python's use in machine learning.
12
00:01:37.325 --> 00:01:47.670
If you enjoy this episode, you can explore further on my new show, The Machine Learning Podcast, that helps you go from idea to production with machine learning. To find out more, you can go to the machine learning podcast.com.
13
00:01:48.695 --> 00:01:58.075
Your host is Tobias Macy, and today, I'm interviewing Max Halford about River, a Python toolkit for streaming and online machine learning. So, Max, can you start by introducing yourself?
14
00:01:58.890 --> 00:02:03.550
Oh, hey there. So I'm Max. Yes. I consider myself as a data scientist.
15
00:02:03.930 --> 00:02:06.190
My day job is doing data science. I
16
00:02:06.635 --> 00:02:08.655
actually measure the carbon footprint of
17
00:02:09.115 --> 00:02:10.095
clothing items.
18
00:02:10.715 --> 00:02:16.540
But I have a wide interest in, you know, technical topics, so be it software engineering or data
19
00:02:16.920 --> 00:02:22.940
engineering, do a lot of open source. My academic background is leaning towards finance and computer science and statistics.
20
00:02:23.545 --> 00:02:26.845
I actually did a PhD in applied machine learning,
21
00:02:27.305 --> 00:02:29.165
which I finished a couple of years ago.
22
00:02:29.465 --> 00:02:31.325
So, yeah, an all around node, basically.
23
00:02:32.840 --> 00:02:35.900
And do you remember how Prisca started working in machine learning?
24
00:02:36.439 --> 00:02:42.614
Kind of, Jess. I was a late bloomer. I got started when maybe when I was 21, 22 when I was at university.
25
00:02:43.474 --> 00:02:49.815
I basically had no idea what machining was, but I started this curriculum that involved that was around statistics.
26
00:02:50.500 --> 00:02:51.880
And we had a course,
27
00:02:52.260 --> 00:02:58.599
which was maybe 2 or 3 hours a week about machine learning, and it did kind of blow my mind. It was
28
00:02:59.435 --> 00:03:01.375
around the time when, well,
29
00:03:02.075 --> 00:03:05.135
machine learning and particularly deep learning was starting to explode.
30
00:03:05.675 --> 00:03:06.175
So
31
00:03:06.780 --> 00:03:10.879
I kinda stopped at university. So I was lucky enough to get a theoretical training.
32
00:03:11.660 --> 00:03:22.045
And in terms of the river project, can you describe a bit more about what it is that you've built and some of the story behind how it came to be and why you decided that you wanted to build it in the first place?
33
00:03:22.560 --> 00:03:23.780
When I was at university,
34
00:03:24.320 --> 00:03:25.940
I received a kind of normal
35
00:03:27.280 --> 00:03:30.420
introduction to a regular introduction to machine learning.
36
00:03:30.785 --> 00:03:32.245
And then I did some internships.
37
00:03:32.705 --> 00:03:35.125
I started PhD after my internships.
38
00:03:35.905 --> 00:03:40.085
And I also did a lot of camera competitions on the side. So I was kind of hooked into
39
00:03:40.770 --> 00:03:41.590
machine learning,
40
00:03:42.370 --> 00:03:46.310
and it always felt to me that something was off. Because
41
00:03:47.250 --> 00:03:52.165
when we were learning machine learning, everything made sense. But then when you get to do in practice,
42
00:03:52.625 --> 00:03:54.725
you often find that, well,
43
00:03:55.345 --> 00:04:08.475
it's not playgrounds. Right? The playground scenarios that they describe at university when you learn machine learning just do not apply in the real world. The real world as well. You have data that's coming in, like a flow of data or every day there's new data or
44
00:04:09.015 --> 00:04:15.730
yeah. There's like an interactive aspect to the world around us, the way the data is flowing. It's not like a CSV file.
45
00:04:16.270 --> 00:04:19.490
Yeah. It just felt like fitting a square peg in a round hole.
46
00:04:19.790 --> 00:04:21.010
So I was always curious
47
00:04:21.390 --> 00:04:23.010
in the back of my mind about
48
00:04:23.435 --> 00:04:23.935
how
49
00:04:24.555 --> 00:04:32.815
you could do online machine learning. Well, I didn't know it was called online machine learning because when I was a kid, I remember growing up and thinking that AI was this kind of
50
00:04:33.419 --> 00:04:37.360
intelligent machine that would keep learning as it went on and as it experienced
51
00:04:37.660 --> 00:04:45.324
the world around it. Anyway, when I started my PhD, I was lucky enough to have a lot of time to read. So I read a lot of papers and blog posts and whatnot.
52
00:04:45.625 --> 00:04:47.805
And I can't remember the exact day
53
00:04:48.185 --> 00:04:52.365
or week that I stumbled upon it, but I just started learning about online machine learning.
54
00:04:52.770 --> 00:04:54.390
Maybe some blog posts or something.
55
00:04:54.930 --> 00:04:56.390
And then it was like a big
56
00:04:56.930 --> 00:05:01.430
explosion in my head, and I was like, wow. This is crazy. Right? This this actually exists.
57
00:05:01.865 --> 00:05:05.165
And I was so curious as to why it wasn't more popular.
58
00:05:05.945 --> 00:05:06.445
And,
59
00:05:06.905 --> 00:05:10.445
you know, at the time, I did a lot of open source as a way to learn,
60
00:05:10.770 --> 00:05:13.990
And so it just felt natural to me to start implementing
61
00:05:14.690 --> 00:05:16.630
algorithms that I've read in papers on
62
00:05:17.010 --> 00:05:17.510
everywhere.
63
00:05:17.865 --> 00:05:20.605
I've just started writing code to learn, basically, just to
64
00:05:21.625 --> 00:05:24.695
confirm what I'd learned and and whatnot. That's just the way I learned.
65
00:05:25.039 --> 00:05:29.300
And it kind of evolved into what is, which is a well, an open source
66
00:05:29.840 --> 00:05:35.564
package that people use. Now if I may expand a little bit, VIVO is actually the merger between 2 projects.
67
00:05:36.185 --> 00:05:38.525
So the first project is called Psyche Multifrow.
68
00:05:39.145 --> 00:05:41.245
It was a package that was developed
69
00:05:41.710 --> 00:05:44.770
before I even got into machine learning. It has roots in
70
00:05:45.150 --> 00:05:46.530
academia in New Zealand,
71
00:05:47.070 --> 00:05:49.330
comes from an old package called Noah in Java.
72
00:05:49.925 --> 00:06:01.690
Anyway, I wasn't not aware of that. On my end, I started to get a package called cream at the time. So in French, creme means cream, and it plays something with incremental, which is another way to say online.
73
00:06:02.310 --> 00:06:03.450
So, yeah, I developed
74
00:06:03.830 --> 00:06:05.690
cream on by myself.
75
00:06:06.525 --> 00:06:12.225
And at some point, it just made sense to reach out to the guys from Psyche Multiflo and to propose a merger.
76
00:06:12.764 --> 00:06:13.264
So
77
00:06:13.720 --> 00:06:16.780
it took us quite a while, but after 9 months
78
00:06:17.480 --> 00:06:20.220
of negotiation and, you know, figuring out the details,
79
00:06:21.095 --> 00:06:23.675
we merged, and we called the new package, Vevor.
80
00:06:24.055 --> 00:06:47.560
You mentioned that it's built around this idea of online machine learning. And in the documentation, you also refer to it as streaming machine learning. I'm curious if you can just talk through what that really means in the context of building a machine learning system and some of the practical differences between that and the typical batch oriented workflow that most folks who are working in ML are go going to be familiar with?
81
00:06:47.940 --> 00:06:52.199
1st, just a recap on machine learning. The whole point of machine learning is to
82
00:06:52.740 --> 00:06:54.039
teach a model
83
00:06:54.635 --> 00:07:05.120
to learn from data and to take decisions. So, you know, monkey see, monkey do. And the typical way you do that is that you fit a model to a bunch of data, and
84
00:07:05.740 --> 00:07:06.639
that's it really.
85
00:07:07.180 --> 00:07:07.680
But
86
00:07:08.460 --> 00:07:11.955
online machining is the equivalent of that, but for streaming data.
87
00:07:12.995 --> 00:07:13.735
So you
88
00:07:14.514 --> 00:07:19.414
stop thinking about data as a file or a table in a database, but you think of it as
89
00:07:19.750 --> 00:07:21.290
a flow of data stream.
90
00:07:21.990 --> 00:07:27.050
So online machine learning, you could call it incremental machine learning, you could call it streaming machine learning.
91
00:07:27.575 --> 00:07:38.190
I mean, I more often see online machine learning being used, although if you Google that, you kind of find these online courses for that online machine learning. So that's, like, kind of cool online.
92
00:07:38.570 --> 00:07:40.030
But anyway, yeah, it's just this
93
00:07:40.570 --> 00:07:44.750
way to say, can I do machine learning but with streaming data? And so
94
00:07:45.210 --> 00:07:49.575
the rule is that an online model is 1 that can learn
95
00:07:50.275 --> 00:07:51.975
1 sample at a time. So
96
00:07:52.355 --> 00:08:00.340
usually, you show a model a whole dataset, and they can work with that dataset. It can calculate the average of the data. It can do a whole bunch of stuff.
97
00:08:00.720 --> 00:08:04.020
But the restriction here with online machining is that
98
00:08:04.625 --> 00:08:23.479
the the model cannot see the whole data. It can't hold it in memory. It can only see 1 sample at a time, and it has to work like that. So it's a restriction. Right? So it makes it harder for the models to learn, but it also has many, many implications. If you have a model that can learn that way, well, you can have a model that just can keep learning as you go on.
99
00:08:23.824 --> 00:08:24.884
Because a regular
100
00:08:25.505 --> 00:08:27.205
machine learning model, once
101
00:08:27.585 --> 00:08:29.604
it's been fitted to a dataset,
102
00:08:30.145 --> 00:08:32.165
you have to retrain it from scratch
103
00:08:32.570 --> 00:08:33.470
if you want to incorporate
104
00:08:34.010 --> 00:08:37.870
new samples into your model. That can be a source of frustration,
105
00:08:38.330 --> 00:08:51.335
and that's why I was calling the square peg in the round hole before. So say you have a model, an online model that is just as performant as a a batch model. Well, you know, if you if you just regardless of performance, accuracy,
106
00:08:52.180 --> 00:08:53.480
that has many implications,
107
00:08:53.940 --> 00:08:58.600
and it actually makes things easier. Because if you have a model that can keep learning,
108
00:08:59.285 --> 00:09:01.545
well, you don't have to, for instance,
109
00:09:01.925 --> 00:09:09.110
schedule retraining of your model. You can just every time you have a new sample arrives, you can just tell your model to learn from
110
00:09:09.410 --> 00:09:10.550
that, and then you don't.
111
00:09:10.850 --> 00:09:15.270
And so that ensures that your model is always as up to date as possible and that has obviously
112
00:09:15.970 --> 00:09:16.950
many, many benefits.
113
00:09:17.275 --> 00:09:18.095
If you think about
114
00:09:18.475 --> 00:09:22.415
people working on the stock market, so trying to forecast
115
00:09:22.715 --> 00:09:25.855
the evolution of a particular stock,
116
00:09:26.390 --> 00:09:29.450
They've actually been doing online machine learning since the eighties
117
00:09:29.830 --> 00:09:30.730
because, obviously,
118
00:09:31.270 --> 00:09:34.650
they have a lot to lose by making all this public. It just never
119
00:09:34.995 --> 00:09:37.975
got into a big thing, and it always stayed in in stock market companies.
120
00:09:38.755 --> 00:09:39.815
So the practical
121
00:09:40.195 --> 00:09:40.695
differences
122
00:09:41.154 --> 00:09:41.815
are that
123
00:09:42.480 --> 00:09:46.340
you are working over stream of data. You're not working over static data set.
124
00:09:46.720 --> 00:09:54.425
This stream of data has an order, meaning that the fact that 1 sample arrives before the other, well, that has a lot of meaning, and that's actually reflecting
125
00:09:54.964 --> 00:10:02.079
what's happening in the real world. In the real world, well, you have data that's surviving in a certain order. Well, if you train your model
126
00:10:02.620 --> 00:10:03.120
offline
127
00:10:03.660 --> 00:10:05.360
on that data, you want to, you know,
128
00:10:06.459 --> 00:10:10.015
process it in the same order. And so that ensures that you are actually
129
00:10:10.635 --> 00:10:11.135
reproducing
130
00:10:12.155 --> 00:10:16.895
the conditions that happen in the real world. Now another practical consideration is that
131
00:10:17.500 --> 00:10:19.360
online learning is much less
132
00:10:20.060 --> 00:10:21.440
popular or predominant
133
00:10:21.819 --> 00:10:38.899
than batch learning, and so a lot less research and software work has been put into online learning. So if you are a newcomer to the field, well, there's just not a lot of resources to learn from. Actually, you could just spend a day on Google, and you probably find all the resources you there are because there's just
134
00:10:39.200 --> 00:10:41.300
not so many of them. There's probably, like,
135
00:10:41.839 --> 00:10:49.075
by memory, just 10 links on Google that you can learn from about online learning. So it's a bit of a niche topic.
136
00:10:49.535 --> 00:10:57.700
In terms of the fact that batch is such a predominant mode of building these ML systems despite the fact that it's not
137
00:10:58.080 --> 00:11:00.260
very reflective of the way that
138
00:11:00.640 --> 00:11:04.900
the real world actually operates. Why do you think that's the case, that
139
00:11:05.425 --> 00:11:11.365
streaming or online machine learning is still such a niche topic and hasn't been more broadly adopted?
140
00:11:11.985 --> 00:11:15.160
Sometimes it feels like I'm trying to teach a new religion,
141
00:11:15.620 --> 00:11:18.360
which feels a bit weird because there's not a lot of us doing it.
142
00:11:18.740 --> 00:11:19.940
So I'm also very
143
00:11:20.580 --> 00:11:27.285
I never try to force people into this. There's obviously many good reasons why batch learning is still done.
144
00:11:27.745 --> 00:11:31.445
And now from a historical point of view, I think it's interesting because
145
00:11:31.960 --> 00:11:41.580
we always used to use statistical models to explain data and not necessarily to predict. So you just have a data set, and you just like to understand,
146
00:11:42.105 --> 00:11:47.404
you know, what variables are affecting a particular outcome. So for instance, if you take linear aggression,
147
00:11:47.945 --> 00:11:49.245
historically, it's been used
148
00:11:49.785 --> 00:11:51.084
to explain the
149
00:11:51.700 --> 00:11:52.200
impact
150
00:11:52.820 --> 00:11:57.000
of global warming on the rise of sea level. But not necessarily to predict
151
00:11:57.700 --> 00:11:58.200
if,
152
00:11:58.754 --> 00:12:07.175
you know, the temperature of the globe was higher, what would be the impact on the sea level. But then someone said, let's use machine learning to predict
153
00:12:07.589 --> 00:12:09.290
outcomes in a business context,
154
00:12:09.910 --> 00:12:13.050
and that's why we have this big event of machine learning.
155
00:12:13.750 --> 00:12:17.050
And we've kind of been using the tools that have been lying around.
156
00:12:17.415 --> 00:12:19.435
So we've been using all these tools
157
00:12:19.975 --> 00:12:21.274
that we used for
158
00:12:22.134 --> 00:12:26.074
to a dataset and explain it, but now we've been using them for predicting. So
159
00:12:26.839 --> 00:12:27.740
these models
160
00:12:28.680 --> 00:12:29.500
are static.
161
00:12:30.040 --> 00:12:34.459
Like, the people who when we started doing linear regression, we never really worried about
162
00:12:34.795 --> 00:12:44.415
streaming data because datasets were small, datasets were static. Well, the Internet didn't even exist, so there was no real notion of IoT or sensors or
163
00:12:44.940 --> 00:12:45.760
streaming data.
164
00:12:46.380 --> 00:12:46.880
So
165
00:12:47.420 --> 00:12:48.640
the fact is that
166
00:12:49.260 --> 00:12:51.760
we've never needed online models.
167
00:12:52.380 --> 00:12:53.120
And so
168
00:12:53.905 --> 00:12:59.365
as a field, you look at the academia and the industry. We're very used to batch learning,
169
00:12:59.905 --> 00:13:03.045
and we're very comfortable with it. There's a lot of good software,
170
00:13:03.650 --> 00:13:06.310
and this is what people are being taught at university. So
171
00:13:07.250 --> 00:13:09.270
I'm not saying that online learning is
172
00:13:09.970 --> 00:13:14.065
necessarily better than batch learning, but I do think that the reasons why
173
00:13:14.925 --> 00:13:20.305
batch learning is so predominant in comparison is because we are too used to it, basically.
174
00:13:20.920 --> 00:13:23.660
And I do think that and I see it every week.
175
00:13:24.120 --> 00:13:25.900
People who are trying to rethink
176
00:13:26.520 --> 00:13:27.020
their
177
00:13:27.640 --> 00:13:29.900
job or their projects and say and say,
178
00:13:30.455 --> 00:13:34.075
maybe I could be doing online learning. It actually makes more sense. So
179
00:13:34.455 --> 00:13:36.555
I think it's a question of habits, really.
180
00:13:36.935 --> 00:13:38.235
For people who are
181
00:13:38.615 --> 00:13:46.980
assessing which approach to to take in their ML projects, what are some of the use cases where online or streaming ML is the
182
00:13:47.545 --> 00:14:04.920
more logical approach or what the decision factors look like for somebody deciding, do I go with a batch oriented process where I'm going to have this large set of tooling available to me, or do I want to use online or streaming ML because the benefits outweigh the potential costs of
183
00:14:05.315 --> 00:14:07.495
plugging into this ecosystem of tooling?
184
00:14:08.035 --> 00:14:13.175
So I'll be honest. I think it always makes sense to start with a batch model.
185
00:14:13.840 --> 00:14:21.940
Why? Because, you know, if you're pragmatic and you actually have deadlines to meet and you just wanna be productive, there's so many good solutions to
186
00:14:22.515 --> 00:14:26.615
train a batch model and deploy it. So, you know, I would just go with that
187
00:14:27.075 --> 00:14:29.255
to start with. And then, yeah, there's
188
00:14:29.750 --> 00:14:37.050
the question of, could I be doing this online? So I think there's 2 cases. There's cases where you need it. And so I have a great example.
189
00:14:37.435 --> 00:14:39.295
So Netflix, when they do recommendations,
190
00:14:39.915 --> 00:14:49.710
you know, you arrive on the website and Netflix recommends movies to you. Netflix actually retrains a model every night or every week, but they have many models anyway. But
191
00:14:50.330 --> 00:14:55.950
they are learning from your behavior to kind of retrain their models to update their recommendations.
192
00:14:56.445 --> 00:14:56.945
Right?
193
00:14:57.725 --> 00:15:05.505
There's a team at Netflix that are working on learning instantly. So if you are scrolling on the Netflix website and you
194
00:15:06.380 --> 00:15:07.920
see your recommendation for Netflix,
195
00:15:08.700 --> 00:15:20.655
the fact that you did not click on that recommendation is a signal that you do not want to watch that movie maybe or that the recommendations will be changed. So if you're able to have a model, for instance, maybe in your browser
196
00:15:21.195 --> 00:15:23.535
that would learn in real time from
197
00:15:24.210 --> 00:15:32.150
your browsing activity and that could update and learn on the fly, that'd be really powerful. And the only way to do that is to have
198
00:15:32.635 --> 00:15:33.935
a model per user
199
00:15:34.395 --> 00:15:35.615
that is learning online.
200
00:15:36.154 --> 00:15:36.815
And so
201
00:15:37.515 --> 00:15:40.795
you cannot just use batch models for that. Yeah. You can't just
202
00:15:41.399 --> 00:15:51.565
every time a user scrolls or ignores a movie, you can't just take all the history of data and fit the model. It would be much too heavy, and it just not practical. So sometimes the only
203
00:15:52.185 --> 00:16:00.350
way forward is to do online learning. But, again, this is quite niche. Like, Netflix recommendations, I mean, obviously, are working reasonably well, I believe, because
204
00:16:00.990 --> 00:16:03.089
they're, you know, just from their market value.
205
00:16:03.470 --> 00:16:07.329
But if you are pushing the envelope, then sometimes you need online learning.
206
00:16:07.845 --> 00:16:13.865
Now another case is when you do not necessarily need it, but you want it because it makes things easier. So
207
00:16:14.404 --> 00:16:16.105
a good example I have is
208
00:16:16.490 --> 00:16:18.830
imagine you're working on the app that categorizes
209
00:16:19.290 --> 00:16:20.510
tickets. So
210
00:16:20.890 --> 00:16:31.985
for instance, on the help support software. So, you know, you go on the website and you're sending a form or you send a message or an email, and you're asking you have some problem maybe with reimbursement on your ultimately, you bought on Amazon.
211
00:16:32.610 --> 00:16:38.150
And then, you know, there's a customer service behind that, human beings that are actually answering those questions.
212
00:16:38.770 --> 00:16:39.270
And
213
00:16:39.805 --> 00:16:47.105
it's really important to be able to categorize each request and put into a bucket so that it gets assigned to the right person.
214
00:16:47.740 --> 00:16:54.720
And maybe the public manager has decided that we need a new category. And so there's this new category,
215
00:16:55.545 --> 00:16:56.765
and your
216
00:16:57.145 --> 00:16:59.325
model is classifying tickets
217
00:16:59.785 --> 00:17:03.565
into 1 of several categories. If you introduce a new category,
218
00:17:04.140 --> 00:17:07.600
it means that you have to retrain the model to incorporate it.
219
00:17:08.060 --> 00:17:12.400
I was in discussion with a company, and they were only able to or their budget
220
00:17:12.774 --> 00:17:19.514
was that they were only able to retrain the model every 3 months. So if you introduce a new ticket, a new category into your system,
221
00:17:19.980 --> 00:17:22.640
the model would only pick it up and predict it
222
00:17:23.500 --> 00:17:28.559
after 3 months. So that sounded kind of insane and, you know, wasn't collected at all.
223
00:17:29.725 --> 00:17:36.305
And I was not aware of the exact details, but it it just seemed too expensive for them to retrain their model from scratch. So
224
00:17:36.780 --> 00:17:41.920
if they were using an online model, well, potentially, that model could just learn the new tickets
225
00:17:42.380 --> 00:17:45.680
on the fly. And, you know, if you just introduced it and people started,
226
00:17:46.225 --> 00:17:49.445
you know, you learn in this feedback loop where you introduce a new category,
227
00:17:49.985 --> 00:17:54.965
people send the email, maybe a human assigns that ticket to a category, so that becomes
228
00:17:55.580 --> 00:17:58.320
a signal for the model. The model picks that up, learns,
229
00:17:58.860 --> 00:18:11.804
and, yeah, it's gonna incorporate that category into its next predictions. So that's a scenario where you don't necessarily need online learning, but actually online learning just makes more sense and makes your your system easier to maintain, to
230
00:18:12.390 --> 00:18:13.370
work with, basically.
231
00:18:14.550 --> 00:18:17.610
There are a number of interesting things to dig into there.
232
00:18:18.310 --> 00:18:24.875
1 of the things that you mentioned is the idea of having a model per user in the Netflix example.
233
00:18:25.655 --> 00:18:29.435
And I'm wondering if you could maybe talk through some of the
234
00:18:29.990 --> 00:18:37.690
conceptual elements of saying, okay, I've got this baseline structure for it. This is how I'm going to build the model. Here is the initial
235
00:18:38.524 --> 00:18:46.945
training set. I've got this model deployed. And now every time a specific user interacts with this model, it is going to learn their specific behaviors
236
00:18:47.325 --> 00:18:47.985
and be
237
00:18:48.480 --> 00:18:50.420
tuned to the data that they are generating.
238
00:18:50.960 --> 00:18:54.100
Would you then take that information and feed that back into
239
00:18:54.640 --> 00:19:00.715
the baseline model that gets loaded at the time that the browser interacts with the website, and just some of the ways to think about
240
00:19:01.175 --> 00:19:01.675
the
241
00:19:02.135 --> 00:19:06.110
potential approaches for how to say, okay. I've got a model,
242
00:19:06.410 --> 00:19:13.550
but it's going to be customized per user and just managing the kind of fan out, fan in topologies that might result from
243
00:19:13.885 --> 00:19:14.385
those
244
00:19:14.685 --> 00:19:15.905
event based interactions
245
00:19:16.685 --> 00:19:19.985
spread out across n number of users or entities.
246
00:19:20.605 --> 00:19:33.715
Now it sounds insane when you say it because to have 1 model per user and, you know, have it deployed on the user's browser or mobile or phone or covers what on Apple Watch, It does kinda sound insane, but it is interesting, I guess.
247
00:19:34.095 --> 00:19:35.615
I don't think there are so many
248
00:19:36.255 --> 00:19:38.915
I mean, I'm not aware of a lot of companies that
249
00:19:39.230 --> 00:19:42.289
would have the justification to actually do this, and
250
00:19:42.669 --> 00:19:46.929
I've never had the occasion to actually work into a setting where I would I would do this.
251
00:19:47.235 --> 00:19:51.255
But I had 1 good example where I was kinda doing some pro bono consulting.
252
00:19:51.794 --> 00:19:53.735
It was this car company where
253
00:19:54.514 --> 00:19:56.455
the onboard navigation system,
254
00:19:56.950 --> 00:20:01.130
they wanted to build a model where they could guess where you're going to. So, basically,
255
00:20:01.590 --> 00:20:06.465
depending on where you left, if you left home in the morning, you're probably going to work. And
256
00:20:07.005 --> 00:20:09.025
they would then use this to, you know,
257
00:20:09.645 --> 00:20:22.200
just give you send you news about your itinerary or things like that. They really needed a model that would just be able to learn online, and they made the bold decision to say, okay. We're going to embed the model
258
00:20:22.555 --> 00:20:25.855
into the car. It's not going to be, like, a central
259
00:20:26.555 --> 00:20:34.640
model that's, you know, hosted on some big server and that the car interaction the actually, the intelligence is actually happening in the car.
260
00:20:34.940 --> 00:20:40.460
And so when you think about that, it's really interesting because it creates a decentralized system. There's not like a single
261
00:20:41.345 --> 00:20:45.605
it actually creates a system where you don't even need the Internet for the model to work. So
262
00:20:45.985 --> 00:20:47.445
there's so many operational
263
00:20:47.745 --> 00:20:54.900
requirements for that. Actually, now that I think of it and I'm talking about cars, I realized that Tesla that's actually what Tesla is doing. Like, they're they're computing
264
00:20:55.280 --> 00:20:56.180
and making decisions
265
00:20:56.800 --> 00:21:00.894
inside the car, you know, doing a bunch of stuff, and they're also communicating with
266
00:21:01.274 --> 00:21:02.735
mother servers. But
267
00:21:03.195 --> 00:21:08.000
the actual computer, they actually have GPUs in the car doing compute with their deep learning models and and what
268
00:21:08.640 --> 00:21:17.140
not. So it's definitely possible to do this. Right? But clearly not something that our company would go would go through or would have the need to to do.
269
00:21:17.524 --> 00:21:24.184
It's interesting also to think about how something like that would play into the federated
270
00:21:24.565 --> 00:21:30.460
learning approach where you have federated models where there is that core model that's being built and maintained
271
00:21:31.080 --> 00:21:36.605
at the core. And then as users interact at the edge, whether it's on their device or in their browser,
272
00:21:36.905 --> 00:21:41.085
it loads a federated learning component that has that streaming ML
273
00:21:41.490 --> 00:21:41.990
capability.
274
00:21:42.370 --> 00:21:43.430
So the model
275
00:21:43.890 --> 00:21:49.990
evolves as the user is interacting with it on their device, and then that information is then sent back to
276
00:21:50.375 --> 00:21:51.195
the centralized
277
00:21:51.495 --> 00:21:57.674
system to be able to feed back into the core model so that you can have these kind of parallel streams of
278
00:21:58.149 --> 00:22:04.649
the model that the user is interacting with as being customized to their behavior at the time that they're interacting with it, but it does still
279
00:22:05.195 --> 00:22:09.375
get propagated back into the larger system so that the
280
00:22:09.755 --> 00:22:10.575
new intelligence
281
00:22:10.955 --> 00:22:12.335
is able to
282
00:22:12.790 --> 00:22:16.810
generate an updated experience for everybody who then goes and interacts with it in the future.
283
00:22:17.350 --> 00:22:19.690
Yeah. That's really, really interesting. So
284
00:22:20.034 --> 00:22:20.855
I think first off,
285
00:22:21.235 --> 00:22:22.695
the the fact is that I
286
00:22:23.315 --> 00:22:29.015
I'm actually still young, and there's so many things that I don't know. And I don't have, like, the technical savvy to be able to
287
00:22:29.400 --> 00:22:34.780
suggest ways forward. But this is obviously things that, you know, I think about.
288
00:22:35.080 --> 00:22:40.635
So these things like Hub Wild, which is a a project from Google, they have a a paper where they discuss these things.
289
00:22:41.095 --> 00:22:51.530
I think that's a really simple thing that if you wanted to do this, if you're the the listener wanting to do something like this, I think there's a simple pattern, which is to maybe once a month have a model that is retrained
290
00:22:52.070 --> 00:22:55.210
and that you just train in batch, and that model is going to be
291
00:22:55.995 --> 00:23:00.255
is gonna be like a hydro. Like, you're gonna copy it, and you're gonna send it to each user.
292
00:23:00.635 --> 00:23:01.295
And then
293
00:23:01.755 --> 00:23:11.970
each copy for each user is going to be able to learn in this whole environment. And for instance, a good idea would maybe to with your model to increase, like, the learning rates so that
294
00:23:12.575 --> 00:23:13.554
every sample
295
00:23:14.015 --> 00:23:16.835
that the user gives you matters a lot. So
296
00:23:17.135 --> 00:23:21.500
for instance, if we take the Netflix example, you would have, you know, your run of the mill
297
00:23:21.960 --> 00:23:24.700
recommendation system model that you would just train in batch,
298
00:23:25.320 --> 00:23:32.625
you know, and you just use all the tools that we in community use. But then you would embed that into each person's browser,
299
00:23:33.005 --> 00:23:35.985
and maybe you do this once a month. And then that model
300
00:23:36.309 --> 00:23:41.210
for each user would be a coffee, a clone, or like a just a separate model now.
301
00:23:41.830 --> 00:23:42.330
And,
302
00:23:42.790 --> 00:23:56.000
you know, it would keep learning in an online manner. So maybe your model was trained in batch initially, but now for each user, it's, yeah, it's actually it's being trained online. So for instance, you can do this with factorization machines that can be trained in batch, but also online.
303
00:23:56.779 --> 00:23:59.600
And, yeah, you would use a high learning rate
304
00:24:00.205 --> 00:24:05.745
so that every sample matters a lot basically. And so you, the user, are tuning your model.
305
00:24:06.285 --> 00:24:10.809
And so I don't know how YouTube does it for instance, but I do imagine they have some sort of
306
00:24:11.429 --> 00:24:19.715
core model. They're just learning how to make good recommendations. But obviously, YouTube, there are some rules that make it so that, you know,
307
00:24:20.095 --> 00:24:34.755
recommendations are tailored to each user. And I don't know if this is done online, and I don't know if it's actually machine learning. It's probably just rules or scores. But, yeah, I think it's a really fun idea to play around with, and I do think that online learning enables this even more.
308
00:24:35.135 --> 00:24:36.115
As far as
309
00:24:37.054 --> 00:24:39.955
the operational model for people who are using
310
00:24:40.340 --> 00:24:42.040
online and streaming machine learning.
311
00:24:42.740 --> 00:24:44.760
If they're coming from a batch background,
312
00:24:45.060 --> 00:24:47.800
they're going to be used to dealing with the
313
00:24:48.294 --> 00:24:58.430
train, test, deploy cycle where I have my dataset, I build this model, I validate the model against the test dataset that I've held out from the training data.
314
00:24:59.050 --> 00:25:11.455
Everything looks good based on the, you know, area under curve or whatever metrics I'm using to validate that model. Now I'm going to put it into my production environment where maybe it's being served as a Flask app or a fast API app.
315
00:25:11.995 --> 00:25:22.460
And then I'm going to monitor it for concept drift, and, eventually, I say, okay. This is no longer performing up to the specifications, so now I need to go back and retrain the model based on my updated datasets.
316
00:25:22.815 --> 00:25:23.715
And I'm wondering
317
00:25:24.174 --> 00:25:30.995
what that process looks like for somebody building a streaming ML model with something like River and how you address things like
318
00:25:31.330 --> 00:25:32.230
concept drift
319
00:25:32.850 --> 00:25:35.990
and, you know, how concept drift manifests in this streaming
320
00:25:36.370 --> 00:25:45.775
environment where you are continually learning and you don't have to worry about, you know, the real world data that I'm seeing is widely divergent from the data that I use to train against.
321
00:25:46.155 --> 00:25:49.695
There's so many things to dig into, and I'll try to give a comprehensive answer.
322
00:25:50.179 --> 00:25:54.440
So first off, it's important to understand that Revo itself is
323
00:25:54.820 --> 00:25:58.280
to online learning what scikit learn is to batch learning. So
324
00:25:58.695 --> 00:25:59.515
it only
325
00:26:00.455 --> 00:26:01.675
desires to be
326
00:26:02.295 --> 00:26:03.275
a machine learning
327
00:26:03.815 --> 00:26:04.315
library.
328
00:26:04.935 --> 00:26:06.955
Right? So it just contains basically
329
00:26:07.370 --> 00:26:08.990
algorithms, routines to
330
00:26:09.530 --> 00:26:10.510
train a model
331
00:26:10.810 --> 00:26:18.565
and to have models that can learn and predict. And what you're going towards to with your question is MLOps. So how does
332
00:26:19.105 --> 00:26:23.285
the life cycle look like for an online model? And so this is always something that
333
00:26:23.640 --> 00:26:25.580
I'm spending a lot of time to look into.
334
00:26:26.120 --> 00:26:30.300
The answer is that the first part of the answer is that online learning
335
00:26:30.920 --> 00:26:31.420
enables
336
00:26:32.355 --> 00:26:33.174
different patterns.
337
00:26:34.115 --> 00:26:40.695
And I believe that these patterns are simpler to reason about. So as you said, you usually start off by
338
00:26:41.309 --> 00:26:45.009
training a model, then evaluating it against a test set, and
339
00:26:45.470 --> 00:26:47.570
maybe going to report to your stakeholders
340
00:26:47.870 --> 00:26:51.165
and show them the performance and guarantee that, you know,
341
00:26:51.625 --> 00:26:53.165
the percentages of full positives
342
00:26:53.785 --> 00:26:58.550
is underneath a certain threshold, then yes, we can diagnose cancer with this model or not.
343
00:26:59.030 --> 00:27:03.130
And yeah, and then you kind of deploy it. Maybe if you get lucky, if you get the approval, and
344
00:27:03.590 --> 00:27:15.595
you sleep you know well or not well at night depending on how much you trust your model. But there's this notion of you deploy a model, and it's like a baby in the world, and this baby is not going to keep learning. So,
345
00:27:16.139 --> 00:27:19.759
you know, it's a lie to believe that if you deploy a batch model,
346
00:27:20.460 --> 00:27:23.200
you're going to be able to just let it,
347
00:27:23.605 --> 00:27:28.025
you know, run by itself. There's actually main things main things that has to happen there. So
348
00:27:28.405 --> 00:27:32.430
the reality is that any machine learning project, you know, any serious project
349
00:27:32.730 --> 00:27:34.030
is never finished.
350
00:27:34.330 --> 00:27:39.070
It's like software, basically. We have to think of machining, projects have software engineering. And obviously,
351
00:27:39.450 --> 00:27:46.415
well, we all know that you never just deploy a feature, a software engineering feature, and just never look at it. You monitor it,
352
00:27:46.715 --> 00:27:50.095
you take care of it while investigating bugs and whatnot. So
353
00:27:50.470 --> 00:27:57.370
batch learning in that sense is a bit it's a bit difficult to work with because obviously you can have if your model is drifting,
354
00:27:57.965 --> 00:28:00.705
so meaning that its performance is dropping because
355
00:28:01.165 --> 00:28:13.540
the data that it's looking at is different than the training set it was trained on, you basically have to be very lucky if you want your model to pick up performance. So you're gonna have to do something about it. And, yeah, you can just retrain it. But
356
00:28:13.914 --> 00:28:18.895
what you do with online learning is that you can have the model just keep learning as you go. So
357
00:28:19.434 --> 00:28:20.335
there is no
358
00:28:21.115 --> 00:28:23.375
distinction between training and testing.
359
00:28:24.000 --> 00:28:27.940
What online learning encourages you to do is to deploy your model
360
00:28:28.560 --> 00:28:30.420
as soon as possible. So
361
00:28:30.800 --> 00:28:31.940
say you have a model,
362
00:28:32.255 --> 00:28:36.515
and it's not being trained on anything, well, you can put into production straight away.
363
00:28:36.975 --> 00:28:39.155
When samples arrive, it's gonna make a prediction.
364
00:28:39.534 --> 00:28:40.195
So maybe,
365
00:28:40.740 --> 00:28:43.320
you know, user arrives on your website, you make recommendations,
366
00:28:43.700 --> 00:28:45.559
that's your prediction. And then
367
00:28:45.860 --> 00:28:48.360
your user is going to click or not on
368
00:28:48.865 --> 00:28:50.485
something you recommended to her,
369
00:28:51.105 --> 00:28:51.605
and
370
00:28:52.065 --> 00:28:54.804
that's gonna be feedback to your for your model to keep training.
371
00:28:55.105 --> 00:28:55.605
So
372
00:28:56.270 --> 00:29:03.010
that is already a guarantee that your model is kind of up to date and kind of learning. And so that's really interesting because
373
00:29:03.790 --> 00:29:05.570
just enables so many good patterns.
374
00:29:06.415 --> 00:29:12.355
You can still monitor your the performance of your model. If the performance of your online model is dropping,
375
00:29:13.510 --> 00:29:16.809
I mean, I haven't seen that yet, but it probably means that your problem is
376
00:29:17.190 --> 00:29:18.410
really hard to solve.
377
00:29:19.190 --> 00:29:26.475
So the really cool thing I stumbled upon was this idea of of test then train. So the idea that
378
00:29:26.775 --> 00:29:28.715
imagine this scenario where you
379
00:29:29.174 --> 00:29:33.130
you have a classification model that is running online. And so what would happen is that
380
00:29:33.990 --> 00:29:35.290
you have your model.
381
00:29:35.670 --> 00:29:38.410
Your model is generating features. So
382
00:29:38.975 --> 00:29:49.240
say the user lives on the website and the features are, what's the time of the day? What's the weather like? What are the top films at the moment? And these are features. And these are features that you have at a certain point in time to
383
00:29:49.880 --> 00:29:55.065
you generate these features. And then later on, when you get the feedback, so did your recommendation
384
00:29:55.465 --> 00:29:57.644
was it a success or not? That's training
385
00:29:57.945 --> 00:29:59.325
data for your your model.
386
00:29:59.705 --> 00:30:01.644
You use the same features that you used
387
00:30:02.024 --> 00:30:02.765
for predictions.
388
00:30:03.409 --> 00:30:05.510
You use those features for training.
389
00:30:05.970 --> 00:30:11.669
And so you can see here that there's a clear feedback loop. The event happens. The user comes on the website.
390
00:30:12.195 --> 00:30:13.655
Your model generates features.
391
00:30:14.355 --> 00:30:15.015
And then
392
00:30:15.395 --> 00:30:24.640
at some later point in time, the feedback arrives. So was the prediction successful or not? Or if so if not, by by how much was the error? And then,
393
00:30:25.100 --> 00:30:29.840
yeah, you can use this feedback, join it with the feature that you generate predicting, and
394
00:30:30.205 --> 00:30:32.865
use that as as training data. So
395
00:30:33.165 --> 00:30:36.785
and you essentially have, like, a small queue or database that's storing
396
00:30:37.245 --> 00:30:37.985
your predictions,
397
00:30:39.029 --> 00:30:40.730
your features, and your
398
00:30:41.270 --> 00:30:44.570
training data, and the labels that make your training data.
399
00:30:45.029 --> 00:30:45.529
So
400
00:30:46.205 --> 00:30:47.905
the big difference here is that
401
00:30:48.684 --> 00:30:49.184
you,
402
00:30:50.044 --> 00:30:51.505
you do not necessarily have
403
00:30:51.885 --> 00:30:59.580
to do a training test phase before deploying your model. You can actually just deploy your model initially, and it just learns online, and then you can monitor.
404
00:31:00.279 --> 00:31:01.740
A really cool thing is that
405
00:31:02.120 --> 00:31:05.875
if you do this, you have a log of people coming on your website.
406
00:31:06.175 --> 00:31:08.434
You're making predictions. You gain features.
407
00:31:09.055 --> 00:31:12.195
People clicking around and interacting with your recommendations.
408
00:31:12.815 --> 00:31:14.035
This creates a log
409
00:31:14.380 --> 00:31:16.080
of what's happening on your website.
410
00:31:16.460 --> 00:31:19.120
And so this log, what's really cool is that you can
411
00:31:19.820 --> 00:31:26.025
offline, afterwards, after the fact, you can process it in the same order it arrives in, and you can replay
412
00:31:26.645 --> 00:31:35.750
what the history of what happened. So it means that if you on the side, when you're redeveloping a model or you want to develop a better model, you can just take this log of events,
413
00:31:36.370 --> 00:31:40.155
run for it, and do this prediction and training dance
414
00:31:40.455 --> 00:31:41.835
the whole life cycle.
415
00:31:42.215 --> 00:31:48.740
You know, you're replaying the feedback loop, and then you have a very accurate representation of how your model would have performed
416
00:31:49.440 --> 00:31:58.865
on that sequence of events. So that's really powerful because the way you're designing your model there is that you have a rough sketch of a model, which you deployed,
417
00:31:59.405 --> 00:32:01.345
then you have a log on that model.
418
00:32:01.965 --> 00:32:02.865
So you know
419
00:32:03.405 --> 00:32:07.299
you can evaluate the performance of that model, but more importantly, you can have
420
00:32:07.679 --> 00:32:08.740
a log of the events.
421
00:32:09.360 --> 00:32:12.580
And then when you're designing the version 2 of your model,
422
00:32:13.120 --> 00:32:13.620
you
423
00:32:14.645 --> 00:32:16.985
have a very reliable way to
424
00:32:17.524 --> 00:32:19.865
estimate how your new model would have performed.
425
00:32:20.404 --> 00:32:24.220
And that's really cool. Because when you are doing train test splits
426
00:32:24.600 --> 00:32:25.580
in batch learning,
427
00:32:26.200 --> 00:32:33.485
that is not representative of the way the way our world. The whole problem is that what you do with train and test is people are spending so much time
428
00:32:33.945 --> 00:32:36.525
making sure that their training test split is correct,
429
00:32:37.065 --> 00:32:40.525
when in fact, even having a good training test split is not
430
00:32:40.940 --> 00:32:46.080
a good proxy of the real world. A good proxy of the real world is to just replay
431
00:32:46.380 --> 00:32:47.120
through history.
432
00:32:47.660 --> 00:32:54.455
So and that's something that you can only do online learning. That's really cool. Now to come to your point about concept drift,
433
00:32:55.235 --> 00:33:02.520
so concept drift is there's many different kind of concept drift, and Chip has a very good on her blog.
434
00:33:02.820 --> 00:33:10.755
What matters really is that concept drift, the result of it is usually that your model is not performing as well. Right? It's gonna be a drop in performance. And so
435
00:33:11.135 --> 00:33:19.850
the first thing you see on your in your monitoring dashboard is that a metric has dropped. And then when you dig into it, you see that maybe there's a class imbalance
436
00:33:20.230 --> 00:33:24.890
or that the correlation between a feature and a class has changed or
437
00:33:25.284 --> 00:33:29.465
something like that. So essentially saying that the data the model has been trained on
438
00:33:29.845 --> 00:33:30.585
is not
439
00:33:30.965 --> 00:33:33.865
representative of the new data that has been seen in production.
440
00:33:34.540 --> 00:33:35.120
But again,
441
00:33:35.660 --> 00:33:37.840
I've only said this a few times, but online models,
442
00:33:38.220 --> 00:33:49.195
which put them in place with the correct camera ops setup, you they are able to learn as soon as possible. So that just guarantees that your model is as up to date as possible. So you're basically really doing the best you can.
443
00:33:49.495 --> 00:33:54.370
So drift is always possible. You can always obviously have a model that's degrading or that's just going haywire.
444
00:33:54.830 --> 00:33:57.090
That's not related necessarily to the
445
00:33:57.390 --> 00:33:59.090
online learning aspect of things.
446
00:33:59.390 --> 00:34:00.130
And so
447
00:34:00.510 --> 00:34:02.850
there are also ways to cope with this. So
448
00:34:03.455 --> 00:34:07.555
for instance, Dan Crankshaw and his team at Berkeley, they developed a system called CLIPr.
449
00:34:08.255 --> 00:34:20.070
It's a kind of an ops tool. It's a research project, but it's it's also it's I think it's been deprecated, but it's the ideas are still there. It's a project where they have a meta model, which is
450
00:34:20.395 --> 00:34:28.335
kind of looking at many models being run-in production and deciding online which model should be used to make a prediction. So it's kind of like a teacher
451
00:34:29.079 --> 00:34:39.205
selecting the best student at a certain point in time and, you know, kind of seeing throughout the year how the students are evolving and, like, which students are getting better or not good.
452
00:34:39.585 --> 00:34:46.980
And so you can do this with bandits, for instance. But yeah. So just to say that there are many ways to deal with concept drift,
453
00:34:47.280 --> 00:34:48.820
and the online models,
454
00:34:49.120 --> 00:34:49.620
again,
455
00:34:50.480 --> 00:34:50.955
help
456
00:34:51.435 --> 00:34:56.735
to cope with content drift in just a way, actually. It just makes sense more so than batch bundles.
457
00:34:57.115 --> 00:35:04.339
And so digging now into River itself, can you talk through how you've implemented that framework and some of the
458
00:35:04.640 --> 00:35:05.140
design
459
00:35:05.440 --> 00:35:09.140
considerations that went into how do I think about exposing this
460
00:35:09.485 --> 00:35:15.825
online learning capability in a way that is accessible and understandable to people who are used to building batch models?
461
00:35:16.420 --> 00:35:19.960
So I like to think of VIVA more as a library than a framework.
462
00:35:20.340 --> 00:35:21.720
If I'm not mistaken,
463
00:35:22.420 --> 00:35:25.435
framework kind of forces you into a certain
464
00:35:26.214 --> 00:35:31.755
behavior or way to do things, and there's an inversion of control where the framework is kind of designing things for you.
465
00:35:32.080 --> 00:35:36.660
So, you know, if you look at Keras and PyTorch, Keras is very much more a framework
466
00:35:37.200 --> 00:35:38.900
in comparison to PyTorch because
467
00:35:39.234 --> 00:35:51.480
PyTorch, for me, the reason why it was successful is that it it kind of gave inversion of control towards the user. You can do so many things in PyTorch, and it's very flexible and doesn't really impose a single way of doing so
468
00:35:52.020 --> 00:35:53.000
we have the ImmanuelRiver.
469
00:35:53.700 --> 00:35:56.360
River, again, is just a library to do
470
00:35:56.735 --> 00:36:01.155
machine learning, online machine learning. But it's it just contains the algorithms. It doesn't really
471
00:36:01.535 --> 00:36:07.020
force you to, you know, read your data in a certain way, or you could use it in a web app. You could use it offline.
472
00:36:07.640 --> 00:36:10.380
You could do it use it on an offline IoT sensor.
473
00:36:10.840 --> 00:36:24.890
Livr is not concerned with that. It's just an hybrid that is agnostic with regards to that. So now to come in to to what Livr is, it is in terms of online machining, it is general purpose. So So it's not dedicated to anomaly
474
00:36:25.270 --> 00:36:34.165
detection or forecasting or classification. It covers all of that. That's the ambition at least. So just to note there is that it's actually really hard to develop and maintain because
475
00:36:34.625 --> 00:36:37.285
other maintainers and I, we are not actually
476
00:36:38.260 --> 00:36:43.260
specialized in different domains, and we kinda have to, you know 1 day, I'm going to be doing
477
00:36:43.700 --> 00:36:44.760
working on the forecasting
478
00:36:45.224 --> 00:36:50.444
module other than the other day I'm gonna work on anomaly detection and it's it's kinda crazy. So
479
00:36:50.744 --> 00:36:55.280
it's still fun. What we do provide is a common interface. So just like scikit learn,
480
00:36:55.660 --> 00:36:56.160
every
481
00:36:56.780 --> 00:37:02.000
piece of the puzzle and river follows a certain interface. So we have transformers, we have regressors,
482
00:37:02.515 --> 00:37:05.255
we have anomaly detectors, of course, we have classifiers,
483
00:37:05.635 --> 00:37:14.600
so binary and multi class. We have forecasting models for time series. And so we guarantee to the user that each model follows a certain
484
00:37:15.220 --> 00:37:15.720
API.
485
00:37:16.100 --> 00:37:20.525
So every model is gonna be able to have a learn method, so it can just learn from new data.
486
00:37:20.985 --> 00:37:23.885
And they usually have a predict method to make prediction.
487
00:37:24.185 --> 00:37:26.685
And so forecasters will have a forecast method.
488
00:37:27.280 --> 00:37:32.340
Anomaly detectors will have a score method, which I must say, anomaly score. And so
489
00:37:33.040 --> 00:37:36.420
the strength of Viva is to, yeah, provide this consistent
490
00:37:37.184 --> 00:37:37.684
API
491
00:37:38.305 --> 00:37:39.684
for doing online machinery.
492
00:37:40.305 --> 00:37:48.910
And it's a bit opinionated because it's well, it just it likes like it learn really. It just says, okay, you're gonna have learn and predict, but that's a reasonable
493
00:37:49.210 --> 00:37:50.190
thing to impose.
494
00:37:50.569 --> 00:37:52.589
And that makes it easier for users to
495
00:37:52.970 --> 00:37:59.055
switch in new models because they have the same interface. And so, again, just to conclude on what I said at the start,
496
00:37:59.355 --> 00:38:01.695
we made the explicit choice to follow the
497
00:38:02.155 --> 00:38:09.440
single responsibility principle in that. Never only manages the machine learning aspect of things and not the deployments and whatnot. And so
498
00:38:09.820 --> 00:38:11.680
if you wanted to use it in production,
499
00:38:12.065 --> 00:38:14.405
see people doing this, you have to worry
500
00:38:15.025 --> 00:38:19.285
about some of the details yourself. Right? If you want to deploy in a in a web app, well,
501
00:38:19.740 --> 00:38:34.950
we do not help at the moment at all. You have to deploy your own web app. As far as the overall design of the framework, you mentioned that it actually started off as 2 separate projects, and then you went through the long process of merging them together.
502
00:38:42.070 --> 00:38:48.454
Actually why the merger between Creme and Scikit Multiflow took a certain time was that although we were both
503
00:38:48.755 --> 00:38:50.055
online learning libraries,
504
00:38:50.720 --> 00:38:54.260
there were some subtle differences, which were kind of important. So
505
00:38:54.800 --> 00:38:55.940
my vision with
506
00:38:56.400 --> 00:38:58.580
Cream at the time and Moover now is that
507
00:38:58.994 --> 00:39:06.775
we should only cater to models which are what I call pure online models, is that in that they can learn from a single sample of data at a time.
508
00:39:07.110 --> 00:39:09.930
But there are also mini batch models, so models
509
00:39:10.550 --> 00:39:16.355
which can learn from streaming data, but in chunks. So like in a mini batch of data. And scikitmultiflow
510
00:39:16.655 --> 00:39:20.275
was kind of doing this, so much like PyTorch and
511
00:39:20.575 --> 00:39:22.675
TensorFlow and you know deep learning models.
512
00:39:23.210 --> 00:39:26.029
And so I kind of had to convince them that
513
00:39:26.569 --> 00:39:28.990
there were reasons why it was just a bit better.
514
00:39:29.690 --> 00:39:30.589
Why? Because,
515
00:39:30.964 --> 00:39:34.265
you know, if you think about again a user infect on the website
516
00:39:34.805 --> 00:39:35.865
or just any
517
00:39:36.165 --> 00:39:38.905
web requests or things that are happening in the real life,
518
00:39:39.250 --> 00:39:40.869
You want to learn as soon as possible.
519
00:39:41.250 --> 00:39:42.869
You don't want to have to wait
520
00:39:43.569 --> 00:39:44.069
for,
521
00:39:44.849 --> 00:39:54.155
you know, 32 samples to arrive to have a batch to be able to feed that to your model. You could obviously, but it just made sense to me to have something simpler where we only
522
00:39:54.615 --> 00:39:58.475
care about pure online learning. Because it means that you don't have to store anything,
523
00:39:59.080 --> 00:40:03.100
you just learn on the fly. And so I guess
524
00:40:03.400 --> 00:40:07.020
the interaction I had with Flaky Multipline kind of confirmed this idea.
525
00:40:07.560 --> 00:40:08.060
And
526
00:40:08.525 --> 00:40:21.360
I guess, you know, they were a bit doubtful when we did the merger because and maybe I was a bit too opinionated, but history proved that it actually made sense, and it's not a decision that we look back on. Like, we're really happy with this now.
527
00:40:21.660 --> 00:40:29.075
So river has, you know, arguably not a success. So it's working. It's alive. It's breathing. It's been going on for
528
00:40:29.455 --> 00:40:34.035
2 years and a half, the project. And so we have a steady intake of
529
00:40:34.529 --> 00:40:43.509
users that are adopting it, and you know we see this from emails we receive, from GitHub discussions and issues, and just general feedback we get. So
530
00:40:44.145 --> 00:40:47.845
the general idea of having a live bid that is only focused towards
531
00:40:48.385 --> 00:40:48.885
ML
532
00:40:49.265 --> 00:40:56.630
and just the algorithms is something that we just gonna keep going with, because it just it looks like it's working, and it looks like this is what people want.
533
00:40:57.010 --> 00:40:58.630
You know, a simple example is,
534
00:40:58.985 --> 00:41:02.365
hey. I want to compute a covariance matrix online.
535
00:41:02.825 --> 00:41:07.885
Well, Livr aims to be the go to library to answer those kind of questions. Right?
536
00:41:08.480 --> 00:41:10.339
But the truth is that people,
537
00:41:10.960 --> 00:41:13.380
they don't just need that. They also need ways to
538
00:41:14.000 --> 00:41:16.980
deploy these models and do MLOps online. So
539
00:41:17.305 --> 00:41:21.885
what we basically did the next steps are for us to build new tools in that direction.
540
00:41:22.185 --> 00:41:23.805
Now we also think that
541
00:41:24.240 --> 00:41:27.060
the initial development of Vivos was a bit fast and furious.
542
00:41:28.320 --> 00:41:33.780
The aim was to implement as many algorithms as possible and, you know, just to cover the wide spectrum of
543
00:41:34.224 --> 00:41:34.724
machine
544
00:41:35.265 --> 00:41:37.525
learning. Now that we've, you know, covered
545
00:41:37.905 --> 00:41:39.444
quite a few topics,
546
00:41:39.984 --> 00:41:44.600
and we also have day jobs. So when I was developing with initially, I was doing a PhD, so I had
547
00:41:45.000 --> 00:41:53.625
ironically, I have more time than now because I have a proper job. But we value our time a bit more, and we're not in this fast and surest mode. We kind of just focus on
548
00:41:54.085 --> 00:41:59.625
picking certain models which are valuable and we see value in and just spending time to implement them properly.
549
00:42:00.005 --> 00:42:04.410
And we also see the final aspect is that we see that people, they don't just want
550
00:42:04.789 --> 00:42:11.035
our user base does not just want algorithms. They also want us to educate them. So they have general questions as towards,
551
00:42:11.495 --> 00:42:20.869
you know, what is online learning, and how do I do it, and how do I decide what model to use, and, you know, all the questions that we're covering in this podcast, basically. So I think there's a huge
552
00:42:21.570 --> 00:42:23.670
need for us to kind of
553
00:42:24.050 --> 00:42:26.550
move into a educational aspect. So
554
00:42:26.885 --> 00:42:37.340
when I was younger, scikit learn was my bible. Like, I would just spend so much time not even using it, not even just using the code, but actually just reading through the documentation because it's just so excellent.
555
00:42:37.880 --> 00:42:38.380
So
556
00:42:38.920 --> 00:42:41.260
obviously, that takes a lot of time, a lot of energy,
557
00:42:41.800 --> 00:42:42.300
people,
558
00:42:42.680 --> 00:42:46.905
contributors, and help, but definitely something towards which we are moving.
559
00:42:47.285 --> 00:42:48.585
In terms of the
560
00:42:48.965 --> 00:42:49.865
overall process
561
00:42:50.805 --> 00:42:54.025
of building the model using something like River,
562
00:42:54.430 --> 00:43:03.490
When people are building a batch model, they end up getting a binary artifact out that is the entire state of that model after it has gone through that training process.
563
00:43:03.975 --> 00:43:06.235
And I'm curious if you can talk to how
564
00:43:06.615 --> 00:43:13.890
River manages the stateful aspect of that model as it goes through this continual learning process, both
565
00:43:14.430 --> 00:43:29.349
in a, you know, sandbox use case where somebody is just playing around on their laptop, but also as you push it into production where maybe you want to be able to use this model and scale out serving it across a fleet of different servers and just some of the
566
00:43:29.730 --> 00:43:35.829
state management that goes into being able to continually learn as new information is presented to it.
567
00:43:36.375 --> 00:43:36.875
So
568
00:43:37.175 --> 00:43:41.675
batch learning the great advantage of batch learning is that once you train your model,
569
00:43:42.055 --> 00:43:43.675
it's essentially a pure function.
570
00:43:44.109 --> 00:43:45.410
There are no side effects.
571
00:43:46.109 --> 00:43:46.609
The,
572
00:43:47.150 --> 00:43:54.525
you know, decision process that's underlying the model is not gonna change. So, you know, you can push the envelope and compile it. You can
573
00:43:54.825 --> 00:43:59.860
pickle it. You can convert it to another format. So that's what lnx does. You
574
00:44:00.240 --> 00:44:04.740
can compile it so that it can run on a mobile device. I mean, it does not need the ability
575
00:44:05.600 --> 00:44:11.445
to train anymore. It's just basically a Python function. Or just it's just a function basically that takes
576
00:44:11.825 --> 00:44:17.125
an input and outputs something. So there's also a good reason why batch learning is predominant.
577
00:44:17.670 --> 00:44:22.250
But with Revver, it's different because online models, they need to keep this ability
578
00:44:22.950 --> 00:44:25.714
to learn. That's what you've been saying. So it's
579
00:44:26.095 --> 00:44:27.954
actually kind of straightforward,
580
00:44:28.494 --> 00:44:28.994
but
581
00:44:29.535 --> 00:44:32.755
the internal representation of most models of river is
582
00:44:33.309 --> 00:44:33.809
fluid,
583
00:44:34.270 --> 00:44:36.849
dynamic. It's usually stored in dictionaries that can
584
00:44:37.150 --> 00:44:39.329
increase and decrease in size. So
585
00:44:39.869 --> 00:44:42.530
imagine you have a new feature that arrives in your stream,
586
00:44:42.955 --> 00:44:55.370
Well, every model never coped with that. It's they're not static. There's a new feature that appears well, they handle it gracefully. So for instance, a linear regression model is just going to add a new weight to its internal
587
00:44:56.150 --> 00:44:59.050
dictionary of weight. Now in terms of civilization
588
00:44:59.510 --> 00:45:00.970
and pickling and whatnot,
589
00:45:01.575 --> 00:45:03.035
river is mostly written
590
00:45:03.494 --> 00:45:05.355
well, basically, river stands on the shoulders
591
00:45:06.215 --> 00:45:06.875
of Python,
592
00:45:07.415 --> 00:45:11.770
very much so. So we do not depend very much on NumPy or Pandas or SciPy.
593
00:45:12.230 --> 00:45:15.690
We mostly depend on Python standard library. We use dictionaries
594
00:45:16.070 --> 00:45:16.650
a lot,
595
00:45:17.135 --> 00:45:17.635
and
596
00:45:17.935 --> 00:45:23.075
that plays really nicely with the standard library. It's very easy just to you can take any Rivo model,
597
00:45:23.375 --> 00:45:28.560
pickle it, and just save it. You can also just dump it to JSON or whatnot.
598
00:45:29.020 --> 00:45:29.520
Also,
599
00:45:29.900 --> 00:45:31.120
the paradigm of
600
00:45:31.455 --> 00:45:35.635
you train a model, you pickle it, and you have an artifact that you can upload anywhere,
601
00:45:35.935 --> 00:45:40.035
it's a bit different with online learning because you would train this differently. You would
602
00:45:40.450 --> 00:45:42.069
maintain your model memory.
603
00:45:42.450 --> 00:45:44.470
So if you have a web server serving your model,
604
00:45:44.930 --> 00:45:49.269
you would not just load the model to make a prediction, you would just keep it in memory
605
00:45:50.115 --> 00:45:53.335
and make a prediction of it because it's in memory, so you don't have to load it anymore,
606
00:45:53.715 --> 00:45:54.535
preload it.
607
00:45:54.835 --> 00:45:59.700
And then when a sample arrives, your model is in memory. You can just make it learn from that. So,
608
00:46:00.080 --> 00:46:02.100
yeah, I think the big difference is that you
609
00:46:02.400 --> 00:46:06.340
hold your model memory rather than picking it to the disk and
610
00:46:06.755 --> 00:46:07.974
loading it when necessary.
611
00:46:08.595 --> 00:46:12.934
In terms of the use of the dictionary as that internal state representation,
612
00:46:13.740 --> 00:46:15.680
As you said, it gives you the flexibility
613
00:46:15.980 --> 00:46:20.000
to be able to evolve with the updates and data.
614
00:46:20.619 --> 00:46:22.880
But at the same time, you have this
615
00:46:23.315 --> 00:46:23.815
heterogeneous
616
00:46:24.355 --> 00:46:26.455
data structure that can be mutated
617
00:46:26.835 --> 00:46:30.135
as the system is in flight, and you don't necessarily
618
00:46:30.515 --> 00:46:31.495
have a
619
00:46:31.799 --> 00:46:37.180
strict schema being applied to it. And I'm just curious if you can talk to the trade offs of
620
00:46:37.559 --> 00:46:39.315
being able to add that flexibility,
621
00:46:39.615 --> 00:46:42.515
but also lacking in some of the validation
622
00:46:42.895 --> 00:46:49.580
and, you know, schema and structure information that you might want in something that's dealing with these volumes of data?
623
00:46:49.960 --> 00:46:52.460
So, yeah, Viva uses dictionaries. So
624
00:46:52.920 --> 00:46:54.940
the advantage of dictionaries are plentiful.
625
00:46:55.664 --> 00:46:58.244
First of all, a really important thing is that
626
00:46:59.184 --> 00:47:02.724
dictionaries are to list what handles data frames are to nonpyras.
627
00:47:03.480 --> 00:47:12.675
So a dictionary has names, and that's really important. So it means that each 1 of your feature actually has a name to it. And I find that hugely important because,
628
00:47:12.975 --> 00:47:17.315
you know, we always see features as just numbers, but they also have names, and that's just really important.
629
00:47:17.990 --> 00:47:20.570
Imagine you have a bunch of features coming in.
630
00:47:20.870 --> 00:47:22.090
Now if that was
631
00:47:22.630 --> 00:47:29.205
a list or a NumPy array, you have no real way of knowing which column corresponds to each variable. If you
632
00:47:29.585 --> 00:47:31.845
switch 2 columns with each other,
633
00:47:32.305 --> 00:47:35.925
that could just be a really silent bug which really affect you. Whereas
634
00:47:36.230 --> 00:47:37.610
if you name each feature,
635
00:47:38.070 --> 00:47:44.810
if the column order changes, well the names of the columns are being permuted too, so you can kind of identify that. So
636
00:47:45.115 --> 00:47:54.734
what's really cool with dictionaries and and that works with River is that the order of the features that you're receiving doesn't matter, because we access every feature by name and not by position.
637
00:47:55.380 --> 00:47:58.039
Dictionaries also allow you are mutable in size. So,
638
00:47:58.339 --> 00:48:03.960
you know, if a new feature arrives or a new feat a feature disappears between 2 different samples,
639
00:48:04.655 --> 00:48:06.195
that just works. So
640
00:48:06.735 --> 00:48:11.155
it's really cool also that dictionaries on when you think about it naturally sparse.
641
00:48:12.500 --> 00:48:16.120
So imagine that on the Netflix projects, the features that you receive
642
00:48:16.580 --> 00:48:18.440
are the name of the user,
643
00:48:19.085 --> 00:48:19.585
semicolon
644
00:48:20.444 --> 00:48:20.944
1
645
00:48:21.805 --> 00:48:23.404
or the dates 1 or
646
00:48:23.885 --> 00:48:29.099
Yeah you can just store sparse information in a dictionary, that's kind of really useful.
647
00:48:29.720 --> 00:48:31.660
There's this robustness principle
648
00:48:31.960 --> 00:48:36.460
that we follow with reverse. So robustness principle is that we are is to be conservative
649
00:48:37.355 --> 00:48:40.494
in what you do, but liberal in what you accept. So
650
00:48:40.955 --> 00:48:42.735
we're very liberal and that accepts
651
00:48:43.755 --> 00:48:51.050
heterogeneous data, as you said. So dictionaries are different in size, dictionaries would have which have different orders or whatnot, but
652
00:48:51.350 --> 00:48:54.010
that is really flexible for users. So a common
653
00:48:54.595 --> 00:48:56.775
use case is to deploy whether in
654
00:48:57.155 --> 00:49:00.455
a web app, and in the web app, you're receiving JSON data
655
00:49:00.915 --> 00:49:06.050
a lot of the time. So the fact that JSON data has a 1 to 1 relationship with Python dictionaries
656
00:49:06.510 --> 00:49:08.210
makes it really easy to integrate
657
00:49:08.670 --> 00:49:14.545
with it into a web app. Whereas if you have a regular batch model, you have to mess about with
658
00:49:15.005 --> 00:49:15.505
casting
659
00:49:15.965 --> 00:49:17.985
the JSON data to a NumPy array,
660
00:49:18.340 --> 00:49:21.240
and, you know, that has a cost actually. It actually has a cost because
661
00:49:21.620 --> 00:49:22.120
although
662
00:49:22.660 --> 00:49:25.240
NumPy, Torch, TensorFlow are,
663
00:49:25.735 --> 00:49:27.595
you know, good at processing matrices,
664
00:49:28.055 --> 00:49:29.915
there's actually a cost that comes with
665
00:49:30.215 --> 00:49:30.715
taking
666
00:49:31.095 --> 00:49:33.835
native data, such as dictionaries, and casting them to
667
00:49:34.480 --> 00:49:38.020
a higher order data structure, such as an NumPy array,
668
00:49:38.400 --> 00:49:43.220
that has a real cost. In a web app where, you know, it's you're reading in terms of milliseconds,
669
00:49:43.714 --> 00:49:44.214
well,
670
00:49:44.755 --> 00:49:48.775
you're spending a lot of your time just converting your JSON data to NumPy.
671
00:49:49.234 --> 00:49:50.214
Whereas with River,
672
00:49:50.675 --> 00:49:56.540
because it consumes dictionaries, well, the data you receive, I don't know if you're coding in Django, Flask, or FastAPI,
673
00:49:56.920 --> 00:50:02.194
the data you receive, your your request is a dictionary. So you don't have to convert the data. It just
674
00:50:02.494 --> 00:50:04.994
runs. So actually, if you if you take a river model,
675
00:50:05.295 --> 00:50:08.515
like a a linear regression river and a linear regression in torch,
676
00:50:08.820 --> 00:50:11.080
it's actually gonna be much faster in river because
677
00:50:11.460 --> 00:50:19.515
there's no conversion cost. Plus the features are names, and plus you just you don't have to worry about the features being mixed or anything. So
678
00:50:19.895 --> 00:50:39.234
it just makes a lot of sense really dictionaries in that sense. Now the pitfalls, obviously, it's it's not perfect. The pitfall is that I kinda disagree that there's a problem with the short dictionaries. I actually think that dictionaries, well, if you wanted to, you could create like, an issue in Python, you can actually use a data class, and you can convert that to a dictionary and feed that into
679
00:50:39.615 --> 00:50:42.595
your model. The data class helps you to create structure.
680
00:50:42.930 --> 00:50:59.355
So I don't think that's really a problem. Quite the contrary, I think. The fact is that also a dictionary can be nested. So maybe your features that you're feeding to the model doesn't have to be a flat dictionary. It's actually Nester's, and that's really cool too. You know, you can have features for your user, you can have features for your page, you can have features for the day, anything.
681
00:50:59.910 --> 00:51:02.250
Things that you cannot necessarily do with a flat
682
00:51:02.550 --> 00:51:08.315
structure such as a data frame or an empire. Anyway, I'm talking about benefits. I should be talking about cons.
683
00:51:08.875 --> 00:51:11.295
But yeah. I guess just the main con of
684
00:51:11.675 --> 00:51:17.309
processing dictionaries is that, you know, if you wanted a VIVA model to process a 1, 000, 000 samples,
685
00:51:17.849 --> 00:51:20.029
it would take much more time than
686
00:51:20.490 --> 00:51:22.029
processing a 1, 000, 000 samples
687
00:51:22.329 --> 00:51:24.625
with pandas data frame or NumPy array.
688
00:51:25.185 --> 00:51:35.500
Because, yeah, the point of Viber is to process is to be half of processing 1 sample at a time, but not necessarily processing a 1000000 samples at a time. But those are 2 different problems. So
689
00:51:36.040 --> 00:51:39.420
although, you know, you take Torch or scikit learn,
690
00:51:39.800 --> 00:51:45.765
their goal is to be able to process offline data really quick. But the goal of VIVA is to process
691
00:51:46.545 --> 00:52:01.445
online data, single samples, as fast as possible. And you're comparing apples and oranges if you you wanna do the comparison. It just doesn't work. So yeah. Actually, you know what? I don't think Hellyville downsized to using dish. It just helps a lot. And and to confirm this, we have a lot of users who
692
00:52:01.745 --> 00:52:07.605
tell us about this. They say, well, it's actually fun to use Rev. I just it just makes sense because it's very close to
693
00:52:07.950 --> 00:52:13.170
the data structures I use in Python. I don't have to introduce a new data structure to my system.
694
00:52:13.790 --> 00:52:15.410
So for somebody who is
695
00:52:15.795 --> 00:52:16.295
using
696
00:52:16.835 --> 00:52:24.775
River to build a machine learning model, can you just talk through the overall process of going from idea through development to deployment?
697
00:52:25.150 --> 00:52:29.890
I'm going to rehash what I said before, but I think the great benefit of online learning
698
00:52:30.510 --> 00:52:38.435
and river is that you can cut the R and D phase. So I've seen so many projects where there's an R and D phase, and
699
00:52:38.895 --> 00:52:44.250
the model, you know, gets validated at some point in time, but there's, like, a real big
700
00:52:44.710 --> 00:52:50.250
gap of time between the start of the r and d phase and the moment when the model is deployed. And
701
00:52:50.575 --> 00:52:53.635
the process of using river and an extreme
702
00:52:53.935 --> 00:52:58.195
model in general is to actually, as I said, deploy the model as soon as possible,
703
00:52:58.850 --> 00:52:59.910
monitor its predictions.
704
00:53:00.490 --> 00:53:00.990
And
705
00:53:01.570 --> 00:53:04.790
it's okay because sometimes that model well, you know, you can
706
00:53:05.250 --> 00:53:10.684
deploy it in production and those predictions do not necessarily have to be served to the user.
707
00:53:11.144 --> 00:53:13.644
So you just make the predictions, you monitor them,
708
00:53:14.150 --> 00:53:21.690
and it creates you a log again of training data and predictions and features. And that's what you call shadow deployment. You
709
00:53:22.145 --> 00:53:23.845
have a model which is
710
00:53:24.385 --> 00:53:35.210
deployed, is making predictions, but those predictions are not being used, you know, to inform decisions or to change or to influence the behavior of users. They just exist for the sake of existing and for monitoring.
711
00:53:35.670 --> 00:53:41.095
1 thing to mention is that once you deploy this model, you have your log of events.
712
00:53:41.395 --> 00:53:43.095
That's the phase where you want to
713
00:53:43.555 --> 00:53:53.680
maybe design a new model. And you're going to have this model replace the existing model in production or coexist with it because you have a meta model or not. So I mentioned that you can
714
00:53:54.220 --> 00:53:58.505
take your log of events, replay it in the order in which it arrived, and
715
00:53:59.285 --> 00:54:00.665
have a good idea of
716
00:54:01.045 --> 00:54:02.985
how well your model would have performed.
717
00:54:03.365 --> 00:54:04.905
That's called progressive validation.
718
00:54:05.680 --> 00:54:09.460
So it's just this idea that if you have your log of events,
719
00:54:09.920 --> 00:54:12.660
every sample you're first gonna make a prediction, and then you're gonna
720
00:54:13.040 --> 00:54:14.980
learn from it. So I have a good example.
721
00:54:15.385 --> 00:54:22.285
There's a data set on Kaggle called the New York Taxi Data Set, and it's basically a log of people asking for a
722
00:54:22.740 --> 00:54:26.599
hailing a taxi for a ride, and they depart from a position,
723
00:54:27.140 --> 00:54:30.520
and they arrive at another position later in time. And so
724
00:54:30.865 --> 00:54:33.924
the goal of a machine learning system in this case could be to
725
00:54:34.305 --> 00:54:37.444
predict how long the taxi trip is going to last. So
726
00:54:38.040 --> 00:54:39.500
when the taxi departs,
727
00:54:40.200 --> 00:54:47.740
you want your model to make a prediction. How long is this taxi ride gonna last? And so maybe that's gonna inform, I don't know, the cost of the
728
00:54:48.265 --> 00:54:53.805
trip, or it's going to help decision makers, you know, rear behind the taxis or I don't know, whatever.
729
00:54:54.585 --> 00:55:01.550
But you can imagine that this is a great feedback loop because you have your model make the prediction, and then later, maybe 18 minutes later or something,
730
00:55:01.930 --> 00:55:15.885
we had the ground truth arrived. So you know how long the taxi trip actually lasts, and that's your ground truth. And then then you can compare your prediction with your model. And that enables progressive validation because you have a a log of events. You have when the taxi trip departs,
731
00:55:16.430 --> 00:55:17.490
what was the value
732
00:55:17.790 --> 00:55:18.770
my model predicted,
733
00:55:19.310 --> 00:55:20.530
what features I used
734
00:55:20.990 --> 00:55:22.210
at prediction time,
735
00:55:22.670 --> 00:55:27.185
and later on, I had to go on through. And so I can just replay
736
00:55:28.765 --> 00:55:32.945
the logs of events for, you know, I don't know, 7 days and
737
00:55:33.680 --> 00:55:37.620
progressively evaluate my model. So I like this taxi example because
738
00:55:38.240 --> 00:55:40.740
it's easy to reason about and, you know, taxis are
739
00:55:41.085 --> 00:55:45.904
easy to understand. But the tax example is really what online learning is about. It's about
740
00:55:46.285 --> 00:55:50.119
this feedback loop between predicting and learning. And just
741
00:55:50.580 --> 00:55:51.480
to remind you,
742
00:55:51.780 --> 00:55:55.640
but how it would be with a batch model is that you would have your taxi date set, and
743
00:55:56.020 --> 00:56:05.785
well, I don't know, you would split your dates in 2, you would have start of the week, end of the week, train your model on the start of the week, evaluate on the rest of the week. Oh, no. The data I trained on
744
00:56:06.325 --> 00:56:08.825
for the start of the week is not represented on the weekend.
745
00:56:09.520 --> 00:56:11.060
Yeah. And it just becomes a bit
746
00:56:11.600 --> 00:56:14.340
weird. It becomes this situation where you're trying to
747
00:56:15.200 --> 00:56:19.734
reproduce conditions in real life, but you're never really sure of it. And you can only really know
748
00:56:20.115 --> 00:56:22.775
how well your batch model is going to do well
749
00:56:23.075 --> 00:56:23.895
in production.
750
00:56:24.595 --> 00:56:27.494
And online learning, it just kind of encourages you to
751
00:56:28.290 --> 00:56:30.950
go for it, to deploy your model straight away and not have to
752
00:56:31.490 --> 00:56:40.175
have this weird r and d phase where you live in a lab. You think you might be right, but then you're not really sure. And yeah, online learning just brings you closer to the reality,
753
00:56:40.635 --> 00:56:41.455
in my opinion.
754
00:56:42.075 --> 00:56:45.295
As you have been developing this project
755
00:56:45.599 --> 00:56:46.819
and helping people
756
00:56:47.119 --> 00:56:50.180
understand and adopt it, what do you see as some of the
757
00:56:50.720 --> 00:56:52.900
conceptual challenges or complexities
758
00:56:53.200 --> 00:56:53.940
that people
759
00:56:54.455 --> 00:57:02.155
experience as they're starting to adapt their thinking to how to build a machine learning model in this online streaming format versus
760
00:57:02.770 --> 00:57:12.905
the batch oriented workflow where they do have to think about the train test split, just the overall shift in the way that they think about iterating on and deploying and developing these models?
761
00:57:13.445 --> 00:57:19.225
It's a hard question, but I think there's 2 aspects. There's the online learning aspect, and then there's the MLOps aspect.
762
00:57:19.900 --> 00:57:21.520
Now in terms of MLOps,
763
00:57:21.900 --> 00:57:23.760
I think I I covered enough, but
764
00:57:24.300 --> 00:57:26.400
it's much like a batch model. You have to
765
00:57:27.184 --> 00:57:33.444
deploy your model, which means maybe serve it behind the web app. As I mentioned, the ideal situation is to have your model
766
00:57:34.420 --> 00:57:37.800
loaded into memory, and that's making prediction and training.
767
00:57:38.180 --> 00:57:39.400
But all that is really
768
00:57:40.180 --> 00:57:41.960
harder to do than to say.
769
00:57:42.425 --> 00:57:51.165
The truth is that there's actually no framework out there which allows you to do this. You could do this yourself, and this is what we see. I mean, we get users who ask us questions
770
00:57:51.810 --> 00:57:57.030
in a bit of context, so on on GitHub or in mails, but and they're asking us, how do I
771
00:57:57.410 --> 00:58:00.550
deploy my model? What should I be doing? And we always give the same answers.
772
00:58:00.944 --> 00:58:01.444
But
773
00:58:01.984 --> 00:58:10.810
the fact is that, you know, we have these users who have basically embraced Viva and they understand it, but then they get into the production phase. And that's not what we're trying to
774
00:58:11.210 --> 00:58:13.150
well, we feel bad because
775
00:58:13.609 --> 00:58:16.030
they are all making the same mistakes in some way,
776
00:58:17.095 --> 00:58:20.395
and Viva is is not there to help them because that's not the purpose of Viva.
777
00:58:20.935 --> 00:58:22.635
So, yeah, there's a lack of
778
00:58:23.390 --> 00:58:26.770
tooling to actually just, you know, deploy an online model.
779
00:58:27.150 --> 00:58:29.329
So, yeah, that's the MO ops aspect.
780
00:58:29.630 --> 00:58:31.569
I think in terms of some online learning,
781
00:58:31.914 --> 00:58:35.775
a big challenge is that not everyone has the luxury to have
782
00:58:36.394 --> 00:58:39.775
a PhD during which you can spend days nights
783
00:58:40.230 --> 00:58:43.130
going through online learning papers and trying to understand it, and
784
00:58:43.589 --> 00:58:47.849
that's what I and others had a chance to do. But a lot of our users,
785
00:58:48.150 --> 00:58:52.224
you know, they see the value of online learning, and they want to
786
00:58:52.605 --> 00:59:00.610
put it into production, but they have deadlines to meet. Right? They have to ship their project in 6 weeks, and they just do not have the time to understand things in detail. So
787
00:59:01.070 --> 00:59:03.970
things like what I just described, progressive validation,
788
00:59:04.430 --> 00:59:07.170
it kind of takes them a bit of time to understand.
789
00:59:07.684 --> 00:59:10.345
And so again, what we need to do is to
790
00:59:10.964 --> 00:59:11.944
spend more time
791
00:59:12.244 --> 00:59:15.704
creating resources or, you know, just diagrams to explain
792
00:59:16.170 --> 00:59:18.910
what online learning is about. And that in terms of
793
00:59:19.450 --> 00:59:20.430
library design,
794
00:59:20.890 --> 00:59:29.235
it's really important. Like, if we wanted to introduce a new method to all our estimators, I would be against that. Like, the whole point of Vivo is to make it as simple as possible
795
00:59:29.695 --> 00:59:39.369
so that people can just, you know, be productive, understand it. So, yeah, I think that just to recapitulate those 2 problems is that people do not necessarily have the
796
00:59:39.750 --> 00:59:42.645
resources to learn about online learning, and then
797
00:59:42.945 --> 00:59:44.964
there are operational problems around
798
00:59:45.665 --> 00:59:48.085
serving these models into production. So
799
00:59:48.464 --> 00:59:52.900
it's kind of like a batch model because you have to search a model behind an API,
800
00:59:53.200 --> 00:59:56.180
you know, and you have to monitor it. And these are things, well,
801
00:59:56.480 --> 00:59:59.859
you know, that are common to a batch model, but there's the added
802
01:00:00.295 --> 01:00:01.994
complexity of having your model
803
01:00:02.454 --> 01:00:04.315
being, you know, maintain a memory
804
01:00:04.615 --> 01:00:06.954
and keep learning and stuff and things that are
805
01:00:07.360 --> 01:00:10.580
basically not common. I mean, if you had to Google it or find something on GitHub,
806
01:00:11.120 --> 01:00:14.500
you you just kind of find these hacky projects, but no
807
01:00:14.955 --> 01:00:34.160
real good library to do that, at least not yet. And 1 of the things that we didn't discuss yet is the types of machine learning use cases that River supports where I'm speaking specifically to things, logistic regressions and decision trees versus deep learning and neural networks. And I'm just wondering if you can talk to the
808
01:00:34.615 --> 01:00:35.434
types of
809
01:00:35.815 --> 01:00:40.635
machine learning approaches that River is designed to support and some of the
810
01:00:41.015 --> 01:00:44.400
reasoning that went into where you decided to put your focus.
811
01:00:45.100 --> 01:00:50.480
Viva, again, is a general purpose library, so it was quite a few things. There are some
812
01:00:51.135 --> 01:00:54.835
cases or flavors of machine learning, which are especially
813
01:00:55.615 --> 01:00:56.115
interesting
814
01:00:56.815 --> 01:01:00.435
when you cast them in an online learning scenario. So
815
01:01:00.950 --> 01:01:02.810
if you're doing anomaly detection,
816
01:01:03.349 --> 01:01:04.650
so for instance, you have
817
01:01:05.030 --> 01:01:09.050
people doing transactions on a in a banking system, so they're making the payments,
818
01:01:09.455 --> 01:01:12.115
And you might want to be doing anomaly detection to detect
819
01:01:12.415 --> 01:01:12.915
fraudulent
820
01:01:13.295 --> 01:01:13.795
payments.
821
01:01:14.575 --> 01:01:16.595
That is very much a
822
01:01:17.055 --> 01:01:18.915
situation where you have streaming data.
823
01:01:19.339 --> 01:01:20.079
And so
824
01:01:20.619 --> 01:01:24.079
in that case, you would like to be doing online anomaly detection. So
825
01:01:24.940 --> 01:01:25.440
we
826
01:01:25.900 --> 01:01:29.695
see that every time we put out a notebook or
827
01:01:30.315 --> 01:01:32.575
a new anomaly detection method,
828
01:01:33.115 --> 01:01:36.335
a lot of people start using it. We start having bug reports and
829
01:01:36.740 --> 01:01:39.880
and whatnot. So it's kind of surprising, but it's a good thing. But,
830
01:01:40.740 --> 01:01:42.680
yeah, I think there are modules
831
01:01:43.140 --> 01:01:49.635
and aspects of which are clearly bring a lot of value to users. So that would be anomaly detection, but also
832
01:01:50.575 --> 01:01:59.430
we have forecasting models. So when you do online forecasting, that just makes sense. But you have sensors which are, I don't know, measuring the temperature of something. People want to do that in real life
833
01:01:59.730 --> 01:02:06.655
in in real time. There's also a good example I have is we have this engineer who's working on the water pipes in in Italy.
834
01:02:07.035 --> 01:02:07.775
He's trying
835
01:02:08.315 --> 01:02:12.990
to predict how much water is going to flow through certain points in his pipeline.
836
01:02:13.450 --> 01:02:13.950
So
837
01:02:14.410 --> 01:02:18.670
he has sensors all over the pipeline, and he's trying to just do a forecasting model.
838
01:02:19.015 --> 01:02:21.595
And so it just makes so much sense for him
839
01:02:21.975 --> 01:02:23.995
to be able to have his model run online
840
01:02:24.295 --> 01:02:25.515
inside the sensors
841
01:02:26.030 --> 01:02:28.050
or inside the IoT systems he's running.
842
01:02:28.510 --> 01:02:30.690
So just all that to say that
843
01:02:31.550 --> 01:02:33.010
there are some more exotic
844
01:02:33.595 --> 01:02:38.895
parts of the river, such as anomaly detection and forecasting, which are not which are probably more value than
845
01:02:39.355 --> 01:02:44.500
the classic models such as, you know, linear regression, classification, regression.
846
01:02:45.200 --> 01:02:48.580
Again, at the start of the show, I I talked about Netflix recommendations.
847
01:02:49.360 --> 01:02:50.660
So we have some
848
01:02:51.120 --> 01:02:51.620
very
849
01:02:52.365 --> 01:02:52.865
basic
850
01:02:53.405 --> 01:02:55.105
bricks to be able to make recommendations.
851
01:02:55.805 --> 01:03:01.505
Well, we have factorization machines, and we have some kind of ranking system so that if you have users and items,
852
01:03:02.030 --> 01:03:02.690
you can
853
01:03:02.990 --> 01:03:05.090
kind of build a ranking of
854
01:03:05.390 --> 01:03:06.930
preferred items for a user.
855
01:03:07.230 --> 01:03:08.850
So we have these kind of exotic
856
01:03:09.470 --> 01:03:11.010
machine learning cases which
857
01:03:11.585 --> 01:03:12.565
provide value,
858
01:03:13.425 --> 01:03:16.325
but require us to spend a lot of time to
859
01:03:17.105 --> 01:03:18.645
work on them basically. So
860
01:03:19.299 --> 01:03:25.400
it's very difficult for me and for other contributors to be specialized in anomaly detection time series forecasting
861
01:03:25.859 --> 01:03:26.359
recommendation.
862
01:03:27.165 --> 01:03:28.865
But, yeah, all this to say that
863
01:03:30.125 --> 01:03:31.265
covers a wide spectrum.
864
01:03:31.724 --> 01:03:37.070
You can do preprocessing. You can extract features. You can do classification, regression, forecasting, anything.
865
01:03:37.550 --> 01:03:42.370
We try to well, because it's online, it's just a bit unique. In your experience of
866
01:03:42.910 --> 01:03:44.770
working with the River Library
867
01:03:45.204 --> 01:03:52.265
and working with end users of the tool, what are some of the most interesting or innovative or unexpected ways that you've seen it used?
868
01:03:52.579 --> 01:04:01.400
Well, unexpected is a good 1. There's 1 thing that comes to mind. We have this person who is a beekeeper. So a person who is, you know, taking care of bees and
869
01:04:01.785 --> 01:04:04.605
I guess they once a week or every 2 weeks, they
870
01:04:04.905 --> 01:04:08.525
go to the beehive and they pick up the honey in the beehive.
871
01:04:09.119 --> 01:04:20.525
And this person has many beehives, and so they don't like to waste their time going into the beehive and actually checking if there's honey or not. So they have a sensor. He has sensors that are in each beehive.
872
01:04:20.905 --> 01:04:21.805
They're kinda measuring
873
01:04:22.425 --> 01:04:25.565
how much honey is in each beehive, and he likes to forecast
874
01:04:25.950 --> 01:04:26.770
how much honey
875
01:04:27.390 --> 01:04:33.630
he is expected to have in, you know, the weeks or months to come based on the weather, based on past data, based on
876
01:04:34.445 --> 01:04:36.385
I don't know, what information he uses.
877
01:04:37.085 --> 01:04:41.185
But really, really just fun just to see this person doing this hackish project
878
01:04:41.539 --> 01:04:44.259
where they just thought it would be fun to use online learning to do it.
879
01:04:44.819 --> 01:04:47.720
And again, it wasn't an IoT context, so
880
01:04:48.180 --> 01:04:49.079
that made sense.
881
01:04:49.395 --> 01:04:53.975
I guess innovative, I was kind of impressed when I heard about this project of having
882
01:04:54.915 --> 01:04:59.309
a, you know, a model within each car to determine where your destination
883
01:04:59.609 --> 01:05:06.910
would be. So I don't know. You wake in the morning, you tell your car, you know, is it Saturday you go to go to the market? Is it a weekday you go to work?
884
01:05:07.215 --> 01:05:08.915
It sounds silly, obviously, but
885
01:05:09.215 --> 01:05:13.155
having this this idea of having 1 model per user is is kind of fun.
886
01:05:13.615 --> 01:05:18.099
The most impactful project I heard about, and I know it's which is being used, is
887
01:05:18.880 --> 01:05:21.619
a situation where this company, they
888
01:05:22.079 --> 01:05:23.859
prevent cyber attacks. So
889
01:05:24.395 --> 01:05:25.775
they monitor this grid
890
01:05:26.155 --> 01:05:27.615
of servers and computers,
891
01:05:28.395 --> 01:05:29.615
and they're monitoring
892
01:05:29.915 --> 01:05:31.535
traffic between machines.
893
01:05:31.910 --> 01:05:33.530
And so they're trying to understand
894
01:05:34.790 --> 01:05:36.730
when some of the traffic is malicious,
895
01:05:37.030 --> 01:05:41.285
and hackers basically trying to get into a system. So you can
896
01:05:42.145 --> 01:05:45.125
detect this by looking at the patterns of
897
01:05:45.505 --> 01:05:49.285
this traffic. Right? And the trick is that
898
01:05:49.970 --> 01:05:55.670
behind the traffic, the malicious traffic is hackers, and they're constantly changing their patterns,
899
01:05:56.530 --> 01:05:59.109
so access patterns to actually not be detected.
900
01:05:59.515 --> 01:06:02.174
And so if you manage to label traffic
901
01:06:02.555 --> 01:06:03.055
as,
902
01:06:03.515 --> 01:06:04.174
you know,
903
01:06:04.555 --> 01:06:07.535
malicious, well, you want your model to keep learning. So
904
01:06:08.069 --> 01:06:10.970
they have the system where they have, like, thousands of machines, and
905
01:06:11.510 --> 01:06:16.650
they have a few machines that are dedicated to just learning from the traffic and in real time,
906
01:06:17.055 --> 01:06:17.555
adapting,
907
01:06:17.855 --> 01:06:18.355
learning,
908
01:06:18.815 --> 01:06:19.315
detecting
909
01:06:19.615 --> 01:06:20.515
anomalous traffic,
910
01:06:20.815 --> 01:06:21.875
sending it to
911
01:06:22.175 --> 01:06:25.474
human beings so that they can actually verify themselves, label it, etcetera.
912
01:06:25.910 --> 01:06:27.370
And so it's really cool to know
913
01:06:27.750 --> 01:06:32.970
that VIVO is being used in that context. Like, it just made so much sense for them to say,
914
01:06:33.315 --> 01:06:39.095
wow. We can actually do this online, and we don't have to retrain. And on batch learning was getting in their way.
915
01:06:39.395 --> 01:06:42.840
They had this system which was going at a 1000 miles an hour, just
916
01:06:43.240 --> 01:06:46.220
hundreds of thousands of data accumulated all the time. And
917
01:06:46.520 --> 01:06:59.095
batch learning was just, you know, it was just, again, annoying for them. They they having a system that will enable them to do all this online just made sense for them. And to know that you can do this at such a high amount of traffic, it was really cool and exciting.
918
01:06:59.570 --> 01:07:00.630
In your own experience
919
01:07:01.010 --> 01:07:08.150
of building the project and using it for your own work, what are some of the most interesting or unexpected or challenging lessons that you've learned in the process?
920
01:07:08.575 --> 01:07:14.355
I think I'm just gonna focus a bit on the human aspect here. But although I have been doing open source,
921
01:07:14.655 --> 01:07:17.395
you know, quite a bit, I've always had this
922
01:07:18.180 --> 01:07:20.120
approach where I probably
923
01:07:20.980 --> 01:07:24.635
work too much on new projects that I make myself rather than on
924
01:07:25.115 --> 01:07:39.430
existing projects. So I just rather just do my things myself rather than contribute just to executive stuff. And it's not always necessarily good, but it's just the way I work. And so Vivo is really the first open source project where I work for other people.
925
01:07:39.744 --> 01:07:41.525
So like probably many people,
926
01:07:41.904 --> 01:07:44.805
a lot of my open source work is I just work on it myself.
927
01:07:45.105 --> 01:07:50.600
And obviously, I you work in companies and where you probably have a review process and you work with other people, but
928
01:07:50.900 --> 01:07:51.880
this is the first
929
01:07:52.580 --> 01:07:55.320
open source project where I really work with a team of people.
930
01:07:55.985 --> 01:08:02.725
And it's fun. It's just really so much fun. Like, just a month ago, we actually got to meet altogether and to
931
01:08:03.105 --> 01:08:04.965
have this, like, informal reunion.
932
01:08:05.480 --> 01:08:06.620
So that was really fun.
933
01:08:07.160 --> 01:08:08.700
And you realize that,
934
01:08:09.320 --> 01:08:18.925
you know, after 3 years, there are ups and downs, and there's moments where you just do not want to work on Liv anymore and you want to you know, you have work, you have friends, girlfriends, whatnot.
935
01:08:19.385 --> 01:08:25.245
And so the only way to subsist as a open source project in the long term is to have multiple people working. So
936
01:08:25.610 --> 01:08:27.150
do an open source, and,
937
01:08:27.690 --> 01:08:34.990
you know, it's not realistic to do it on your own if you want something to be successful and to actually have an impact in the long term. So
938
01:08:35.515 --> 01:08:37.295
it's actually really important to
939
01:08:37.915 --> 01:08:38.415
just
940
01:08:38.875 --> 01:08:45.670
be nice and to have people around you who help you. And although not everyone contributes as much as I do or core maintainers
941
01:08:46.450 --> 01:08:57.655
do, people help a lot and they make things alive. Like, it's always a joy when I open an issue on GitHub, and I see that someone from the community has answered the question, and I don't have to do anything. It helps tremendously.
942
01:08:58.435 --> 01:09:09.425
Yeah. We've already talked a bit about some of the situations where online learning might not be the right choice, but for the case where somebody is going to use an online streaming machine learning approach,
943
01:09:10.065 --> 01:09:15.045
What are the cases where River is the wrong choice and maybe there's a different library or framework that would be better suited?
944
01:09:15.505 --> 01:09:30.335
Well, yeah. Again, honestly, I think that online learning is the wrong choice in 95% of cases. Like, you do not want to make the mistake to think that your problem is a online problem. You probably, most of the time have a batch problem that you can solve with a batch library.
945
01:09:30.715 --> 01:09:36.494
You know, I mean, scikit learn now, if you open it and you just learn it, it's always going to work reasonably well. So
946
01:09:37.020 --> 01:09:44.880
sometimes I would just go for that. 1 thing we do get a lot is people asking how you can do deep learning with river. So they want to train deep learning models online.
947
01:09:45.485 --> 01:09:51.664
So the answer is that we do have a sister library that is called Torchriver, and it's dedicated towards
948
01:09:52.445 --> 01:09:54.145
training Torch models online.
949
01:09:54.750 --> 01:09:58.690
So but again, that is a bit finicky at the moment and still need some work being done on it.
950
01:09:59.070 --> 01:10:05.875
But yeah. If you want to be doing deep learning and you want to be working with images and sound and, you know, structured data,
951
01:10:06.415 --> 01:10:11.315
River is not the right choice even online, and you probably have to be looking at PyTorch.
952
01:10:11.910 --> 01:10:19.450
As you continue to build and iterate on the river project, what are some of the things you have planned for the near to medium term or any
953
01:10:25.585 --> 01:10:34.560
roadmap. So it's a Notion page with a list of stuff we're working on. That 1 mostly has a list of algorithms to implement, and it's mostly there
954
01:10:34.940 --> 01:10:35.440
to,
955
01:10:35.900 --> 01:10:36.480
you know,
956
01:10:36.785 --> 01:10:41.605
make people know what we're working on and to encourage new contributors to work on something.
957
01:10:42.625 --> 01:10:49.040
So the few contributors we have just pick what they want to work on and, you know, just in general, all their preference.
958
01:10:49.580 --> 01:10:52.800
So for instance, me this summer, I've decided to work on
959
01:10:53.265 --> 01:10:55.685
online covariance matrix estimation. So
960
01:10:56.145 --> 01:10:59.685
if you actually learn online covariance matrix, it's kinda useful because
961
01:11:00.100 --> 01:11:09.655
it's very useful in financial trading. And if you have an inverse covariance matrix that you can estimate online, that unlocks so many other algorithms, such as Bayesian linear aggression,
962
01:11:10.915 --> 01:11:17.655
elliptic envelope method for unknown detection, Gaussian processes, and whatnot. So I think I'm still in the
963
01:11:18.050 --> 01:11:21.190
nitty gritty details of implementing algorithms and not necessarily
964
01:11:21.490 --> 01:11:23.110
applying them to stuff.
965
01:11:23.410 --> 01:11:26.230
I'm kinda counting on users to to do the applications.
966
01:11:27.005 --> 01:11:35.025
It just feels right at the moment. Now 1 thing I'm working on in the mid to long term is Beaver. So eventually, I want to try to
967
01:11:35.370 --> 01:11:40.350
spend less time in River and work on a tool I'm building called Beaver. So Beaver is
968
01:11:40.730 --> 01:11:43.070
a tool to deploy and maintain
969
01:11:43.764 --> 01:11:47.065
online learning models. So essentially an MLOps tool for
970
01:11:47.525 --> 01:11:49.625
an MLOps tool for online learning.
971
01:11:50.005 --> 01:11:50.744
So it's
972
01:11:51.140 --> 01:11:54.920
in its infancy, but it's something I've I've been thinking about a lot.
973
01:11:55.300 --> 01:11:57.480
So I recently gave a talk on it in Sweden.
974
01:11:57.940 --> 01:12:06.265
I've sketched a blog post and some slides where I tried to describe what it's going to look like. But the goal of this project is to create a
975
01:12:07.050 --> 01:12:10.190
very simple, user friendly tool to deploy a model,
976
01:12:10.650 --> 01:12:12.670
and I'm hoping that that is going to encourage
977
01:12:13.425 --> 01:12:24.470
people to actually use river and to use online learning because they're gonna say, hey. Okay. I can learn, but I can also just deploy the model and, you know, and both tools play nicely together. So
978
01:12:24.850 --> 01:12:30.150
yeah, the future of Vivo is to have Beevor and to have this reference tool to deploy online models.
979
01:12:30.565 --> 01:12:33.065
It's not going to be catered just towards River.
980
01:12:33.365 --> 01:12:37.945
The goal is to be able to, you know, run it with any model that can learn online.
981
01:12:38.480 --> 01:12:52.405
Well, for anybody who wants to get in touch with you and follow along with the work that you're doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest barrier to adoption for machine learning today.
982
01:12:52.864 --> 01:12:55.764
I'm always impressed by how much the field is maturing.
983
01:12:56.530 --> 01:13:03.750
I think that there's a clear separation now between regular machine learning, like business machine learning, I like to call it, and deep learning.
984
01:13:04.204 --> 01:13:06.945
I think those 2 fields are becoming my separate fields.
985
01:13:07.324 --> 01:13:11.985
So I've kind of stayed away from deep learning because I just not my
986
01:13:12.390 --> 01:13:13.210
cup of tea, but
987
01:13:13.750 --> 01:13:21.290
some very interesting in business machine learning. So getting things that I call it. And I think I'm impressed by how much the community
988
01:13:22.005 --> 01:13:26.425
has evolved in terms of knowledge. People are the average ML practitioner
989
01:13:27.045 --> 01:13:30.180
today is just so much more proficient than 5 years ago.
990
01:13:30.660 --> 01:13:32.920
And I think it's a big question of education
991
01:13:33.300 --> 01:13:34.120
and tooling.
992
01:13:34.900 --> 01:13:37.800
The tricky thing about an ML model when it's not deterministic,
993
01:13:38.340 --> 01:13:39.000
and so
994
01:13:39.375 --> 01:13:43.635
it's difficult to guarantee that its performance over time is going to be good,
995
01:13:44.175 --> 01:13:49.800
and let alone certify the model or convince stakeholders that they should adopt it. So
996
01:13:50.500 --> 01:13:53.560
in the real world, you don't just deploy a model and cross your fingers.
997
01:13:54.100 --> 01:13:56.840
So although we've gone past the
998
01:13:57.485 --> 01:13:57.985
test
999
01:13:58.845 --> 01:14:04.385
and R and D phase of the model, we are still not there in terms of deploying the model. And so
1000
01:14:04.720 --> 01:14:06.260
the reality is that there's
1001
01:14:06.560 --> 01:14:08.260
usually a feedback loop
1002
01:14:08.560 --> 01:14:10.420
where you monitor your model and
1003
01:14:10.880 --> 01:14:13.460
possibly retrain it, be online or,
1004
01:14:13.765 --> 01:14:16.425
you know, offline retraining. It doesn't matter.
1005
01:14:16.805 --> 01:14:21.225
And so I don't think we're really good at that right now. I don't think that we have great tools to
1006
01:14:21.605 --> 01:14:25.590
have human beings in the loop who work hand in hand with
1007
01:14:25.970 --> 01:14:29.590
machine learning models. So I think that tools like Progyny,
1008
01:14:30.050 --> 01:14:31.350
which is a tool to
1009
01:14:31.815 --> 01:14:41.275
have a user work hand in hand with an ML system by labeling data that the model is unsure about, they're crucial. They're game changers because they create real systems where
1010
01:14:42.330 --> 01:14:43.870
you care about,
1011
01:14:44.810 --> 01:14:47.710
you know, new data coming in, retraining your model,
1012
01:14:48.010 --> 01:14:49.390
having a human validate
1013
01:14:49.755 --> 01:14:54.715
predictions, stuff like that. So I think we have to move away from only having models that are
1014
01:14:55.115 --> 01:15:00.410
only having tools that aren't destined towards training model, but we also need to get better tools that,
1015
01:15:00.950 --> 01:15:05.830
you know, encourage you to monitor your model, to keep training it, to work with it, to
1016
01:15:06.845 --> 01:15:12.945
yeah, again, just treat machine learning as software engineering and not just as some research project.
1017
01:15:13.485 --> 01:15:38.489
Alright. Well, thank you very much for taking the time today to join me and share the work that you've been doing on River and helping to introduce the overall concept of online machine learning. It's definitely a very interesting space, and it's great to have tools like River available to help people take advantage of this approach. So thank you for all of the time and effort that you and the other maintainers are putting into the project, and I hope you enjoy the rest of your day. Oh, thank you. Thanks for having me. It was great.
1018
01:15:40.835 --> 01:15:53.920
Thank you for listening. Don't forget to check out our other shows, the Data Engineering Podcast, which covers the latest on modern data management, and the Machine Learning Podcast, which which helps you go from idea to production with machine learning. Visit the site at pythonpodcast.com
1019
01:15:54.940 --> 01:16:02.695
to subscribe to the show, sign up for the mailing list, and read the show notes. And if you learned something or tried out a project from the show, then tell us about it. Email hostspythonpodcast.com
1020
01:16:04.195 --> 01:16:05.015
with your story.
1021
01:16:05.360 --> 01:16:10.260
And to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.