ACERCA DE ESTE EPISODIO
Visit our site to listen to past episodes, support the show, join our community, and sign up for our mailing list.
SummaryIan Ozsvald and Emlyn Clay are co-chairs of the London chapter of the PyData organization. In this episode we talked to them about their experience managing the PyData conference and meetup, what the PyData organization does, and their thoughts on using Python for data analytics in their work.
Brief Introduction- Hello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.
- Subscribe on iTunes, Stitcher, TuneIn or RSS
- Follow us on Twitter or Google+
- Give us feedback! Leave a review on iTunes, Tweet to us, send us an email or leave us a message on Google+
- Join our community! Visit discourse.pythonpodcast.com for your opportunity to find out about upcoming guests, suggest questions, and propose show ideas.
- I would like to thank everyone who has donated to the show. Your contributions help us make the show sustainable. For details on how to support the show you can visit our site at pythonpodcast.com
- Linode is sponsoring us this week. Check them out at linode.com/podcastinit and get a $20 credit to try out their fast and reliable Linux virtual servers for your next project
- I would also like to thank Hired, a job marketplace for developers and designers, for sponsoring this episode of Podcast.__init__. Use the link hired.com/podcastinit to double your signing bonus.
- Your hosts as usual are Tobias Macey and Chris Patti
- Today we are interviewing Ian Ozsvald and Emlyn Clay about their work with PyData London, a group within the PyData organization. PyData London represents the largest Python group in London at ~2850 members, they hold regular monthly meetups for ~200 members at AHL near Bank and a yearly conference for around ~300 members. Last year, they and their sponsors raised over £26,000 to sponsor the development of core numerical libraries in Python.
On Hired software engineers & designers can get 5+ interview requests in a week and each offer has salary and equity upfront. With full time and contract opportunities available, users can view the offers and accept or reject them before talking to any company. Work with over 2,500 companies from startups to large public companies hailing from 12 major tech hubs in North America and Europe. Hired is totally free for users and If you get a job you’ll get a $2,000 “thank you” bonus. If you use our special link to signup, then that bonus will double to $4,000 when you accept a job. If you’re not looking for a job but know someone who is, you can refer them to Hired and get a $1,337 bonus when they accept a job.
- Introductions
- How did you get introduced to Python? – Chris
- What is the PyData organization, how does PyData London fit into it and what is your relationship with it? – Tobias
- In what ways does a PyData conference differ from a PyCon? – Tobias
- Does PyData do anything in particular to encourage users from disciplines that might not be aware of how much our community has to offer to choose the Python suite of data analysis tools? – Chris
- You have both spent a good portion of your careers using Python for working with and analyzing data from various domains. How has that experience evolved over the past several years as newer tools have become available? – Tobias
- For someone who is just getting started in the data analytics space, what advice can you give? – Tobias
- How can conferences like PyData help strengthen the bonds and synergies between the Python software community and the sciences? – Chris
- There are a number of different subtopics within the blanket categorization of data science. Is it difficult to balance the subject matter in PyData conferences and meetups to keep members of the audience from being alienated? – Tobias
- Data science is a young field and we’ve yet to see lots of examples of the successful use of data. How are London-based companies using data with Python? – Ian
- Is there a Python data science library you think needs a little love? – Emlyn
- Tobias
- Chris
- Ian
- Emlyn
The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA
EN ESTE EPISODIO
MOSTRAR NOTAS 🔗
TRANSCRIPCIÓN 🔗
NOTE
Transcription provided by Podhome.fm
Created: 7/6/2024 4:19:03 PM
Duration: 3791.163
Channels: 1
1
00:00:14.094 --> 00:00:15.635
Hello, and welcome to podcast.init,
2
00:00:16.240 --> 00:00:24.615
The podcast about Python and the people who make it great. You can subscribe to our show on Itunes, Stitcher, or TuneIn Radio, or you can add our RSS feed to your pod catcher of choice.
3
00:00:25.095 --> 00:00:37.030
You can follow us on Twitter or Google Plus, and please give us feedback. Leave us a review on iTunes to help other people find the show. Send us a tweet or an email. Leave us a message on Google Plus or in our show notes, or you can also join our new community. Visitdiscourse.pythonpodcast.com
4
00:00:38.850 --> 00:00:43.615
for your opportunity to find out about upcoming guests, suggest questions, and propose show ideas.
5
00:00:44.235 --> 00:00:48.570
I would like to thank everyone who has donated to the show. Your contributions help us make the show sustainable.
6
00:00:48.970 --> 00:00:52.270
For details on how to support the show, you can visit our site at python podcast.com.
7
00:00:53.610 --> 00:00:56.590
Linode is sponsoring us this week. You can check them out at linode.com/
8
00:00:57.405 --> 00:01:02.385
podcast in it and get a $20 credit to try out their fast and reliable Linux virtual servers for your next project.
9
00:01:02.925 --> 00:01:08.280
I would also like to thank Hired, a job marketplace for developers and designers, for sponsoring this episode of podcast.onnet.
10
00:01:08.740 --> 00:01:09.880
Use the link hired.com/podcastonnet
11
00:01:10.979 --> 00:01:12.360
to double your signing bonus.
12
00:01:12.899 --> 00:01:21.725
Your host as usual are Tobias Macy and Chris Patty. Today, we are interviewing Ian Oswald and Emlyn Clay about their work with PIData London, a group within the PIData Organization.
13
00:01:22.170 --> 00:01:39.354
PyData London represents the largest Python group in London at about 2, 850 members. They hold regular monthly meetups for 200 members at AHL, near bank, and a yearly conference for around 300 members. Last year, they and their sponsors raised over £26, 000 to sponsor the development of core numerical libraries in Python.
14
00:01:39.990 --> 00:01:52.395
So, Ian and Emlen, could you please introduce yourselves? Ian, why don't you go first? I'll be happy to introduce myself, but I'm just gonna correct you there. So, PyDays London is not just London's largest Python user group, but it's the UK's largest
15
00:01:53.015 --> 00:02:11.004
Python user group. And I think, Emlyn, tell me, I I think we're Europe's largest Python user group as well. I guess we're I'm not quite sure. I would say, I gave Tobias this little intro so that he could he could finish it. And I didn't wanna blow our horn too much. It's been a crazy 2 years building this meetup, building this group.
16
00:02:11.705 --> 00:02:20.020
And yeah. No. We are. We're Europe's biggest. We're certainly certainly Europe, the UK is and I think Europe's. Yeah. I'm gonna claim Europe until somebody tells me I'm wrong. We're not the US',
17
00:02:20.320 --> 00:02:26.475
the New York Python user group is definitely a lot larger than us. Yes, sir. But, on the world stage, we're not doing bad. So,
18
00:02:28.055 --> 00:02:32.610
who am I? So my name is Ian Oswald. I've been working in data science
19
00:02:33.090 --> 00:02:36.870
for 15 years. I'm an O'Reilly author with High Performance Python,
20
00:02:37.170 --> 00:02:38.150
and with Emlyn,
21
00:02:38.450 --> 00:02:44.515
we co chaired the first conference 3 years ago, Pilate to London, and then we got the meetup out of that.
22
00:02:45.135 --> 00:02:47.155
And I'm a long time Pythonister
23
00:02:47.614 --> 00:02:49.955
and c plus plus programmer before that,
24
00:02:50.255 --> 00:02:51.795
and international speaker.
25
00:02:52.560 --> 00:02:53.780
Emlyn, how about you?
26
00:02:54.400 --> 00:02:59.620
Right. So, yeah. My name is Emlyn Clay. I'm a bit of a chameleon in that my background is in pharmacology,
27
00:03:00.080 --> 00:03:03.095
which is sort of a lot of wet science, doing drug discovery.
28
00:03:03.635 --> 00:03:04.614
I got into,
29
00:03:05.954 --> 00:03:11.575
doing software sort of, tinkering about 10 years ago. I've been doing it professionally for about 6 years.
30
00:03:12.180 --> 00:03:15.720
Yeah. I, I co chair PyLab London, with Ian here,
31
00:03:16.260 --> 00:03:19.799
and I use Python in anger. I have many production systems now up there.
32
00:03:20.260 --> 00:03:20.760
And
33
00:03:21.724 --> 00:03:25.584
as well as Python, I write in another a number of other languages,
34
00:03:25.965 --> 00:03:28.944
as well, but it's certainly 1 that's close to my heart.
35
00:03:30.340 --> 00:03:32.440
So how were you each introduced to Python?
36
00:03:33.300 --> 00:03:33.960
Hey, Chris.
37
00:03:34.500 --> 00:03:35.000
So
38
00:03:35.460 --> 00:03:39.925
I started using Python in my first real job back,
39
00:03:40.405 --> 00:03:41.465
around 1999.
40
00:03:42.645 --> 00:03:43.525
It was somewhere around,
41
00:03:44.085 --> 00:03:45.465
2, 001 or so.
42
00:03:45.845 --> 00:03:57.065
At the time I was senior programmer for a c plus plus group in an artificial intelligence research company. And I was really proud of the speed of our c plus plus code doing things like, logistics optimization
43
00:03:57.365 --> 00:04:04.220
work. And 1 day I was given some Python code, a sax parser written in Python for some HTML,
44
00:04:05.000 --> 00:04:11.345
and asked to improve this parser for 1 of the NLP's, the natural language processing teams in our French office.
45
00:04:11.745 --> 00:04:13.365
And I remember looking at it disdainfully
46
00:04:13.745 --> 00:04:20.580
and thinking why would I need to write in this ridiculous scripting language, and I've got the power of c plus plus behind me and a team of 5 working with me.
47
00:04:21.060 --> 00:04:21.720
And then
48
00:04:22.259 --> 00:04:24.759
learning Python in the space of a day
49
00:04:25.139 --> 00:04:27.319
learning how to improve the sax parser,
50
00:04:27.620 --> 00:04:31.365
I suddenly realized I was more productive in a day's learning
51
00:04:31.825 --> 00:04:35.525
than I was after 5 years as a senior programmer with c plus plus,
52
00:04:35.825 --> 00:04:48.199
for text processing. And that was a bit of a wake up call. And then since then, I've never looked back. I use Python for almost all of my work, almost exclusively in the last 5 years for data science,
53
00:04:48.505 --> 00:04:51.965
falling back every now and again to use c plus plus when necessary,
54
00:04:52.745 --> 00:04:59.219
hardly touching any other languages and just just really, it's like a an over 10 year love affair with the Python language.
55
00:05:00.000 --> 00:05:00.979
And, Emlen?
56
00:05:01.840 --> 00:05:02.560
Right. So,
57
00:05:03.504 --> 00:05:08.965
I like to consider the first the time I got introduced to Python was the day that Excel broke.
58
00:05:09.824 --> 00:05:14.400
So at the time, I was analyzing doing some signal processing on ECG data.
59
00:05:14.940 --> 00:05:17.920
Unlike most life scientists, we don't really have any,
60
00:05:18.460 --> 00:05:20.720
formal training in doing any programming.
61
00:05:21.355 --> 00:05:23.755
So I was trying to find anything to sort of,
62
00:05:24.235 --> 00:05:28.095
get it working well. And I've been toying around with MATLAB for a little while,
63
00:05:28.795 --> 00:05:31.560
but didn't have access to the, computer,
64
00:05:31.940 --> 00:05:34.840
down the computer labs at King's College in London.
65
00:05:35.620 --> 00:05:36.919
And Python was available,
66
00:05:37.245 --> 00:05:42.865
so I could put it on my computer and play with it there, and that kind of, you know, spawned the interest there.
67
00:05:43.645 --> 00:05:47.039
So I started using it to do, biomedical analysis,
68
00:05:47.500 --> 00:05:51.439
statistical analysis on on various datasets I had access to.
69
00:05:52.974 --> 00:06:03.350
Yeah. So I started off with scripting, you know, just in that sort of way, and then from there, I mean, I, you know, the web got more and more impressive, so then span out into, into those sorts of things.
70
00:06:04.130 --> 00:06:10.150
But Python's always been particularly good as sort of a scientist toolbox because all the numerical libraries are particularly strong.
71
00:06:10.825 --> 00:06:14.445
And it feels a lot like MATLAB, so, you can almost forget
72
00:06:14.905 --> 00:06:17.005
that you are you are using it for that purpose.
73
00:06:17.545 --> 00:06:19.085
So yeah, that's how I got introduced.
74
00:06:20.370 --> 00:06:22.630
And what is the PyData Organization?
75
00:06:22.930 --> 00:06:27.190
And how does PyData London fit into it? And what is your relationship with it?
76
00:06:28.075 --> 00:06:41.120
Okay. So that's an interesting question, and I'm I'm not entirely sure what's, what is the PI Data Organization. I know that it started several years ago. It's a relatively young organization. It started back in 2012 in the USA.
77
00:06:42.460 --> 00:06:45.120
And, out of that it spawned a number of conferences. There have been 14 to date, I believe, and
78
00:06:49.965 --> 00:06:52.465
the first 1 that was run-in Europe was run by
79
00:06:53.060 --> 00:06:59.960
Emlen and myself and several of our colleagues 3 years ago. And so we hosted the first in Europe, last year there were a couple in Europe, and this year there will be 5 in Europe, and 5 in America. So it's growing rapidly.
80
00:07:04.685 --> 00:07:07.905
Emlyn, how would you define the organization of PI Data?
81
00:07:08.605 --> 00:07:15.080
Yeah. So I mean, you and I were there, I think, at the second ever 1, and that was attached to PyCon,
82
00:07:15.539 --> 00:07:16.599
US in 2013.
83
00:07:17.379 --> 00:07:28.505
Mhmm. This is where you and I met. I mean, I think they had 1 before that at the, the Google campus. Mhmm. And this was, as far as we're aware, it was started up as a as a counterpoint to,
84
00:07:29.470 --> 00:07:33.330
PyCon, which was more sort of general all sorts of topics that touch on Python.
85
00:07:33.630 --> 00:07:36.930
And PyData was focused on people who were using it for,
86
00:07:37.470 --> 00:07:37.970
doing
87
00:07:39.585 --> 00:07:41.365
data analysis, signal processing,
88
00:07:42.705 --> 00:07:46.725
statistics, things of that nature. Things that were much more sort of a data driven problem,
89
00:07:48.040 --> 00:07:52.860
And all of it, I guess, built around NumPy, which in some ways, NumPy is kind of its own
90
00:07:53.480 --> 00:07:54.540
entire ecosystem
91
00:07:54.840 --> 00:07:56.700
on top of in inside of Python.
92
00:07:58.085 --> 00:08:03.625
So, yeah. I think that's kind of what it was. It was trying to make it so that it was very much about numerical computation,
93
00:08:04.725 --> 00:08:05.785
parts than Python.
94
00:08:06.245 --> 00:08:07.289
And it's slightly,
95
00:08:07.830 --> 00:08:08.889
it sits alongside
96
00:08:09.270 --> 00:08:13.610
the older SciPy conference and the European Euro SciPy series,
97
00:08:13.975 --> 00:08:33.904
which are I would argue a little bit more academically focused and Pydata sits a little bit more on the industrial side. Both have industrial and academic speakers along, but I think both have a slightly different focus. I think it's quite nice to be able to mix the 2 sides, but accepting that, perhaps it's more interesting for a series to have a lot of academic contribution,
98
00:08:34.365 --> 00:08:37.950
and others for industrialists to share their their their,
99
00:08:38.510 --> 00:08:46.205
achievements and their problems, and the success stories of actually getting things shipped. Because that that alone is really quite a hard technical challenge,
100
00:08:46.585 --> 00:08:48.205
in a changing data world.
101
00:08:50.905 --> 00:08:51.405
And
102
00:08:51.785 --> 00:08:54.470
taking a brief diversion, you guys mentioned that
103
00:08:54.949 --> 00:09:06.824
the meetup has grown to a pretty significant number of members. And I'm just wondering if you can give some insight into how you managed to build and expand that community and how you manage to sort of keep it together and keep people interested.
104
00:09:08.084 --> 00:09:10.649
So we have I'm just looking at the screen now, 2, 872
105
00:09:12.070 --> 00:09:16.250
members. I'm really proud of that number. I'm proud every time it goes up, and I get very excited.
106
00:09:16.790 --> 00:09:17.110
I,
107
00:09:17.785 --> 00:09:18.345
I scraped,
108
00:09:18.825 --> 00:09:21.885
the growth numbers, and we generated a graph for the meetup.
109
00:09:22.265 --> 00:09:24.285
And then we have a graph that shows,
110
00:09:25.305 --> 00:09:27.325
I don't have to describe this. 2 lines,
111
00:09:27.710 --> 00:09:29.970
1 goes up at a certain rate until Christmas,
112
00:09:30.590 --> 00:09:37.405
a year a year and a bit back, and then it grows linearly, but a faster rate. We don't know what changed. It just it grew faster.
113
00:09:39.145 --> 00:09:40.205
Now exactly how
114
00:09:40.665 --> 00:09:54.685
we've helped make it grow, I'm not quite sure. Emlen, what are your thoughts on, why why the audience is growing so well? We know exactly how we've made this grow. So well, I think 1 of the big thing that attracts people, to Pilates London is that it's unashamedly,
115
00:09:55.065 --> 00:09:57.645
sort of, intermediate advanced level stuff.
116
00:09:58.105 --> 00:10:02.445
So it's really interesting talks whether, you know, the speakers are
117
00:10:02.990 --> 00:10:10.130
completely at liberty just to sort of go, here's the stuff I'm doing and here's all this interesting, you know, these ways fits together,
118
00:10:10.514 --> 00:10:23.529
which I think creates a fantastic driving force which has helped bring all these excellent speakers in. We've always made it sort of light and fun. We've we've had excellent sponsors. They've always been great at providing, you know, pizza and beer. It's very social
119
00:10:24.230 --> 00:10:24.730
and,
120
00:10:25.350 --> 00:10:32.425
I think that's that's kind of what sort of pushed it along. I think around us though, there's been the the general macro environment is that,
121
00:10:33.045 --> 00:10:35.225
Python has just become more and more important,
122
00:10:35.550 --> 00:10:36.610
for data scientists.
123
00:10:37.070 --> 00:10:40.450
And, data science as a job has become,
124
00:10:41.630 --> 00:10:44.210
furiously more in demand, especially in London
125
00:10:44.514 --> 00:10:45.954
over the past, you know,
126
00:10:46.595 --> 00:10:47.495
18 months,
127
00:10:48.195 --> 00:10:59.860
or so. There is there is certainly something interesting around the timing. Around 6 years ago at a hacker news event in London, I stood up on stage at the end and said to about the 500 people in the audience,
128
00:11:00.480 --> 00:11:15.480
I said, hey, I want to organize some kind of data related meetup. Who's with me? And about 30 hands gingerly went up. And I looked around and thought, wow, that's not not a huge number of people given 500 people who are coming to this Hacker News event.
129
00:11:16.019 --> 00:11:30.790
And then a couple of years later some data science groups started in London, and ours started about a year later. And we've I think there are 12 related data science events now from machine learning and deep learning through to visualization and Kaggle competitions.
130
00:11:31.250 --> 00:11:38.595
So there's a whole range of events now that have all sprung up and grown in the last 3 years. There's definitely a a local timing event in London. I don't really know,
131
00:11:38.975 --> 00:11:44.355
I don't know quite how that works. I know with the meetup, we're very keen to have a meetup every month
132
00:11:44.800 --> 00:11:48.560
sponsored by a good host, and so we've had Pivotal at the beginning.
133
00:11:49.279 --> 00:11:53.300
So they got us started, and then list the the fashion recommender
134
00:11:54.194 --> 00:12:06.200
us for another 6 months, and then we moved to AHL, a hedge fund, who've been sponsoring us for over a year now. And as Emeline says, we only have good speakers. We have no product sales pitches. It's all about good technical,
135
00:12:06.660 --> 00:12:08.760
data storage, the highs and the lows.
136
00:12:09.140 --> 00:12:12.040
We have a good diverse speaker set and audience.
137
00:12:12.894 --> 00:12:14.595
And, yeah, we we try to encourage
138
00:12:15.375 --> 00:12:17.714
interesting stories and a nice environment
139
00:12:18.334 --> 00:12:21.955
facilitated by beer and pizza, and it seems to work terribly well.
140
00:12:22.639 --> 00:12:26.180
Yeah. The rise in data science as a
141
00:12:26.639 --> 00:12:27.699
particular discipline
142
00:12:28.240 --> 00:12:28.740
definitely.
143
00:12:30.985 --> 00:12:42.640
I've definitely seen that happening in a lot of different places, and I can definitely see how that would contribute to the growth of a data oriented meetup group. And, also, as yours you're mentioning how there are a number of somewhat related,
144
00:12:43.260 --> 00:12:46.000
more niche meetups, but I think that having a
145
00:12:46.334 --> 00:12:48.915
sort of unifying group where people can
146
00:12:49.375 --> 00:12:51.555
get a, you know, cross discipline
147
00:12:51.935 --> 00:12:54.834
view of what other people are working on is definitely
148
00:12:55.240 --> 00:13:05.235
very interesting. I know that whenever I, you know, go to any meetups, it's very interesting seeing what people in other domains and other particular, subcategories of my, of
149
00:13:05.695 --> 00:13:07.154
my work are up to.
150
00:13:07.855 --> 00:13:16.460
So you mentioned that you are co chairs of the PyData London conference. And I'm wondering in what ways does a PI Data conference differ from a Picon?
151
00:13:19.400 --> 00:13:20.780
Oh, that's quite good. So
152
00:13:22.445 --> 00:13:29.425
generally, yeah. A lot of it focuses more on here is the particular data problem I tried to solve and here is the implementation I used.
153
00:13:29.820 --> 00:13:33.360
And I know that there are a lot of PyCon talks that are about that as well.
154
00:13:34.060 --> 00:13:35.040
But I guess PyCon
155
00:13:35.580 --> 00:13:37.839
focuses on so many broad things,
156
00:13:38.265 --> 00:13:43.404
that generally it's sort of a bigger mixed bag, whereas there's generally a sort of a a common narrative.
157
00:13:43.785 --> 00:13:53.459
Do you do you reckon, Ian, there's a common sort of pattern that the speakers, they come along and it's, you know, here's my problem, here's how we've implemented it, and here's what we're going on to in future, which,
158
00:13:54.315 --> 00:13:57.615
you know, is is kind of a common framework to all the talks that are done with PyData.
159
00:13:58.555 --> 00:14:02.735
Right. But the the focus is definitely at the PyData simply around the data
160
00:14:03.130 --> 00:14:05.470
rather than, for example, web development
161
00:14:05.850 --> 00:14:08.110
or back end system support.
162
00:14:08.810 --> 00:14:12.045
I certainly remember when I attended a EuroPython
163
00:14:12.345 --> 00:14:13.485
in 2, 010.
164
00:14:14.585 --> 00:14:15.325
I remember
165
00:14:15.705 --> 00:14:36.475
going to all the different tracks and listening to a lot of the talks there, and then getting to the end of the conference and being a little bit grumpy. And I spoke to a colleague and said, you know, hey. It's it's great that we've got quite so many talks on Django and general web development, but surely people are doing stuff with the data in their databases, not just putting it on the screen in in pretty boxes.
166
00:14:37.320 --> 00:14:57.240
And then, my friend said, well, you know, there's a community conference here that the way you fix this is that you do something technical next year on data. Oh 0, crikey. Yeah. Actually, it's probably time that I I did something and gave back rather than just consuming these tools that are provided for me. So the next year I proposed a high performance Python tutorial.
167
00:14:58.180 --> 00:15:04.005
I didn't know how to pitch it, and so, I crammed in as much as I could. And it turns out that,
168
00:15:04.404 --> 00:15:17.870
I thought I had 3 hours. I had actually 2 and a half hours of breaks, and so I crammed about 6 hours of material into 3 hours. I didn't let my students out of their room, and I did my bit to to make sure that everyone learned lots of high performance computing,
169
00:15:18.410 --> 00:15:20.750
as a precursor to dealing with data.
170
00:15:21.290 --> 00:15:30.654
And then after that, I sat back and thought, oh, I wonder if there are other Python conferences that deal with data rather than just just everything like PYcons and, the Europythons.
171
00:15:31.310 --> 00:15:35.570
And that's when I discovered Eurospine. So Eurospine is a bit more academically focused,
172
00:15:36.029 --> 00:15:40.275
and that's been running for, I think, 8 years now, moving throughout Europe.
173
00:15:40.995 --> 00:15:51.400
And then the Pydata series sprung up, in the last, 3 or 4 years. And definitely, we all of the Pydata, all the videos are online. You can find them, at PIData.org.
174
00:15:52.260 --> 00:15:57.255
All the videos are there, and they all focus on data processing. I'd say that's the core difference to
175
00:15:57.795 --> 00:15:58.295
Pycoms.
176
00:15:59.635 --> 00:16:06.200
And have you noticed a difference in terms of the sort of the the feel of the audience or the,
177
00:16:06.579 --> 00:16:12.920
meaning the demographics or just the general populace of people who attend a PyData conference versus a PyCon?
178
00:16:13.755 --> 00:16:18.654
Oh, that's an interesting 1. Yeah. Right. So Evelyn, what do you think? Right. So I guess,
179
00:16:19.435 --> 00:16:22.735
certainly we we don't feel like the audience has changed
180
00:16:23.550 --> 00:16:27.090
inside the group, that's been quite nice. So along the lines of,
181
00:16:27.790 --> 00:16:32.450
where 1 of the things that is the continued strength of Play Data is that the mix of people there
182
00:16:32.975 --> 00:16:34.035
is, is fantastic.
183
00:16:35.295 --> 00:16:37.875
Compared to a oh, compared to a PyCon,
184
00:16:38.655 --> 00:16:39.235
I guess
185
00:16:41.779 --> 00:16:48.519
what do you reckon? Like the sort of, almost the job spec of a person. There's a lot more, of course, there's a lot more data science people.
186
00:16:49.855 --> 00:16:59.290
Invariably, we have a lot of people who are, sort of former scientists who have come into programming. That, I think, is more common than having sort of career programmers,
187
00:17:00.470 --> 00:17:09.255
I think as our sort of mix. Yeah. I'd agree. And we certainly have, maybe a third of our audience here are data engineers. So people who deal with the data plumbing,
188
00:17:09.555 --> 00:17:14.695
getting in, storing it, and serving it up in efficient ways. Sure. I know that we have,
189
00:17:15.360 --> 00:17:16.820
I'm pretty sure our audience
190
00:17:17.120 --> 00:17:21.380
is sort of 40% PhD, 40% master's degree,
191
00:17:21.840 --> 00:17:22.915
and then the remainder
192
00:17:23.315 --> 00:17:26.695
are probably they probably have some level of higher education.
193
00:17:27.715 --> 00:17:30.375
I think comparing that to general Python conferences,
194
00:17:30.940 --> 00:17:32.480
you will have a a wider
195
00:17:32.780 --> 00:17:36.960
mix of academic backgrounds and a far wider mix of industrial,
196
00:17:37.820 --> 00:17:39.280
backgrounds. I know it's
197
00:17:39.904 --> 00:17:43.445
1 of the PyCons I went to in the US several years ago.
198
00:17:43.745 --> 00:18:00.665
I was incredibly impressed just to see lots of parents walking around with their children. And they brought their children because they were Raspberry Pi hackathons. And so you could turn up with your child, and then the child would start to learn programming in Python on a Raspberry Pi, and then they take it at home afterwards. And so not scientific at all, but educating
199
00:18:01.045 --> 00:18:03.145
the the next most available generation.
200
00:18:03.685 --> 00:18:09.230
And I thought that was just that, you know, blimey, that was very different to a typical scientifically led conference.
201
00:18:11.130 --> 00:18:12.510
And do you think that
202
00:18:12.890 --> 00:18:13.765
by virtue of
203
00:18:14.885 --> 00:18:16.665
a bit of more shared background
204
00:18:17.125 --> 00:18:22.425
in industry causes any difference in the social dynamics? Like, do you notice people at PI Data Conferences
205
00:18:22.965 --> 00:18:23.465
being
206
00:18:25.490 --> 00:18:33.190
do do you see them having an easier time sort of picking up conversations with their peers versus somebody at a PyCon because of the difference in background? Or
207
00:18:33.490 --> 00:18:42.215
I I think so. Yeah. I definitely I definitely know that's the case. I think because we're we're somewhat because we're, sort of, in many ways, a more specialized Python conference towards
208
00:18:42.610 --> 00:18:44.309
data analysis and data problems.
209
00:18:45.250 --> 00:18:49.350
Generally it's a good assumption that most of the people you're talking to have a familiarity
210
00:18:49.730 --> 00:18:50.230
with
211
00:18:50.615 --> 00:18:52.075
all the mix of things, statistics,
212
00:18:52.375 --> 00:18:58.075
machine learning. I mean, you can roughly assume that anyone you talk to at a pi data knows a bit about machine learning,
213
00:18:58.615 --> 00:18:59.434
which is nice.
214
00:19:00.049 --> 00:19:06.470
So yeah. No. That that has certainly been when it's very easy and fluid to get talking to people. We encourage a very,
215
00:19:07.650 --> 00:19:08.150
conversational,
216
00:19:08.725 --> 00:19:10.424
you know, discourse driven,
217
00:19:11.285 --> 00:19:11.785
style.
218
00:19:13.445 --> 00:19:28.855
You know, we even have we even have heckling in the meetings. Right? I mean we actually It's all about the heckling. Oh, I don't know whether this is just a UK thing, but there's something about having it so that it's really interactive. It's not just you know, turn up, sit down, watch the watch the show. It's, you know,
219
00:19:29.155 --> 00:19:33.335
thank you for coming along, you know, have a pizza and a beer, but get in the conversation. Right?
220
00:19:33.990 --> 00:19:35.770
Know, it's all about giving back and
221
00:19:36.150 --> 00:19:38.870
and that's kind of that's that. That's that, I guess, is,
222
00:19:40.070 --> 00:19:57.060
yeah. That's that's, I would say, is, yeah, kind of the the focus of what we've got there. Certainly 1 of the more difficult things at a general Python conference will be that you turn around, you say, hello, my name is Ian. Who are you? And you start a conversation with somebody and then you discover that you have a very, very different background. So
223
00:19:57.440 --> 00:19:57.940
possibly,
224
00:19:58.400 --> 00:20:09.660
say, in a PyCon US with 2, 000 500 people, it's quite possible you've got no shared background at all, no shared work interests. You're both at the conference for your own reasons, and it's quite hard to get some kind of overlap.
225
00:20:10.040 --> 00:20:15.825
And that's lovely and, it's interesting, but it's it's hard to find people who really care passionately about your core
226
00:20:16.285 --> 00:20:35.544
subject. And we have a similar problem with our Pydata meetups in that we have 200 people turning up every month. So our our regular meetups are the size of a small conference every month for free. So that's a heck of an achievement. But we do have people turning up with a high performance background or a natural language processing background or a time series econometric background,
227
00:20:36.404 --> 00:20:45.820
or they're data engineers, and they don't really care about what you do with the data. They just want to store it. So So 1 thing we started to experiment with is sections of the room during the breaks being designated.
228
00:20:46.760 --> 00:20:53.315
That corner, that's for visualization. And that corner, that's for machine learning. And then the bar, that's for general conversation. And that way,
229
00:20:53.695 --> 00:20:59.850
the the 200 people have a better chance of finding people who want to talk about their particular subject right now.
230
00:21:00.590 --> 00:21:03.930
And I I think that seems to be working really quite nicely.
231
00:21:04.470 --> 00:21:09.195
So I I think it's interesting that, the same problem happens at a different scale even at a conference,
232
00:21:09.575 --> 00:21:10.315
like PyCon.
233
00:21:11.895 --> 00:21:22.260
So does PyData do anything in particular to encourage users from disciplines that might not be aware how much our community has to offer to choose the Python speed of data do you have data analysis tools?
234
00:21:23.520 --> 00:21:26.080
I guess we we sort of do. Yeah. So we've,
235
00:21:26.885 --> 00:21:31.225
there are plenty of people who come along to Pydata who are part of the, the r user groups.
236
00:21:31.925 --> 00:21:36.200
I think we did a we did a plug for 1 of their conferences. Didn't we? Yes. Absolutely.
237
00:21:36.900 --> 00:21:42.235
The the IRRRL conference. Yeah. We plugged the last couple of them. And they've plugged us as well. And they've plugged us as well.
238
00:21:42.795 --> 00:21:47.775
You know, we chat to the big O people, who are, I mean, they use a fair chunk of Python.
239
00:21:48.235 --> 00:21:52.015
But then again, they tend to also be sort of, you know, high more high performance focused.
240
00:21:53.419 --> 00:21:56.460
Yeah, weirdly, I mean, we sometimes get comments in the meetup,
241
00:21:57.020 --> 00:22:01.039
there was a few meetups ago, someone put, there wasn't any Python discussed in this particular
242
00:22:01.340 --> 00:22:02.255
talk. And
243
00:22:02.795 --> 00:22:11.135
yeah, while we are we are sort of focused on Python because it's a popular language in this space and it's a good language to sort of, you know, put up on a slide and and read and understand.
244
00:22:11.970 --> 00:22:22.085
It's it's really it's very interesting because once you get to a certain level with with like if you need extra performance, then yeah, you need to drop down into c or maybe siphonize your code or something. Right?
245
00:22:22.865 --> 00:22:29.160
Or for others, there's compatibility things like you have to get it running on the JVM because that's, you know, the setup they've got.
246
00:22:29.800 --> 00:22:39.635
So there is a fair bit of overlap that occurs naturally anyway with the people who come along. So I'd like to think that we sort of we pivot around both the Python bit and the data bit,
247
00:22:40.255 --> 00:22:41.555
of our of our group.
248
00:22:41.935 --> 00:22:42.175
So,
249
00:22:42.815 --> 00:22:58.665
we don't have anything in particular. Sorry. Well, 1 of the points we made for the conference is that Python is definitely the hook, and so we want the majority of speakers to be talking about the use of Python. But it's not just about Python. We need people who are using R or Julia or SPSS or MATLAB
250
00:22:59.370 --> 00:23:07.070
or any of or Java, I guess, for the big data side. People coming into the community with other ideas, other impressions, other backgrounds, other use cases
251
00:23:07.404 --> 00:23:17.690
so that we can have a nice cross discipline sharing of experience at the conference. Otherwise, you you end up in danger of being in a little echo chamber in your own closed ecosystem,
252
00:23:18.070 --> 00:23:28.525
everyone patting each other on the back and saying, oh, yes. Well, you've got the best techniques and kind of ignore everybody else. And, you know, that's that's not how the rest rest of the world works. So we do our best to be very inclusive there.
253
00:23:29.305 --> 00:23:41.915
And, yeah, sure, we have had people complain that, that, 3 months in a row, we've had other languages being discussed alongside Python. How about just an old Python night? And yeah. Sure. We normally just have Python nights, for our monthly meetups.
254
00:23:42.295 --> 00:23:45.115
But I'm very keen to get other people in from other backgrounds.
255
00:23:45.655 --> 00:23:48.909
Spread the word and see who comes in with a contrarian opinion.
256
00:23:49.450 --> 00:23:54.110
Hopefully, upsets, some of the natural order of things and changes some of the the normal thinking,
257
00:23:54.490 --> 00:24:13.270
because that's how new ideas get tested, and that's what science is all about. I think that's fantastic. And and I've I've certainly, if you've been listening or or even if you haven't on this podcast, I've been a huge proponent of people need to break out of their sandbox and look at what other communities are doing, you know, because Absolutely.
258
00:24:13.715 --> 00:24:20.775
I I feel very strongly that, you know, I came from Ruby and learned Python and said, oh my god. You guys have built some amazing toys over here.
259
00:24:21.155 --> 00:24:28.669
So, you know, and vice versa. So I I think it like, being it's great to be a Python fan, but
260
00:24:29.049 --> 00:24:35.665
be aware of what else is happening outside the boundaries of the Python community. It's the only way we're gonna continue to learn and evolve.
261
00:24:36.605 --> 00:24:47.970
Yeah. Oh, definitely. Definitely. I mean, I don't know whether it's, whether we're just fortunate with the people we have or maybe it's the whole the focus on being or focusing on data or that a lot of our members are sort of ex scientist.
262
00:24:48.385 --> 00:24:55.345
But everyone is very much is very critical of everything. Everything they use, all the tools they use and things of that nature. So,
263
00:24:56.065 --> 00:24:56.965
yeah, I think
264
00:24:57.620 --> 00:24:59.000
the sort of the
265
00:24:59.300 --> 00:25:08.405
the reason why Ian and I sort of try and encourage that in our group, is simply because that's that's sort of the background we come from. But you're quite right, it could I mean, there's there's
266
00:25:08.865 --> 00:25:10.645
there's no winners in a flame war.
267
00:25:11.265 --> 00:25:11.765
So
268
00:25:12.980 --> 00:25:18.120
it it makes a lot of sense to be able to discuss which which thing is better, what makes it objectively better.
269
00:25:19.300 --> 00:25:23.625
Because I don't know, I don't think I don't know whether you can make it entirely as a pure Python guy.
270
00:25:24.005 --> 00:25:33.530
I think there is something to be said for if you can just if you're a little bit polyglot, like have 1 language you're really strong in and then know a smattering of the others. You can be a lot more dangerous
271
00:25:33.990 --> 00:25:35.930
than than just knowing 1 language.
272
00:25:37.030 --> 00:25:41.290
Yeah. And it's very rare that you come across a code base that has only 1
273
00:25:41.675 --> 00:25:45.295
language in it or not necessarily a single code base. But, you know, a particular
274
00:25:45.915 --> 00:26:03.635
application environment, you know, even if you're just doing web development, you're guaranteed to be running up against Java script. Or if you're doing data you know, big data analytics, you're bound to come in to come across Java at some point. Or, you know, if you're doing a lot of distributed computing, chances are you're gonna come across something either Java or Erlang.
275
00:26:04.095 --> 00:26:05.635
So being able to at least
276
00:26:06.815 --> 00:26:21.585
understand how to parse those languages to be able to figure out what's going on when you come across an edge case is definitely important. So even if you don't necessarily write anything in other languages, having some familiarity in what other languages are capable of and why they're used in those cases and
277
00:26:21.885 --> 00:26:27.460
how they operate under the covers a little bit is definitely valuable and makes you much more valuable as an engineer.
278
00:26:28.080 --> 00:26:32.580
Sure. Sure. I mean, to to Chris's point about, yeah, the crossover with something like Ruby.
279
00:26:32.960 --> 00:26:39.205
I think there's a fair there's there's there's a fair chunk of that involved with, with things like Chef. Right? And while,
280
00:26:39.745 --> 00:26:41.925
you'll have, Python developers who themselves,
281
00:26:42.865 --> 00:26:44.065
aren't probably using,
282
00:26:44.490 --> 00:26:48.990
Ruby as their main scripting language for data analysis. They're probably using something like Chef or Puppet,
283
00:26:49.690 --> 00:26:54.835
to, you know, to orchestrate all the servers because that becomes a concern when you're plumbing things together.
284
00:26:56.174 --> 00:27:00.034
Yeah. Absolutely. That's my day job actually, writing Python and Chef.
285
00:27:00.660 --> 00:27:03.640
Very good. Very good. I'm I'm more of an Ansible man myself.
286
00:27:04.580 --> 00:27:07.640
But, I have had on occasion to play with Chef and Puppet.
287
00:27:08.025 --> 00:27:15.885
Yeah. Need to look at the interval. Sorry, Tobias. I was just gonna say I'm a salt stack. I'll do when whenever I can be. Oh.
288
00:27:17.960 --> 00:27:28.115
So you've both spent a good portion of your careers using Python for working with and analyzing data from various domains. And I'm wondering how that experience has evolved over the past several years as newer tools have become available.
289
00:27:29.455 --> 00:27:32.274
Oh, well, so I've been using Python for,
290
00:27:32.735 --> 00:27:41.980
I think, 14 years, something like that now. So what would you what would you have started on here? What version is you oh, it's it's and, oh,
291
00:27:42.365 --> 00:27:45.585
I don't know. What was what was Python 14 years ago?
292
00:27:45.885 --> 00:27:49.024
I have no idea. I I know Wasn't even a major number, mate.
293
00:27:49.565 --> 00:27:55.630
It was probably like 0 No. No. No. It was well it was well established. There was at least 1 book in the local Waterstones bookstore.
294
00:27:56.410 --> 00:28:02.005
I remember because, I I just thought, well, it's it must be legitimate because there's at least 1 book out for it.
295
00:28:03.025 --> 00:28:06.085
Now, I know that that was that was sometime around
296
00:28:06.509 --> 00:28:09.330
the unification of the math libraries,
297
00:28:10.429 --> 00:28:10.929
into
298
00:28:11.230 --> 00:28:13.490
NumPy by Travis Oliphant.
299
00:28:14.335 --> 00:28:21.395
So that was the beginning of the scientific community coming together with a shared base layer of numeric tools.
300
00:28:22.160 --> 00:28:27.140
And it was you know, that was way back before pandas and scikit learn and the like. I guess,
301
00:28:27.919 --> 00:28:29.934
1 thing I've seen evolve is the the the stack,
302
00:28:30.495 --> 00:28:33.715
massively improving. So we had NumPy with, homogeneous,
303
00:28:34.255 --> 00:28:36.434
arrays of data and then pandas,
304
00:28:36.815 --> 00:28:37.875
allowing heterogeneous,
305
00:28:38.654 --> 00:28:40.150
arrays to be put together,
306
00:28:40.450 --> 00:28:42.150
in an Excel like spreadsheet.
307
00:28:42.450 --> 00:28:44.710
And scikit learn making it super easy,
308
00:28:45.170 --> 00:28:47.350
to do your machine learning. And, matplotlib
309
00:28:47.650 --> 00:28:50.525
growing up, quite nicely in to a very powerful,
310
00:28:51.065 --> 00:28:54.365
a little bit fiddly at times, but a very powerful plotting library.
311
00:28:55.545 --> 00:29:09.125
I think 1 of the to me, 1 of the most interesting things has been the disconnect with some of the clients. So I run my own consultancy, and I've been consulting for over 10 years. So I've spoken with a lot of clients. And when they see data science and Python,
312
00:29:09.745 --> 00:29:10.565
on the rise,
313
00:29:10.865 --> 00:29:14.245
they they kinda match that with a lot of the publicity,
314
00:29:14.545 --> 00:29:18.240
partly from the big data world and partly from companies like IBM.
315
00:29:18.620 --> 00:29:23.440
And they begin to assume that because it's AI in data science,
316
00:29:23.980 --> 00:29:25.705
probably because, say, Google
317
00:29:34.424 --> 00:29:35.565
conversational level,
318
00:29:36.090 --> 00:29:45.105
freely available, and they can predict human buying habits, and you could knock out a prototype in the space of a week to improve spending patterns and advertising.
319
00:29:45.805 --> 00:30:01.544
And I find it I kinda it's kind of interesting. It's, it's a it's a bit of a tricky conversation to calm clients down there, but it's lovely that clients have bought into the fact that Python and the data science community is so well established that this kind of thing must be possible by now, surely.
320
00:30:02.005 --> 00:30:05.945
So, that's an interesting evolution, I think, from tools used by scientists
321
00:30:06.245 --> 00:30:11.330
through to clients, assuming that some of these things just kind of solved problems now when and they're not yet.
322
00:30:12.110 --> 00:30:14.930
I would say I have forgotten the the question initially,
323
00:30:15.710 --> 00:30:17.410
positive. Yeah. Was I monologuing?
324
00:30:18.115 --> 00:30:24.375
Oh, dear. So I'm just wondering how in your in the length of your career of using Python for data analysis,
325
00:30:24.835 --> 00:30:26.535
I'm wondering how your experience
326
00:30:26.900 --> 00:30:30.040
has evolved over the past several years as newer tools have become available.
327
00:30:31.060 --> 00:30:36.120
So 1 of the big experiences that I've had is that my continued use of R
328
00:30:36.875 --> 00:30:37.375
decreases,
329
00:30:38.394 --> 00:30:44.654
over time. I think about the time when, you know, pandas came in and gave us a really decent data frame object,
330
00:30:45.100 --> 00:30:48.860
that was a big use case for for things that I was doing. We we're just organizing,
331
00:30:49.660 --> 00:31:00.225
clinical trial data sets, things of that nature. If you're doing a sort of a small trial that, you know, you've got a couple gigabyte file, that kind of thing, it can all fit in memory. You can if you can do it with something like Pandas, it's
332
00:31:00.525 --> 00:31:06.080
very very quick and very expressive to get stuff done. Matplotlib has become easy and easier to use,
333
00:31:07.180 --> 00:31:07.680
and
334
00:31:08.300 --> 00:31:11.040
let's see. Seaborn came out a few years ago,
335
00:31:11.425 --> 00:31:15.445
and that just just by importing Seaborn, all of your graphs got prettier,
336
00:31:15.745 --> 00:31:18.565
which was a really nice effect I found.
337
00:31:18.945 --> 00:31:22.529
So it's generally I find myself relying less and less on that. I
338
00:31:23.149 --> 00:31:31.285
for for all of my sins have done lots and lots of MATLAB, because of signal processing of you know, biomedical signals and things of that nature.
339
00:31:31.665 --> 00:31:33.925
And I've become less and less,
340
00:31:34.545 --> 00:31:39.570
attached to it, but I still, you know, I'm still stuck on there. So I see a fair chunk of room,
341
00:31:40.110 --> 00:31:43.650
where I can move on in future. But broadly, what's happened is, I've been homogenizing
342
00:31:44.110 --> 00:31:45.410
onto the Python platform,
343
00:31:46.285 --> 00:31:51.345
because it's been expanding to fill my needs, which has been, it's been lovely to see.
344
00:31:51.645 --> 00:31:53.505
And I've only been using it for
345
00:31:55.159 --> 00:31:56.620
6 years or so,
346
00:31:57.000 --> 00:32:07.195
so much shorter time, but in that time, you know, it's become very strong indeed for for what I need it for. I guess that point about homogenizing onto the Python stack is the 1 of the key ones here. So
347
00:32:07.495 --> 00:32:13.909
rather than having languages that are specialized at certain points, of the the scientific problem spectrum.
348
00:32:14.289 --> 00:32:21.215
With Python, we get a language and a toolset that means we can go from taking data kind of from anywhere
349
00:32:21.675 --> 00:32:25.935
and then doing the necessary monging on that data to turn it into something useful,
350
00:32:26.235 --> 00:32:30.320
modeling it, visualizing it, exporting it, and putting it into production,
351
00:32:30.700 --> 00:32:49.870
under monitoring, deployed wherever you need it, and you can do it all in 1 language. And that means then we although we may not have the best language and the best libraries for any 1 particular part of that process, you've only got 1 language to keep in mind. So you've got a a bunch developers who are just dealing with 1 conceptual model, 1 language, 1 syntax,
352
00:32:50.250 --> 00:33:03.920
and that really eases development and deployment and debugging and support and all of those things. And I think really that's the the main strength that I see of Python evolving in the data science world. You can kind of do all of it just in the 1 ecosystem.
353
00:33:07.534 --> 00:33:14.015
And so for somebody who's just getting started in the data analytics space, what advice can you each give about, you know, how to
354
00:33:15.100 --> 00:33:22.745
you know, what what are the best things to learn first, how to establish themselves, or, just any general advice as far as, you know, doing data analytics
355
00:33:23.225 --> 00:33:30.045
itself. Oh, sure. Well, I'm gonna dive in there because because I've given a couple of keynotes talking about this kind of thing.
356
00:33:30.585 --> 00:33:30.825
So,
357
00:33:31.545 --> 00:33:33.470
it was my honor to give a couple of keynotes,
358
00:33:33.770 --> 00:33:35.870
PyCon Ireland, PyCon Sweden,
359
00:33:36.410 --> 00:33:37.150
and the Budapest
360
00:33:37.610 --> 00:33:44.805
Business Intelligence Conference over the last year and a half. And 1 of the things I talked about was helping people, particularly engineers,
361
00:33:45.745 --> 00:33:50.470
move on with developing data science products and getting successful products shipped.
362
00:33:51.090 --> 00:33:51.830
And so
363
00:33:52.290 --> 00:34:11.470
1 of the things that I think people get a bit hung up about, and forget is that you can go an awful long way with some really simple data science work. And by really simple, I mean, drawing graphs and doing a bit of filtering on your data. There's so much you could do. That's partly why Tableau is so powerful. You can do an awful lot just by drawing your data
364
00:34:11.770 --> 00:34:16.035
and cleaning it up. So you want clean data that you can visualize and explain,
365
00:34:16.415 --> 00:34:24.755
and then maybe you can do a bit of statistical modeling or some machine learning on it using tools like Pandas and scikit learn. Then you do some more visualization.
366
00:34:25.440 --> 00:34:31.140
And then when you've got something that's really robust, then maybe you want to get it deployed and shipped out there.
367
00:34:31.520 --> 00:34:34.020
And if you want to practice, then something like
368
00:34:34.444 --> 00:34:57.125
Kaggle, the machine learning competition site, that's an ideal place because there are lots of competitions from simple to quite complex with an open forum full of solutions and discussions about how people have improved things. So there's quite a wealth of material out there which helps people get started. But the 1 of the biggest things I've seen people get hung up on is the need to do something really, really clever and cutting edge like a deep learning spark distributed solution,
369
00:34:57.589 --> 00:35:13.435
when actually you can get away with putting it in a small data frame in RAM and running a logistic regression classifier on it and visualizing it in 2 dimensions. And maybe it turns out that's just as good. And if that's a robust solution, then great. Just do that.
370
00:35:15.415 --> 00:35:16.700
Yeah. I've definitely,
371
00:35:17.000 --> 00:35:21.420
particularly in the big data space, seen a lot of mention of people saying how,
372
00:35:22.280 --> 00:35:35.589
you know, people are far people are far too likely to get carried away in the approach to a given problem where they see they see a data problem. They say, oh, I'm gonna throw Hadoop at it when all you really need is Pandas and your laptop. So
373
00:35:36.470 --> 00:35:48.065
Well, it's worth definitely remembering that, I mean, my laptop has 4 cores and 16 gigabytes. But I can go to Amazon, and for a couple of bucks an hour, I can rent a machine with 50 or 60 cores
374
00:35:48.490 --> 00:35:50.350
and hundreds of gigabytes of RAM.
375
00:35:51.130 --> 00:35:56.670
And as long as my problem fits into that, and it turns out you fit an awful lot of data into a couple of 100 gigabytes,
376
00:35:57.235 --> 00:36:14.635
then I have no deployment problems. I don't need to hire someone to run a Spark deployed environment for me. I can iterate really quickly. And because data is not being thrown around amongst the number of machines, I get a solution back really quickly. So I, as an individual or in a small team, can iterate really quickly on my ideas.
377
00:36:15.015 --> 00:36:26.450
And I I'm a huge proponent of small to medium data just keeping it on 1 machine and going as far as you can with that before having to worry about a big data solution and the complexity that that involves?
378
00:36:27.230 --> 00:36:37.875
Well, I think this is a subproblem of the general problem that we we all face as technologists and that we see shiny toys and we wanna play with them. Right? I mean, it really is that simple, isn't it?
379
00:36:39.135 --> 00:36:50.665
Yeah. I can actually agree on that point. Well, right. No 1 got fired for suggesting a big data solution from 1 of the big providers, but it doesn't mean that it's you know, pragmatically as an engineer, it doesn't mean that it's the right solution.
380
00:36:52.405 --> 00:36:58.985
So how can conferences like PyData help strengthen the bonds and synergies between Python software community and the sciences?
381
00:37:00.650 --> 00:37:03.309
I I think it's it gives them it gives them that niche,
382
00:37:03.690 --> 00:37:05.150
to, to work in.
383
00:37:06.225 --> 00:37:08.885
I mean speaking of, you know, to the fact that we've
384
00:37:09.185 --> 00:37:11.685
such a large proportion of the PI data members
385
00:37:11.985 --> 00:37:12.385
are,
386
00:37:13.025 --> 00:37:15.525
are are either scientists or ex scientists
387
00:37:15.980 --> 00:37:17.200
who are working in industry.
388
00:37:18.540 --> 00:37:26.645
It gives them an opportunity to focus on their particular problem domain with Python just being, you know, the tool that we all sort of, group around.
389
00:37:28.145 --> 00:37:37.210
And I think out of any of the languages, Python has, I think, the strongest, especially for your casual programmer. Speaking as someone who comes from, life sciences where
390
00:37:37.589 --> 00:37:44.125
there are very few people who actually do a lot of programming. A lot of our problems are, you know, heavily computational. They're more,
391
00:37:44.505 --> 00:37:47.645
sort of wet work stuff, you know, getting the actual sensor
392
00:37:47.945 --> 00:37:52.519
on properly and observing a, you know, very large sort of macro change.
393
00:37:54.099 --> 00:38:01.355
Then, yeah, Pylator helped a great deal in that respect. And before that, I mean, Pearl did a great job. I mean, if you're looking at, in the bioinformatics
394
00:38:01.655 --> 00:38:04.234
space, Perl's still very much a king in that area.
395
00:38:04.535 --> 00:38:08.955
But I'd argue for, sort of, anyone who's got a casual interest in doing some programming,
396
00:38:09.320 --> 00:38:12.380
there isn't much better language to advise than Python.
397
00:38:12.760 --> 00:38:13.500
So consequently,
398
00:38:14.120 --> 00:38:17.260
yeah, PlayData is is very much that bridge,
399
00:38:17.835 --> 00:38:25.615
to to help it go 2 ways. You have people who are career programmers who wanna learn more about problem domains, and you got people with problem domains who wanna learn more about programming.
400
00:38:27.050 --> 00:38:28.569
Certainly, we see at the conference,
401
00:38:28.890 --> 00:38:32.030
a lot of industrial users of Python being,
402
00:38:32.490 --> 00:38:46.460
they kind of exist under the term data engineer at the moment. So regular engineers who know how to move data around, and then they're trying to do more interesting things with their data. But if they haven't got a data science background, so they haven't got a PhD or masters in those kind of,
403
00:38:46.940 --> 00:38:47.839
numerous subjects,
404
00:38:48.140 --> 00:38:57.275
then maybe they're wondering, how do they go and do this? But at the conference, they can meet other industrialists who are doing this, and certainly they can meet academics who are working on these kind of problems.
405
00:38:57.655 --> 00:39:03.289
So the conference gives, gives a great ground for cross pollinating different ideas and different solutions.
406
00:39:03.589 --> 00:39:23.510
But rather than just having problems that are presented as being interesting problems that could be solved, instead you've got a driver of, hey, we're failing. We've got this huge dataset. We wanna do something with it, and it's not working. Could someone help us? We've got some money. We've got a strong desire to do something. Can someone come and help us with the problem? I think that's a really interesting driver to bring people together.
407
00:39:23.890 --> 00:39:28.950
And certainly via the PI data, we've been reaching out to other groups here in the UK.
408
00:39:29.645 --> 00:39:34.065
1 of my particular pushes this year is into the Royal Statistical Society.
409
00:39:34.445 --> 00:39:40.820
So the RSS is a couple of 100 years old. We have a couple of members of the RSS in our PI Data group, but not many.
410
00:39:41.200 --> 00:40:04.275
And I know that the Royal Statistical Society have been trying in the last year to get more involved in the big data community regardless of language, just just large datasets in general, to see if they can find interesting ways to apply some of their knowledge into these newer industrial areas that are growing up. And so I'm reaching out into the RSS and trying to get, members to come along and speak and attend the conference,
411
00:40:04.735 --> 00:40:09.395
as another way of just bringing in people with another diverse different set of backgrounds, different set of opinions,
412
00:40:09.775 --> 00:40:12.640
to come and join the conversation, see if we can get some idea sharing
413
00:40:13.920 --> 00:40:19.715
going on. And so there are a number of different subtopics within the blanket categorization of data science. Is it
414
00:40:21.795 --> 00:40:24.935
to conferences and meetups to keep members of the audience from being alienated?
415
00:40:25.955 --> 00:40:28.775
Yeah. No. That that is a really tricky area.
416
00:40:29.760 --> 00:40:31.140
So we did a survey,
417
00:40:31.520 --> 00:40:34.420
well, I say I say we. Ian sent out the survey,
418
00:40:34.880 --> 00:40:37.280
to all of our members to get a, you know, handle on,
419
00:40:38.275 --> 00:40:45.494
the kind of things that they they wanna see more of or what they're involved in. And why? Because we're we're data geeks. We want surveys. We want all the answers.
420
00:40:45.880 --> 00:40:56.555
So we had I mean, we know we know something like a third of the people there. It's London. Right? There's a lot of people involved in the financial services. So people who are in quant finance or you're doing the data plumbing behind it,
421
00:40:57.015 --> 00:40:58.315
all that kind of jazz.
422
00:40:58.615 --> 00:41:06.589
But almost a third of the, responses were sort of other, So it's a really long tail to all these different, you know, domains that they're in.
423
00:41:07.289 --> 00:41:16.825
So, yeah, it's it's kinda tricky. I mean, what we I guess what we keep doing is we keep relying on the fact that the meetup board gives us, you know, commentary back as to how they think we're doing,
424
00:41:17.340 --> 00:41:18.860
they they rate the,
425
00:41:19.340 --> 00:41:25.360
the the meeting to see how they felt it went. And we do our best to, you know, reply to that and, you know, adjust it accordingly.
426
00:41:25.835 --> 00:41:31.615
I guess because we are fairly data driven, we can also sort of categorize in the rest of it the the talks that we've done.
427
00:41:32.075 --> 00:41:53.180
And that gives us a fairly good idea as to whether we're sort of, you know, hitting all the, the the various, areas. But I guess even if someone isn't in a particular problem area, they can still take something home from it. Like, Yeah. Yeah. Emily, you're making this sound terribly principled. I don't think we're being that principled about it. We just keep changing the tune every month. We just keep trying to get different people in different topics
428
00:41:53.480 --> 00:42:14.240
talking. And I think it's that it's that diversity in the talks that matters. Right. But what we don't do, like, for instance, I do a a fair chunk of stuff in in the biomedical area, but I don't deliberately try and get lots of bike and biomedical speakers up on stage. And similarly, you do a lot of machine learning. Right? Mhmm. So it's, you know, we don't see it dominated by any of the 1 particular forces.
429
00:42:15.020 --> 00:42:15.520
So,
430
00:42:16.315 --> 00:42:27.760
yeah. No. Right. But but part of the reason for that is there are other groups in London. I mean, there's there's 12 data related machine learning type groups right now. There's 1 specialized in deep learning. There's another specialized in visualization.
431
00:42:28.060 --> 00:42:29.520
There's 2 or 3 in visualization.
432
00:42:29.900 --> 00:42:37.234
There's another 1 on text analytics. And so we've got groups that really go deep into each of those areas, and I think part of the role of PIData
433
00:42:37.775 --> 00:42:43.954
is to say that with Python and some other languages, you can do an awful lot of work with data, and here's a wide variety
434
00:42:44.260 --> 00:42:47.720
of examples of how data's being used. And, hey,
435
00:42:48.020 --> 00:43:02.250
all the speakers are intelligent. The audience is intelligent. Everyone's gonna pick up something new every month, so come along and learn something new. And at worst, you have a beer with some friendly, interesting people, and you go to the pub afterwards. So I think that's that's the that's the main thing that happens.
436
00:43:03.589 --> 00:43:11.815
So data science is still a young field, and we've yet to see lots of examples of the successful use of data. How are London based companies using data with Python?
437
00:43:12.914 --> 00:43:21.610
So I've definitely got ideas there. But, Emlen, do you have anything you wanna throw in? So particular companies have been using it? I mean, I guess we have the speakers who've been up sort of recently,
438
00:43:22.150 --> 00:43:33.355
give us a good idea of how they've been using it. So you know, List who there's often you know 6 or 8 of those engineers coming along to the meet ups every single time. 1 of our committee members works at List.
439
00:43:33.780 --> 00:43:36.280
They use it for doing all sorts of really fancy,
440
00:43:37.220 --> 00:43:38.040
image processing
441
00:43:38.500 --> 00:43:40.685
analysis to try and match up, you know,
442
00:43:41.165 --> 00:43:46.785
if this blouse is on this website, is it the same blouse as the 1 on that website and those sorts of things. So that's fascinating.
443
00:43:47.565 --> 00:43:56.590
They the the chaps at Deliveroo who is basically like a UK version of the number of sort of, fast, I say fast food restaurant food delivery services.
444
00:43:57.435 --> 00:44:03.055
They use it to to to plan their fleet so that they can, you know, get optimized delivery
445
00:44:03.515 --> 00:44:07.055
for, you know, products which spoil in a very short period of time.
446
00:44:07.740 --> 00:44:09.680
Yeah, there are there are so many
447
00:44:10.140 --> 00:44:12.220
who are in London and doing this that,
448
00:44:12.780 --> 00:44:17.115
yeah, could reel off more and more and more. But even So I think it's worth, maybe mentioning,
449
00:44:17.654 --> 00:44:20.635
our current host for the meetup, AHL, the hedge fund.
450
00:44:21.095 --> 00:44:27.990
They're an interesting story in that several years ago, they had a heterogeneous environment. They had people using
451
00:44:28.529 --> 00:44:32.365
R and Python and MATLAB and a whole mix of research tools.
452
00:44:32.825 --> 00:44:33.885
And they centralized
453
00:44:34.185 --> 00:44:41.910
on Python, and they hosted the London Financial Python user group years ago ago, after they've switched into this pure Python mode.
454
00:44:42.210 --> 00:44:46.630
And the reason they've, cited for switching to this Python mode is simply
455
00:44:47.035 --> 00:44:48.255
by having the researchers
456
00:44:48.795 --> 00:44:55.580
and the engineering team working with the same language, they can quickly ship and iterate and improve upon their trading approaches.
457
00:44:56.140 --> 00:45:04.240
So by losing a little bit of flexibility in all of the possible ways they could be developing new code, they get rid of a lot of, friction around reimplementation
458
00:45:04.775 --> 00:45:09.835
of ideas that will then go out to engineering. And so they can quickly get things deployed, test,
459
00:45:10.135 --> 00:45:17.130
keep it if it works. If it breaks, then take it out again, and improve upon it. So I really like the fact that they're a traditional
460
00:45:17.590 --> 00:45:18.090
quant
461
00:45:19.030 --> 00:45:23.535
hedge fund, but they've centralized on just Python as their their main tool.
462
00:45:23.835 --> 00:45:28.315
And the other example I'll cite, is channel 4. So this is,
463
00:45:29.035 --> 00:45:31.695
I cite these as a bit of an unusual example.
464
00:45:32.880 --> 00:45:35.940
So we wouldn't normally think of a British broadcaster,
465
00:45:36.640 --> 00:45:50.720
television broadcaster, being data driven. And so channel 4, it's a public services broadcaster. So it's got a remit from the government to provide, certain requirements. It's not a purely commercially driven outfit, but it is commercially driven, so it's ad driven.
466
00:45:51.339 --> 00:45:52.080
And so,
467
00:45:52.460 --> 00:46:11.030
I'm working with these guys at the moment. And there's some interesting projects going on there. 1 is using classifiers to better understand the audience to improve ad targeting. It's a bit like the Google model of having, a good understanding of your audience to target the right kind of ads. But as a result, they're able to increase their ad sale premiums by 30 to 50%
468
00:46:11.474 --> 00:46:18.375
because they're targeting people and they're able to tell the advertisers more about the the audience that they're going to be advertising against.
469
00:46:18.930 --> 00:46:22.630
They're using unsupervised classification methods to better understand
470
00:46:23.010 --> 00:46:29.035
the diversity of the audience and the viewing habits and when they watch, where they watch, what kind of devices they're using,
471
00:46:29.495 --> 00:46:30.875
and also personalization.
472
00:46:31.335 --> 00:46:47.805
So recommending the right kind of shows. So rather than just having the video on demand being this kind of catch up service that you go to if you miss a program, instead maybe it can turn into more of a destination. So a bit more like a Netflix driven site. So in London, as Emlyn says, we have a lot of finance companies,
473
00:46:48.185 --> 00:46:52.400
and I love the fact that media companies and the fashion companies Emily mentioned,
474
00:46:52.700 --> 00:46:56.320
are getting more into the use of data to drive all of their decisions.
475
00:46:57.020 --> 00:47:00.535
And 1 thing we encourage with the PyData meetups and the conferences
476
00:47:00.915 --> 00:47:12.800
is to have companies stand up and say, hey. We're in this niche over here. You haven't heard of us. You probably didn't think we've got a data problem, but we have got a data problem. Here's how we're solving it, and these are our problems. This is where we'd like some feedback to.
477
00:47:13.100 --> 00:47:20.005
And I think that, really helps to knit the ecosystem together in London, and it makes it much friendlier for companies to realize,
478
00:47:20.625 --> 00:47:22.964
realize that they can use their data productively,
479
00:47:23.265 --> 00:47:33.020
take lessons from other domains, and encourage their management that they can see that in other domains, people are solving these hard problems. So why don't we solve those problems too and use our data more productively?
480
00:47:35.234 --> 00:47:42.295
So is there a particular Python data science library that you think needs a little extra love and attention or 1 that you think
481
00:47:42.710 --> 00:47:44.970
or what 1 that doesn't exist that you think should?
482
00:47:46.710 --> 00:47:50.970
Oh. Now let's see. I mean, this always comes down to a personal bug there, but I always think,
483
00:47:51.590 --> 00:47:53.404
sci fi signal could be
484
00:47:53.785 --> 00:47:56.765
it could learn a lot basically from the from the MATLAB
485
00:47:57.144 --> 00:47:57.644
like,
486
00:47:59.144 --> 00:48:00.525
like function APIs.
487
00:48:01.970 --> 00:48:12.125
Just because it does lots it does a lot of the complicated signal processing tasks and it's pretty good at that, but it doesn't make sort of easier, sort of more convenience functions available. So,
488
00:48:12.505 --> 00:48:21.920
oh, man, I should just go and do it, right? I mean, I I always say this to people, and, I never actually step up and do it, so, looks like I shall get my pull request on,
489
00:48:22.460 --> 00:48:25.675
and get involved. But that would be definitely 1 for me.
490
00:48:25.975 --> 00:48:28.315
Ian, can you think of any that you just
491
00:48:28.855 --> 00:48:43.315
Right. Something something I've asked about, in in talks and the keynotes, in the last couple of years is I want to see more tools that help us clean data, particularly text data, because text data is often broken, hard to mark up. It comes badly encoded.
492
00:48:43.695 --> 00:48:54.240
And then most people who deal with it don't have strong training in natural language processing. And they just want a tool where they can chuck in some unstructured text, and it comes out in some kind of marked up useful clean
493
00:48:55.005 --> 00:48:55.505
JSONified
494
00:48:55.965 --> 00:48:56.465
dictionary
495
00:48:56.845 --> 00:49:02.625
like way. And those tools don't exist. So I would love to see some tools that make it easier to work with unstructured
496
00:49:02.925 --> 00:49:03.260
text.
497
00:49:03.740 --> 00:49:07.520
It's a rapidly evolving area, and, Word2Vec from, in GenSim,
498
00:49:08.220 --> 00:49:09.440
that's helping a lot.
499
00:49:09.740 --> 00:49:12.000
But, yeah, certainly, there's room there. And,
500
00:49:12.655 --> 00:49:13.395
in general,
501
00:49:13.855 --> 00:49:20.035
I think 1 thing that all the projects could do is, and, I will use as example statsmodels and, pymc3,
502
00:49:20.930 --> 00:49:22.150
is better documentation.
503
00:49:22.450 --> 00:49:23.910
So these products have documentation,
504
00:49:24.450 --> 00:49:40.080
but documentation is always something in the open source world in general that people forget about. We consume the documentation. It's documentation. It's quite hard to write good documentation. People don't really go and put the time into that. And so I said this, if there's anything that anyone out there listening wants to improve,
505
00:49:40.540 --> 00:49:46.880
go and put in a pull request, or at least go and file a bug report and say, hey. This documentation's a bit wrong, or it's just lacking.
506
00:49:47.185 --> 00:49:50.325
There's there's 1 line, and it doesn't explain anything. That's unhelpful.
507
00:49:50.705 --> 00:49:55.525
But maybe if you said this instead, this would be helpful, and maybe someone could turn this into an improvement,
508
00:49:55.950 --> 00:49:57.010
against the documentation.
509
00:49:57.550 --> 00:50:07.275
All the projects are crying out for documenters. And if you've never contributed to an open source project ever, then 1 of the easiest things you can do is go and improve the documentation.
510
00:50:07.655 --> 00:50:15.740
You get loads of feedback on it, and everyone will love you for it. Because normally, people like writing the code, not documentation. So please go and help improve the documentation.
511
00:50:18.920 --> 00:50:24.655
So before we move on, is there anything that we didn't cover you that you think we should have? Or anything that you wanna bring up?
512
00:50:26.315 --> 00:50:29.340
I think I think that covers a great deal of it.
513
00:50:30.300 --> 00:50:39.055
I'm trying to think. I mean, it's, yeah. No. I I think that covers a great deal about what Pi data does and sort of the the data science
514
00:50:39.615 --> 00:50:45.555
state of play in, in London. And I guess is a, you know, microcosm for the other centers of the world.
515
00:50:46.175 --> 00:50:46.675
Yeah.
516
00:50:47.055 --> 00:50:49.200
Yeah. I'm pretty happy with that. Alright.
517
00:50:49.740 --> 00:50:57.375
So for anybody who wants to get in touch with either of you or follow what you're up to, what would be the best way for them to do that? Ian, how about you go first?
518
00:50:57.755 --> 00:50:59.214
So for me,
519
00:50:59.515 --> 00:51:04.255
I blog and I've got a Twitter account, and both are my name. So that's ianoz
520
00:51:06.060 --> 00:51:09.200
v a l d, and so that's my Twitter account and ianoswald.com
521
00:51:09.900 --> 00:51:10.800
is my blog.
522
00:51:11.900 --> 00:51:17.994
And if you just Google for my name you'll find my past talks and the like. I'm reasonably well represented in Google now.
523
00:51:19.575 --> 00:51:20.875
And what about you, Evelyn?
524
00:51:21.335 --> 00:51:24.690
I think the easiest way to get hold of me is probably on Twitter. My Twitter handle
525
00:51:24.990 --> 00:51:25.730
is atemlynclay.
526
00:51:27.470 --> 00:51:31.090
1 of those lucky ones that has a rare enough name that no 1 had my handle.
527
00:51:31.675 --> 00:51:33.135
Yeah. No 1 has the Ausfords.
528
00:51:34.475 --> 00:51:35.455
Well, that too.
529
00:51:37.115 --> 00:51:42.810
And I I own embling clay dotco.uk, but there's nothing on it. So, no use going there.
530
00:51:44.230 --> 00:51:50.570
Alright. So we will move on to the picks. For my picks today, my first 1 is a tool called Xscape.
531
00:51:51.335 --> 00:51:57.674
And what it does is it's a Linux utility that lets you remap your keys to have different functionality.
532
00:51:58.680 --> 00:52:10.125
So you just write out a config file for it, and it will remap your whatever key you want to whatever other key you want. I personally use it for being able to let my caps lock key be an escape key.
533
00:52:10.585 --> 00:52:23.480
So I use e max and so I I was gonna ask I was gonna ask. That smelled like an e max thing to do. Yeah. Because I I've done that back in the day, and you train your fingers to do it, and then as soon as you use somebody else's computer,
534
00:52:23.994 --> 00:52:25.674
everything gets hard. Yes. So,
535
00:52:26.634 --> 00:52:40.395
Yeah. So in in my I use KDE for my desktop, and so I've remapped the caps lock key to be control when it's used as a modifier. But with Xscape, I can also use it as an escape key when I'm not using it as a modifier. So it does double duty.
536
00:52:42.775 --> 00:52:47.515
And so my other pick for today is going to be the key base file system.
537
00:52:48.090 --> 00:52:52.430
So they recently put out a brief blog post and announcement about
538
00:52:52.730 --> 00:52:59.214
the new tool that they added in. So with the newest version of the keybase utility, it will actually map a network
539
00:52:59.515 --> 00:53:00.974
drive onto your computer
540
00:53:01.434 --> 00:53:02.575
that is encrypted
541
00:53:02.954 --> 00:53:04.895
using your GPG keys.
542
00:53:05.400 --> 00:53:07.819
And so what that does is it lets you
543
00:53:08.440 --> 00:53:11.579
put a you know, put files into directories
544
00:53:12.125 --> 00:53:14.785
based on the IDs of other users of Keybase.
545
00:53:15.485 --> 00:53:34.154
And even if they're not a user of Keybase yet, you can, for instance, put it, you know, put it out a file path that includes their maybe their Twitter username so that when they do sign up for Keybase and associate their Twitter ID with their account, they will then automatically gain access to those files that you put on the Keybase file system. So it's a very, very
546
00:53:42.575 --> 00:53:44.355
people check that out and take a look.
547
00:53:45.055 --> 00:53:50.595
Keybase is still invite only, but that being said, I've got something like a 100 invites remaining.
548
00:53:50.974 --> 00:53:54.320
So for anybody who wants to get an account and check it out,
549
00:53:54.960 --> 00:53:57.860
just get in touch with us by their via Twitter or
550
00:53:58.160 --> 00:53:58.640
get,
551
00:53:59.040 --> 00:54:10.000
log sign up for our discourse forum and leave a post there. Just send us a message about what you like most about the show, and I will follow-up and get you an invite. So, Chris, what about you?
552
00:54:11.020 --> 00:54:18.925
Let's see. So my first pick first of all, I just wanna say Keybase dot io is is awesome. Those folks are doing such a great job making
553
00:54:19.385 --> 00:54:19.865
crypto,
554
00:54:20.185 --> 00:54:21.405
crypto tools
555
00:54:21.705 --> 00:54:29.220
accessible to the every man. I think it's fantastic, and I can't wait till they go public so I can trumpet them from the highest mountain top.
556
00:54:30.240 --> 00:54:40.195
So, my first pick is a book by Iain M. Banks called The Player of Games. It's the second in the his culture series of novels.
557
00:54:40.655 --> 00:54:47.740
And, all I have to say about it is I love these stories. They're amazing. And this is 1 of the only science fiction futures
558
00:54:48.040 --> 00:54:50.940
that I want to live in. Like, if I could push a button
559
00:54:51.240 --> 00:54:59.664
and beam myself into that future being a member of the culture, I would sign up for it in a heartbeat. I wanna live there and do that. It's it's just in a really
560
00:55:00.180 --> 00:55:03.960
interesting, interesting world having to do with the evolution of humankind
561
00:55:04.580 --> 00:55:06.040
and the evolution of machines
562
00:55:06.500 --> 00:55:08.040
in a very sort of nondark,
563
00:55:08.340 --> 00:55:08.840
non,
564
00:55:09.155 --> 00:55:12.855
you know, doom, they're taking over kinda way. It's it's great stuff.
565
00:55:13.234 --> 00:55:15.815
My next My culture universe is fabulous.
566
00:55:16.330 --> 00:55:25.230
I just recently reread accession. So, yeah, I'd I'd double up, your recommendation there. Anything by banks and the the culture universe is amazing. It's a really good utopia.
567
00:55:26.455 --> 00:55:32.880
I I need to read some of this other stuff that's not culture, but I'm I'm enjoying the culture book so much that I'm just starting there. Yeah. Just hoof them up.
568
00:55:34.319 --> 00:55:36.500
My next pick is a game called Undertale.
569
00:55:36.880 --> 00:55:44.115
This game is not what it appears to be. It deserves a second look. It looks like your standard kind of RPG ish kind of thing with an odd,
570
00:55:44.595 --> 00:55:45.495
combat mechanic,
571
00:55:45.875 --> 00:55:51.869
but it really is not. It is an exploration in morality, and it is totally worth a look.
572
00:55:54.089 --> 00:56:08.270
That's that's an is it annoying laugh that I hear there? No. No. No. An exploration of a galaxy. Yeah. It's through computer game. I love the idea of that. Yeah. Yeah. It's it's definitely worth checking out. You guys should definitely give it a look. And my 3rd and final pick is a movie called The Big Short.
573
00:56:08.730 --> 00:56:11.550
I was really impressed with this film. It's
574
00:56:11.850 --> 00:56:15.470
beautifully crafted. The acting is great. The writing is is phenomenal.
575
00:56:16.845 --> 00:56:17.665
It it takes
576
00:56:18.365 --> 00:56:22.625
it does a really excellent job of explaining some fairly complex concepts
577
00:56:23.140 --> 00:56:25.720
in a way that is both incredibly accessible,
578
00:56:26.020 --> 00:56:40.145
incredibly funny, and it just sort of shocks your brain into into being receptive. It's a really good movie. I I highly highly recommend. Oh, I agree. I mean, I watched that a little while ago. My wife and I devoured all of the things that were coming up for nomination.
579
00:56:41.650 --> 00:56:52.105
And, no, I thought The Big Short was 1 of the better ones. I I very much enjoy a good technical film. I mean, I don't work in the financial sector and I I know a little bit about the stock market and how it works, but this was,
580
00:56:52.485 --> 00:57:00.310
this felt really quite meaty as to what it was doing. And then of course the thing that gets you towards the end is that nothing's been learned. It's just amazing
581
00:57:00.690 --> 00:57:02.710
how, you know, all these things have occurred.
582
00:57:03.330 --> 00:57:10.085
And, yeah, they did an exceptional job at explaining, you know, how the system got itself into such a bender.
583
00:57:10.625 --> 00:57:14.164
So, yeah. No. I'd I'd simply echo that. That's a great film to go see.
584
00:57:14.545 --> 00:57:16.510
So Ian, what picks do you have for us?
585
00:57:16.890 --> 00:57:24.765
Right. I'm gonna go for 4 and I'll try and keep them brief. So Emlyn mentioned Seaborn, the visualization library. If you're using Python for visualization
586
00:57:25.065 --> 00:57:34.920
and you're using Pandas to represent your data, then you have to look at seaborn, s e a b o r n. It can consume a pandas data frame and it will do things like,
587
00:57:35.380 --> 00:57:50.980
spit out box plots and strip plots, of your columns of data with neat labeling and sensible color schemes and heat maps and violin plots and kernel density estimates and all sorts of things really really easy like in a 1 or 2 lines. Definitely go for seaborne for visualization.
588
00:57:52.960 --> 00:57:54.340
I'm working on a project,
589
00:57:54.880 --> 00:57:56.020
to try to understand
590
00:57:56.535 --> 00:57:58.795
allergic reactions through machine learning,
591
00:57:59.175 --> 00:58:05.310
and it's something I'm covering on my blog. So this is a it's a it's a pick because I'm interested in this, and I'll be talking about this at conferences.
592
00:58:05.870 --> 00:58:10.290
If anyone out there has a background around understanding allergic rhinitis
593
00:58:10.910 --> 00:58:22.035
and allergies in general and is interested in the idea of machine learning to help citizen science your way through to solving this, then I would love your feedback, and you'll find that mentioned on my blog.
594
00:58:23.420 --> 00:58:28.559
I've got a book which is a slightly contentious pick. This is Rui Miguel Forte's
595
00:58:29.099 --> 00:58:30.799
Mastering Predictive Analytics
596
00:58:31.099 --> 00:58:40.775
in R. So it's deliberately an R book, not a Python book. I'm looking at it as kind of the a secret hidden manual to Python stats models.
597
00:58:41.200 --> 00:58:42.500
So R and statsmodels,
598
00:58:43.040 --> 00:58:46.500
so statsmodels actually is a subset of all of these things available in R.
599
00:58:46.960 --> 00:58:47.859
In this book,
600
00:58:48.625 --> 00:58:50.005
Miguel is covering
601
00:58:50.305 --> 00:58:51.045
a statistician's
602
00:58:51.425 --> 00:59:03.680
view upon machine learning using R. And you can just take it sideways and go straight across into statsmodels, which I don't believe has a a book available at the moment and you can use the same techniques in scikit learn. So if you want a statistician's
603
00:59:04.059 --> 00:59:10.535
view upon doing predictive analytics, I really recommend mastering predictive analytics in R. And finally,
604
00:59:11.075 --> 00:59:12.615
if you're visiting London,
605
00:59:13.075 --> 00:59:18.490
I recommend you take a walking tour of London with unreal city audio.
606
00:59:18.790 --> 00:59:24.010
This is a small niche group of actors who will take you on an interactive tour of London
607
00:59:24.435 --> 00:59:25.095
with props.
608
00:59:25.635 --> 00:59:29.815
Props will include people leaping out from behind buildings to give you a discourse,
609
00:59:30.355 --> 00:59:59.795
and they will include things like on the coffee house tour, you will have original 17th century bitter coffee to drink. There's a chocolate house tour where you get get original 17th century hot chocolate. And at the moment, we've just missed it, and we're gonna be going for 1 of the ones coming up, a medieval wine tour, where 1 of the goals is to get you a little bit tipsy whilst teaching you history of the center of London. So unreal city audio is a massive, massive, shout out those guys. If you come to London and you want a really good tour, go to Unreal City Audio.
610
01:00:00.415 --> 01:00:01.315
I love them.
611
01:00:07.480 --> 01:00:09.100
Alright, and what about you, Emlen?
612
01:00:09.560 --> 01:00:17.214
Right, well I, yeah. I've only got I've only got the 1 pick. This is based on, some, stuff I was playing with today. So I love making
613
01:00:17.515 --> 01:00:19.375
slideshows in the IPython notebook.
614
01:00:19.914 --> 01:00:25.060
Even if they're not really necessarily much about code, I just like making it inside of that notebook.
615
01:00:25.440 --> 01:00:28.740
And I was playing around today with the the template flag so I could customise,
616
01:00:29.119 --> 01:00:32.625
what was going on. And because it uses the lovely ginger templating,
617
01:00:33.645 --> 01:00:34.145
syntax,
618
01:00:35.724 --> 01:00:40.305
you can do phenomenal things. Like if you've ever wanted to get really involved in making
619
01:00:40.800 --> 01:00:50.565
sort of slightly different animations, really exciting slides where, you know, the graph animates slightly or, you know, being able to blend bits of like d 3 into a presentation.
620
01:00:51.744 --> 01:00:53.525
I'm I'm up for doing a,
621
01:00:53.825 --> 01:00:56.145
a pitch to some private equity people. And,
622
01:00:57.050 --> 01:01:02.750
I was doing it so that the whole application that we've recently done is inside the IPython notebook that you can play with.
623
01:01:03.290 --> 01:01:08.055
And so I found the template flag to be just fantastic for dealing with bits around that.
624
01:01:09.555 --> 01:01:10.775
So yeah. I mean,
625
01:01:11.155 --> 01:01:15.930
definitely have a look. If you haven't ever made a slideshow in anything other than PowerPoint,
626
01:01:16.470 --> 01:01:19.770
put down the PowerPoint right now and get IPython
627
01:01:20.230 --> 01:01:21.370
and use the nbconvert,
628
01:01:22.184 --> 01:01:26.184
command line tool to turn it into a reveal, reveal dot j s,
629
01:01:27.065 --> 01:01:27.565
slideshow
630
01:01:28.025 --> 01:01:31.500
and, and profit. Everyone will think you're the coolest guy,
631
01:01:31.800 --> 01:01:32.540
in the world.
632
01:01:33.640 --> 01:01:34.140
Fact.
633
01:01:37.960 --> 01:01:45.628
Emily, do you have a blog post or something right, so I really should. I really should have a blog.
634
01:01:46.199 --> 01:01:49.740
I do not. Right. So I really should. I really should have a blog.
635
01:01:50.200 --> 01:01:56.765
I do not. So I will do, but I if you do need somebody to look at I was looking at, Damien Alvada's,
636
01:01:57.465 --> 01:02:00.285
blog and he's 1 of the core contributors to IPython.
637
01:02:00.665 --> 01:02:01.120
He
638
01:02:01.520 --> 01:02:18.765
went before I was doing a, a conference talk at Pioneer London a few years ago, I tweeted out, oh, no. This thing is broken. And Damien Arvada got on the case and fixed it for me just before the conference. So he's an absolute hero. He's got a whole bunch of good stuff there. So, yeah. I'll put it in your show notes, for all the lovely,
639
01:02:19.480 --> 01:02:20.540
all the lovely listeners.
640
01:02:22.120 --> 01:02:33.935
Alright. Well, we definitely appreciate the both of you joining us and taking time out of your day to tell us more about your PyData Organization in London and both of your experience in using Python for your own work. So
641
01:02:34.315 --> 01:02:34.974
I definitely
642
01:02:35.515 --> 01:02:42.640
learned a lot and, I appreciate your time. Oh, great. Tobias. Chris, thank you very much for having us. This has been a lot of fun. It has been. Thank you.