이 에피소드에 관해
Visit our site to listen to past episodes, support the show, join our community, and sign up for our mailing list.
SummaryWriting tests is important for the stability of our projects and our confidence when making changes. One issue that we must all contend with when crafting these tests is whether or not we are properly exercising all of the edge cases. Property based testing is a method that attempts to find all of those edge cases by generating randomized inputs to your functions until a failing combination is found. This approach has been popularized by libraries such as Quickcheck in Haskell, but now Python has an offering in this space in the form of Hypothesis. This week, the creator and maintainer of Hypothesis, David MacIver, joins us to tell us about his work on it and how it works to improve our confidence in the stability of our code.
Brief Introduction- Hello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.
- Subscribe on iTunes, Stitcher, TuneIn or RSS
- Follow us on Twitter or Google+
- Give us feedback! Leave a review on iTunes, Tweet to us, send us an email or leave us a message on Google+
- Join our community! Visit discourse.pythonpodcast.com for your opportunity to find out about upcoming guests, suggest questions, and propose show ideas.
- I would like to thank everyone who has donated to the show. Your contributions help us make the show sustainable. For details on how to support the show you can visit our site at pythonpodcast.com
- Linode is sponsoring us this week. Check them out at linode.com/podcastinit and get a $20 credit to try out their fast and reliable Linux virtual servers for your next project
- Open Data Science Conference on May 21-22nd in Boston. 20%
- Your hosts as usual are Tobias Macey and Chris Patti
- Today we are interviewing David MacIver about the Hypothesis project which is an advanced Quickcheck implementation for Python.
- Introductions
- How did you get introduced to Python? – Chris
- Can you provide some background on what Quickcheck is and what inspired you to write an implementation in Python? – Tobias
- Are there any ways in which Hypothesis improves on the original design of Quickcheck? – Tobias
- Can you walk us through the execution of a simple Hypothesis test to give our listeners a better sense for what Hypothesis does? – Chris
- Have you had trouble getting people to use Hypothesis? How has adoption been? – David
- What does this sort of testing get you that conventional testing doesn’t? – David
- Why do you think this sort of testing hasn’t caught on in the Python world before? – David
- Are there any facilities of the Python language that make your job easier? Are there aspects of the language that make this style of testing more difficult? – Tobias
- What are some of the design challenges that you have been presented with while working on Hypothesis and how did you overcome them? – Tobias
- Given that testing is an important part of the development process for ensuring the reliability and correctness of the system under test, how do you make sure that Hypothesis doesn’t introduce uncertainty into this step? – Tobias
- Given the sophisticated nature of the internals of Hypothesis, do you find it difficult to attract contributors to the project? – Tobias
- A few months ago you went through some public burnout with regards to open source and Hypothesis in particular, but circumstances have brought you back to it with a more focused plan for making it sustainable. Can you provide some background and detail about your experiences and reasoning? – Tobias
- What’s next for Hypothesis? – Chris
- Tobias
- Chris
- David
The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA
이 에피소드 내
노트보기 🔗
전사 🔗
NOTE
Transcription provided by Podhome.fm
Created: 7/6/2024 4:35:16 PM
Duration: 2821.286
Channels: 1
1
00:00:13.905 --> 00:00:15.445
Hello, and welcome to podcast.init,
2
00:00:16.065 --> 00:00:24.390
the podcast about Python and the people who make it great. You can subscribe to our show on Itunes, Stitcher, TuneIn Radio, or add our RSS feed to your pod catcher of choice.
3
00:00:24.690 --> 00:00:37.850
You can also follow us on Twitter or Google Plus, and please give us feedback. You can leave a review on Itunes to help other people find the show. Send us a tweet or an email. Leave us a message on Google Plus or on our show notes. And And you can also join our community. Visit discourse.pythonpodcast.com
4
00:00:39.190 --> 00:00:45.050
for your opportunity to find out about upcoming guests, suggest questions, propose show ideas, and follow-up with past guests.
5
00:00:45.715 --> 00:00:50.375
I would like to thank everyone who has donated to the show. Your contributions help us make the show sustainable.
6
00:00:50.754 --> 00:00:53.894
For details on how to support the show, you can visit our site at python podcast.com.
7
00:00:55.270 --> 00:00:58.090
Linode is sponsoring us this week. Check them out at linode.com/podcast
8
00:00:59.030 --> 00:01:03.770
in it and get a $20 credit to try out their fast and reliable Linux virtual service for your next project.
9
00:01:04.365 --> 00:01:09.185
I'd also like to announce that the open data science conference is happening in Boston on May 21st, 22nd.
10
00:01:09.725 --> 00:01:13.985
If you look at our show notes, you can find a discount code to get 20% off your tickets.
11
00:01:14.509 --> 00:01:22.850
Your host as usual are Tobias Macy and Chris Patti. And today, we're interviewing David McKeever about the hypothesis project, which is an advanced quick check implementation for Python.
12
00:01:23.305 --> 00:01:25.164
David, could you introduce yourself, please?
13
00:01:26.024 --> 00:01:29.085
So, hi. Yes. I'm David McKeever. I wrote hypothesis,
14
00:01:29.465 --> 00:01:39.830
and that's probably most most of what Python people know me for as I was relatively unknown in the community beforehand. But previously, I tend to do sort of back end data engineering stuff.
15
00:01:40.475 --> 00:01:46.895
I was just gonna ask, what kinds of data engineering work did you do? Is it mainly big data plumbing kind of thing or
16
00:01:47.450 --> 00:01:50.430
sort of database management? Just out of curiosity.
17
00:01:51.610 --> 00:02:03.905
It was a mix of things. So usually what happened is I got hired to do something interesting like recommendations or so on, and then discovered that the data pipeline for the system I was supposed to be working on was such a mess that I ended up being the 1 to fix it instead.
18
00:02:04.365 --> 00:02:05.744
Isn't that always the way?
19
00:02:06.079 --> 00:02:06.579
Unfortunately,
20
00:02:06.960 --> 00:02:09.300
yeah. So how did you get introduced to Python?
21
00:02:09.920 --> 00:02:14.955
So what actually happened was I got introduced to Python simply because I got a job writing it.
22
00:02:15.515 --> 00:02:19.935
I'd previously been writing Ruby and previously a bunch of other languages before that.
23
00:02:20.715 --> 00:02:22.575
But when I changed jobs,
24
00:02:24.320 --> 00:02:26.500
the company I was moving to was using Python.
25
00:02:26.800 --> 00:02:29.140
And so I figured I should probably learn it.
26
00:02:29.520 --> 00:02:34.375
So can you provide some background on what QuickCheck is and what inspired you to write an implementation in
27
00:02:34.675 --> 00:02:41.015
Python? Okay. So what QuickCheck is, at a high level, it's basically randomized testing. It comes from a programming language called Haskell.
28
00:02:41.390 --> 00:02:43.410
And the original original version,
29
00:02:43.790 --> 00:02:44.450
it was
30
00:02:45.070 --> 00:02:45.890
less about
31
00:02:46.750 --> 00:02:53.895
or they thought about it less about as a form of testing and more a way of doing formal methods that was a lot less hard than doing formal methods.
32
00:02:54.835 --> 00:02:55.335
But
33
00:02:56.435 --> 00:03:02.520
the idea caught on, and it's been ported to a variety of languages with varying degrees of success.
34
00:03:03.140 --> 00:03:03.640
And
35
00:03:04.260 --> 00:03:13.355
particularly in the hypothesis incarnation, it's really much more like conventional unit testing with some additional magic rather than its methods incarnation.
36
00:03:13.815 --> 00:03:16.635
The story of why I prototyped Python is a little funny
37
00:03:16.990 --> 00:03:19.650
in that originally I did it to learn Python.
38
00:03:19.950 --> 00:03:23.650
It was basically just a little prototype that I wrote back in 2013,
39
00:03:24.225 --> 00:03:24.725
and
40
00:03:25.265 --> 00:03:30.885
I needed Python projects. I'd had good luck with quick check-in its Scala check implementation beforehand,
41
00:03:31.505 --> 00:03:33.295
and it didn't seem to be anything
42
00:03:33.680 --> 00:03:36.560
very satisfying for Python. So I thought I would just have a play around and see what it
43
00:03:37.360 --> 00:03:45.685
how hard it would be. And the original prototype wasn't very good, but the problem is that neither were any of the other things similar to it for Python. And
44
00:03:45.985 --> 00:03:51.680
so with the Lightning role, I didn't really work on it for a while. And then I sort of got, out of the job I was in in 2014.
45
00:03:52.220 --> 00:03:58.480
And beginning of 2015, I had some time between jobs. I realized people were using this hypothesis thing that I had written 2 years ago,
46
00:03:58.925 --> 00:04:06.760
because there wasn't anything better. And it kind of annoyed me that there wasn't anything better. So I decided to take some time and just improve it and try and,
47
00:04:07.300 --> 00:04:10.840
make it a more useful project. And it sort of got away from me.
48
00:04:11.220 --> 00:04:15.080
So are there ways in which Hypothesis improves on the original design of QuickCheck?
49
00:04:15.905 --> 00:04:18.165
1 of the ways is simply that it's not working at Haskell,
50
00:04:18.705 --> 00:04:20.885
which makes it a lot more usable for,
51
00:04:22.145 --> 00:04:25.729
a much wider variety of tasks. Like, in some sense, the
52
00:04:26.030 --> 00:04:28.849
objective of a hypothesis is to make this sort of testing mainstream.
53
00:04:29.150 --> 00:04:37.004
But along the way, I had to figure out a lot of useful functionality that isn't present or is sort of differently present in the original QuickCheck.
54
00:04:37.465 --> 00:04:38.444
Probably the biggest,
55
00:04:39.145 --> 00:04:43.490
feature of Hypothesis over QuickCheck is that it has an example database.
56
00:04:43.950 --> 00:04:45.810
So what happens is that
57
00:04:46.430 --> 00:04:48.610
when you have a failing example,
58
00:04:49.550 --> 00:04:50.050
hypothesis
59
00:04:50.865 --> 00:04:52.245
will save it in its database
60
00:04:52.545 --> 00:04:59.285
so that when you rerun the test, it will rerun with the previous failing example rather than having to generate a new 1.
61
00:04:59.670 --> 00:05:04.570
Then crew check has some functionality for that in more recent versions, but
62
00:05:05.110 --> 00:05:06.890
hypothesis is a lot more advanced.
63
00:05:07.590 --> 00:05:08.330
There also
64
00:05:09.225 --> 00:05:13.565
a bunch of things that it does to make more things work out of the box
65
00:05:14.025 --> 00:05:14.525
and
66
00:05:14.985 --> 00:05:18.765
to improve sort of data quality and exam and the final example
67
00:05:19.450 --> 00:05:26.830
quality compared to original quick check. I also think it's a better design from the point of view of the user interface of the library and
68
00:05:27.405 --> 00:05:30.145
how people interact with it, but that may be a matter of taste.
69
00:05:31.324 --> 00:05:32.705
And are there
70
00:05:33.164 --> 00:05:36.785
significant differences between it and, you mentioned ScalaCheck
71
00:05:37.230 --> 00:05:38.610
as another implementation.
72
00:05:39.150 --> 00:05:43.330
I'm just curious what your impression is in terms of So there's quite a lot of difference between it and ScalaCheck.
73
00:05:43.710 --> 00:05:54.255
Sorry, I didn't mean to interrupt you there. So Hypothesis isn't very closely based on ScalaCheck. And 1 of the big differences between Hypothesis and either ScalaCheck or QuickCheck is that Python is dynamically typed and
74
00:05:54.729 --> 00:05:56.990
Haskell and Scala are statically typed.
75
00:05:57.849 --> 00:05:59.710
And in Haskell and Scala, the
76
00:06:00.490 --> 00:06:04.110
the way you generate data is very closely tied to the type system.
77
00:06:04.525 --> 00:06:06.305
So you don't have
78
00:06:06.845 --> 00:06:09.824
necessarily custom generators that you're composing in
79
00:06:10.125 --> 00:06:14.145
the way that you do in hypothesis. You instead say, I want a list of integers. Give me a list of integers.
80
00:06:18.030 --> 00:06:21.250
That's probably the major difference from the outside. The,
81
00:06:22.030 --> 00:06:30.205
the way it interacts with test frameworks and test runners is all very Python specific, so it's obviously quite from Scala or Haskell.
82
00:06:30.745 --> 00:06:31.245
And
83
00:06:31.920 --> 00:06:34.660
there's a bunch of additional functionality in terms of
84
00:06:34.960 --> 00:06:35.460
how
85
00:06:36.160 --> 00:06:39.780
you compose data and how you build things up that are quite different.
86
00:06:40.645 --> 00:06:48.585
Personally, when I was writing it, I thought they were quite similar, but I've since been having some, quick chat users trying to use hypothesis and getting
87
00:06:48.910 --> 00:06:54.770
confused. So I think there's probably more of a difference than I realized at the time. Do you think that there's value in dissociating
88
00:06:55.150 --> 00:06:59.345
hypothesis from the term quick check and using the more generic term of property based testing?
89
00:06:59.965 --> 00:07:04.945
I do often do that. Yeah. Particularly from an implementation point of view, it's actually very different from quick check.
90
00:07:05.430 --> 00:07:12.330
But the problem is that what happens when people are looking explicitly for this sort of thing is that they inevitably Google for Python QuickCheck. So
91
00:07:12.815 --> 00:07:22.595
it's hard to disassociate too much on it. Right. So can you walk us through the execution of a simple hypothesis test to give our listeners a better sense for what hypothesis does?
92
00:07:23.440 --> 00:07:23.940
Sure.
93
00:07:24.400 --> 00:07:32.260
So you've got a hypothesis based test, and it's got this given decorator that specifies how to generate some values for some arguments.
94
00:07:32.665 --> 00:07:40.639
And what will happen is that or when the test runs, it will call your test function multiple times with arguments provided from that
95
00:07:41.360 --> 00:07:42.340
from hypothesis's
96
00:07:42.880 --> 00:07:52.694
data generation. So the first thing it does is it looks in the database, and it says, do you have any saved examples for this test? If it does, then it pulls the arrows out of the database and tries those. And
97
00:07:52.995 --> 00:07:53.735
if things
98
00:07:54.275 --> 00:07:57.735
if any of the previously saved examples fail, then it
99
00:07:58.240 --> 00:07:59.620
will skip to next phase.
100
00:08:00.080 --> 00:08:06.580
But if they all pass, then what happens is it moves on to a generation phase where it's essentially picking randomly generated examples
101
00:08:07.074 --> 00:08:10.375
from a distribution that's been provided by the strategy implementation
102
00:08:10.914 --> 00:08:19.790
and tries to find 1 of these that causes the test to fail. And if neither of the previous step those 2 steps have caused a test failure, then it stops.
103
00:08:20.250 --> 00:08:22.990
But if any of them have caused a test failure,
104
00:08:23.354 --> 00:08:28.655
then it takes that failure, and it tries to prune it down. Because the problem with randomly generated examples
105
00:08:29.115 --> 00:08:32.700
is that they're often huge and messy, and they contain lots of extraneous detail
106
00:08:33.640 --> 00:08:39.340
that doesn't actually matter for failing the test. Like, you might have an a 1, 000 element list, but all that actually matters
107
00:08:40.405 --> 00:08:44.265
is non empty and contains 1 non 0 number or something like that.
108
00:08:44.805 --> 00:08:51.260
So it just sort of basically hacks bits off until it can't find anything to remove from the example that would still cause the test to fail.
109
00:08:52.120 --> 00:08:52.620
And
110
00:08:53.000 --> 00:08:59.525
at that point, it then says, okay. Here's the smallest example. First of all, I'll save that in my database so that when you rerun the test,
111
00:09:00.545 --> 00:09:00.865
the,
112
00:09:01.345 --> 00:09:02.565
we'll get that 1 again.
113
00:09:03.080 --> 00:09:10.940
And then it reruns the test 1 final time to sort of let the test failure bubble up and basically look like it had just not
114
00:09:11.805 --> 00:09:13.345
picked exactly the right example,
115
00:09:13.725 --> 00:09:18.545
from the very beginning. Does that make sense? That that does. And that's really neat. It almost seems like
116
00:09:18.889 --> 00:09:22.910
in addition to the randomized property testing aspect of the tool,
117
00:09:23.290 --> 00:09:24.510
it sort of performs
118
00:09:24.970 --> 00:09:26.350
some of the processes
119
00:09:27.154 --> 00:09:32.935
that a good QA tester would perform when sort of, like, ruggedizing a test suite anyway,
120
00:09:33.475 --> 00:09:35.574
in terms of finding the minimal
121
00:09:36.230 --> 00:09:36.730
possible,
122
00:09:37.110 --> 00:09:46.125
the the you know, the most minimal failure case for a given test. It's very neat that it has that kind of intelligence built in. It sounds like it could, in that sense, be a much more effective
123
00:09:47.385 --> 00:09:50.045
testing tool for developers than standard tests.
124
00:09:51.250 --> 00:09:52.530
That's definitely the goal,
125
00:09:53.010 --> 00:09:53.910
to some extent.
126
00:09:54.530 --> 00:10:03.904
So the phrase I sometimes use is thinking with the machine, where basically what's happened is that hypothesis gets to be a sort of little external brain for your testing process,
127
00:10:04.365 --> 00:10:04.865
where
128
00:10:05.324 --> 00:10:10.770
hypothesis is not itself very smart, but it is very fast. And sort of the combination of
129
00:10:11.310 --> 00:10:16.850
a experienced human tester plus hypothesis doing a lot of the heavy lifting for you ends up being,
130
00:10:17.584 --> 00:10:29.100
like, a a much larger and much more experienced QA department simply because there's sort of a force multiplier between the 2. I haven't talked to many people who work directly in QA about using Hypothesis, but it's something I'm hoping to do more because
131
00:10:29.480 --> 00:10:34.454
I think that there's probably quite a lot of value to them as well as to normal developers.
132
00:10:35.235 --> 00:10:35.735
Definitely.
133
00:10:36.355 --> 00:10:43.415
As a matter of fact, just before we did the podcast, I was talking to 1 of my coworkers who's in test automation here,
134
00:10:43.860 --> 00:10:44.760
about Hypothesis.
135
00:10:45.220 --> 00:10:48.520
So, hopefully, we can use this episode as a way to
136
00:10:48.900 --> 00:10:52.440
shine more light on on it and get it to be used by a wider
137
00:10:53.144 --> 00:10:54.285
audience of people.
138
00:10:54.985 --> 00:10:55.964
That would be excellent.
139
00:10:56.745 --> 00:11:02.285
So have you had any trouble getting people to use Hypothesis? And I'm wondering how the adoption curve has been.
140
00:11:02.890 --> 00:11:13.245
It's actually mostly been much better than I would have expected. It's very weird having a project which most people just say really nice things about. There have been some people who just seem very uncertain how to get started,
141
00:11:13.625 --> 00:11:33.275
and I'm trying to sort of improve documentation and writing around that to give people easy entry points into Hypothesis. But generally speaking, when people do get started, they sort of rave enthusiastically about it, which is really gratifying. I mean, it's still early days. I think 1 of the consequences of the fact that I've done so much full time work on Hypothesis is that a lot of the Python community
142
00:11:33.895 --> 00:11:40.139
sort of sees it as having just come out of left field where it basically just jumped dropped this completed project on them. So
143
00:11:41.399 --> 00:11:48.705
it's not got as many people as do not necessarily expect for sort of the level of readiness it currently has, but we've got,
144
00:11:49.485 --> 00:11:54.065
Mercurial is using it. PyPI is using it. Those are probably the 2 big names, but
145
00:11:54.370 --> 00:12:10.305
there's a whole pile of little projects using it at this point. I know a bunch of companies are using it. So pretty good for the most part. It definitely hasn't hit the sort of test saturation something some of the bigger tools like pie dot test or coverage have, of course. But,
146
00:12:11.610 --> 00:12:14.510
I'm hopeful that it'll get there at some point in the next couple of years.
147
00:12:15.530 --> 00:12:22.365
And what does the integration story look like for using Hypothesis along with some of those other tools, like, that you mentioned?
148
00:12:23.225 --> 00:12:31.150
So mostly really good. I sort of made it originally because I was lazy, but it turned out to be an amazingly good idea, is that Hypothesis is not a test runner.
149
00:12:31.610 --> 00:12:39.705
It doesn't have any testing framework built into it. You can just use it with whatever test framework you like. So it works basically out of the box with py.test,
150
00:12:40.085 --> 00:12:41.465
with nose, with,
151
00:12:42.725 --> 00:12:48.000
unit test if you really must. And as far as I know, basically every test runner.
152
00:12:48.460 --> 00:12:58.905
The only things I know who have some problems is that it doesn't play very well with asynchronous tests. We've got some people using it with asyncio, and it does more or less work. But with Twisted's,
153
00:12:59.925 --> 00:13:01.945
trial test runner, it's a bit problematic.
154
00:13:02.325 --> 00:13:02.825
And
155
00:13:03.290 --> 00:13:27.339
the combination of Pypi plus coverage plus hypothesis is 1 that I generally recommend people avoid simply because every sort of every pair of things in this this triple have slight problems with each other. I mean, PyPI and hypothesis don't really help with each other. It's more that it's not as fast as you would expect given how much faster PyPI usually is. But high client coverage and hypothesis and coverage both have a bit of struggle sometimes.
156
00:13:27.640 --> 00:13:37.775
The other problem that I think people sometimes have is that letting hypothesis generate your coverage for you without some explicit examples to see this can be a bit of a problem because it makes your coverage random.
157
00:13:38.440 --> 00:13:48.154
Yeah. I can see how that would be the case, but given the fact that the input data is somewhat randomized. So it would potentially pick some arbitrary code paths to execute. So it's interesting.
158
00:13:48.695 --> 00:13:57.529
Normally, it's pretty good at getting as a 100% of coverage reachable by the test. But if you're running enough tests, then you're gonna be unlucky on 1 of them, basically.
159
00:13:58.550 --> 00:14:01.930
What does this sort of testing get you that conventional testing doesn't?
160
00:14:02.709 --> 00:14:20.800
So the major thing is that it is a huge effort saving because you can write this sort of testing with semi conventional testing, like the pytest parameterized stuff. But the problem is that you then have to write all the examples out by hand. And this is both tedious,
161
00:14:21.420 --> 00:14:24.640
which means you're not gonna do it nearly the degree that you should
162
00:14:25.245 --> 00:14:29.425
And also error prone that it's very easy to be forgetful. And having
163
00:14:29.885 --> 00:14:33.505
the computer basically take care of that for you means that you can concentrate
164
00:14:33.885 --> 00:14:34.385
on
165
00:14:35.190 --> 00:14:37.370
testing some much more high level things
166
00:14:39.190 --> 00:14:43.290
and basically have it figure out all the edge cases that you forget.
167
00:14:44.045 --> 00:14:44.545
So,
168
00:14:45.085 --> 00:14:45.585
typically,
169
00:14:46.045 --> 00:14:46.945
people's experience
170
00:14:47.404 --> 00:14:51.105
using hypothesis that it won't get it won't let them get away with things because
171
00:14:51.710 --> 00:15:01.365
it will try every edge case, even once they've they forgot when writing both the tests and the original code. So, basically, it's just an automated way of being much more than you
172
00:15:01.665 --> 00:15:03.445
would necessarily otherwise have been during your testing.
173
00:15:04.065 --> 00:15:04.565
Interesting.
174
00:15:05.025 --> 00:15:09.650
Does it constrain the ways in which you would generally write a unit test in order to be compatible with,
175
00:15:10.130 --> 00:15:10.630
how
176
00:15:10.930 --> 00:15:11.430
hypothesis
177
00:15:11.810 --> 00:15:12.310
executes?
178
00:15:12.770 --> 00:15:17.430
Just thinking in terms of the, you know, for instance, like setting up of mocks and and so forth.
179
00:15:18.210 --> 00:15:20.855
Mostly, it doesn't do much. You do have
180
00:15:21.475 --> 00:15:41.964
to use hypothesis' own setup and teardown hooks if you want something on setup and teardown to work simply because you need to do it before and after each example, rather than before and after the entire test function. The other thing that can sometimes happen is that hypothesis isn't very forgiving if your tests are slow because you're running your tests many times, so it tends to
181
00:15:42.745 --> 00:15:45.725
exacerbate that sort of problem. It will tend to
182
00:15:46.390 --> 00:15:53.770
encourage you to write small, fast tests, but that's not a bad thing in and of itself. Although it does cause some problems for Django users simply because
183
00:15:54.514 --> 00:16:10.680
the Django test runner is relatively slow and has to do a lot of database setup and teardown. Sorry. Did that answer the question? And yes. No. It absolutely did. I bet you that's also problematic for developers who tend to sort of have lots of database dependencies and and who tend to
184
00:16:11.075 --> 00:16:12.615
blur the lines between
185
00:16:12.995 --> 00:16:25.870
unit testing and integration testing. I bet hypothesis is a a tough pill to swallow for them. It can be. And then to some other sense sorry. To some degree, I'm also 1 of those developers in that I don't necessarily believe there is a hard and fast line between
186
00:16:26.250 --> 00:16:31.915
unit testing and integration testing. But it's not like you can't do that under hypothesis. It's just that it's
187
00:16:32.475 --> 00:16:36.255
what tends to happen is that once you get a failing integration test, you end
188
00:16:36.870 --> 00:16:42.970
up wanting to write a failing unit test to isolate that case so you can work with it more easily, which,
189
00:16:43.430 --> 00:16:44.810
again, is no bad thing.
190
00:16:45.555 --> 00:16:49.574
So why do you think this sort of testing hasn't caught on in the Python world before?
191
00:16:51.475 --> 00:16:54.055
So the major reason is that it was really hard to write.
192
00:16:54.790 --> 00:16:58.570
It turns out that a lot of, the design of the original QuickJack has
193
00:16:58.949 --> 00:17:01.610
very Haskell shaped assumptions built through it. And
194
00:17:02.115 --> 00:17:04.695
so Python isn't statically typed. Python isn't immutable.
195
00:17:04.995 --> 00:17:05.495
And
196
00:17:06.035 --> 00:17:16.060
both of these things turned out to be quite hard to make it work well. And it does all work well out of the box now, but that basically required a lot of work on the part of designing it. And
197
00:17:16.440 --> 00:17:22.605
even without those, sort of the baseline level of difficulty implement in good quick check is really high. And
198
00:17:22.985 --> 00:17:24.765
so what you saw prior to hypothesis
199
00:17:25.145 --> 00:17:31.580
or not even prior to hypothesis, prior to my the hypothesis reboot is that there were about 8 or 9 different abandoned projects on PyPI,
200
00:17:31.960 --> 00:17:33.659
which were basically doing,
201
00:17:34.519 --> 00:17:35.419
similar thing.
202
00:17:35.720 --> 00:17:37.259
And, basically, a combination
203
00:17:37.639 --> 00:17:40.924
of luck and timing meant that I had
204
00:17:41.225 --> 00:17:42.365
enough time to
205
00:17:42.745 --> 00:17:46.300
basically plunge into the problem and brute force my way through it. So,
206
00:17:46.760 --> 00:17:52.780
the result was that hypothesis managed to overcome the hump, and we, by virtue of a couple of months, saw it work. And
207
00:17:53.395 --> 00:18:00.615
then once you've got the basic tooling working and it worked well enough that people can use it, at that point, people can get excited about it. Whereas previously,
208
00:18:01.200 --> 00:18:12.465
if they had tried any of the things that were lying around, most likely, they would have gone, this is a really nice idea, but it sort of makes my life harder in practice. So for example, the, the shrinking tends to be the part at which
209
00:18:13.405 --> 00:18:18.625
most projects, for them to do this, fail. Because it turns out that making that work well is really quite difficult. And
210
00:18:19.190 --> 00:18:24.250
so you would see most of the previous attempts at it sort of got to the generation point.
211
00:18:24.630 --> 00:18:25.690
And then
212
00:18:26.070 --> 00:18:31.575
when it failed, the output output was almost incomprehensible, so people wouldn't really like to use it.
213
00:18:31.955 --> 00:18:38.700
That's really interesting. And it's also really interesting that you sort of picked a problem space that many people had attempted, but no 1 had really
214
00:18:39.240 --> 00:18:41.500
succeeded at in a truly workable way.
215
00:18:41.880 --> 00:18:47.455
And congratulations for choosing to climb a fairly tall mountain and and actually getting there safely.
216
00:18:47.995 --> 00:18:51.274
Thank you. I'm not necessarily sure I would say it was a good decision, but the
217
00:18:52.620 --> 00:19:08.605
Yeah. Thank you. I'm not necessarily sure I would say it was a good decision. But it worked out pretty well in the end. I think the like I said, it's not necessarily that I have some unique ability to implement good quick checks. It's simply that this sort of thing takes time, and most people don't have the ability to
218
00:19:08.905 --> 00:19:10.860
take 3 months to basically scratch a niche.
219
00:19:11.820 --> 00:19:13.040
That is all too true.
220
00:19:13.420 --> 00:19:25.695
So are there any facilities of the Python language that make your job easier? And are there aspects of the language that make this style of testing more difficult? And as a corollary to that, I'm curious if the typing module in Python 35
221
00:19:26.555 --> 00:19:30.889
would provide any hooks for you to be able to tie into that to be able to potentially
222
00:19:31.669 --> 00:19:32.169
automatically
223
00:19:32.549 --> 00:19:34.090
determine the type signatures
224
00:19:34.470 --> 00:19:46.070
for the quick check implementation. Because I know that as part of the test setup, you need to tell it sort of what general types are accepted in order to get it to generate an appropriate set of data for it.
225
00:19:46.770 --> 00:19:53.665
So in terms of features of Python that have made my job easier, not many, to be honest. Decorators are very nice, and I use them extensively.
226
00:19:54.525 --> 00:20:00.705
And that's probably the big thing that let me make hypothesis quite so test framework framework agnostic,
227
00:20:01.070 --> 00:20:06.929
in that it means that I can just expose functions that that do what I want and mirror the original functions. And
228
00:20:07.870 --> 00:20:08.690
most other
229
00:20:09.105 --> 00:20:13.445
attempts to do this in other languages don't have quite that level of test framework of agnosticism.
230
00:20:13.905 --> 00:20:19.629
Features that make my life harder, I can definitely list more of this. But the big 1 is just
231
00:20:20.090 --> 00:20:21.950
simply how complicated the
232
00:20:22.490 --> 00:20:22.990
Python,
233
00:20:24.330 --> 00:20:25.549
argument path and conventions
234
00:20:25.850 --> 00:20:32.695
are. If you think they're simple, then that's probably because you have never tried to emulate them. And sort of the differences between
235
00:20:33.155 --> 00:20:35.890
positional arcs and keyword arguments and the ability to mix
236
00:20:36.850 --> 00:20:37.430
them and
237
00:20:38.050 --> 00:20:41.110
how these are then used by the different testing frameworks
238
00:20:41.410 --> 00:20:43.510
cause it caused me a lot of problems.
239
00:20:43.815 --> 00:20:50.955
And at this point, most of the solutions are just nicely packaged up in 1 or 2 functions inside of hypothesis, so it doesn't get any grief now. But
240
00:20:51.280 --> 00:20:52.820
every time I need to,
241
00:20:53.760 --> 00:21:00.020
implement a new code that has to deal with this in a different way, I get sad all over again. There was a third part to your question.
242
00:21:00.345 --> 00:21:05.245
I was curious if the typing module and type annotations in Python 3.5
243
00:21:05.865 --> 00:21:09.885
would provide any way for you to sort of automatically generate the
244
00:21:10.400 --> 00:21:14.580
signature that Hypothesis uses for determining what types of,
245
00:21:14.960 --> 00:21:16.020
inputs to generate.
246
00:21:16.640 --> 00:21:16.800
Right.
247
00:21:17.680 --> 00:21:18.180
So
248
00:21:18.575 --> 00:21:30.600
people keep asking me that, and it's totally understandable because of the original quick check's relationship to typing. But I actually don't think this is a good idea because if you look at hypothesis strategies,
249
00:21:31.220 --> 00:21:32.040
they are
250
00:21:32.500 --> 00:21:50.350
much more fine grained than types. And almost all tests that people write, I find they want to do something a little bit custom because you don't really just want to list the strings. You want a list of strings with at least 1 element where the individual strings can only fit in 3 bytes, like UTF-eight or something like that. And
251
00:21:50.735 --> 00:22:02.419
you want to possibly filter some stuff out. And hypothesis strategies give you a huge variety of knobs to Twiddle and ways to compose them that aren't really present in,
252
00:22:03.279 --> 00:22:06.500
any sort of nondependently typed type system. The
253
00:22:06.880 --> 00:22:08.565
other big problem is that,
254
00:22:08.965 --> 00:22:16.549
I'm not a huge fan of the typing machine. I forgot to be honest. I think that it's got quite a ways to go before it's a useful piece of software. And
255
00:22:17.410 --> 00:22:23.190
I'm a little worried about it's inclusion in Python 3.5 because I think you're gonna need to change a bunch of things before it becomes useful.
256
00:22:24.195 --> 00:22:29.255
But at some point, I'm sure I will do something to integrate along those lines. It's just not quite there yet.
257
00:22:29.715 --> 00:22:36.279
Sure. And what are some of the design challenges that you've been presented with while working on hypothesis and I'm wondering how you overcame them?
258
00:22:36.659 --> 00:22:42.039
So the major design challenges have been simply trying to figure out how to deal with
259
00:22:42.915 --> 00:22:54.200
the extreme flexibility of Python. So you you've got a lot of issues where people can't. So So for example, 1 of the things that hypothesis relies on is being able to replay tests,
260
00:22:54.659 --> 00:22:55.880
and you
261
00:22:56.340 --> 00:22:58.545
have to deal with the fact that,
262
00:22:59.105 --> 00:23:00.165
a function
263
00:23:00.465 --> 00:23:10.480
could do something arbitrarily different the, second time it's been called, or it could have mutated its arguments the first time, or it could basically be asserting that you're never calling in the same argument twice.
264
00:23:10.860 --> 00:23:19.522
And so there's a lot of sort of defensive things that have to happen every, basically every time that hypothesis calls out to someone else's function. There's been a lot of work in terms of, implementations of Python. And the
265
00:23:20.658 --> 00:23:21.158
major
266
00:23:22.295 --> 00:23:22.920
way I've ever
267
00:23:23.560 --> 00:23:24.860
implementations of Python.
268
00:23:25.800 --> 00:23:28.540
And the major way I've overcome that is simply by using Travis,
269
00:23:28.920 --> 00:23:31.100
which has been really good for,
270
00:23:31.865 --> 00:23:35.865
letting me run and observe an Apple CI jobs and observe a number of different,
271
00:23:36.265 --> 00:23:47.750
Python version and not an observed number of operating systems, but 2 operating systems, and then AppFare for Windows. So, Adam, I'd like to say that there were sort of deep interesting theoretical challenges that,
272
00:23:48.309 --> 00:23:48.809
were
273
00:23:49.455 --> 00:23:50.835
things I had to overcome.
274
00:23:51.375 --> 00:23:53.554
But the reality is that those are the easy part,
275
00:23:53.934 --> 00:23:58.034
at least for me, because that's the part that I find interesting and enjoy working on.
276
00:23:58.390 --> 00:24:01.530
And what sort of really surprised me is just the simple math
277
00:24:02.390 --> 00:24:03.130
of, detail
278
00:24:03.510 --> 00:24:07.610
work and sort of boring grunt work that 1 has to do in order to
279
00:24:08.794 --> 00:24:11.695
make all these sort of things into usable production software.
280
00:24:12.715 --> 00:24:16.654
And given that testing is an important part of the development process for ensuring
281
00:24:17.530 --> 00:24:24.590
the reliability and correctness of the system under test, I'm wondering how you make sure that hypothesis doesn't introduce uncertainty into that step.
282
00:24:25.415 --> 00:24:30.315
So hypothesis does introduce uncertainty into that step, but it sort of does it in the direction
283
00:24:30.775 --> 00:24:33.995
in that basically the hypothesis principle is no false positives.
284
00:24:35.040 --> 00:24:37.220
Every time that your test fails with hypothesis,
285
00:24:38.400 --> 00:24:39.380
it's a real failure.
286
00:24:39.920 --> 00:24:45.184
It's it's not sort of the bad kind of randomness in test, where a test can fail flakily.
287
00:24:45.965 --> 00:24:46.465
It's
288
00:24:46.924 --> 00:24:55.160
every failure is a real failure. Then you've got the example database, which I mentioned earlier, which means that every failure is not just real, but it's also replayable.
289
00:24:55.620 --> 00:24:56.120
So
290
00:24:56.420 --> 00:25:03.825
a bug won't go away just because you've had bad luck with a random number generator this time. You actually have to fix the bug. And
291
00:25:04.365 --> 00:25:07.345
there is a sort of explicitly feeding hypothesis with
292
00:25:07.780 --> 00:25:13.160
handpicked examples that you want to want it to use each time to make that more reliable.
293
00:25:13.780 --> 00:25:20.554
And you can also put hypothesis in deterministic mode if you really want, in in which case it just fixes the seed each time and,
294
00:25:21.095 --> 00:25:23.034
reruns the same test each time. So
295
00:25:23.335 --> 00:25:25.195
with all of these basically hypothesis,
296
00:25:25.495 --> 00:25:26.715
it will find
297
00:25:27.330 --> 00:25:29.990
everything that you would have found through the manual testing process.
298
00:25:30.290 --> 00:25:42.760
And the only uncertainty that really remains is basically how much additional stuff will it find. So you can't guarantee that on every single run, hypothesis will find every single bug that hypothesis could find. And I know some people have
299
00:25:43.160 --> 00:25:52.435
had hypothesis happily running for months, and then suddenly it said, hey. These particular 2 floating point numbers no longer work. And, of course, they've not worked all along. It's just that,
300
00:25:53.215 --> 00:25:55.395
it didn't find them before now. And
301
00:25:55.855 --> 00:26:02.790
to some extent, that's a real problem, and it does introduce uncertainty into the testing process. But it doesn't introduce more uncertainty
302
00:26:03.250 --> 00:26:03.750
than
303
00:26:04.130 --> 00:26:05.430
having users does.
304
00:26:05.735 --> 00:26:14.154
And hypothesis just essentially, in this case, becomes another user of your system, who every now and then just drops to you, a line saying, hey, so I found this bug.
305
00:26:15.200 --> 00:26:19.460
And is there a parameter that you can tune to be able to
306
00:26:20.000 --> 00:26:21.780
increase the number of
307
00:26:22.320 --> 00:26:24.340
trials that QuickCheck will run through?
308
00:26:24.985 --> 00:26:29.565
Hypothesis has so many parameters. It's a bit silly. But, yes, 1 of them does do that.
309
00:26:30.505 --> 00:26:32.585
It's by default, it's running,
310
00:26:32.905 --> 00:26:33.885
200 examples
311
00:26:34.720 --> 00:26:37.380
and will time out your tests if,
312
00:26:38.320 --> 00:26:49.075
it takes more than a minute to run. The reality is that, a minute is ridiculous overestimate of how much how long most of these tests run. It's simpler to have some value that you probably be hitting, but
313
00:26:50.029 --> 00:26:52.110
makes it not get into an infinite loop.
314
00:26:52.830 --> 00:27:10.730
I think it's more common for people to turn down the numbers rather than to turn off the numbers, but I typically run it with a thousand rather than 200 simply because I do like a bit more thorough testing. And I think people who are using the Django integration typically turn it down to 50 simply because of the aforementioned speed of the Django test runner.
315
00:27:11.350 --> 00:27:17.770
So given the sophisticated nature of the internals of hypothesis, do you find it difficult to attract contributors to the project?
316
00:27:18.145 --> 00:27:22.565
To some degree. 1 of the nice things about hypothesis is that it's relatively well factored.
317
00:27:23.025 --> 00:27:23.525
So
318
00:27:23.985 --> 00:27:25.684
it's much easier for people to
319
00:27:26.290 --> 00:27:30.790
contribute new strategies to the strategy library without really understanding
320
00:27:31.170 --> 00:27:36.775
how the sort of the core kernel of code of generation and shrinking works because
321
00:27:37.235 --> 00:27:39.255
things should be built on top of other things.
322
00:27:39.795 --> 00:27:42.375
There's sort of lots of little internal libraries that
323
00:27:42.870 --> 00:27:50.650
you can more or less work on independently if you find a problem with them. So there have been quite a few contributions around sort of that sort of periphery. The internals,
324
00:27:50.985 --> 00:27:56.365
yeah, no 1 worked on them other than me. And that sort of I think as far as today is by design,
325
00:27:56.665 --> 00:27:57.165
that
326
00:27:57.625 --> 00:28:09.310
right now, a lot of the hypothesis internals are almost a research project. Like, they're very much production software, but they're production software where I'm still figuring out a lot of the theory. I'm still figuring out improvements.
327
00:28:09.965 --> 00:28:13.905
So it's quite hard for other people to come into that
328
00:28:14.365 --> 00:28:19.320
more because of the fact that we'll move out from under them rather than because of the level of sophistication.
329
00:28:19.620 --> 00:28:41.450
And for the moment, I'm okay with that. Once it stabilizes a bit more, I will be writing up a lot of stuff about how everything works and how we're trying to make it more approachable to people so that people are over the meaning of it. And it's also it's actually not that large. So if I were to stop changing things and basically get hit by a bus tomorrow, then someone else could easily figure out how it works. It's just that
330
00:28:41.750 --> 00:28:48.625
right now, no 1 has a very good reason to because I'm likely to change the answer to how it works next month. No 1 has so far.
331
00:28:49.405 --> 00:29:00.990
A few months ago, you went through some public burnout with regards to open source and hypothesis in particular, but circumstances have brought you back to it with a more focused plan for making it sustainable. I'm wondering if you are
332
00:29:01.405 --> 00:29:10.625
comfortable with providing some background and detail about your experiences and reasoning behind all of that. So the reality is I'm still a bit burned out and I'm sort of working against my own best interests.
333
00:29:11.510 --> 00:29:18.410
Part of the problem is that shortly after doing public burnout, I had this really great idea that I just couldn't resist implementing. But fortunately,
334
00:29:18.870 --> 00:29:22.255
also around this time, I got a few contracts related to hypothesis,
335
00:29:22.955 --> 00:29:24.015
doing some training,
336
00:29:24.475 --> 00:29:25.855
doing some custom development.
337
00:29:27.320 --> 00:29:31.740
A lot of the use of hypothesis material was done by me,
338
00:29:32.040 --> 00:29:33.179
paid by a client,
339
00:29:33.880 --> 00:29:45.640
which was nice. So I've managed to sort of make enough money this year that I'm not feeling a complete sucker. But if you do know anyone who wants more hypothesis training or contracting, that will always help because I wouldn't really regard
340
00:29:46.020 --> 00:29:53.115
hypothesis as sustainable right now. And and this is sort of a problem. I think it's it's not just me who has this problem. I know that,
341
00:29:53.675 --> 00:30:00.955
potentially, most of the Python projects struggle quite a lot with this. I don't think anyone has ever paid Ned to work on coverage.
342
00:30:01.450 --> 00:30:06.110
I know that the PyPI project has some good commercial customers, but
343
00:30:06.490 --> 00:30:07.710
I think that
344
00:30:08.010 --> 00:30:15.695
they could use a lot more and they could use a lot more funding given how amazing PyPI is. And this is sort of a problem across open source where,
345
00:30:16.394 --> 00:30:21.500
a lot of the things we build on, we basically go, hey. This works. Great. Let me build my stuff on it.
346
00:30:21.960 --> 00:30:31.304
And then money doesn't really fund back or flow back from the people using it to the people who are watching the 1st place. And even though I'm
347
00:30:31.684 --> 00:30:37.145
doing business on top of hypothesis, to some extent, that's still very true because the stuff that is most useful
348
00:30:37.550 --> 00:30:45.810
to the Python community at large and also the stuff I most want to be working on is really improving hypothesis functionality, improving hypothesis internals,
349
00:30:46.365 --> 00:30:47.585
and generally
350
00:30:47.885 --> 00:30:49.825
making it a more useful tool. And
351
00:30:50.125 --> 00:30:52.145
that's sort of the 1 bit that no 1 is paying.
352
00:30:53.165 --> 00:30:57.909
People are paying me to teach them how to use better. They're paying me to use hypothesis,
353
00:30:59.330 --> 00:31:01.510
to test things for them. But
354
00:31:02.945 --> 00:31:09.845
making hypothesis work itself work better is just not something that anyone so far has proven very interested investing in.
355
00:31:10.225 --> 00:31:11.684
And in the long run,
356
00:31:12.010 --> 00:31:14.590
I don't know how sustainable that is, but for the moment,
357
00:31:15.610 --> 00:31:20.270
I'm interested enough in it that I'll take what I can get in terms of cash flow and
358
00:31:20.575 --> 00:31:21.554
see how things go.
359
00:31:22.575 --> 00:31:23.475
Yeah. The
360
00:31:23.855 --> 00:31:30.700
money in open source discussion is definitely 1 that has been going on for a while and has not yet reached any useful conclusion.
361
00:31:31.080 --> 00:31:34.780
I do know that they're probably going to mispronounce her name, but Niafra
362
00:31:35.080 --> 00:31:35.580
has
363
00:31:36.295 --> 00:31:39.915
produced a series of blog posts and also done an interview with the folks at the changelog
364
00:31:40.215 --> 00:31:41.115
on this topic,
365
00:31:41.495 --> 00:31:42.315
trying to
366
00:31:42.960 --> 00:32:01.929
bring it a little more out into the open in the general conversation, you know, even bringing it outside of, people who work in open source and in the day to day. So definitely 1 that's interesting and 1 worth keeping an eye on. And, also a sort of outlier from the general trend is the Jupyter project. Recently received, I think,
367
00:32:02.309 --> 00:32:04.010
$6, 000, 000 in funding. So
368
00:32:04.630 --> 00:32:07.130
definitely a an example where,
369
00:32:08.070 --> 00:32:12.955
someone at least saw the utility in providing some money to the project, but it,
370
00:32:13.515 --> 00:32:16.414
certainly could do with being a more general trend.
371
00:32:16.929 --> 00:32:20.150
I do get the impression that the scientific and Python community
372
00:32:20.610 --> 00:32:23.330
are a little bit more on the ball,
373
00:32:23.650 --> 00:32:28.905
than a lot of the rest of the Python community. And there is more funding for this tool and for the,
374
00:32:29.305 --> 00:32:31.245
for this side of the ecosystem.
375
00:32:31.785 --> 00:32:43.310
I mean, there's NumFocus, for example, who do a lot of NumPy and sci fi related funding. I don't know how much money they actually have provided to them by commercial customers, but it's at least, it at least exists.
376
00:32:43.845 --> 00:32:45.625
And presumably, I know,
377
00:32:46.005 --> 00:32:47.225
Jupyter is
378
00:32:48.245 --> 00:32:51.090
really heavily used in the scientific community. So
379
00:32:51.570 --> 00:32:55.110
so that's sort of my guess. This is where a lot of this comes from, but I may just be making that up.
380
00:32:55.490 --> 00:33:02.325
There are some glimmers of light at the end of the tunnel, though. Right? We've had a couple of guests on here. Jessica McKellar from the Python Foundation
381
00:33:02.705 --> 00:33:03.924
as just a for instance,
382
00:33:04.304 --> 00:33:05.044
they support
383
00:33:05.424 --> 00:33:15.155
various open source projects that are important to the Python community, so you might wanna talk to those folks, or those folks might wanna talk to you if any of our listeners are Python Foundation
384
00:33:15.535 --> 00:33:18.515
folks with their hands in the purse strings. And, also,
385
00:33:18.815 --> 00:33:21.715
we just had a really interesting conversation around this
386
00:33:22.120 --> 00:33:23.320
with the,
387
00:33:23.640 --> 00:33:28.060
gentleman behind the Read the Docs project. Tobias, do you remember his name? Eric Holscher.
388
00:33:28.440 --> 00:33:34.385
Eric Holscher. Thank you so much. I was gonna say it also came up briefly in our conversation with, Maciej Pielkovski
389
00:33:34.685 --> 00:33:36.930
discussing R Python and the PyPI project.
390
00:33:37.730 --> 00:33:38.230
Absolutely.
391
00:33:38.850 --> 00:33:41.270
With regards to the conversation with Eric,
392
00:33:41.570 --> 00:33:44.630
he brought up the idea of open source projects
393
00:33:45.355 --> 00:33:46.095
as libraries,
394
00:33:46.795 --> 00:33:56.240
as opposed to I mean, not like in the technical sense, but in the sense of everybody contributes to the public library because it's considered to be a an important resource for the community.
395
00:33:56.620 --> 00:33:57.680
In the same way,
396
00:33:57.980 --> 00:34:02.960
all the companies that use open source software should be contributing back to those projects because
397
00:34:03.805 --> 00:34:06.145
those libraries, those open source projects
398
00:34:06.525 --> 00:34:08.305
are enabling the company to
399
00:34:08.845 --> 00:34:12.465
get higher velocity, and it's essentially valued for free.
400
00:34:12.960 --> 00:34:17.059
So if you have any corporations or big companies that are using
401
00:34:17.519 --> 00:34:18.019
Hypothesis,
402
00:34:18.559 --> 00:34:27.655
it might be a reasonable thing to sort of say, hey. Do you like the library? Would you like to see future development happen? Then would you be willing to fund it or something like that?
403
00:34:28.355 --> 00:34:39.785
So thinking in order, so the Python Software Foundation, first of all, their funding work is really good, and I'm not saying anything against them in this regard. But what tends to happen as I understand it, and
404
00:34:40.165 --> 00:34:45.145
maybe I've got the wrong end of the stick, is that what they will typically do is offer you grants for
405
00:34:45.470 --> 00:34:47.170
a specific thing. So
406
00:34:47.630 --> 00:34:51.250
if there were a particular piece of functionality that I wanted to add to Hypothesis,
407
00:34:51.630 --> 00:34:54.644
then I could apply for a grant to work on that piece of functionality.
408
00:34:55.025 --> 00:35:05.250
And they don't really fund sort of ongoing work on projects like this, which is a perfectly reasonable decision, but sort of means that it's hard to use them as a route to sustainability.
409
00:35:05.710 --> 00:35:13.725
In terms of corporate funding, the big problem that I tend to see, and this isn't just a problem with open source funding, this is also a problem I have with
410
00:35:14.025 --> 00:35:16.525
trying to sell into corporates, is that
411
00:35:16.985 --> 00:35:26.200
the people who are most invested in these things aren't the people with their hands on the purse strings. So it's relatively hard as a developer working for a corporate to go,
412
00:35:26.505 --> 00:35:30.905
I have this great open source library I'm using. I really want to support it. I will,
413
00:35:31.785 --> 00:35:32.925
let me give them money
414
00:35:33.225 --> 00:35:36.980
because the developer doesn't have any access to the budget. And
415
00:35:37.280 --> 00:35:44.240
so what you often see is sort of the closer you get to the money, the further you get away from being able to see why you might want on this,
416
00:35:44.945 --> 00:35:46.645
this tool you're using downstream.
417
00:35:46.945 --> 00:35:50.565
And that's sort of part of why I've tried to go down the training route is because
418
00:35:50.865 --> 00:35:55.630
it tends to be much more clearer to people with access to budget that they want to spend their training budget
419
00:35:56.010 --> 00:35:57.470
than it does
420
00:35:58.170 --> 00:36:04.365
to them that they want to throw money at guy who, as far as they know, isn't doing anything for
421
00:36:04.665 --> 00:36:06.525
them. So what's next for Hypothesis?
422
00:36:07.305 --> 00:36:11.700
There are a couple of different directions I can go with Hypothesis right now. 1 of
423
00:36:12.080 --> 00:36:16.900
the big things, which I maybe shouldn't admit in the Python podcast, is that I'm looking into
424
00:36:17.280 --> 00:36:19.220
how to port hypotheses to other languages.
425
00:36:19.555 --> 00:36:20.055
So
426
00:36:20.595 --> 00:36:28.055
in the last couple of months, I've been doing some work on a project that I originally named conjecture. In the end, it sort of has been rolled back into hypothesis,
427
00:36:28.380 --> 00:36:29.200
which is
428
00:36:29.500 --> 00:36:38.585
trying to make the internals less complicated in a way that makes them much easier to make them work in a new language. So I'm looking into I'm thinking about this sort of tool in,
429
00:36:39.205 --> 00:36:47.040
in the context of Java or I got a lot of people asking me about GAR, which is actually 1 of the more annoying languages to port it to for, reasons.
430
00:36:47.740 --> 00:36:51.280
And so at some point, I'm gonna start looking into that, at least partly because
431
00:36:52.315 --> 00:36:55.355
other languages are often a lot better about paying for tooling than,
432
00:36:55.755 --> 00:37:00.335
private community necessarily is. So it's much easier, for example, sell proprietary tools,
433
00:37:01.420 --> 00:37:11.065
in the C plus plus or Java world. And I still have an open source core for all of these things, but it's really gives me a few more options in terms of sustainability front. The other thing is that recently, I just,
434
00:37:11.705 --> 00:37:17.565
sort of I've had effort for 3 months. I suddenly have this brainwave. Usually, it's the result of reading someone else's work, and
435
00:37:17.960 --> 00:37:18.460
I've
436
00:37:18.760 --> 00:37:28.535
been doing a whole bunch of reading about, formal language theory and induction of regular languages. And I've suddenly been going, oh, my god. I've got so many amazing ideas for how to improve hypothesis.
437
00:37:28.995 --> 00:37:38.290
So that's sort of that's the fun direction, which is sort of very impractical from a business point of view. But I'm I'm certainly going to end up spending a bunch of time in it already because
438
00:37:38.590 --> 00:37:41.250
I'm my own worst enemy as far as money is concerned.
439
00:37:41.550 --> 00:37:42.050
And
440
00:37:42.725 --> 00:37:51.120
related to that, there's been sort of I've probably been promising it for about 6 months now, but there's been a long running discussion about how to improve hypothesis
441
00:37:51.820 --> 00:37:55.040
by using coverage information to try and get it to,
442
00:37:55.500 --> 00:37:58.365
target interesting corners of your code because it can track things.
443
00:37:58.825 --> 00:38:09.630
Sort of the research directions I'm thinking about right now should have bearing on that and have bearing on how to make that work better. So if you're if you're not a Python user, then what's next and is most exciting is obviously
444
00:38:10.170 --> 00:38:15.310
the other language stuff. If you're a Python user, it's basically getting better data and being able to find,
445
00:38:15.975 --> 00:38:19.755
magically difficult corners of your code that you wouldn't expect hypothesis to be able to find.
446
00:38:20.295 --> 00:38:25.480
Is there anything that our listeners can do to help the project? Like, where could you use contributions effectively?
447
00:38:25.960 --> 00:38:30.869
So the big thing that is very easy for people to get started on, and I do get those contributions on, is documentation improvements are always great. Because
448
00:38:35.155 --> 00:38:36.695
hypothesis documentation is
449
00:38:37.075 --> 00:38:37.974
better than adequate.
450
00:38:38.275 --> 00:38:41.575
In places, it's even pretty good, but it's a relatively
451
00:38:42.035 --> 00:38:46.350
complicated project, which is quite unfamiliar to people. So writing about hypothesis
452
00:38:47.050 --> 00:38:48.990
here in the context of the documentation,
453
00:38:50.490 --> 00:38:54.035
or in just on their own blogs or whatever is always welcome
454
00:38:54.495 --> 00:38:55.395
and always gratefully
455
00:38:55.855 --> 00:39:08.750
received. Writing new sources of data generation is always useful, and it's very easy for this you to do this in your own project and for the bits that you want to generate for your own stuff and then try and contribute them back upstream. I'm generally
456
00:39:09.425 --> 00:39:12.165
very welcome to people adding stuff to the hypothesis.strategies
457
00:39:12.865 --> 00:39:37.170
module. The other thing is simply getting more projects tested using it. If you have a project that's not using Hypothesis and you would like it to be better tested, just start using Hypothesis for it, both simply because I like it when there's more quality software out there and testing it with hypothesis will almost certainly find bugs in it. But also, the more people use it, the better feedback I get and sort of the better idea I have about what the library needs to improve.
458
00:39:38.430 --> 00:39:45.744
So before we move on, is there anything else that we didn't ask you that you think we should have or anything that you wanna bring up? Nothing comes to mind. Yeah.
459
00:39:46.045 --> 00:39:54.440
Okay. So for anybody who wants to follow you and keep up to date with what you're up to, what would be the best way for them to do that? So I'm available on Twitter as Doctor McKeever,
460
00:39:55.059 --> 00:39:57.799
and I have my blog at drmciever.com.
461
00:39:59.175 --> 00:40:04.315
There is also a tiny letter, which I'm very much overdue to updates this month, which is at tinyletter.com/youguesseditdrmcgaver.
462
00:40:06.950 --> 00:40:09.210
That one's more specifically how office is focused.
463
00:40:09.990 --> 00:40:15.770
Great. So with that, I will take us into the picks. And for my pick today, I'm going to choose Typeform,
464
00:40:16.325 --> 00:40:17.065
which is
465
00:40:18.085 --> 00:40:19.464
a form building service
466
00:40:19.765 --> 00:40:23.464
that lets you put together some very nice looking surveys.
467
00:40:23.925 --> 00:40:24.425
And
468
00:40:24.840 --> 00:40:36.494
if you are on 1 of their paid plans then it also supports logic branching so that depending on the answers that somebody makes to 1 question you can put them down a path of a few different sets of other questions and
469
00:40:36.795 --> 00:40:38.255
just a lot of really great functionality.
470
00:40:38.714 --> 00:40:41.055
It works great on mobile. It has
471
00:40:41.435 --> 00:40:42.015
a good
472
00:40:43.150 --> 00:40:51.970
good way to view the analytics associated with it to see how many people viewed it, what kinds of devices they were on, what their answers were. And it also has Zapier integrations
473
00:40:52.335 --> 00:40:55.075
to be able to pipe that data out to other services.
474
00:40:55.615 --> 00:41:06.890
So I've been using it for a couple of different surveys that I've sent out, So I'll put those in the show notes as well. Most specifically, I put together a survey for listeners of this podcast. So if you have any feedback
475
00:41:07.515 --> 00:41:12.575
that you wanna give, you can fill that out. And also, I put together a survey for
476
00:41:14.210 --> 00:41:18.630
trying to do some market research on continuous integration and people's experience with that
477
00:41:19.410 --> 00:41:22.130
to get some ideas for a project that I've been,
478
00:41:22.529 --> 00:41:44.234
started working on. So if anybody wants to fill that out, that'd be great as well. So I'll put all that in the show notes. And with that, I'll pass it to you, Chris. Thanks, Tobias. I I just wanted to say, the CI form, I just filled it out today, and I was very, very impressed. Like, I was thinking actually as I was filling it out, well, I wonder what he's using because this site is really, really slick. I love the the UX. It's it's great stuff.
479
00:41:46.140 --> 00:41:50.080
So my first pick is a kind of brainless wonder,
480
00:41:50.940 --> 00:41:55.845
game that I've been enjoying lately, video game, for iOS devices called Seashine.
481
00:41:56.305 --> 00:42:04.280
It's a perfect my brain is fried. I'm commuting. I just wanna enjoyably pass a little bit of time kind of thing. Basically, you're a jellyfish,
482
00:42:04.900 --> 00:42:06.840
and you're swimming around in the ocean,
483
00:42:07.540 --> 00:42:37.395
trying not to get killed or eaten, and it it uses the the the touch interface of the tablet beautifully. You swim by sort of making little swishes with your finger. It's it's gorgeous. The graphics are great. The sound design is is really good. It's just an awful lot of fun, and it's free. I'm sure they, at some point in time, try to commodit you know, monetize with something or other, but I haven't hit it so far. My next pick is actually something that we had on the show quite a number of episodes back, but I've been really sort of getting into it lately. Check. Io,
484
00:42:38.530 --> 00:42:39.590
it's a website,
485
00:42:40.370 --> 00:42:42.530
where it, basically, it's a it's a Python,
486
00:42:42.850 --> 00:42:44.855
problem solving site where you get
487
00:42:45.494 --> 00:42:55.810
practice problems in Python, and they've gamified the whole thing, and you can publish your solutions and have them reviewed by other people. And I just I've really been enjoying I've I've been, lately trying to undertake,
488
00:42:56.670 --> 00:43:05.395
improving my algorithmic problem solving skills, And Check. Io has really just sort of made it a a real pleasure and sort of given me the incentive to want
489
00:43:05.855 --> 00:43:07.555
to go further and push harder
490
00:43:07.859 --> 00:43:15.240
in terms of, you know, getting the next level and whatever the case may be. And the interface is really nice, and it's also super easy to use an IDE
491
00:43:15.540 --> 00:43:19.065
to, you know, work on your on your code. They make it really nice.
492
00:43:20.165 --> 00:43:21.625
My last pick is,
493
00:43:22.485 --> 00:43:24.505
a former coworker of both Tobias,
494
00:43:24.965 --> 00:43:25.705
and mine,
495
00:43:26.370 --> 00:43:27.190
Mike Kudermarsh,
496
00:43:27.490 --> 00:43:34.710
currently works for, Product Hunt in San Francisco, has written this excellent series of articles called the junior developer series,
497
00:43:35.045 --> 00:43:39.785
which is a kind of bland sounding name, but a really phenomenal series of articles,
498
00:43:40.565 --> 00:43:41.065
containing
499
00:43:42.060 --> 00:43:43.760
lots of sort of really interesting
500
00:43:44.300 --> 00:43:52.805
kind of off the beaten path tips for people starting out in the industry. Not the usual, like, okay, learn this, don't learn that crud. It's more like,
501
00:43:53.185 --> 00:44:03.089
this is how you work effectively with your coworkers. This is how you, you know, can work in a way that will have you be liked by your teammates. This is how to generate good pull requests.
502
00:44:03.470 --> 00:44:15.615
This is how to, you know, all of that kind of sort of, like, tips that are not covered elsewhere in the standard. Here's how to get started doing dev kind of articles that are so, you know, rampant around the intertubes these days.
503
00:44:16.560 --> 00:44:22.500
I've heard some criticize it for having too many emojis, but I think that if you disregard the article series for,
504
00:44:22.855 --> 00:44:27.515
you know, just over Mike's love of of the emoji, then I think you're doing yourself a disservice.
505
00:44:28.135 --> 00:44:30.780
Push forward and read it anyway. It's excellent stuff.
506
00:44:31.100 --> 00:44:34.080
And that's it for me. David, what do you have for us for picks?
507
00:44:34.860 --> 00:44:42.795
So I've got 3 picks. The first is a recent discovery, and the 2 are sort of long running papers for me. The first is
508
00:44:43.175 --> 00:44:45.115
Make It Stick by Peter Brown,
509
00:44:45.415 --> 00:44:48.810
which is a really good book about learning theory
510
00:44:49.110 --> 00:44:57.994
and in particular about sort of what the science supports in terms of how we learn and through practical advice for how to actually apply it in your day to day life.
511
00:44:58.295 --> 00:44:58.795
So,
512
00:44:59.575 --> 00:45:07.790
this is the sort of thing I personally love reading about, but it's also got some really quite useful tips about how to how to make it stick.
513
00:45:08.650 --> 00:45:09.150
And,
514
00:45:09.450 --> 00:45:15.230
I would recommend checking that. The 2 classic favorites from me are there's a service I use called Beeminder,
515
00:45:16.634 --> 00:45:19.835
and I frequently have to tell people I'm not affiliated with them because,
516
00:45:20.474 --> 00:45:22.894
and not even on commission, unfortunately.
517
00:45:24.060 --> 00:45:34.365
But basically, it is a great way of sort of building good habits and essentially committing to yourself to do things that you would otherwise have days or a week.
518
00:45:34.905 --> 00:45:37.565
And the final thing is simply an author recommendation.
519
00:45:38.650 --> 00:45:42.030
It's sort of a running joke that whenever anyone asks me for fiction recommendations,
520
00:45:43.450 --> 00:45:45.470
I listen to them very carefully, and I go,
521
00:45:47.005 --> 00:45:53.665
but the author you really want to read is, Lewis Macmaster Bujold, and in particular, her Borkoskin books, which are
522
00:45:54.330 --> 00:45:56.110
relatively light space opera,
523
00:45:56.650 --> 00:45:58.750
particularly in the original ones. But
524
00:45:59.130 --> 00:46:01.550
in the later ones, when she has
525
00:46:02.315 --> 00:46:05.055
figured out basically that they will let her keep
526
00:46:05.595 --> 00:46:11.295
publishing these regardless of what she writes as long as it's in the universe, and she started telling a much wider variety of stories
527
00:46:12.140 --> 00:46:15.040
in this universe, and some have been really good.
528
00:46:15.660 --> 00:46:21.645
All right. Well, we really appreciate you taking the time out of your day to join us and tell us more about Hypothesis
529
00:46:22.025 --> 00:46:33.030
and how it came to be and how people can take advantage of it as well as your experiences in building it. So I appreciate that, and I hope you enjoy the rest of your day. Yes. And you. Thank you very much for having me. Thanks. Bye bye.