درباره این اپیزود
In this episode, Lane talks to Alex DeBrie, author of the DynamoDB book. Today's talk covers various aspects such as DynamoDB's comparison with Amazon S3, its benefits, use cases, constraints, and cost considerations, while also covering other AWS and Google Cloud services. Alex also shares his insights into his journey of writing the book on DynamoDB and touches on topics like access patterns, secondary indexes, and billing modes. Alex also shares his professional experiences, including consulting vs freelancing, thoughts of entrepreneurial aspirations, and gives helpful advice for those that are considering pursuing a similar career.
Learn back-end development - https://boot.dev
Listen on your favorite podcast player: https://www.backendbanter.fm
Alex's Twitter: https://twitter.com/alexbdebrie
Alex's Website: https://www.alexdebrie.com
- (00:00) - Introduction
- (01:27) - Who is Alex DeBrie?
- (02:39) - What is DynamoDB?
- (04:15) - EC2 instance
- (05:50) - Amazon S3
- (06:25) - DynamoDB is more like S3
- (07:40) - Difference between DynamoDB and S3
- (08:20) - What do we mean when we say NoSQL
- (10:08) - BigQuery and BigTable
- (12:31) - Some of DynamoDB's benefits
- (13:15) - When to use DynamoDB
- (15:58) - Constraint of number of connections
- (18:06) - DynamoDB is a multi-tenant service
- (19:21) - How does DynamoDB shake up against something like MongoDB
- (22:22) - DynamoDB is opinionated, but it provides good results consistently
- (25:54) - You can only do certain things in DynamoDB, but they are guaranteed to be fast
- (26:42) - Relational Databases - Theory vs Practicality
- (31:08) - How Alex came to write a book about DynamoDB
- (32:15) - What happens when SQL runs, depends heavily on the system underneath
- (33:57) - DynamoDB doesn't have a query planner
- (36:08) - Access patterns
- (38:04) - Use case for Secondary Indexes
- (39:43) - Costs of DynamoDB
- (40:45) - Billing modes for DynamoDB
- (45:26) - Provisioning and planning for expenses
- (48:40) - Super Mario 64 Hack
- (49:34) - What Was Alex's Last Full Time Job
- (51:02) - Consulting vs Freelancing
- (52:23) - Does Alex see himself going back to a Full Time Job?
- (53:07) - Does Alex have any entrepreneurial urges?
- (54:01) - What you should think about before jumping into freelance/consulting
- (56:01) - Authority in the consulting world
- (57:11) - Where to find Alex
یادداشت ها را نشان دهید 🔗
رونوشت 🔗
Lane
You wrote the Dynamo DB book like you're the Dynamo DB guy.
00:00:03:16 - 00:00:05:07
Alex
I kind of fell into it by accident.
00:00:05:08 - 00:00:29:17
Lane
Do you ever see yourself going back to a full time job?
00:00:29:19 - 00:00:52:02
Lane
Welcome back to another episode of the Back in Banter podcast today. I'm joined by Alex Debris. Now, Alex, I think I first became familiar with like the stuff that you're working on through the serverless framework, which we'll talk about. But primary reason I invited you on stage to talk about Dynamo DB, which is a database that's tightly coupled with like the IWC ecosystem that I've never used.
00:00:52:06 - 00:00:58:20
Lane
So hopefully I can ask a bunch of interesting questions, will, you know, we'll be able to discover some things we haven't talked about yet on the pod?
00:00:58:21 - 00:01:10:17
Alex
Gotcha. Well, you already have a few knocks against you there despite admitting you haven't even tried out Dynamo. But now it's it's very niche and interesting and different and yeah a lot of people have it so. But that's great. This'll be a great conversation.
00:01:10:19 - 00:01:27:11
Lane
I hope it'll be fine. I should be able to do a lot of compare and contrast against like the databases that I am more familiar with, like Mongo, Elasticsearch, PostgreSQL. We'll see how that, how that goes. But to start, do you want to take just a second and introduce yourself? Tell us who you are, what you've worked on in the past, that kind of stuff.
00:01:27:15 - 00:01:54:11
Alex
Sure. So, yeah, my name's Alex Debris. I, I think sort of started being in the community and things like that when I worked at Service Framework, as you're mentioning. So I worked for service Framework for two and a half years and just I got involved in the IWC and service community from that as part of that, started working with Dynamo Hub a lot and sort of like actually fell into I went deeper and deeper into Dynamo and I wrote a website about it and then started doing talks and wrote a book about it.
00:01:54:13 - 00:02:07:20
Alex
And then now for the last four years, yeah, for years I've been mostly an independent consultant, Dynamo, Serverless, IWC, just different things like that. So that's mostly what I work on these days, of course.
00:02:07:22 - 00:02:11:18
Lane
So you're actually at the moment, like on your own doing the consulting on.
00:02:11:20 - 00:02:21:02
Alex
On my own, Yeah. I've been on my own for yeah, most of, most of the last four years. I had like a brief five, five month stint in the middle of that. But yeah, mostly on my own for the last five, four years.
00:02:21:07 - 00:02:40:11
Lane
Okay. Wow. We might have to touch on that as well, because we've talked about consulting on this pod before, but like, I haven't really had a guest that's done too much of it, especially in recent years. So we might touch on that as well. That would be fun. Okay, so let's dive right into Dynamo. Like, first of all, what is Dynamo?
00:02:40:11 - 00:02:44:23
Lane
Our listeners are familiar with PostgreSQL. Maybe they've heard of Mongo, but what is Dynamo DB?
00:02:45:00 - 00:03:03:15
Alex
So I like the short phrase I say is it's a fully managed NoSQL database from a U.S. NoSQL database. People kind of know and understand, although I think it's I think it's actually pretty unique from a lot of the other ones. But like, one thing that's interesting is like it's it's from yes, you can only use it on IWC because of that.
00:03:03:15 - 00:03:26:12
Alex
It's fully managed in a way that most databases aren't fully managed. Right. Where it's not like you're sort of picking a server size and server class and different things like that, and they're, you know, they're managing it for you somewhat, but also like, you know, you're still in charge of certain things, maybe some backups or always configuring backups or different things like that, like Dynamo really is more fully managed for you.
00:03:26:14 - 00:03:45:19
Alex
It's more like S3 or ESXi than it than it is like then. Yes. And it's a little different from S3 because like, S3 is like, Hey, you can just start throwing as much data as you want into it. Whereas Dynamo, you know, you can have provisioned capacity. We have to set like how much, how much you can do per second, things like that or you can do on the main.
00:03:45:21 - 00:04:11:20
Alex
But I'd say it's more like those, it's, it's a multitenant service where like, you know, behind this, like they're giving you an end point and behind the scenes they've got like a bunch of different microservices that are working together and sort of everyone's data and metadata and all that stuff is using the same shared infrastructure like in S3 rather than RDF, where it's like if you spin off a Postgres instance, theoretically someone could take you to a data center and like point to a database and say like your RDF instance lives there.
00:04:11:20 - 00:04:15:06
Alex
That's like not true with Dynamo, right? It's split spread across a bunch of machines.
00:04:15:10 - 00:04:33:02
Lane
Okay, so you throw out a bunch a bunch of products that are like, just like recap a little bit. That's like a way back to first principles. So like, let's say I have an easy to instance and it's just like this VM running in the cloud. It's like this box in some of this data center. I could install a database on that box like a my sequel or PostgreSQL.
00:04:33:03 - 00:04:50:22
Lane
This to me would be like the least serverless thing ever. This is the most server awful thing that I could have, right? And I got to manage that box and I got to like, install the stuff and I've got upgrade dependencies and I got to do my own backups when ATO is launched. This branding was kind like this is a managed database and you know, you use that term as well.
00:04:50:22 - 00:05:16:24
Lane
I think it's a great way to describe it, but like and it is like you buy this already s box and now instead of just like a vanilla Linux server, you get this Postgres database, you get a way to connect to it. Versioning is done for you automatically. You can like check some boxes to get backups automated, but like you could still go find your RTC database in a native boost data center right like it exists in a server somewhere.
00:05:17:01 - 00:05:48:03
Alex
Yeah, it's like a specific thing. You're connecting to a specific thing. Like they will give you, you know, an IP address or like some sort of long domain name and you're connecting to that specific server directly, like you're saying, like they're doing the upgrades on that specific server directly and things like that. Or if that specific server had some hardware failure or something like that, like you would notice that as a particular user, even though that's like a managed service, like you specifically as a user, are going to have problems that maybe the person next to you wouldn't have.
00:05:48:09 - 00:05:50:19
Alex
They migrate you to a new server failover or things like that.
00:05:51:00 - 00:06:15:24
Lane
And S3 is like more like the fully serverless thing of like at least my experience, this three is like I'm given some URL to point to, but it's not like the URL of a specific VM. It's just like the URL of S3 and I just tell it which like key value pair. I want to store things and like shove all of this data into this key and then I'll, I'll come back and get it later.
00:06:16:00 - 00:06:24:21
Lane
Backups or I'm assuming there's a ton of redundancy that's just built in like you don't have to worry about that yourself. So this dynamo, it sounds like Dynamo is more like S3.
00:06:24:24 - 00:06:55:17
Alex
And it was more like S3 where it's like a region wide service and you're going to be connecting to the service rather than to a specific server or you're sort of like data plane operations, right? So you'll be connecting to dynamite TV dot U.S. one dot Amazon eight of US dot com or whatever like that doing those operations and then there's you know there's load balancers that are for to request routers and then they'll sort of look up the metadata for your specific data and then figure out like which storage nodes and which specific servers I need to go to to read data or write data or things like that.
00:06:55:17 - 00:07:15:10
Alex
But it's like a more multitenant region wide service that ABC is truly, fully managing for you rather than a more server based one. Even a managed server like RTC Postgres or my sequel or even like Document DB, or if you use MongoDB Atlas or anything like that.
00:07:15:15 - 00:07:32:01
Lane
Feels like maybe the like non program or analog would be almost like using a SAS product or like a like a Google drive where it's like when you use Google Drive, you're not really concerned about where exactly the data is getting stored. You just kind of upload it to Google Drive and trust that later you'll be able to get it off of Google Drive.
00:07:32:01 - 00:07:40:19
Alex
So yeah, yeah, exactly. I think that's a that's a great comparison. I get Google Drive is more S3 like you know directly but like it's in that same sort of vein I would say Yeah, yeah.
00:07:40:19 - 00:08:01:14
Lane
So dynamo like that is a key difference between Dynamo and S3. S3 is like kind of raw. We call it object storage, which is kind of like a different version of file storage, but it's like pretty raw. You just like shoving stuff into it, uploading files, like for example on boot def, we use, we use the GCP alternative, but it's the same exact thing is S3, which is like we're uploading images.
00:08:01:14 - 00:08:18:12
Lane
So like your profile images, they go into that object storage and we serve them from there. So with Dynamo, it's more like a database. It's like we'd compare it to a mongo or a PostgreSQL. And I think you said earlier was this is, this is no sequel. Me That feels like the most useless term ever. It's just not sequel.
00:08:18:12 - 00:08:20:18
Lane
Like, what do we mean when we say no sequel?
00:08:20:20 - 00:08:48:03
Alex
Yeah, exactly. So yeah, I had that phrase earlier and I talked about the operational aspects and sort of that stuff, but like the sort of modeling and developer and user aspects of it. So it's a now we could do so it's a no SQL database. Again, like you're saying, bad terms sort of came out in the, in the Twins where like, you know, everyone's using traditional R.T. BMS relational databases and then you had this sort of new wave of other type of databases.
00:08:48:05 - 00:09:11:06
Alex
MongoDB probably being the definitely being the most popular of those. But you also have like Cassandra, weirdly, Elasticsearch is probably in there, which is like its own funky thing. There's there's not a lot of like, you know, what makes it NoSQL, I'd say probably like the the main throughline for all of them would be like easier to shard your data and split it across multiple machines would be some of it.
00:09:11:06 - 00:09:34:04
Alex
And given that relaxing some of the traditional relational features you have, whether that be transactions, whether that's like schema enforcement type stuff, foreign keys into your joints, like giving up some of those type of features in order to often in order to like shard more easily, which means, you know, theoretically scale more easily and things that can be web scale.
00:09:34:06 - 00:10:03:03
Alex
Exactly. All that sort of stuff. Been weird now I think, over the last ten years, because I think a lot of like the a lot of the NoSQL ones have moved a lot closer to their relational database counterparts. Right. They're adding joins they're adding aggregations, they're adding transactions, different things like that. And Dynamo has transactions, too. So I think a lot of honestly, like a lot of the databases, whether it be MySQL, Postgres, Mongo elastic, I think those are more similar to each other than Dynamo.
00:10:03:03 - 00:10:08:19
Alex
I think dynamos like way off on it's island in like a different realm of of things generally.
00:10:09:00 - 00:10:17:18
Lane
It also very different from like BigQuery because BigQuery is also like this kind of serverless managed thing in the Google world.
00:10:17:20 - 00:10:27:10
Alex
Yeah, and I always get confused. Okay, so big query. Big query is more like the old AP one, right? Like a little more like like, Oh, that like aggregation type queries. Is that right?
00:10:27:12 - 00:10:38:16
Lane
It's pretty rare to use. And so, like, now that you say that, I'm like, okay, yeah, obviously things are pretty different because but with BigQuery you wouldn't really use it for like a transactional application. You're more like running analytics on it.
00:10:38:16 - 00:10:55:07
Alex
Yeah, exactly. But, but, but Google does have is a big table and I think that's actually guess fairly close to Dynamo. So big tables like closer Dynamo I would say like operationally both big table and big query if I understand correctly, I think are more similar to Dynamo, whereas like more multi-tenant type stuff and you're not provisioning servers.
00:10:55:09 - 00:11:01:14
Alex
I'm not percent on that. So I don't do a ton in the Google world, but my big table data model wise I think is going to be closer to Dynamo.
00:11:01:15 - 00:11:28:22
Lane
Experience with big table. Well, I want to hear you talk a little bit more about Dynamo, because I'm not I'm definitely not an expert on Dynamo, but like in just to like set the tone for like the comparison I have done big table. We did this really dumb thing at this company I was at. This is like this is, this is fucking terrible idea We used big table as like our CRUD database for just like generic crud services and big table has some weird strengths on it.
00:11:28:22 - 00:11:49:21
Lane
At my opinion. Do not lend well to like a startup trying to build CRUD application because it's a very key value store. And but the weird thing is that the keys you can kind of sort them but like they're Lexa graphically sorted. So all of the conversations we had at this company around like database model was just like, get the key right.
00:11:49:21 - 00:12:00:10
Lane
Like we have to make the keys in the proper order so that we can sort correctly. And if you eff up the key, the only thing you can do is like copy the data into a new key structure.
00:12:00:12 - 00:12:19:12
Alex
Oh, I going to. So you could have just substituted big table there and enter Dynamo like that about Oh God. Okay save It's like all right exact same thing. Yeah And that's why like Dynamo is super weird there there there are pros and cons time. I'm like I love Dynamo for for certain things. But then there is like, I don't begrudge anyone that doesn't use it because it is weird and different.
00:12:19:12 - 00:12:24:17
Alex
And especially like if you're talking about startup flexibility type world, it can be restrictive and limiting in that sense.
00:12:24:17 - 00:12:31:21
Lane
So this feels like crazy, right? Like that's why Google uses big tables, like run. Google is like it's like the scale is enormous.
00:12:31:23 - 00:12:59:07
Alex
I mean, like the nice thing about Dynamo is like you're just trading off a different you're getting these sort of like operational and scalability benefits and especially like predictability, like basically once you sort of model it in a certain way, you know that like, hey, it doesn't matter how many concurrent queries are being run or if someone's running a giant aggregation query or if we have way more users or way more data or anything like that, you know, you're going to get the same predictable performance for a given request.
00:12:59:07 - 00:13:15:00
Alex
When you have ten users as when you have 100,000 users, when you have a million or ten, you know, a billion users, whatever, it's going to be predictable in that sense. But it does you lose a ton of the flexibility that you get from that relational model then. And so I get the upsides of that as well.
00:13:15:04 - 00:13:38:13
Lane
Let me ask this question. After I used Big table, I would not recommend using it as like the main CRUD store for like a typical SAS application. Right? But I can definitely see use cases for it either like when your credit application gets so large that you just need the scale benefits or you have like a specific part of your application that has specific data needs, right?
00:13:38:13 - 00:13:57:06
Lane
Like maybe it's not the main crud part of your application like users and groups and organizations billing, but maybe it's like, you know, we're piping in social media data and we have like this massive amount of data that we just need to store somewhere and we kind of need to maybe like sort by timestamp or something. When do you think about using Dynamo?
00:13:57:06 - 00:13:58:19
Lane
What are the use cases that are, that are big?
00:13:58:19 - 00:14:32:13
Alex
Yeah, I like, I think that like small focus one is is really are like a lot of data but like pretty fairly simple model like you're saying key value or like you know groups of data that you need to order by timestamp. That's really good. I see people use it for sort of everything and you absolutely can. I think one of the reasons Dynamo took off, the reason I got in Dynamo was because it worked so well with serverless applications, especially in the early days, where, I mean just like a lot for a lot of reasons, if you were using a lambda, it did not work well to use a more like server fold database,
00:14:32:13 - 00:14:49:06
Alex
like as my school mongo. Again, anything that had like a specific server that you could point to and part of that was around connection limits, right? A lot of those because they're a single server, they're often limited in how many connections they can have to them. Right? And my cycle's gotten better at that Postgres has gotten better at that.
00:14:49:06 - 00:15:04:07
Alex
But it's still like usually you don't want to have a ton of those, which usually you solve with like connection pooling in your in your application layer or something like that. But if you have Lambda functions now, it's like one request per function. Each one is doing a connection like you can't hold those open. So that was a problem.
00:15:04:07 - 00:15:21:16
Alex
I'd also say like, you know, often you want to have those, those database resources in a VPC. Again, since it's a specific server, you want to sort of protect it from outside Internet traffic and things like that. It used to be a real pain to put Lambda function that a VPC like serious cold starts and just like an operational pain that's gotten better as well.
00:15:21:21 - 00:15:47:00
Alex
And then also just like some provisioning and permission stats where it's a lot easier to use Dynamo than it is to use a relational database with that. So a lot of people did that and a lot of people built, you know, their applications on on top of Dynamo, including the core of their stuff. I will say like the core of your application that like you're saying, users and groups, it has like many too many relationships, you have flexibility, you have people changing their usernames or emails or things like that, like more mutable data that can be tricky.
00:15:47:00 - 00:15:58:05
Alex
And that's like where you might want to think about it a little bit and see what your appetite is for, for something like that. So it sort of depends on like what constraints you want, what operational profile you want, different things like that.
00:15:58:05 - 00:16:17:19
Lane
I love that you brought up the constraint of number of connections because at least when I was a new developer, I assumed that usually with a database that the things you'd run into most quickly are like running out of CPU on the database or running out of memory or running out of disk. In fact, I think I really overestimated the problem of disk.
00:16:17:21 - 00:16:52:12
Lane
Disk tends to be like not a problem at all. Yeah, you just add more disk. Like that's not like CPU memory. The thing that like kind of caught me by surprise was how often the number of connections is actually the problem. In fact, we just ran into this on boot dev a few months ago. What was happening was in our so our back is written and go and we'd have like, you know, it's a CRUD application, it's a JSON API, we'd have a handler and in the spirit of optimization, his handler would make like let's say four different SQL queries, each query would be ran in parallel.
00:16:52:12 - 00:17:15:23
Lane
So we were spawning a go routine for each one and making a database query in each go routine. What that does under the hood is open for connections to the database. Right? We were using like the built in connection pooling and stuff, but when we started to have more traffic, like we were hitting our I mean post-crisis like default connection limit on like Google's already has equivalent cloud storage or classical, it's like 100 connections.
00:17:15:23 - 00:17:16:23
Lane
It's like say yeah.
00:17:17:00 - 00:17:28:00
Alex
To fail right. Yeah. To 50 or 6500 like where it starts to like really even throw a blob. Although I was talking to someone yesterday and he says it can easily do thousands now so we'll see. But yeah, it is like much lower than you would expect.
00:17:28:05 - 00:17:57:21
Lane
Like I probably could go update some configs, right? Like I'm using whatever the defaults are at the moment, but like defaults are there for a reason, right? Like, you know, you don't want to like overuse your hardware or whatever and get actual crashing crash loops going on. Yeah, but I was just kind of blown away because, you know, our CPU and our RAM usage were selects crazy low because our queries were fairly efficient and it was just like, you know, we have 15 people on the site at the same time they're hitting these handlers that are opening four connections each like, yeah, blows out.
00:17:57:23 - 00:17:58:19
Alex
Hit those issues.
00:17:58:19 - 00:18:06:04
Lane
Yeah. It's a real, real problem that yeah, it sounds like you know using Dynamo or I'm guessing like Spanner or BigQuery or Big Table or whatever, it takes care of some of that.
00:18:06:06 - 00:18:23:19
Alex
I mean I think the big thing with, with Dynamo the reason it doesn't have that those issues is again it's a multitenant service. You're not hitting a single server, you're hitting a fleet of load balancers that can terminate connections that then go to like request routers. And those are both like stateless layers that they can just scale out infinitely and all that sort of stuff.
00:18:23:19 - 00:18:39:08
Alex
So then it's like you don't need to worry about overloading your individual service, whereas like Postgres, you know, they're going to, they're going to create a new process for each connection you have to it, and that requires resources on the server. And that's why they want to want you to limit the number of connections you make to it.
00:18:39:09 - 00:18:58:05
Lane
That makes perfect sense. Okay, cool. So kind of contrasted it a little bit, especially like against some of the SQL databases like Post Grasser or RTX, I guess in particular which are listening like this is A.W. s managed version of skill databases. So I think is it just stress my sequel or these are.
00:18:58:05 - 00:19:05:18
Alex
Oh, they have, they have Oracle, they have SQL Server and they've recently support for DB to have you use DB to ever.
00:19:05:18 - 00:19:08:13
Lane
Have not. I forget about Microsoft in Oracle someday.
00:19:08:13 - 00:19:21:22
Alex
Yeah, yeah, yeah. As well. But yeah, every event this year they add support for for DB too, which I imagine that's just like a bunch of migration, you know, moving on prem to the cloud use cases. They want to support that probably not a lot of people like spinning up new DB two instances these days.
00:19:21:24 - 00:19:25:21
Lane
So how does it shake up against something like Mongo?
00:19:25:23 - 00:19:52:10
Alex
I think they get grouped together a lot and they're actually quite different. I would say like the comparison I always have for like Mongo versus Dynamo is like Mongo is libertarian and Dynamo is authoritarian, so Manga manga lets you do anything well, lets you write any query you want. And, and again, let's, let's go back like a lot of these NoSQL databases, they're splitting your data across shards or partitions or something like that.
00:19:52:10 - 00:20:22:21
Alex
But if you have, you know, a 100, 100 gig database, maybe you have, you know, three shards or three partitions and each one holds 30 gig or something like that. Right? 33 gigs. So something like that. They're splitting across these different storage nodes and they're often going to be splitting by some sort of key. And ideally that key is used in your query, just sort of, you know, So when a request comes in, it can just say, Oh, I just need to go to this one shard and do that query there rather than hitting all the shards to take, you know, that sort of query.
00:20:22:23 - 00:20:42:08
Alex
Mongo will enforce that. Mongo like if you write a query that needs to go across all your things, it's just going to that query will come in, it'll go to all three of your shards or all ten of your shards or how many you have for fetch. Those intermediate results there go back to the router is going to like combine those results and figure out what you want and then send you the final result.
00:20:42:08 - 00:21:02:03
Alex
Right. Which means if you're doing if like most of your queries are like that, it means that horizontally scaling isn't going to help you as much. Because if you add another shard on there, well, all those shards are still getting all those queries. So like continuing the further shard, you're still going to have like this baseline level of usage that's going to all the shards.
00:21:02:03 - 00:21:21:09
Alex
You're just not going to get as much bang for your buck there. Dynamo, on the other hand, is like going to religiously enforce that. You include what's called a partition key in your data access patterns. And that partition key is how they're splitting your data. So now when you have a request come into Dynamo, it hits that request router.
00:21:21:14 - 00:21:43:23
Alex
That request router can say, okay, this partition key lives on this specific partition, I'm going to forward it there. And now only that partition is involved in servicing that request. And that means like as your database grows, they can add more partitions and you get essentially like linear scalability, right? Because because you're not doing multi partition reads or writes generally for that sort of thing.
00:21:44:01 - 00:22:03:06
Alex
You can use Mongo like dynamo and only write queries and model your data. So it's only hitting a single shard and sort of getting that linear scalability or mongo is going to let you do whatever you want and do these sort of like scatter gather type queries where they're hitting all the shards in parallel to like fetch and handle that specific result.
00:22:03:06 - 00:22:22:02
Alex
And now you're not getting all those like horizontal scalability benefits that, that, you know, no sequel is sort of aimed at, but it is giving you a lot more flexibility. Whereas like Dynamo, if you're access pattern change or if you have a query that that needs to span across your database or across like larger chunks of data, it's a lot harder to you're sort of like manually doing it yourself in different ways.
00:22:22:08 - 00:22:45:10
Lane
I don't know, like this analog is or this analogy is not going to be perfect, but I think it'll it'll maybe get us somewhere, not like, you know, dynamically type language like Python, kind of just do whatever you want and then like, maybe it works. Yeah, Yeah. Whereas like in a statically typed language, like O or TypeScript or C or whatever, forced by the compiler to like, adhere to certain ways of writing your code.
00:22:45:12 - 00:23:04:00
Lane
And then if it compiles like it's, you know, you're kind of eliminate these, this whole set of problems when you run the thing I experience working in elastic and Mongo is yeah. Like you can kind of like write any query. I guarantee to be fast you have to run it and see how fast it's going to be. It sounds like Dynamo might be more similar to like we've already talked about Big Table.
00:23:04:00 - 00:23:21:03
Lane
I think another analog might be Firebase, which is like you have this really restrictive way that you're allowed to query the data, but when you do query it that way, it's kind of guaranteed to be fast, right? You're not going to have this like I don't know who to the end blow up on the efficiency of the query.
00:23:21:05 - 00:23:36:20
Lane
Like you said, the tradeoff there is you have to be extremely strict about a rigorous about how you structure the data, because if you want to change it later, at least and I don't want to put words into your mouth about Dynamo, but like my experience, the big tables, like you really have to copy the data into a new format.
00:23:36:22 - 00:23:37:02
Lane
Yeah.
00:23:37:02 - 00:23:54:14
Alex
Like if you want to change the primary key of your data of an item in Dynamo, you can't just update the primary key. You do have to copy that new newly new item that Dynamo does have like secondary indexes where you can like, you know, add a secondary index and basically re index that data in a different way that allows for different sort of read based access patterns on that.
00:23:54:16 - 00:24:15:12
Alex
And like just side point here, one thing that's interesting about Dynamo and Mongo there is like if you add a secondary index in Dynamo, it's going to copy your data onto different partitions, right? For that secondary index, like in a different shape, it's going to do that asynchronously. So when you write an item, Dynamo is going to write it to your main table with that, that main primary key for that.
00:24:15:12 - 00:24:33:16
Alex
But it's also going to asynchronously replicate it to these to your secondary indexes, which are like on different partitions elsewhere. So again, when you create that secondary index, you're getting that same thing where it's like, okay, which partition is this on? I can go right there and get it very quickly and you still get that like consistent performance on that for that secondary index.
00:24:33:18 - 00:24:49:24
Alex
Whereas Mongo, Mongo also lets you set up additional indexes on your data and if you happen to index on something that doesn't correlate with your, your shard key, you know it's going to be indexed locally on each shard. But now when you do a query, that's when it has to fan out and hit all those different features. And though each reader indexes to do that.
00:24:49:24 - 00:25:22:19
Alex
So like dynamo sort of like forces you into that and so that's interesting like that. So to go back to your like question you asked, unlike the dynamic versus type or static languages, that's like interesting in some sense, but I don't want to overstate what Dynamo does there. I do think it eliminates a class of problems. It eliminates the problem of I issued a query to Elasticsearch, and sometimes it takes 10 milliseconds and sometimes it takes a minute, you know, and it's like the same query or just like a different parameter or a different data shape in the background.
00:25:22:19 - 00:25:51:07
Alex
Like you don't have that problem at all. With Dynamo, every query you do to Dynamo is going to be like within a fairly bounded time window, so you'll get a response quickly and it might even be an error response saying, Hey, you're overloading our database, you're throttled rather than just be like a super slow query. Right? But the one thing I would say is like sometimes because of how Dynamo is structured, like if you have an access pattern where you just by nature need to read a lot of records and filter them out or something like that, that access pattern as a whole can take a while because maybe you need to issue a lot
00:25:51:07 - 00:26:07:05
Alex
of requests to Dynamo to satisfy that access pattern. So it's sort of like the individual to request a dynamo is always going to be sort of time bounded within this very narrow window. But it could mean your entire operation that you're doing is so can be slow. And I do want to be kind of.
00:26:07:05 - 00:26:10:14
Lane
Forced to do multiple requests to get the data that you want.
00:26:10:18 - 00:26:27:05
Alex
Potentially. Yeah. And the reason I make that distinction is sometimes people write me a dynamo so fast, like you can do whatever you want with it. It's like, No, you can't do whatever you want with it. You go and do certain things and because you can only do those certain things, those certain things are guaranteed. Be fast. But if you have other things that don't fit into that and like now, you need to compose lots of those things together.
00:26:27:06 - 00:27:05:18
Lane
I recently wrote a course on school for boot dev, and one of the interesting things about SQL and it's not even well, it is, it is kind of I want to mix up like SQL, the language and like kind of the theory behind relational databases, but like the theory behind relational databases is actually interesting. It's really like academic and theoretical in the sense that like you can structure your data in, in such and such a way and we are going to like the word we use is normalized or we're going to normalize the data so that we're like never storing things twice and everything can be related in a proper and sane way.
00:27:05:24 - 00:27:20:15
Lane
And then we can we are like once the data is modeled in that relational way, we can query it however we want and it's like insanely flexible, right? Like as long as you model the data correctly upfront, being super flexible. And the nice thing about that when you're building like web apps, for example, is you get the data model right?
00:27:20:20 - 00:27:43:16
Lane
And now your application logic like what you want to actually happen for the end user. You could change that very quickly. Like one day it's doing X and then your customers complain you like you just change it to Y. You don't have to like go re-architect your database. Problem with that is like, it's all really great in theory, but the practicality of like keeping those queries that go into this standard model of data fast, it's actually like really, really hard.
00:27:43:16 - 00:28:14:01
Lane
And like these, these big databases like PostgreSQL, my school and even like Cockroach was trying to like do it horizontally, scaling a hard time. And so my understanding is like the philosophy behind something like Dynamo is like, well, okay, we get that. Like, you know, the theory behind data based normalization super useful. But in the practical sense, we've seen how slow it can be for certain use cases and is that like an explicit tradeoff in your mind of like, we're doing this for practical reasons, we're doing this for speed or in this scale.
00:28:14:01 - 00:28:32:21
Alex
Like you're saying like that flexibility is nice, but then it comes with like this sort of unknown cost or like it depends like on how sort of diligent you are about using the freedom that you have. Right? Because the problem with like a any, any relational query is it's like sort of unbounded in terms of how many rows you can go across.
00:28:32:21 - 00:28:52:16
Alex
Right. If you do like, hey, I want to find like for you which users have watched the most amount of video or complete the most lessons, right? And now you need a go probably to join across some tables, you know, do a group buy on that user ID and do like a some on or account on the number of rows they've done or maybe some on the total video time, things like that, which if you have ten rows, it has to read ten rows.
00:28:52:16 - 00:29:13:24
Alex
If you now have 10 million rows now it has to read 10 million rows and there's no way around that. So like there's this unbounded nature of it, especially when you're talking about aggregations or potentially joins different things like that that can just give you like significant variability in your performance, right? Dynamo doesn't have that probably because Dynamo doesn't do aggregations.
00:29:13:24 - 00:29:33:10
Alex
You can't do any aggregations in Dynamo. Like you just have to read the data that can do it. It also limits on how much data you can read in a single request. So that's one mag of data. So you just can't do anything more on that. So I mean, one thing is like it sort of pushes those constraints upfront for you and like makes you think about those and just playing around them in a way that you could do in a relational database as well.
00:29:33:10 - 00:30:02:07
Alex
But people just often don't because they have the freedom to sort of do whatever they want with it, you know? And the other thing like with like related to that is Dynamo doesn't have a query platter, right? Like you're basically getting just direct access to items and sort of like the, the indexes themselves rather than having like this indirection layer of the query planner, it's doing a lot of that work for you and I like that was something that's kind of eye opening for me when I came to Dynamo because I didn't really understand that much how how databases worked.
00:30:02:07 - 00:30:32:01
Alex
I knew school pretty well and could do that stuff, but didn't like understand what was going on. And one thing I really like about Dynamo is it just like gives you a better intuition on how databases work now and like even if you go back to a relational database or Elasticsearch or something like that, you just get a sense, right, Ooh, that's going to be an expensive thing because I know how much data it's touching or potentially touching and the transformation that has to do and just like all those sorts of different things and gives you a, it gives you like a little bit of a feel for that, even if you don't end up
00:30:32:05 - 00:30:34:05
Alex
sort of going all in on, on Dynamo.
00:30:34:07 - 00:30:52:05
Lane
Okay, Yeah, that makes sense. So I want to circle back to query planning in just a second. But I also just realized, like, I don't think I did your original intro justice. Like people don't realize like you wrote the Dynamo DB book, like you're the Dynamo DB guy. We don't, we don't skimp on the quality of the guest on this show.
00:30:52:11 - 00:31:06:20
Lane
So you're not just hearing some random person that's like use Dynamo. DB Talk about Dynamo would be like Alex Debris wrote Dynamo DB book. Go check it out. I'll drop a link to the description below and then also plug it again at the end. But I just want to make sure that like.
00:31:06:22 - 00:31:21:12
Alex
Yeah, thanks. I appreciate it. Mike Yeah, it's kind of weird. I kind of fell into it by accident and I don't have like a deep systems training. I actually went to law school and like, I'm like a self-taught developer overall, so I don't have like, that deep thing. But just like, again, sort of fell into it in the service world.
00:31:21:16 - 00:31:41:05
Alex
Some videos by a guy named Rick Houlihan that used to work for Dynamo Networks for Mongo and then just talked with him, talked with some other people and sort of just like kept going down that dynamo route. And the one thing I liked about Dynamo is it is very understandable. I feel like you need to learn like four or five rules and you learn how it works and you can apply it in all these different ways.
00:31:41:07 - 00:31:45:20
Alex
And for whatever reason, a lot of people don't like to learn those those four or five rules, and it keeps me busy.
00:31:45:20 - 00:31:47:05
Lane
I guess learning is hard.
00:31:47:07 - 00:32:07:02
Alex
Yeah, exactly. So it's like very predictable and just like the amount of like the number of factors that that go into any one dynamo query or data model, just like way fewer than, than a relational database. We have to worry about how many connections, how big like things they're doing, how big my buffer pool for doing flip queries on the hour or something like that.
00:32:07:02 - 00:32:15:13
Alex
That's slowing everything down, just like a bunch of different factors there that sort of don't apply to the Dynamo world and help can help you predict performance in a nice way. So I like that aspect of it.
00:32:15:14 - 00:32:29:15
Lane
I want to I want to just circle back really quick. That skill builder, that query builder part that you mentioned, because that's really, really fascinating to me when I write a skill query and ask you a query, I'm writing this language structured query language, and the purpose of language is to give me this like really flexible way, like.
00:32:29:15 - 00:32:41:14
Lane
EXPRESS how I want to extract data from a system. But a lot of things support ask you out post-arrest my sequel sequel like I don't know there's like sequel layers on top of even like no sequel no.
00:32:41:16 - 00:32:43:14
Alex
DVD support sequel unbelievably.
00:32:43:16 - 00:33:04:08
Lane
Yeah right Like you just got to write like sequel all the time. What actually happens when the sequel runs is so dependent on the system underneath. And I think sometimes we don't realize that because, like, you know, when you're writing a language like, see pretty close to the metal kind of kind of know what's going on, you're compiling for a specific architecture on a machine.
00:33:04:12 - 00:33:22:05
Lane
Like obviously it's still higher level than like assembly, but whatever. But like, you know, like, would you compare it to Python? Like you really get kind of what's going on at the machine level as well? I think it's like actually this really high abstraction layer that sometimes we forget because I think sometimes we compare as well. It's like an or M, like a sequel is like the lower level thing that we're doing.
00:33:22:05 - 00:33:43:11
Lane
But really there is this entire layer of abstraction that's going on and figuring out what's going on underneath the skull is exactly like you said. You need this like query planning tool. So for example, boo dev, when we have a query that we see is consistently taking a long time, say like 2 seconds to complete, we'll go in and we'll look at the query planner for the tool and see like what's actually happening, which indexes are being hit, right?
00:33:43:11 - 00:33:58:02
Lane
What's the bigger complexity of the different stages of the square that's being run? Because it's actually not obvious from the sequel itself. Sometimes it sounds like with Dynamo, you, you maybe don't don't need that query planner. Is that what I want there just like.
00:33:58:03 - 00:34:25:16
Alex
Isn't a query planner. So like with a relational database and this is why I say Mango's ongoing Elasticsearch are actually like more similar to Postgres in my sequel than they are. A dynamo is like both of them have a query planer layer where like you're going to say what data you want back and then it's going to pass that there's something in between the physical data and you and it's going to parse that out and say like, okay, based on that, it has three these three filters, it has these two joins and things like that.
00:34:25:18 - 00:34:45:06
Alex
It might look at some table statistics that it's been maintaining about like cardinality of different values in the table or things like that or like how many values are in this table to decide which kind of joint to do, when to do that, join, when to apply the filters, all these different things and sort of choosing how to do that and then going in and executing that particular thing.
00:34:45:06 - 00:34:58:14
Alex
Right? So there's this layer between you on that and based on that, it can be hard to understand, like what are its statistics bad? And it's sort of using the wrong index. Maybe it doesn't know enough about my data and or there's this quirk of my data that would make it run faster if it sort of knew about that.
00:34:58:14 - 00:35:14:21
Alex
But it just doesn't know about it. Or like, is it using the wrong type of drawing or is it just like it's hitting too much data? I wrote sort of a bad query and didn't realize what it's doing there with Dynamo. There's there's not there's not a query later layer, right. Like the API, the API you're doing are like basically direct access to your index.
00:35:14:22 - 00:35:36:09
Alex
You have like your key value type operations, get item put item update item delete item. You know, if you have a tree index, you understand that's operating on a single item, you know, it's a log in to find that particular item and then you just operate on it. And then there's a query operation that basically says START it started this record and read a continuous set of items till till this record till this is true, right.
00:35:36:09 - 00:35:53:15
Alex
Or until it hit ten items or until you had a mag or whatever conditions you want to have on that. And that again is just like reading a B index. So it's like it's more like direct access to a B tree than it is like this intermediate layer that's saying like, okay, there are these five different indexes that I sort of need to work.
00:35:53:15 - 00:36:08:24
Alex
And because of that, you don't get sort of like index intersection or joins or like some nice fancy things that you can your planner can do. But because that you do get like that the predictability aspect of that that again and that makes Dynamo attractive.
00:36:09:01 - 00:36:26:16
Lane
That scan operation that you just mentioned seemed to me to be kind of like the unique thing about Big Table and I was using it would be like, you get like this order and look up to a row and then we'd like scan some number of rows, usually, like you said, until like the key changes to some value and like read all of those rows out.
00:36:26:18 - 00:36:33:20
Lane
Yeah. Is that usually timestamps or is it usually like go look up a thing by a key? And then we want like the date range from here to here very.
00:36:33:20 - 00:36:45:06
Alex
Often timestamps that's can the big one maybe alphabetical type things for for certain things but like a timestamped is going to be the most common one for that and like you'll be surprised about how many of your access patterns.
00:36:45:06 - 00:36:46:01
Lane
Are.
00:36:46:03 - 00:36:54:15
Alex
Manipulating individual record or retrieve a continuous set of records using the single filter like this user's most recent ten things or.
00:36:54:15 - 00:36:55:00
Lane
Or by.
00:36:55:01 - 00:37:13:11
Alex
Things like it. So like Dynamo is not flexible by but you'd be spread like the more sort of apps you build these registrars. Hey, these like the same five patterns I keep using over and over and I can actually sort of plan for those and, and model my dynamo correctly for that. In a lot of cases, I'm.
00:37:13:11 - 00:37:31:13
Lane
Thinking of like the user's table on dev or something and like, maybe this is an unfair comparison because maybe you wouldn't throw a user's table in there. Let me think of something else. Like, okay, like lessons. Like we've got the concept of lessons, right? And like as a user, you complete lessons and there's a few different timestamps that are like on that record.
00:37:31:15 - 00:37:46:07
Lane
One is when the lesson was completed at and one was is like when the solution was viewed as like if the student kind of cheated, quote unquote, and like viewed the the solution, if you want to order by both of those, sometimes.
00:37:46:09 - 00:37:50:11
Alex
For a particular customer or across larger groups, either.
00:37:50:13 - 00:38:04:14
Lane
Single player like single user who different timestamps on the same record, you're going to want to order by both of those. Is there an easy way to store that data once and due to different order, buy queries, or am I basically restricted to like copying this data into two different rows?
00:38:04:16 - 00:38:27:16
Alex
So that's where secondary indexes will work. Like so you can have your main table and maybe you have partition key of user ID that's going to split your data across these different partitions and sort of grouping by that value. And then you have what's called a sort key maybe on completed at right. And now if you go to your main table and you say, Hey, give me Lane's records me the most, the ten most recent ones he's completed, that's going to hit on our main table.
00:38:27:18 - 00:38:55:00
Alex
And now and now you say, Oh, but I also want to sort a different times by solution view. That one, you can create a secondary index where the partition key again is that user ID, but now the sort key on the secondary index is that solution view that thing and get that. So now again, it's going to what's going to do when you write that record to your table that's going to write it to your main table and then asynchronously replicate it to your secondary index as well and allow you to query on that sort of pattern.
00:38:55:00 - 00:38:55:06
Alex
Right.
00:38:55:06 - 00:38:57:08
Lane
And it's like a managed copy.
00:38:57:10 - 00:38:58:16
Alex
Yes, exactly.
00:38:58:16 - 00:39:03:06
Lane
It's like a copy that the system takes care of for you rather than like your application code having to do that.
00:39:03:08 - 00:39:21:06
Alex
Yeah. So in some sense, you know, it's like an index in a relational database in some ways, you know, it gives you different ways to access your data. I'd say the big difference would be, hey, it's going to be on different infrastructure somewhere else that you then query directly using that secondary index. But it feels the same for you.
00:39:21:06 - 00:39:25:05
Alex
You only have to write that record once and it's copying it for you.
00:39:25:07 - 00:39:43:02
Lane
That's like that. Implementation details abstracted away from you as the use of okay, that makes a lot of sense to me. One last question I have about Dynamo before I do want to talk a little bit about your consulting work over the last few years. I think that'll be really interesting, but so last question, what Dynamo is like?
00:39:43:06 - 00:39:52:11
Lane
How do you think about the cost of Dynamo and like when is it cheaper to use it and what are the common gotchas for it being more expensive?
00:39:52:16 - 00:40:16:06
Alex
It's it's hard with to think about cost because like it is a managed service. And so, you know, people talk about like total cost of ownership and stuff like that. And if you don't have someone sort of like managing this for you, how much are you saving that way or things like that? I would say like one thing that's nice about Dynamo is you can scale your capacity up and down quite easily.
00:40:16:06 - 00:40:30:03
Alex
Like you can do that very regularly. And so if you have a spike of traffic, you can scale it up to handle that. If it goes down, you can scale that back down, which you couldn't do in a relational database, right? Usually you need to sort of provision for the peak and that's where it is. And then you have like super low utilization.
00:40:30:03 - 00:40:45:20
Lane
Oh, this is interesting to me. I have a I have a question. So so you actually I was imagining it is this like magic control plan where you didn't have to think about infrastructure at all. You have to kind of at least somewhat think about the provisioning for your like cluster or whatever.
00:40:45:21 - 00:41:05:07
Alex
So there's two billing modes for Dynamo. One is the original one is the provision one where just the Dynamo billing is like I think in itself a basically in dynamo, you're going to get charged based on read capacity units and write capacity units. So like if you're reading four kilobytes of data, that's a request. If you're writing a kilobyte of data, that's a right capacity unit.
00:41:05:07 - 00:41:28:14
Alex
If you write five kilobytes of data, they're charging five write capacity units, right? So it's all sort of based in in those sorts of things. The provision mode is you say in advance this is how much capacity I want on a per second basis for my table, and that could be five write units and 2000 read units. And that can change all the time throughout the day as you need to.
00:41:28:20 - 00:41:39:23
Alex
We just add that to whatever you want it to be so like that. You're sort of like managing the capacity for that. If you happen to exceed it on a per second basis, you'll just get a 500 back. It'll say, Hey, you've, you're being throttled right here.
00:41:39:23 - 00:41:45:10
Lane
So you feel like I'm not spending more than this. This is like the amount I want to spend. Give me that. Okay, cool.
00:41:45:11 - 00:42:12:00
Alex
Yeah, the the newer one. This is 2019, I think is basically on demand mode where you are paying per individual read and write to Dynamo TV in like a a true like fully service way basically like you know if you have zero traffic you're not paying for anything, you're paying for data storage. If you have data storage, but you can scale down all the way to zero and it's not going to charge you anything for that scales up pretty rapidly.
00:42:12:00 - 00:42:42:01
Alex
Like basically the scaling, the way it works is like it'll scale essentially instantaneously to two times your previous peak and you won't see any any your previous peak, right? So if you have like peak of traffic comes in and you're doing 5000 reads per second and then it goes back down to 200 and stays there for six months and then you come back and do 5000 again, it'll have that size, it'll go up to 10,000 really fine if it, if it doubles, if it goes beyond doubling in like a very short period of time, we're talking like, you know, 5 to 30 minutes.
00:42:42:03 - 00:42:56:18
Alex
Then it'll take a little bit of time to sort of scale up past that. And so that on demand mode is like now it's totally set it and forget it, right? You don't have to sort of think about that at all. It's just going to charge you per read and write that you do to Dynamo. In that case.
00:42:56:20 - 00:42:59:08
Lane
I'm guessing on demand is more popular.
00:42:59:10 - 00:43:27:02
Alex
It depends on it depends on sort of who you are. Like if you're in sort of the serverless world of smaller, smaller apps and servers, I think most people are okay that the cost of dealing with that is just like not worth it. But if you're like there are people that are doing millions of requests per second against Dynamo and have petabytes of data against Dynamo and in that case it probably makes sense to have someone think a little bit in set up auto scaling, especially, especially if you have like fairly predictable like day time patterns, right?
00:43:27:02 - 00:43:47:15
Alex
Where usage goes up during the day, goes down overnight, like manage sort of them scaling. They're basically the the threshold for of the eurozone threshold of like should you do on demand first provisioned is if you can do better than 17% utilization of your capacity. It makes sense to do provisioned that'll be cheaper sounds kind of low 17% you like that.
00:43:47:15 - 00:43:55:24
Alex
But it's like if you look at your database utilization over the entire day or things like that, it's actually like, great, a lot of like single digit, you know, I.
00:43:55:24 - 00:44:02:13
Lane
Get concerned any time, any databases over 50% on any metric. Yeah, exactly. Like he feels like 20%.
00:44:02:13 - 00:44:28:08
Alex
Yeah, exactly right. So then it's like, okay, how like, how much time is it worth to even think about that sort of problem to eke out some utilization versus like set it on brand and not even worry about it at all at that point? So then like for a lot of folks, I'm like, I always recommend on demand, you know, unless it's like I'm always like, Hey, figure out what your dynamo bill is going to be roughly and figure out if I cutting that in half would be super meaningful to you.
00:44:28:08 - 00:44:41:12
Alex
And if it's not, then like I would say, don't even worry about it because it's it's not going to matter, you know, because some people have like $200 a month dynamo bills and like if you save $100 a month by now you have to like manage your provision through. But like that's just not worth it, you know.
00:44:41:13 - 00:44:51:07
Lane
Would you like to like a new developer or like an indie developer that's like building a site app? Like, yeah, saving $200 a month feels huge. Like there's like 20 Netflix subscriptions or something. I don't know. Whatever.
00:44:51:11 - 00:44:56:20
Alex
Your dynamo is not going to be $200 a month unless you have like some decent traffic on it. You Know. Right. So yeah.
00:44:57:00 - 00:45:16:01
Lane
Yeah. But, but no, the point I'm making is just like I'm thinking of like where boot Def is now. Like I'm employing several full time engineers. It's like my database costs at the moment are on that order. It's like a couple hundred dollars a month and it's like the last thing I'm worried about, right? If it creeped up to a couple thousand dollars a month, then I'd be concerned maybe and think about we could maybe change some things.
00:45:16:01 - 00:45:19:18
Lane
But you're comparing to, like, salaries. Those amounts don't don't matter too much.
00:45:19:18 - 00:45:26:11
Alex
Yeah, totally. So I, I forgot what road we were going down before we took, like, this building side road. Well, we talking about you remember.
00:45:26:12 - 00:45:43:21
Lane
Well, no, no, I just said that Yeah. I'd had questions about like, are you kind of plan for expenses, Right. Yeah. And how you manage that. That was interesting because you talked about provisioning and I was kind of under the impression that it was an understanding of S3, for example, or cloud storage. It's like there is no such thing as provisioning.
00:45:43:21 - 00:45:50:01
Lane
I don't have to worry about it at all. I'm seeing the API, right? It sounds like it's just a slight like just a little more. It's in that.
00:45:50:02 - 00:46:08:22
Alex
It's either way, like you can do that on demand thing and you can also do this trick with on demand that basically like gives you a fake peak of like 40,000 reads and writes per second, which like most people are never going to come close to you. And then I'm usually scaling back down and that means you're basically in sort of S3 world, right, where you never have to worry about scaling and getting throttled and stuff like that.
00:46:08:22 - 00:46:17:04
Alex
You're just doing pay per use like you would. S3 So like you can do that if you want to. For if you're more cost conscious or things like that, then you can think about it.
00:46:17:04 - 00:46:23:04
Lane
It sounds like such a hack. Like why would if any business is able to do that, like why wouldn't they just be like the out of the box option?
00:46:23:05 - 00:46:44:09
Alex
I think they actually kind of don't like doing it because I don't love when people do it because actually dynamos partitioning your data under the scenes and it's it's part it's in choosing how many partitions you have under the scenes. There are like a couple of different factors. It's like how much storage do you have? Like basically every ten gigabytes of storage, they're going to add a new two partition or is based on how much throughput you have.
00:46:44:09 - 00:47:00:01
Alex
So even if you have like a fairly small table, but it's like super hot, it's going to have more partitions just to spread that across more partitions. So what you're doing when you're scaling it up to 40,000 reads in and scaling it back down is they're making a ton of these tiny little partitions all over the place. They're just hanging around forever.
00:47:00:03 - 00:47:04:09
Alex
And if you're doing like tiny amounts of traffic, like now, you have to sort of keep these partitions away.
00:47:04:09 - 00:47:05:10
Lane
Yeah, they probably hate that.
00:47:05:10 - 00:47:06:09
Alex
And like, yeah.
00:47:06:12 - 00:47:07:23
Lane
Ups their costs and stuff.
00:47:07:23 - 00:47:24:09
Alex
Yeah, a little bit. I've sort of asked them like, is that a big annoyance for you? And they said it's actually not that big of a deal because like, you know, their what they have is like these giant storage node, just a fleet of storage nodes, right? And each storage node is holding hundreds of partitions from all different customers, all that things.
00:47:24:09 - 00:47:41:19
Alex
And they're sort of moving those storage nodes around based on like, Hey, this one doesn't have that much actual storage packet with one that's like little denture storage or maybe doesn't have as many requests. So they're like doing that balancing. So if you have a partition, that base has no data and no requests, like it's not hurting them that much, but it's like it's just a lot of metadata they're keeping around.
00:47:41:19 - 00:47:48:22
Lane
This is like, this sounds like it's probably like an undocumented product feature that honestly could go away if they wanted it to.
00:47:48:22 - 00:48:24:03
Alex
Like, it's definite. Yeah. Yeah. It's like there are no guarantees on how many partitions Dynamo is going to make behind the scenes, whether those are like going to like currently they don't recombine partitions in like if they split a partition that split forever, but they could theoretically like there are no guarantees on that. Basically what it is is like a quirk of people reading the documentation very carefully, doing some testing and then realize, okay, if you do like, basically the trick is creating your new table as provisioned mode, setting it to 40,000 reads and writes, which now is like your fake peak that you've ever hit, and then you immediately switched to on demand mode.
00:48:24:03 - 00:48:35:06
Alex
So now you're not actually paying for those 40,000 reads and writes for like a long time, just like a few seconds. Now you switch to on demand mode and now you have a bajillion partitions, but you're only paying for reads and rates as they come in.
00:48:35:08 - 00:48:40:23
Lane
Get to engineering. Hacks are so fun. Like I don't know if you ever watch. Did you ever play Super Mario, 60?
00:48:41:00 - 00:48:43:21
Alex
Oh man, I play with that with my kids right now. It's so much fun.
00:48:43:23 - 00:49:01:04
Lane
So good. For some reason, I'm like, tick tock at YouTube. I just get I get recommended like speed runs of surprise. If you people you can like reverse you can reverse long jump backwards into a door enough times that like your speed vector gets so high that it's like go through the door and this is like one of those things.
00:49:01:04 - 00:49:10:04
Lane
It's like, yeah, you just kind of like play with it long enough that you figure out what's going on behind the scenes. That part of the API contracts, right? But like you can, you could get these, these things going. I'm not going to like recommend that.
00:49:10:04 - 00:49:12:05
Alex
That's how do people find that That's amazing.
00:49:12:05 - 00:49:12:23
Lane
I know.
00:49:12:23 - 00:49:19:22
Alex
No, that's more amazing than like someone reading the docs carefully. It's like, what are you doing that you're just backflipping about it, you know? But whatever. Yeah.
00:49:19:24 - 00:49:40:03
Lane
Okay. Amazing. Thank you so much. That was like going over Dynamo. Not only was interesting to me because I haven't really used it very much, but I think it'll be awesome for the audience. I want to switch gears for just maybe like ten or 15 minutes, just a little bit of this part about your experience consulting. What was your last full time job before you started the freelancing stuff?
00:49:40:03 - 00:49:56:22
Alex
So my last full time job before I started freelancing was it sort of was framework. Basically. I quit that in January 2020 and I was like, I'm going to write that down. What mean books I like started to kind of write it while I had a full time job and I'm just like, all my creative energy is going towards my job and I have nothing left to do the writing.
00:49:56:22 - 00:50:13:17
Alex
So I couldn't do that. I thought there was a market for it and I'm like, I'm going to quit and do this. So I quit in January 2020, and most of the next like four months was writing the book. I really that in April, which was like two weeks after, like COVID, shut everything down. And I'm just like, Is anyone going to spend any money?
00:50:13:17 - 00:50:34:14
Alex
Like, I just like, totally blow this whole thing like that that way? Fine. So first couple months for about the book and then like there was book sales, there was like sort of training an additional help on that sort of thing. And then actually like in August of that year, I did join a company is called Steady Work there for like five or six months and then decided I want to get back into just like being on my own.
00:50:34:14 - 00:50:47:23
Alex
I sort of like that independence once I had gotten that taste. So have now been full time since full time, independent since January 2021 and on that. So most I'd say most of the last four years with that little little stint in the middle there.
00:50:48:00 - 00:51:05:17
Lane
Notice you use the word consulting most like, say, less than five years of experience developers who are kind of striking out on their own. They probably won't call themselves consultants. They call themselves like freelancers, and they're like kind of hopping into a project, writing a bunch of code, usually for like a very small company. How does what you're doing with consulting, is that different?
00:51:05:18 - 00:51:28:20
Alex
Yeah, I would say especially like the first couple of years I was consulting, it was more advisory consulting, training type stuff. So like no, almost no hands on code for a while, whereas like, hey, we, we want to use Dynamo. Would you help us understand Dynamo, which could be a training workshop? It could be we have this specific data, all that want to design.
00:51:28:20 - 00:51:46:15
Alex
And would you just help us walk through the process of how to design that data model and think about it? Because with Dynamo, you have to do a lot more upfront design to like design your data model. You were talking about database normalization earlier and it's like with a relational database, you sort of like normalize your data first and then think what are the queries I need to access my data with Dynamo?
00:51:46:15 - 00:52:04:21
Alex
It's like, what are the queries I need to access my data and how do I design my table to handle those queries? Right? So it's a lot of upfront design and just like getting people to understand that that shift. So that was a lot of it. One thing that I, I like that because it's like very flexible and I see a lot of different things and it's like helping people that way.
00:52:04:21 - 00:52:23:00
Alex
But I didn't get to do as much hands on code, which is like what I really like doing. I missed that aspect of it. So now I've taken on more like hands on, hands on work the last year, year and a half like that. So now it's like a balance of like some advisory staff, some hands on hands on code.
00:52:23:00 - 00:52:25:05
Lane
If you see yourself going back to a full time job.
00:52:25:05 - 00:52:48:20
Alex
It's like I was hanging out with my friend the other day. It's it's hard to imagine being independent for like 40 years. But I also like and I really like the flexibility and independence of being on my own. I don't know. I want to be independent as long as I can and like my wife loves it and like, you know, we have kids and like if they have a school break or, you know, over the summer we like go on road trips somewhere or something like that very easily.
00:52:48:20 - 00:53:06:24
Alex
I like that aspect of it. So now the biggest thing I tell people like this is just like working as part of a team and like the camaraderie there and also like building towards something bigger than yourself, like the the joy when it takes off and like all that stuff. Those feelings I miss when I'm not like working as part of a team as much.
00:53:07:02 - 00:53:11:13
Lane
Any other entrepreneurial urges like you start your own like product?
00:53:11:14 - 00:53:34:02
Alex
Yeah, potentially. I don't think I'm like, I don't want to excel badly, but like, I don't think I'm like super great at, like the 0 to 1 type phase of just like purely starting it from zero and doing it myself. So potentially, I also think it just gets harder as you get like more established in your life. Like we have four kids and like, you know, house and all the things.
00:53:34:02 - 00:53:53:13
Alex
It's just like now it's like it's hard to start from zero and like, work your way up to like, you know, ramen profitable and all that sort of thing. So it's just like a harder bar there. So maybe like, I like the idea of it, but I haven't had anything that like because my attention in that way I thought would be worthwhile to go do that yet.
00:53:53:17 - 00:53:56:09
Lane
Crazy project ideas Yeah, not yet.
00:53:56:09 - 00:53:56:16
Alex
Yeah.
00:53:56:17 - 00:54:13:02
Lane
Cool. We'll see. We'll see what what comes the next couple of years. Sounds exciting. What's like the one takeaway if there's people listening to this pod that are thinking about either freelancing or consulting, what's like the big thing they should think about before jumping into that world?
00:54:13:05 - 00:54:34:15
Alex
How will you get clients? Think about how you will get clients and specifically like you want to get clients that want to work with you because you're Lane and not because you are a developer that knows Python or something like that, that like someone that wants to work with you specifically because that gives you some amount of like pricing power and just like selectivity of just like making sure it's the right fit for you and things like that.
00:54:34:17 - 00:54:49:02
Alex
I think you want to try to avoid early billing. I hate our billing. I said I was a lawyer. I hated that was like my least favorite part of being a lawyer was like keeping track of my hours. Like I might, you know, school kid or something like that. Like, I just did not like tracking all my hours.
00:54:49:02 - 00:55:04:17
Alex
And I think you don't want to track your hours as a consultant either, because it ruins your incentive to to get better. If you get more efficient, you actually lose money, which is, which is just silly. But then you also kind of feel bad or it's just like a bad situation. So figure out how you're going to get clients.
00:55:04:17 - 00:55:21:24
Alex
I think on the Internet, that's so a lot of ways easier than ever, but also harder than ever because you're competing with a lot of people, but you can definitely write or do projects or do things where people are like, Man, I want to work with that specific person. They know this. I think they can do a good job or you can have evidence of just like I've built this before, I can do it for you as well.
00:55:21:24 - 00:55:42:10
Alex
But like, that's what you want to do is like figure out some, some way to get clients. And mine has been mostly inbound based on like my own service and stuff that I've written over the years and also clients that I've worked with. But my brother in law does consulting. He's like machine learning guy and he does a lot of word of mouth or attending like local things in the area where people have a computer vision problem.
00:55:42:12 - 00:55:57:05
Alex
And I realized, like he's one of a few people in the area that can help him with that particular problem. So there are different ways to get clients. But but think really hard about how you're going get clients, because I've seen a few people out and then they have like no way to get clients and they don't want to go like pound the pavement or anything like that.
00:55:57:05 - 00:56:01:01
Alex
And it's just like really hard then, you know, like it's really hard in that sense.
00:56:01:01 - 00:56:16:21
Lane
In the freelancing world, I feel like there's not as much of a like necessary components around authority. But in the consulting world, to me it sounds like it's very important was the book was publishing the book a really key part of your ability to, to make a transition into consulting?
00:56:16:22 - 00:56:34:16
Alex
Yeah, Yeah. And it was already like written a bunch of dynamo stuff online. I had like spoken at Reinvent and a few other conferences about Dynamo before doing the Dynamo book. So I had some authority there and then it sort of just like solidified me another way and just like it helps, that dynamo is like a niche area.
00:56:34:16 - 00:56:55:18
Alex
Like if you wrote a book on my sequel therapy, jillion other books and things on my sequel, no one had written anything on Dynamo at all. So it just sort of helped in that particular sense. And so I do feel like I got lucky, just like right place, right time on on some of that stuff. But I do think that now writing online and then having a bigger a bigger piece like a book, I think that can help you for sure.
00:56:55:18 - 00:57:11:13
Lane
Fantastic. Well, thanks again for coming on. This has been a lot of fun. I always love when we do an episode on something entirely new. I mean, we're getting up there, there's don't actually know. This is going to be about episode three, but yeah, like it's, you know, it's fun to do some original stuff but as a last is the last thing.
00:57:11:16 - 00:57:15:20
Lane
Do you want a plug? Is there anything where's the book? Where can people find you? That kind of stuff.
00:57:15:20 - 00:57:34:11
Alex
Yeah, I'd say if you want to book Dynamo audible.com, you can find that out. You can find me on Twitter. Alex B debris or my, you know, website, Alex, Broadcom, things like that. But most of all, just reach out if you want to talk about Dynamo. I kind of rambled, I think in different spots today. But if you want like a do you want to talk about or understand like when Dynamo can be right for you or not?
00:57:34:11 - 00:57:48:12
Alex
Or like, what are the sort of pros and cons of things like that? I try to be fair on that. That aspect, I, I do love Dynamo, so there's that aspect of it, but I do try to be fair about like think about trade offs and realizing it's not right for every situation. So feel free to reach out any time.
00:57:48:12 - 00:57:54:07
Alex
Like, Yeah, thanks for having me on. There's been a lot of fun to get to chat with you on the podcast, so yeah.
00:57:54:09 - 00:57:56:17
Lane
Thanks, Alex. Yeah, we'll talk to you later, man.
00:57:56:19 - 00:57:56:24
Alex
Thanks.