Geoff Huston 0:00
You know, history makes liars of us all, and you're right. That
was the thinking when networks were unreliable. TCP kind of added
strength to it, but if networks were basically reliable, things
were fine. And then we put the Internet on mobile phones, and the
loss rate, at least initially, was awesome. It was one of the few
systems that had a loss rate of one in every 10 packets never made
it. It was pretty bad, and you know what collapsed? TCP.
George Michaelson 0:37
You're listening to Ping, a podcast by APNIC discussing all things
related to measuring the Internet. I'm your host George Michaelson.
This time, I'm talking again to Geoff Huston from APNIC Labs in
his regular monthly spot on ping. A lot of Geoff's experimental
measurement system is looking at behaviors in the DNS. Geoff has
noticed that for the huge, multi-million unique DNS queries that
he generates in clients that are coming back to the APNIC Lab's
DNS servers-an astonishingly high number-are retransmissions. In
fact, it's virtually double the expected query load for 150
million unique DNS labels that Geoff generates that will never be
seen by another client. He's seeing 270 million DNS queries. Many
of them are an almost immediate repeat. That is a lot of repeated
traffic. What's going on? Is this an instance of the tragedy of
the commons, or is it simply a fact of life in the modern Internet.
Geoff, welcome back to Ping. What should we talk about this time?
Geoff Huston 1:46
Hi, George. You know, it's a pleasure to be back. I'm just
actually back from a tiny bit of travel, and I I went to a couple
of meetings where the DNS was on the agenda.
George Michaelson 1:55
Ha ha!
Geoff Huston 1:56
I was at a meeting of the Operations and Research Group, DNS OARC,
and later on the following week, I went to a meeting of the sister
RIR organization of APNIC, the Ripe organization, which had their
their meeting also in Edinburgh, and there was a DNS session there
too. And I want to talk about things DNS, but
George Michaelson 2:17
right, let's do it. DNS, it
Geoff Huston 2:19
is a little bit of economics. Do you know about the tragedy of the
commons?
George Michaelson 2:23
Well, you know, it's a phrase that I think everyone's very
familiar with. But knowing a phrase and understanding exactly what
it means is kind of subtly different. So, why don't you give me a
rundown?
Geoff Huston 2:36
Well, it kind of works on the theory. The commons were evidently,
I'm like I wasn't around, and neither was anyone else, dear
listener. An aspect of medieval England where a collection of
houses, a village, would have a common area that nobody owned, and
that everyone could graze their animals.
George Michaelson 2:56
So nobody owned it, but everybody shared it.
Geoff Huston 3:00
Shared it. They could use it without charge. It wasn't anybody's
land. It was everyone's land. It was the commons, and and the
problem was it only worked if everyone was restrained in the
number of beasts they sent out on the commons going eat grass.
Because as soon as someone got greedy and sent out, you know,
more, the place just went to rack and ruin because all the grass
got eaten. It turned to mud. You know, bad things happen, and it
actually led to, according to the economists who studied this, a
subsequent move of the fencing of the commons to try and sort of,
if you privatized it, people would look after it better. But
there's a generic kind of principle about unconstrained human
behavior [George: yeah], free resources tend to get exploited to
the point of overconsumption and degradation.
George Michaelson 3:49
It feels a little bit like a version of the prisoner's dilemma
that the only good answer in the prisoner's dilemma is for them
both to do the thing that's mutually beneficial, and if either one
of them tries to get their own safety secured. It stuffs up the
whole thing. The Commons. Everybody's got to be a good player. And
the minute there's one bad apple that takes more than their free
rights, the whole thing goes to pot.
Geoff Huston 4:13
So, free resources get abused. They just do. Free money. Free
whatever. And so there's this aspect of this applied to
the Internet's domain name system.
George Michaelson 4:25
Well, I'm not paying for it, so there's a quality that makes it
feel to me very free.
Geoff Huston 4:31
Well, if I run an authoritative server where I answer queries that
other users have made, so if you want to know about potaroo.net.
It's my server that's on the hook. Does anyone pay me?
George Michaelson 4:44
So, answer here means specifically be the authoritative answer
because there might be many other DNS players along the path who
feel like they are air quotes providing the answer, but ultimately
somebody has to have asked you, Geoff, in order. For that to
happen,
Geoff Huston 5:01
right? Someone needs to have sort of seeded the DNS with the
original sets of you know data, and that's actually an act that
doesn't get compensated by by the people who actually are making
the queries. Queries are free, and if you observe from that, well,
do people over query because they're free? That's a really
interesting question.
George Michaelson 5:25
Yeah,
Geoff Huston 5:25
you see, the DNS works and it's really efficient because it uses
the UDP protocol, the datagram protocol, the protocol that says
every packet is an adventure. If I send you a UDP data packet. You
may or may not get it. UDP doesn't care. You see, this actually,
oddly enough, and let me digress a tiny little bit, was actually
one of the issues with the early forays into packet networking in
the 1960s. We had been used to a form of networking which, when we
moved from analog to digital signals. We regarded the signal we
were switching, even digital, as a consistent stream of packets,
and we ended up switching the most incredibly expensive thing we
could possibly switch: time. So that if you had a common resource
switch between, say, eight people for the first eighth of a
second, it was user A, and then the next eighth, user B, etc. In
other words, you tried to sort of fair share amongst competing
people by actually switching on on a common view of the time,
which is interesting. But what it meant was that the network was
carrying the integrity of the conversations. It didn't matter how
clever or dumb you were when you attached your speaker and
microphone to this network. The network itself would preserve that
digital data as it switched it through the system. Circuit switch
systems. When we went to packetization, there was a debate, and it
was a furious debate. The debate was in a packet switch network.
The packets don't go all the way from A to B. It's actually a
series of of hops. There are intermediate switches, and so you eat
the first switch that then passes it to the second switch, then
the third, then the fourth. Now it doesn't really matter how you
define the sequence of switches, whether it's predetermined,
whether it's you know like the Internet dynamically discovered,
doesn't really matter. But the issue is when one intermediate
system hands a packet to the next intermediate system, does it
care that the next one received it? Ooh, big question. Because if
you care X 25 then you actually get a whole sequence of A sends it
to B, B acts it back to A, B sends it to C, C acts it back to A.
What does this mean? Well, if someone pops a packet in to this
system, it will either come out really reliably, or the interior
of the network will go, oh my God! All hope is lost. Let me
destroy the circuit. Let me give you an indication of failure.
George Michaelson 8:05
Because every hop along the path, people are saying, "Did you get
it? And the person on the other end is saying, "No. In which case,
you have the burden to retransmit it. So bounded in time, it's
going to get there, or the whole thing's going to say, "We've got
a problem",
Geoff Huston 8:21
right? So X 25 based networks, which were really popular for a
small window of time in the g 1980s worked on this principle of
reliable hop by hop transmission. One of the networks that I did
work on, DEC net, and a venerable protocol called DDCMP, and it
was actually the same kind of thing. All the intermediate systems
effectively explicitly acknowledge successful receipt of a packet,
and so the thing injecting packets didn't have to care. Once you
passed the packet to the network, you could assume that it was
going to come out the other end. What you couldn't assume was
speed, because you know, if someone had stuffed it up, they'd sit
there and try and try and try for quite some time before they gave
up on the entire conversation.
George Michaelson 9:08
So, if we bring this back to an Internet context and the idea
of Internet protocol, the actual packets that are the Internet
protocol packets, whether they be four or six, they're kind of
stuck in the weeds on this, they're not saying absolutely must be
on the network like DDCMP or X 25 They're actually saying don't
care what I run on piece of string, giant fiber optic 100 gigabit
channel, totally up to you. I'm just a packet, mate.
Geoff Huston 9:35
Well, that was the genius, and I would say of the Internet that
looking at everyone else who was busy doing reliable, reliable
switching and effectively synthesizing a circuit out of these hop
by hop packet elements, the Internet said, "Nah, every packet is an
adventure. It's not the network's problem to identify and fix it.
George Michaelson 9:57
I think every packet is an adventure. Needs to be on a T-shirt. I
think that's brilliant.
Geoff Huston 10:02
That's right. We're off on a big adventure. Some of us might not
make it. Whoops. And that was kind of it. The argument was, and it
got encapsulated in what what we call the end to end principle,
that it was left to the two machines at either end who were
communicating to figure out if any packets that hadn't made it,
and repair the damage.
George Michaelson 10:23
Right,
Geoff Huston 10:24
and you kind of go, well, that was just IP, and you go, well, no,
that was Ethernet.
George Michaelson 10:28
No,
Geoff Huston 10:28
it was Ethernet. It was a whole generation of successor data
protocol, you know, designs that effectively said, well, this is a
packet system. I don't actually care about the reliability of the
packets. That's up to the communicating systems at either end. All
I care about is just pushing it through, and if I fail, neh; let's
just move on with our lives, right?
George Michaelson 10:49
So we don't actually routinely construct applications and services
into IP packets. We've got a layering model, and there's a choice
moment coming, which I think brings us back to DNS and UDP. How do
you make something on top of this have the behaviors you need to
know is this unreliable or reliable?
Geoff Huston 11:09
Yes, a lot of stuff use TCP. This reliable transport. Now TCP has
overheads because there's no network state, but there really is
state at either end. Instead of I to send you a packet, I have to
go through a dance. Hi, George. I've got a number. Here's the
number I'm thinking. The packet is going to be called a
synchronized packet, so it's a SYN.
George Michaelson 11:30
Hi, Geoff. Here's an ACK for that. And by the way, here's a
number I made up that is my number that you should remember
when I tell you things.
Geoff Huston 11:39
Ah, George, I see you've received my packet and sended my number
back number back to me, and I've got your number. Let me
acknowledge your number. This is the third part of the three-way
handshake, and now we both have sequencing numbers that the others
using.
George Michaelson 11:55
Terribly polite protocol, isn't it, Geoff? So polite. Oh, thank
you for sending me that packet.
Geoff Huston 12:01
Oh, thank you. Here's my first data packet. Now I know I said
magic number 10. Well, it's got 10 bytes in it, so the new magic
number will be 20.
George Michaelson 12:11
Thank you, Geoff. I got that magic number 20. Here's my answer
with magic number two.
Geoff Huston 12:16
Ah, I see that you've received magic number 20. I will delete my
copy because you have it. I don't need to resend it. Thank you,
George. Oh, and by the way, let me acknowledge 2. You see how it
kind of moves. That instead of doing hop by hop, did you get it?
Did you get it? Did you get it? The two end systems play the did
you get it dance by advancing their respective sequence numbers.
The cost. Well, it's a three-way hands shape to start with.
George Michaelson 12:44
So there's a bit of delay coming in.
Geoff Huston 12:46
This whole thing about let's exchange sequence numbers works
perfectly fine if we've got stuff to say to each other. But George
only wanted to say hi.
George Michaelson 12:56
Short conversation.
Geoff Huston 12:58
That's two characters. That's that's two bytes, so it's kind of
SYN SYN ACK ACK " hi" ACK FIN FIN ACK FIN. You know, that's a lot
of packets to go " hi"
George Michaelson 13:10
"hi" yeah.
Geoff Huston 13:12
And when the approach came to think about implementing the domain
name system and its resolution as an Internet protocol, it was very
quickly realized that there were many, many points of
conversation, many points of authority potentially, but each
conversation would be tiny. Hi, do you know about dub dub
dub.apnic.net
George Michaelson 13:34
No,
Geoff Huston 13:34
done. You know that's it. I don't need handshake handshake. You
know, blah blah blah. It's just you know, fine. Tell me yes or no.
I don't care.
George Michaelson 13:44
I think tied up in this is sort of a subsidiary story that if we
were talking about a medium that was so atrociously noisy, one
packet in two never got through. You might be very uncomfortable
with that class of behavior. But we are living in a world where
the behavior assumption was most packets get through, and you'll
need to say, "I didn't get that", or you'll need as a sender to
say, "Looks like you didn't get that. I think I'll do it again".
Is kind of down in the threshold where you're not really paranoid
about checking every single time. It mostly works.
Geoff Huston 14:19
You know, history makes liars of us all, and you're right. That
was the thinking when networks were unreliable. TCP kind of added
strength to it, but if networks were basically reliable, things
were fine. And then we put the Internet on mobile phones, and the
loss rate, at least initially, was awesome. It was one of the few
systems that had a loss rate of one in every 10 packets never made
it.
George Michaelson 14:42
Wow!
Geoff Huston 14:42
It was pretty bad. And you know what collapsed TCP.
George Michaelson 14:46
Seriously, because
Geoff Huston 14:47
if the loss rate is high enough, it never gets. I suppose what I'd
call a full head of steam. It always just sputters. Hi, I's one
packet. Because the way TCP works is that if you send me an ACK
for a packet. I'm going. Well, that's fine. Here's two packets,
and you send me the ACKs for those. And I go, well, George, this
is cooking. Here's 4.
George Michaelson 15:08
I'll ACK that.
Geoff Huston 15:08
8, 16,
George Michaelson 15:10
I'll ACK that.
Geoff Huston 15:11
You know, TCP has this system, which ironically is called slow
start, which is anything but. [George: Yeah] it's actually
exponential start, and it just zooms forward. If I've got a lot to
say, TCP works brilliantly as long as as long as the loss rate
isn't bad. If the loss rate's around two 3% max, TCP is okay
George Michaelson 15:32
because you get the second part of the TCP behavior, which is
under loss. The back off is extreme.
Geoff Huston 15:38
Oh, it's very extreme. Oh, a packet loss! Panic! Ring the fire
alarms! You know, ring the bell. Let's let's back right off. And
so, oddly enough, you might have thought the TCP was sort of
perfect for crappy networks. It's actually very imperfect. And
what actually became really quite neat was UDP, because
applications expected damage. It's kind of yeah, right? You know,
lost a packet. It's not a big drama, and it was UDP was always
intended for short transactions. Although some people abused it
and started doing weird and wacko things using gigabytes of data
over UDP, but they were loonies.
George Michaelson 16:14
Well, they were loonies then, Geoff. But if we look at modern
parlance and QUIC using UDP for megabytes of data has come back
into fashion, but we've drifted slightly here. The point is, UDP
was the way to say I'm not going to invest in a layer to do all of
that reliability checking for me. I'm going to fling packets out
with the least overhead, particularly when I'm asking questions
like DNS questions. What's the address of this name?
Geoff Huston 16:42
Single packet question, single packet answer. And if packets get
damaged, it's left to the DNS application to figure it out and ask
again somewhere else.
George Michaelson 16:52
Ah, now you just said ask again.
Geoff Huston 16:55
Well, yes, because if you don't get an answer, what are you meant
to do? Go and sulk? I'm sorry to get answer. Nobody likes me. It
doesn't work like that.
George Michaelson 17:05
It's not the " Eeyore" network, is it?
Geoff Huston 17:08
Oh, absolutely not. I need an answer. I'm going to get one. I'm
going to keep on asking til I get one. And so it's persisted under
loss, right? Because that's the way we designed it. The
application has to sort of make up for a network that isn't
perfect. Okay,
George Michaelson 17:25
right
Geoff Huston 17:25
now I said before queries are free. So when do you re query?
George Michaelson 17:31
Now I'm a good person, Geoff. I want to go to Internet heaven, so
I'm going to say the right thing to do is to re query when you're
really sure you have to, and not re query any other time. There
you go. I'm a good player in the Commons.
Geoff Huston 17:48
Well, you've also read RFC 1536 that said you'll wait for about
five seconds.
George Michaelson 17:53
That's a long time.
Geoff Huston 17:54
Oh, continents have drifted in five seconds. You know, attention
span. Entire generations of attention spans have come and gone.
[George: Squirrel] exactly. But that was the advice, you know. Be
patient, wait. Don't splutter packets into the network. And of
course, since then, the networks have got bigger. But at the same
time, we have got way more impatient. We are extremely impatient
about this. And so, what's the best way to get an answer quickly,
George Michaelson 18:24
Vincent? Twice, Vincent twice.
Geoff Huston 18:27
Well,
ask, ask, ask, ask, ask. You know,
George Michaelson 18:29
are you there yet? Are we there yet? Are we there yet? Are we
there yet?
Geoff Huston 18:33
Yeah,
I want to drive from Canberra to Sydney. Okay, how do I get there
fast?
George Michaelson 18:36
We
there yet?
Geoff Huston 18:37
Sent five cars, but I'm only in one of them. Sometimes it doesn't
work.
George Michaelson 18:43
That's the million man month statement about you can get a baby in
under nine months if you get nine women to carry a baby, isn't it?
Geoff Huston 18:51
That's the same kind of fallacy, you know. But oddly enough, in
the DNS, it's not quite a fallacy. You see,
George Michaelson 18:57
what?
Geoff Huston 18:58
Well, the thing is, if this stuff is unreliable, and let's let's
go to our magic 10% Then 10% of the time, I'm not going to get an
answer, but I'm going to wait for five seconds. So 10% of folk
will have a five second delay, and they'll ask again. 10% of them
will not get an answer for another five seconds. So 1% of folk
will be waiting for at least 10 seconds for an answer, and 10% of
them won't get it, etc. etc.
George Michaelson 19:24
Well, in a network of 2 billion users, 1% is an astronomically
large number.
Geoff Huston 19:30
Well, it is, and the whole issue comes that maybe that's the wrong
strategy.
George Michaelson 19:35
Yeah.
Geoff Huston 19:36
So one thing you can do is not wait five seconds. I will work out
reasonable approximation is for when you should respond. A packet
can circle round the globe in a second. In fact, you can do it for
half a second if you're lucky, and on you know the right pieces of
cable. So I'll wait for a second because if you haven't answered
it,
George Michaelson 19:54
yeah.
Geoff Huston 19:54
In fact, damn it, 300 milliseconds.
George Michaelson 19:57
You know what? I would feel like this. Was still in the class of
good behavior in the commons. If you're genuinely likely to get an
answer sooner than that, waiting 300 feels like a pretty good fit
for human centric delay, the experience of waiting for a thing to
happen, and the reality sometimes you have to do it again. I could
be happy in that world, and I'd love to believe it's another short
episode, and you're telling me you went out and looked, and that's
exactly what you saw because we're good citizens in the commons of
the DNS. That's true, isn't it, Geoff?
Geoff Huston 20:29
I'm sorry. Yeah, right. So we went and looked. We went and looked
for a day using this ad based experiment, and we sent out a large
number of names, oddly enough, because we don't want caching to
happen in the DNS, which glosses over everything. All these names
are unique, but that's easy. You just make a long name and you
know add some variable fields every time. So on this day, we did
150 million unique names that got asked for by users, and they're
unique, and they're unique in a couple of ways. One, the name only
exists just once in time. That's it. Secondly, the only the only
systems that can answer that DNS query are ones we run. So we see
the query because we're the only people that are authoritative,
and no one can cache it because it's unique.
George Michaelson 21:18
You kind of brought the cache word to the table, and I think we're
probably going to want to come back and talk about that in another
episode because there is so much in cache. Cache is fascinating.
Geoff Huston 21:30
If it wasn't for the cache, there wouldn't be the DNS. But that's
a talk for a different day.
George Michaelson 21:34
But the thing is, you've constructed a world with lots of unique
labels where caching simply can't apply because they've never
existed in the system before.
Geoff Huston 21:43
Yep, 150 million of them in 24 hours. We were very busy,
George Michaelson 21:46
good-sized experiment.
Geoff Huston 21:48
These all over the planet, but we received 250 million. In fact,
270 million actual queries. So 150 million becomes 270 million in
the DNS.
George Michaelson 22:00
Well, it's not quite double, but it's a hell of a lot closer to
double than a small percentage are repeated. That's most queries
are repeated, Geoff.
Geoff Huston 22:12
On average, most queries are repeated. On average, the DNS is
going " to be sure, to be sure", and on average, it's it's sending
out the query a second time. Now, one theory is the Internet's far
worse than anyone ever suspected. They've all wasted their
obligatory 300 milliseconds or whatever their time out might be.
They haven't got an answer in time, and they're requery.
George Michaelson 22:33
Well, I just want to put my hand up here and say we're recording
this using a system that is transmitting live audio and video, and
effectively, is like a two to 300 kilobit stream continuously from
me to you and you to me. And if the level of behavior in that
system demanded every single packet that we are doing was sent
twice, I'm fairly comfortable we'd be seeing artifacts like delay
or break up on this video, Geoff. That's what we used to see in
old noisy systems when things could not reliably get data through.
So if you and me can reliably send hundreds of kilobits of data to
do this video and audio recording, I'm not buying that the DNS is
suffering a noisy network and HAS to retransmit. You've just said
it IS retransmitting, but you haven't said it HAD to do it.
Geoff Huston 23:23
Well, we go a bit further because we're seeing all these queries,
and so what we actually do is then go. Well, a duplicate query has
the same query name a dot b dot c and the same query type a record
quad a whatever because we've got a recording of every query we
see, why don't we time the interval between the first query and
the duplicate query? Why don't we actually see how long people are
waiting? How patient?
George Michaelson 23:50
That's quite interesting. So you don't actually need their clock.
You just have to have your own clock recording the interval
between seeing it once and seeing it again.
Geoff Huston 24:00
Well, I just record the time when I see a query. Go back over the
logs and go. Well, that was the duplicate of the first query.
What's the time difference between the first and the second?
George Michaelson 24:10
And if you're using a system like DNS cap or TCP dump pcaps, you
actually automatically get a packet timer, don't you? An arrival
time, which means you can do the packet interval, either from a
log or from a data stream.
Geoff Huston 24:23
Pcaps, pcap's is an extraordinary amount of data that wade
through. I'm using logs. It's good enough for for this jazz. So
yeah, the log says, you know, here's a time interval between
queries. And we've said before, this is not time out. This is not
a crappy network. So what sort of time interval do people like to
do because as I said, half of all the traffic we see, almost half,
are duplicates. So of the duplicates, just looking at the ones
that weren't the first, but it's a query that happened before,
right?
George Michaelson 24:55
Yeah,
Geoff Huston 24:55
45% of all those duplicates come to us within 10. Milliseconds of
the original. Oh my God!
George Michaelson 25:03
Just before we drill down there, this is quite a skewed
distribution, isn't it? I mean, a small number of these duplicates
probably arrive a significant time after, like more than 300
milliseconds. Yeah.
Geoff Huston 25:14
Well, a small number might well be the first time didn't make it.
I'm twiddling my thumbs and I'm doing a DNS timeout, and I'm
trying again.
George Michaelson 25:22
Real loss.
Geoff Huston 25:23
Real
loss will cause retransmission.
George Michaelson 25:25
Now to come back to the astonishment moment, almost half of the
retransmitted packets are 10 milliseconds or less. [Geoff: Less]
Holy cow, Geoff. Holy cow.
Geoff Huston 25:39
Well, let's up the clock a little bit and go 20 milliseconds, 20
thousandths of a second, two hundredths of a second. That's not a
timeout. No, 60-5% of all queries occur almost like we'd say tap
tap, and it's the same query.
George Michaelson 25:56
I am astounded, Geoff. I mean, to me, that shriek's deliberately
coded.
Geoff Huston 26:01
Okay, so let's dig a little deeper because it's always fun to dig
a little deeper. Do you know about happy eyeballs, dear listeners?
And George, maybe folk do.
George Michaelson 26:11
Got a kind of vague feel. It's an odd thing, isn't it? It's this
idea that we need to construct a way to handle how do we make six
happen if four and six both exist, it's in that space, isn't it?
Geoff Huston 26:23
It's in that space, and it used to be when I'm opening up a TCP
session that a Linux operating system would wait for 108 seconds
before saying that's not going to work. So if you didn't know if
you could open it using V4 or V6, and you tried in say v4 the TCP
part of the operating system would say, "Hey, I'm working on this
three way handshake. Leave it to me. Oh, that SYN didn't make it.
Let me time out and try again. Oh, that SYN didn't make it, and
it'd take 108 seconds to go. I don't think it's working. Let's try
V6
George Michaelson 26:59
I have a story for you from the deep dark days of the
early Internet, when people still used the file called host. txt,
and it took so long to serially check line by line that file to
find where I was coming from that when I used the protocol called
FTP to try and download code from London, getting it from a
machine at the University of Maryland, the 30-second timer on the
login would time out and disconnect me before it had worked out
who I was.
Geoff Huston 27:30
Steam-powered computers, George. They were wonderful. Turn off the
steam. Go to electronics. Oh, you British! My apologies to all the
British listeners. I didn't mean it. This whole idea, though, of
duplicating became a technique where instead of waiting for
conventional timers, you do try and speed up the process. So what
we called happy eyeballs was you actually try a connection in v4
but impatiently. In fact, you do it in six first, and if that
doesn't work, you do a rapid switch over to b4
George Michaelson 28:05
You kind of set two.
Geoff Huston 28:07
It's a two horse race.
George Michaelson 28:08
Two horse race. It's not wait for one then the other. You fire
them off with a just little bit of an advantage.
Geoff Huston 28:15
You handicap the one you prefer, and if you kind of go, oh well,
if you've implemented both protocols over the same paths and it's
roughly the same time. Whoever you handicap will win the race. So
I handicap six. I send out the two horses at the same time. Six
has a slight favor. You'll make a V6 connection if you can, but
you'll rapidly fall back to v4 [George: r ight] and so you go.
That's good.
George Michaelson 28:36
Putting this into the context of the DNS, I'd love to say, Geoff,
it must be that you're measuring happy eyeballs, then. That's
where you're going, isn't it?
Geoff Huston 28:43
No, not really. I'm trying to find an explanation as to why almost
two thirds of all the duplicates occur within 20 milliseconds.
Someone is doing this, bang bang, and happy eyeballs is part of it
because certainly in the DNS, if I kind of want to make sure, then
testing V4 and V6 is not a bad strategy. But there's more to it
than that. In theory, I'm meant to have names with multiple
authoritative name servers, and if I want to increase the
resilience, a very clever recursive resolver, one that's making
the queries, would go, "Ah, Geoff, I see you've got four name
servers out there that can answer my question. Let me send the
same question to all four at once". Queries are free. I'll just
take the first one to answer.
George Michaelson 29:28
Right, but Geoff, that wouldn't look like if you air quote it
duplicates unless you measured across the whole system. Whereas
you're at a single point from what you've been.
Geoff Huston 29:37
So , that's
where the argument breaks down. So let's see what other clever
things people do. One of them, which we have noticed in a number
of ISPs, is they're not really sure about their own DNS
infrastructure. So, well, it's been in a dusty cupboard for 10
years. There's a bit of mold and bit rot growing on the shazzy.
Might not be working.
George Michaelson 30:01
The perversity of the Internet, where you trust the remote party
more than your own equipment that you bought and set up to run
service.
Geoff Huston 30:10
Google's public DNS has an awesome reputation. Cloudflare's all
one has an awesome reputation. So tell you what, let me send the
query to one of these open resolvers, and at the same time, work
diligently hard to resolve the name myself. If I get an answer
from the open resolver, I'm done. Otherwise, I'll keep on working.
Yeah, not a bad strategy.
George Michaelson 30:30
I could believe in this strategy, Geoff. But again,
Geoff Huston 30:33
you could believe it. That would occur.
George Michaelson 30:35
I bring back to the table that would not, in strict sense,
normally look like a duplicate to you would it because you would
see different originators asking you the same question,
Geoff Huston 30:46
same name. A duplicate is the same name, not the same asker.
George Michaelson 30:50
Ah, ah ha.
Geoff Huston 30:52
So v4 and V6 if they're asking for the same name, are duplicates.
Cloudflare or Google and My Resolver are duplicates if they're
asking for the same query type, the same query number.
George Michaelson 31:04
Right. So this begins to sound like a possible collection of
reasons you would see twice.
Geoff Huston 31:10
An
optimization. Yeah. You know, send off four race horses down
different tracks. Whichever one comes back, I'll rub with it. You
know, the v4 racehorse, the V6 the one to Google, the one to
Cloudflare, the one to open DNS, whatever queries are free. Just
send them all out. Take the first answer you get. You're done.
Sounds good.
George Michaelson 31:26
Mostly got a few if but maybes. But yeah, I could be there.
Geoff Huston 31:30
So 1/5 of all the duplicates come from the same IP address. Jaw
dropping moment.
George Michaelson 31:36
Wait, I'm sorry. 45% of all the duplicates are rapid fire.
Geoff Huston 31:40
Yes,
George Michaelson 31:40
and 1/5 of the duplicates
Geoff Huston 31:43
of the rapid fire duplicates of the ones that are tap tap come
from the same IP address.
George Michaelson 31:49
I'm not using two people to ask this question for me. I'm asking
it twice. Wow,
Geoff Huston 31:54
what's your
problem, George? Sorry, what?
George Michaelson 31:57
Didn't hear you. What did you say?
Geoff Huston 31:58
Yeah, what? Sorry, what?
Sorry, what? I'm in a group. Sorry, what?
George Michaelson 32:02
You know, entering the room as a tragedy of the commons. I'm
sitting here, frankly, aghast, Geoff, because this is taking a
service that is notionally free, but actually has a publicly
distributed cost. It is not free to provide the DNS either as an
authority or as a public resolver. It's a cost, and it's a packet
cost in terms of congestion on the wire. And you've just said to
me there could be two or three or four or 10 duplications to get
an answer from a system that we already agree is overwhelmingly
reliable, even if you're using the UDP shout into the void
protocol,
Geoff Huston 32:41
60% of the duplicates are only single duplicates. If I say one or
two duplicates, so in other words, a total of three packets, I get
up to 80% So most of the time, it really is tap tap, and it's so
bizarre. Not tap tap tap.
George Michaelson 32:55
But
isn't that an abuse of the system?
Geoff Huston 32:57
Well, from my selfish perspective, as the resolver, I just want
the answer, dude, and I'm incredibly impatient. So someone else
has the cost to bear. Tough on them.
George Michaelson 33:06
Ah,
Geoff Huston 33:07
it's the selfish. It's the tragedy of the commons all over again.
And I'm sitting there going, "Wow, where is this?" Well, we know
where it is because we know who's asking.
George Michaelson 33:18
Oh dear, name and shame time. That comes up so often in these
measurements, doesn't it?
Geoff Huston 33:23
It does, but I'll be vague. I'll be general. The hot spots appear
to be where there is, if you will, network infrastructure that may
not be so crash hot. So oddly enough, the Indian subcontinent,
including Bangladesh, we see a lot of duplicates. Double the rate
coming out of the U.S. and Europe. That doesn't mean the U.S. and
Europe don't have it. They do, but the rate is much higher in our
servers that serve the Indian subcontinent, and oddly enough, in
China and Hong Kong in that area, where again the duplicate rate
is pretty much double what we see in other parts of the world.
George Michaelson 33:57
So China, we know, has a national strategic view of DNS in the
context of Internet, and it is extremely common for people to use
intermediary devices because they can satisfy national delivery
goals by buying legitimated systems that do the things the Chinese
state considers necessary for safe Internet. So, for China, my
spidey sense is saying this sounds like intermediaries.
Geoff Huston 34:25
It's hard to say, George, because in our experiment we don't get
to see what happened to start this off. We know an ad got placed
in a browser. We know that, but we can't trace the DNS from that
machine all the way through to the one machine we can see the
machine that is asking our server the question. So between the
user and that final sort of hopping off point, it's all black to
us. We can't tell. But it does strike me as amazing that almost
one half of the DNS are ghost cars. They're nothing. They're not
real. They're just duplicates, and most of those duplicates occur
almost back to back with the original query. And so the answer
set, the answer set, is going out to all of them, duplicate and
original. [George: Yeah], it's not as if I don't answer the
duplicate. I've already answered one.
George Michaelson 35:17
No,
Geoff Huston 35:17
we answer everything.
George Michaelson 35:18
I'm kind of wishing there was the facility in the packet of DNS
under this condition of behavior that allowed a flag to be put in
the question, saying this is the second time I've asked. Because
the thing you don't know is the origination order of these two
packets. You know the arrival order. You can't actually know which
of the two was.
Geoff Huston 35:38
You can't tell, and I can't tell. But if it's the same IP address
and all that's varied is the source port, does it matter which was
the original?
George Michaelson 35:46
No, not really. I mean, it's kind of irrelevant.
Geoff Huston 35:49
Exactly. So, in some ways, the most bizarre behavior, where it's
almost as if the switch is duplicating packets. Now, that used to
happen, and I remember seeing it in about 1990, that certain
switches from a certain really cheap vendor decided that when life
was tough, instead of really switching one packet, they'd turn one
packet into two for free.
George Michaelson 36:11
Because the world always gets better when you put more stuff on
the network, right?
Geoff Huston 36:15
More packets equals more joy, you know. But those days,
thankfully,
George Michaelson 36:21
every packet's an adventure, Geoff.
Geoff Huston 36:23
Even when it gets cloned, but you know they were cloned packets.
These are point of difference. It's it's a DNS source port that's
actually changed.
George Michaelson 36:32
Wow, Geoff, this is really fascinating. I think you might need to
move to other forms of experimentation, like doing packet trace on
the client side resolver to see if you can actually see two packet
queries being initiated out of some system.
Geoff Huston 36:47
Oh, at some point we resort back to understanding which resolvers
are doing this and understanding who runs them, and actually
dropping them a note saying, "Hi, we've noticed something rather
bizarre about the DNS. Yeah, you are running. Who's vendor's code?
Are you using? What version is it? How have you customized it?"
George Michaelson 37:06
So, dear listener, if you happen to know the reason that this
behavior is now ubiquitous in the global Internet, Geoff would love
to know.
Geoff Huston 37:15
Oh,
I would love to know. So much of the Internet happens without us
looking, and even if you were running a system that did it, do you
know? Nobody knows.
George Michaelson 37:24
Oh God, I wouldn't have a clue.
Geoff Huston 37:26
No, no
one looks down at this detail. Quite frankly, even if you run an
authoritative name server trying to find these rapid duplicates,
you've actually really got to look hard. And so you kind of
wonder, George, what else is out there of the Internet? Complete
swap. The only difference is no one's bothered looking for it.
[George: Yeah] you know it's out there, just waiting to be
uncovered.
George Michaelson 37:49
Folks, there's a world of measurement out there. We should all be
looking at this stuff. That's really fascinating, Geoff. You've
written this one up on your blog.
Geoff Huston 37:56
It's
coming up in the next couple of days. I'm just in the final
process of crossing the eyes, dotting the Ts, or whatever you do
with I's and T s.
George Michaelson 38:04
By the time this one goes to air, I'm sure you'll have it online,
so we'll include it in the web page with it. That was really
great, Geoff. Thanks.
Geoff Huston 38:11
Thanks, George. Cheers. Till next time.
George Michaelson 38:14
If you've got a story or research to share here on Ping, why not
get in contact by email to ping at APNIC. net or via the APNIC
social media channels. Also, remember the measurement at APNIC.
net mailing list on Orbit is there to discuss and share relevant
collaborative opportunities, grants and funding opportunities,
jobs or graduate placings, or to seek feedback from the community
on your own measurement projects, be sure to check out the APNIC
website for all your resource and community needs. Until next
time.
You know, history makes liars of us all, and you're right. That
was the thinking when networks were unreliable. TCP kind of added
strength to it, but if networks were basically reliable, things
were fine. And then we put the Internet on mobile phones, and the
loss rate, at least initially, was awesome. It was one of the few
systems that had a loss rate of one in every 10 packets never made
it. It was pretty bad, and you know what collapsed? TCP.
George Michaelson 0:37
You're listening to Ping, a podcast by APNIC discussing all things
related to measuring the Internet. I'm your host George Michaelson.
This time, I'm talking again to Geoff Huston from APNIC Labs in
his regular monthly spot on ping. A lot of Geoff's experimental
measurement system is looking at behaviors in the DNS. Geoff has
noticed that for the huge, multi-million unique DNS queries that
he generates in clients that are coming back to the APNIC Lab's
DNS servers-an astonishingly high number-are retransmissions. In
fact, it's virtually double the expected query load for 150
million unique DNS labels that Geoff generates that will never be
seen by another client. He's seeing 270 million DNS queries. Many
of them are an almost immediate repeat. That is a lot of repeated
traffic. What's going on? Is this an instance of the tragedy of
the commons, or is it simply a fact of life in the modern Internet.
Geoff, welcome back to Ping. What should we talk about this time?
Geoff Huston 1:46
Hi, George. You know, it's a pleasure to be back. I'm just
actually back from a tiny bit of travel, and I I went to a couple
of meetings where the DNS was on the agenda.
George Michaelson 1:55
Ha ha!
Geoff Huston 1:56
I was at a meeting of the Operations and Research Group, DNS OARC,
and later on the following week, I went to a meeting of the sister
RIR organization of APNIC, the Ripe organization, which had their
their meeting also in Edinburgh, and there was a DNS session there
too. And I want to talk about things DNS, but
George Michaelson 2:17
right, let's do it. DNS, it
Geoff Huston 2:19
is a little bit of economics. Do you know about the tragedy of the
commons?
George Michaelson 2:23
Well, you know, it's a phrase that I think everyone's very
familiar with. But knowing a phrase and understanding exactly what
it means is kind of subtly different. So, why don't you give me a
rundown?
Geoff Huston 2:36
Well, it kind of works on the theory. The commons were evidently,
I'm like I wasn't around, and neither was anyone else, dear
listener. An aspect of medieval England where a collection of
houses, a village, would have a common area that nobody owned, and
that everyone could graze their animals.
George Michaelson 2:56
So nobody owned it, but everybody shared it.
Geoff Huston 3:00
Shared it. They could use it without charge. It wasn't anybody's
land. It was everyone's land. It was the commons, and and the
problem was it only worked if everyone was restrained in the
number of beasts they sent out on the commons going eat grass.
Because as soon as someone got greedy and sent out, you know,
more, the place just went to rack and ruin because all the grass
got eaten. It turned to mud. You know, bad things happen, and it
actually led to, according to the economists who studied this, a
subsequent move of the fencing of the commons to try and sort of,
if you privatized it, people would look after it better. But
there's a generic kind of principle about unconstrained human
behavior [George: yeah], free resources tend to get exploited to
the point of overconsumption and degradation.
George Michaelson 3:49
It feels a little bit like a version of the prisoner's dilemma
that the only good answer in the prisoner's dilemma is for them
both to do the thing that's mutually beneficial, and if either one
of them tries to get their own safety secured. It stuffs up the
whole thing. The Commons. Everybody's got to be a good player. And
the minute there's one bad apple that takes more than their free
rights, the whole thing goes to pot.
Geoff Huston 4:13
So, free resources get abused. They just do. Free money. Free
whatever. And so there's this aspect of this applied to
the Internet's domain name system.
George Michaelson 4:25
Well, I'm not paying for it, so there's a quality that makes it
feel to me very free.
Geoff Huston 4:31
Well, if I run an authoritative server where I answer queries that
other users have made, so if you want to know about potaroo.net.
It's my server that's on the hook. Does anyone pay me?
George Michaelson 4:44
So, answer here means specifically be the authoritative answer
because there might be many other DNS players along the path who
feel like they are air quotes providing the answer, but ultimately
somebody has to have asked you, Geoff, in order. For that to
happen,
Geoff Huston 5:01
right? Someone needs to have sort of seeded the DNS with the
original sets of you know data, and that's actually an act that
doesn't get compensated by by the people who actually are making
the queries. Queries are free, and if you observe from that, well,
do people over query because they're free? That's a really
interesting question.
George Michaelson 5:25
Yeah,
Geoff Huston 5:25
you see, the DNS works and it's really efficient because it uses
the UDP protocol, the datagram protocol, the protocol that says
every packet is an adventure. If I send you a UDP data packet. You
may or may not get it. UDP doesn't care. You see, this actually,
oddly enough, and let me digress a tiny little bit, was actually
one of the issues with the early forays into packet networking in
the 1960s. We had been used to a form of networking which, when we
moved from analog to digital signals. We regarded the signal we
were switching, even digital, as a consistent stream of packets,
and we ended up switching the most incredibly expensive thing we
could possibly switch: time. So that if you had a common resource
switch between, say, eight people for the first eighth of a
second, it was user A, and then the next eighth, user B, etc. In
other words, you tried to sort of fair share amongst competing
people by actually switching on on a common view of the time,
which is interesting. But what it meant was that the network was
carrying the integrity of the conversations. It didn't matter how
clever or dumb you were when you attached your speaker and
microphone to this network. The network itself would preserve that
digital data as it switched it through the system. Circuit switch
systems. When we went to packetization, there was a debate, and it
was a furious debate. The debate was in a packet switch network.
The packets don't go all the way from A to B. It's actually a
series of of hops. There are intermediate switches, and so you eat
the first switch that then passes it to the second switch, then
the third, then the fourth. Now it doesn't really matter how you
define the sequence of switches, whether it's predetermined,
whether it's you know like the Internet dynamically discovered,
doesn't really matter. But the issue is when one intermediate
system hands a packet to the next intermediate system, does it
care that the next one received it? Ooh, big question. Because if
you care X 25 then you actually get a whole sequence of A sends it
to B, B acts it back to A, B sends it to C, C acts it back to A.
What does this mean? Well, if someone pops a packet in to this
system, it will either come out really reliably, or the interior
of the network will go, oh my God! All hope is lost. Let me
destroy the circuit. Let me give you an indication of failure.
George Michaelson 8:05
Because every hop along the path, people are saying, "Did you get
it? And the person on the other end is saying, "No. In which case,
you have the burden to retransmit it. So bounded in time, it's
going to get there, or the whole thing's going to say, "We've got
a problem",
Geoff Huston 8:21
right? So X 25 based networks, which were really popular for a
small window of time in the g 1980s worked on this principle of
reliable hop by hop transmission. One of the networks that I did
work on, DEC net, and a venerable protocol called DDCMP, and it
was actually the same kind of thing. All the intermediate systems
effectively explicitly acknowledge successful receipt of a packet,
and so the thing injecting packets didn't have to care. Once you
passed the packet to the network, you could assume that it was
going to come out the other end. What you couldn't assume was
speed, because you know, if someone had stuffed it up, they'd sit
there and try and try and try for quite some time before they gave
up on the entire conversation.
George Michaelson 9:08
So, if we bring this back to an Internet context and the idea
of Internet protocol, the actual packets that are the Internet
protocol packets, whether they be four or six, they're kind of
stuck in the weeds on this, they're not saying absolutely must be
on the network like DDCMP or X 25 They're actually saying don't
care what I run on piece of string, giant fiber optic 100 gigabit
channel, totally up to you. I'm just a packet, mate.
Geoff Huston 9:35
Well, that was the genius, and I would say of the Internet that
looking at everyone else who was busy doing reliable, reliable
switching and effectively synthesizing a circuit out of these hop
by hop packet elements, the Internet said, "Nah, every packet is an
adventure. It's not the network's problem to identify and fix it.
George Michaelson 9:57
I think every packet is an adventure. Needs to be on a T-shirt. I
think that's brilliant.
Geoff Huston 10:02
That's right. We're off on a big adventure. Some of us might not
make it. Whoops. And that was kind of it. The argument was, and it
got encapsulated in what what we call the end to end principle,
that it was left to the two machines at either end who were
communicating to figure out if any packets that hadn't made it,
and repair the damage.
George Michaelson 10:23
Right,
Geoff Huston 10:24
and you kind of go, well, that was just IP, and you go, well, no,
that was Ethernet.
George Michaelson 10:28
No,
Geoff Huston 10:28
it was Ethernet. It was a whole generation of successor data
protocol, you know, designs that effectively said, well, this is a
packet system. I don't actually care about the reliability of the
packets. That's up to the communicating systems at either end. All
I care about is just pushing it through, and if I fail, neh; let's
just move on with our lives, right?
George Michaelson 10:49
So we don't actually routinely construct applications and services
into IP packets. We've got a layering model, and there's a choice
moment coming, which I think brings us back to DNS and UDP. How do
you make something on top of this have the behaviors you need to
know is this unreliable or reliable?
Geoff Huston 11:09
Yes, a lot of stuff use TCP. This reliable transport. Now TCP has
overheads because there's no network state, but there really is
state at either end. Instead of I to send you a packet, I have to
go through a dance. Hi, George. I've got a number. Here's the
number I'm thinking. The packet is going to be called a
synchronized packet, so it's a SYN.
George Michaelson 11:30
Hi, Geoff. Here's an ACK for that. And by the way, here's a
number I made up that is my number that you should remember
when I tell you things.
Geoff Huston 11:39
Ah, George, I see you've received my packet and sended my number
back number back to me, and I've got your number. Let me
acknowledge your number. This is the third part of the three-way
handshake, and now we both have sequencing numbers that the others
using.
George Michaelson 11:55
Terribly polite protocol, isn't it, Geoff? So polite. Oh, thank
you for sending me that packet.
Geoff Huston 12:01
Oh, thank you. Here's my first data packet. Now I know I said
magic number 10. Well, it's got 10 bytes in it, so the new magic
number will be 20.
George Michaelson 12:11
Thank you, Geoff. I got that magic number 20. Here's my answer
with magic number two.
Geoff Huston 12:16
Ah, I see that you've received magic number 20. I will delete my
copy because you have it. I don't need to resend it. Thank you,
George. Oh, and by the way, let me acknowledge 2. You see how it
kind of moves. That instead of doing hop by hop, did you get it?
Did you get it? Did you get it? The two end systems play the did
you get it dance by advancing their respective sequence numbers.
The cost. Well, it's a three-way hands shape to start with.
George Michaelson 12:44
So there's a bit of delay coming in.
Geoff Huston 12:46
This whole thing about let's exchange sequence numbers works
perfectly fine if we've got stuff to say to each other. But George
only wanted to say hi.
George Michaelson 12:56
Short conversation.
Geoff Huston 12:58
That's two characters. That's that's two bytes, so it's kind of
SYN SYN ACK ACK " hi" ACK FIN FIN ACK FIN. You know, that's a lot
of packets to go " hi"
George Michaelson 13:10
"hi" yeah.
Geoff Huston 13:12
And when the approach came to think about implementing the domain
name system and its resolution as an Internet protocol, it was very
quickly realized that there were many, many points of
conversation, many points of authority potentially, but each
conversation would be tiny. Hi, do you know about dub dub
dub.apnic.net
George Michaelson 13:34
No,
Geoff Huston 13:34
done. You know that's it. I don't need handshake handshake. You
know, blah blah blah. It's just you know, fine. Tell me yes or no.
I don't care.
George Michaelson 13:44
I think tied up in this is sort of a subsidiary story that if we
were talking about a medium that was so atrociously noisy, one
packet in two never got through. You might be very uncomfortable
with that class of behavior. But we are living in a world where
the behavior assumption was most packets get through, and you'll
need to say, "I didn't get that", or you'll need as a sender to
say, "Looks like you didn't get that. I think I'll do it again".
Is kind of down in the threshold where you're not really paranoid
about checking every single time. It mostly works.
Geoff Huston 14:19
You know, history makes liars of us all, and you're right. That
was the thinking when networks were unreliable. TCP kind of added
strength to it, but if networks were basically reliable, things
were fine. And then we put the Internet on mobile phones, and the
loss rate, at least initially, was awesome. It was one of the few
systems that had a loss rate of one in every 10 packets never made
it.
George Michaelson 14:42
Wow!
Geoff Huston 14:42
It was pretty bad. And you know what collapsed TCP.
George Michaelson 14:46
Seriously, because
Geoff Huston 14:47
if the loss rate is high enough, it never gets. I suppose what I'd
call a full head of steam. It always just sputters. Hi, I's one
packet. Because the way TCP works is that if you send me an ACK
for a packet. I'm going. Well, that's fine. Here's two packets,
and you send me the ACKs for those. And I go, well, George, this
is cooking. Here's 4.
George Michaelson 15:08
I'll ACK that.
Geoff Huston 15:08
8, 16,
George Michaelson 15:10
I'll ACK that.
Geoff Huston 15:11
You know, TCP has this system, which ironically is called slow
start, which is anything but. [George: Yeah] it's actually
exponential start, and it just zooms forward. If I've got a lot to
say, TCP works brilliantly as long as as long as the loss rate
isn't bad. If the loss rate's around two 3% max, TCP is okay
George Michaelson 15:32
because you get the second part of the TCP behavior, which is
under loss. The back off is extreme.
Geoff Huston 15:38
Oh, it's very extreme. Oh, a packet loss! Panic! Ring the fire
alarms! You know, ring the bell. Let's let's back right off. And
so, oddly enough, you might have thought the TCP was sort of
perfect for crappy networks. It's actually very imperfect. And
what actually became really quite neat was UDP, because
applications expected damage. It's kind of yeah, right? You know,
lost a packet. It's not a big drama, and it was UDP was always
intended for short transactions. Although some people abused it
and started doing weird and wacko things using gigabytes of data
over UDP, but they were loonies.
George Michaelson 16:14
Well, they were loonies then, Geoff. But if we look at modern
parlance and QUIC using UDP for megabytes of data has come back
into fashion, but we've drifted slightly here. The point is, UDP
was the way to say I'm not going to invest in a layer to do all of
that reliability checking for me. I'm going to fling packets out
with the least overhead, particularly when I'm asking questions
like DNS questions. What's the address of this name?
Geoff Huston 16:42
Single packet question, single packet answer. And if packets get
damaged, it's left to the DNS application to figure it out and ask
again somewhere else.
George Michaelson 16:52
Ah, now you just said ask again.
Geoff Huston 16:55
Well, yes, because if you don't get an answer, what are you meant
to do? Go and sulk? I'm sorry to get answer. Nobody likes me. It
doesn't work like that.
George Michaelson 17:05
It's not the " Eeyore" network, is it?
Geoff Huston 17:08
Oh, absolutely not. I need an answer. I'm going to get one. I'm
going to keep on asking til I get one. And so it's persisted under
loss, right? Because that's the way we designed it. The
application has to sort of make up for a network that isn't
perfect. Okay,
George Michaelson 17:25
right
Geoff Huston 17:25
now I said before queries are free. So when do you re query?
George Michaelson 17:31
Now I'm a good person, Geoff. I want to go to Internet heaven, so
I'm going to say the right thing to do is to re query when you're
really sure you have to, and not re query any other time. There
you go. I'm a good player in the Commons.
Geoff Huston 17:48
Well, you've also read RFC 1536 that said you'll wait for about
five seconds.
George Michaelson 17:53
That's a long time.
Geoff Huston 17:54
Oh, continents have drifted in five seconds. You know, attention
span. Entire generations of attention spans have come and gone.
[George: Squirrel] exactly. But that was the advice, you know. Be
patient, wait. Don't splutter packets into the network. And of
course, since then, the networks have got bigger. But at the same
time, we have got way more impatient. We are extremely impatient
about this. And so, what's the best way to get an answer quickly,
George Michaelson 18:24
Vincent? Twice, Vincent twice.
Geoff Huston 18:27
Well,
ask, ask, ask, ask, ask. You know,
George Michaelson 18:29
are you there yet? Are we there yet? Are we there yet? Are we
there yet?
Geoff Huston 18:33
Yeah,
I want to drive from Canberra to Sydney. Okay, how do I get there
fast?
George Michaelson 18:36
We
there yet?
Geoff Huston 18:37
Sent five cars, but I'm only in one of them. Sometimes it doesn't
work.
George Michaelson 18:43
That's the million man month statement about you can get a baby in
under nine months if you get nine women to carry a baby, isn't it?
Geoff Huston 18:51
That's the same kind of fallacy, you know. But oddly enough, in
the DNS, it's not quite a fallacy. You see,
George Michaelson 18:57
what?
Geoff Huston 18:58
Well, the thing is, if this stuff is unreliable, and let's let's
go to our magic 10% Then 10% of the time, I'm not going to get an
answer, but I'm going to wait for five seconds. So 10% of folk
will have a five second delay, and they'll ask again. 10% of them
will not get an answer for another five seconds. So 1% of folk
will be waiting for at least 10 seconds for an answer, and 10% of
them won't get it, etc. etc.
George Michaelson 19:24
Well, in a network of 2 billion users, 1% is an astronomically
large number.
Geoff Huston 19:30
Well, it is, and the whole issue comes that maybe that's the wrong
strategy.
George Michaelson 19:35
Yeah.
Geoff Huston 19:36
So one thing you can do is not wait five seconds. I will work out
reasonable approximation is for when you should respond. A packet
can circle round the globe in a second. In fact, you can do it for
half a second if you're lucky, and on you know the right pieces of
cable. So I'll wait for a second because if you haven't answered
it,
George Michaelson 19:54
yeah.
Geoff Huston 19:54
In fact, damn it, 300 milliseconds.
George Michaelson 19:57
You know what? I would feel like this. Was still in the class of
good behavior in the commons. If you're genuinely likely to get an
answer sooner than that, waiting 300 feels like a pretty good fit
for human centric delay, the experience of waiting for a thing to
happen, and the reality sometimes you have to do it again. I could
be happy in that world, and I'd love to believe it's another short
episode, and you're telling me you went out and looked, and that's
exactly what you saw because we're good citizens in the commons of
the DNS. That's true, isn't it, Geoff?
Geoff Huston 20:29
I'm sorry. Yeah, right. So we went and looked. We went and looked
for a day using this ad based experiment, and we sent out a large
number of names, oddly enough, because we don't want caching to
happen in the DNS, which glosses over everything. All these names
are unique, but that's easy. You just make a long name and you
know add some variable fields every time. So on this day, we did
150 million unique names that got asked for by users, and they're
unique, and they're unique in a couple of ways. One, the name only
exists just once in time. That's it. Secondly, the only the only
systems that can answer that DNS query are ones we run. So we see
the query because we're the only people that are authoritative,
and no one can cache it because it's unique.
George Michaelson 21:18
You kind of brought the cache word to the table, and I think we're
probably going to want to come back and talk about that in another
episode because there is so much in cache. Cache is fascinating.
Geoff Huston 21:30
If it wasn't for the cache, there wouldn't be the DNS. But that's
a talk for a different day.
George Michaelson 21:34
But the thing is, you've constructed a world with lots of unique
labels where caching simply can't apply because they've never
existed in the system before.
Geoff Huston 21:43
Yep, 150 million of them in 24 hours. We were very busy,
George Michaelson 21:46
good-sized experiment.
Geoff Huston 21:48
These all over the planet, but we received 250 million. In fact,
270 million actual queries. So 150 million becomes 270 million in
the DNS.
George Michaelson 22:00
Well, it's not quite double, but it's a hell of a lot closer to
double than a small percentage are repeated. That's most queries
are repeated, Geoff.
Geoff Huston 22:12
On average, most queries are repeated. On average, the DNS is
going " to be sure, to be sure", and on average, it's it's sending
out the query a second time. Now, one theory is the Internet's far
worse than anyone ever suspected. They've all wasted their
obligatory 300 milliseconds or whatever their time out might be.
They haven't got an answer in time, and they're requery.
George Michaelson 22:33
Well, I just want to put my hand up here and say we're recording
this using a system that is transmitting live audio and video, and
effectively, is like a two to 300 kilobit stream continuously from
me to you and you to me. And if the level of behavior in that
system demanded every single packet that we are doing was sent
twice, I'm fairly comfortable we'd be seeing artifacts like delay
or break up on this video, Geoff. That's what we used to see in
old noisy systems when things could not reliably get data through.
So if you and me can reliably send hundreds of kilobits of data to
do this video and audio recording, I'm not buying that the DNS is
suffering a noisy network and HAS to retransmit. You've just said
it IS retransmitting, but you haven't said it HAD to do it.
Geoff Huston 23:23
Well, we go a bit further because we're seeing all these queries,
and so what we actually do is then go. Well, a duplicate query has
the same query name a dot b dot c and the same query type a record
quad a whatever because we've got a recording of every query we
see, why don't we time the interval between the first query and
the duplicate query? Why don't we actually see how long people are
waiting? How patient?
George Michaelson 23:50
That's quite interesting. So you don't actually need their clock.
You just have to have your own clock recording the interval
between seeing it once and seeing it again.
Geoff Huston 24:00
Well, I just record the time when I see a query. Go back over the
logs and go. Well, that was the duplicate of the first query.
What's the time difference between the first and the second?
George Michaelson 24:10
And if you're using a system like DNS cap or TCP dump pcaps, you
actually automatically get a packet timer, don't you? An arrival
time, which means you can do the packet interval, either from a
log or from a data stream.
Geoff Huston 24:23
Pcaps, pcap's is an extraordinary amount of data that wade
through. I'm using logs. It's good enough for for this jazz. So
yeah, the log says, you know, here's a time interval between
queries. And we've said before, this is not time out. This is not
a crappy network. So what sort of time interval do people like to
do because as I said, half of all the traffic we see, almost half,
are duplicates. So of the duplicates, just looking at the ones
that weren't the first, but it's a query that happened before,
right?
George Michaelson 24:55
Yeah,
Geoff Huston 24:55
45% of all those duplicates come to us within 10. Milliseconds of
the original. Oh my God!
George Michaelson 25:03
Just before we drill down there, this is quite a skewed
distribution, isn't it? I mean, a small number of these duplicates
probably arrive a significant time after, like more than 300
milliseconds. Yeah.
Geoff Huston 25:14
Well, a small number might well be the first time didn't make it.
I'm twiddling my thumbs and I'm doing a DNS timeout, and I'm
trying again.
George Michaelson 25:22
Real loss.
Geoff Huston 25:23
Real
loss will cause retransmission.
George Michaelson 25:25
Now to come back to the astonishment moment, almost half of the
retransmitted packets are 10 milliseconds or less. [Geoff: Less]
Holy cow, Geoff. Holy cow.
Geoff Huston 25:39
Well, let's up the clock a little bit and go 20 milliseconds, 20
thousandths of a second, two hundredths of a second. That's not a
timeout. No, 60-5% of all queries occur almost like we'd say tap
tap, and it's the same query.
George Michaelson 25:56
I am astounded, Geoff. I mean, to me, that shriek's deliberately
coded.
Geoff Huston 26:01
Okay, so let's dig a little deeper because it's always fun to dig
a little deeper. Do you know about happy eyeballs, dear listeners?
And George, maybe folk do.
George Michaelson 26:11
Got a kind of vague feel. It's an odd thing, isn't it? It's this
idea that we need to construct a way to handle how do we make six
happen if four and six both exist, it's in that space, isn't it?
Geoff Huston 26:23
It's in that space, and it used to be when I'm opening up a TCP
session that a Linux operating system would wait for 108 seconds
before saying that's not going to work. So if you didn't know if
you could open it using V4 or V6, and you tried in say v4 the TCP
part of the operating system would say, "Hey, I'm working on this
three way handshake. Leave it to me. Oh, that SYN didn't make it.
Let me time out and try again. Oh, that SYN didn't make it, and
it'd take 108 seconds to go. I don't think it's working. Let's try
V6
George Michaelson 26:59
I have a story for you from the deep dark days of the
early Internet, when people still used the file called host. txt,
and it took so long to serially check line by line that file to
find where I was coming from that when I used the protocol called
FTP to try and download code from London, getting it from a
machine at the University of Maryland, the 30-second timer on the
login would time out and disconnect me before it had worked out
who I was.
Geoff Huston 27:30
Steam-powered computers, George. They were wonderful. Turn off the
steam. Go to electronics. Oh, you British! My apologies to all the
British listeners. I didn't mean it. This whole idea, though, of
duplicating became a technique where instead of waiting for
conventional timers, you do try and speed up the process. So what
we called happy eyeballs was you actually try a connection in v4
but impatiently. In fact, you do it in six first, and if that
doesn't work, you do a rapid switch over to b4
George Michaelson 28:05
You kind of set two.
Geoff Huston 28:07
It's a two horse race.
George Michaelson 28:08
Two horse race. It's not wait for one then the other. You fire
them off with a just little bit of an advantage.
Geoff Huston 28:15
You handicap the one you prefer, and if you kind of go, oh well,
if you've implemented both protocols over the same paths and it's
roughly the same time. Whoever you handicap will win the race. So
I handicap six. I send out the two horses at the same time. Six
has a slight favor. You'll make a V6 connection if you can, but
you'll rapidly fall back to v4 [George: r ight] and so you go.
That's good.
George Michaelson 28:36
Putting this into the context of the DNS, I'd love to say, Geoff,
it must be that you're measuring happy eyeballs, then. That's
where you're going, isn't it?
Geoff Huston 28:43
No, not really. I'm trying to find an explanation as to why almost
two thirds of all the duplicates occur within 20 milliseconds.
Someone is doing this, bang bang, and happy eyeballs is part of it
because certainly in the DNS, if I kind of want to make sure, then
testing V4 and V6 is not a bad strategy. But there's more to it
than that. In theory, I'm meant to have names with multiple
authoritative name servers, and if I want to increase the
resilience, a very clever recursive resolver, one that's making
the queries, would go, "Ah, Geoff, I see you've got four name
servers out there that can answer my question. Let me send the
same question to all four at once". Queries are free. I'll just
take the first one to answer.
George Michaelson 29:28
Right, but Geoff, that wouldn't look like if you air quote it
duplicates unless you measured across the whole system. Whereas
you're at a single point from what you've been.
Geoff Huston 29:37
So , that's
where the argument breaks down. So let's see what other clever
things people do. One of them, which we have noticed in a number
of ISPs, is they're not really sure about their own DNS
infrastructure. So, well, it's been in a dusty cupboard for 10
years. There's a bit of mold and bit rot growing on the shazzy.
Might not be working.
George Michaelson 30:01
The perversity of the Internet, where you trust the remote party
more than your own equipment that you bought and set up to run
service.
Geoff Huston 30:10
Google's public DNS has an awesome reputation. Cloudflare's all
one has an awesome reputation. So tell you what, let me send the
query to one of these open resolvers, and at the same time, work
diligently hard to resolve the name myself. If I get an answer
from the open resolver, I'm done. Otherwise, I'll keep on working.
Yeah, not a bad strategy.
George Michaelson 30:30
I could believe in this strategy, Geoff. But again,
Geoff Huston 30:33
you could believe it. That would occur.
George Michaelson 30:35
I bring back to the table that would not, in strict sense,
normally look like a duplicate to you would it because you would
see different originators asking you the same question,
Geoff Huston 30:46
same name. A duplicate is the same name, not the same asker.
George Michaelson 30:50
Ah, ah ha.
Geoff Huston 30:52
So v4 and V6 if they're asking for the same name, are duplicates.
Cloudflare or Google and My Resolver are duplicates if they're
asking for the same query type, the same query number.
George Michaelson 31:04
Right. So this begins to sound like a possible collection of
reasons you would see twice.
Geoff Huston 31:10
An
optimization. Yeah. You know, send off four race horses down
different tracks. Whichever one comes back, I'll rub with it. You
know, the v4 racehorse, the V6 the one to Google, the one to
Cloudflare, the one to open DNS, whatever queries are free. Just
send them all out. Take the first answer you get. You're done.
Sounds good.
George Michaelson 31:26
Mostly got a few if but maybes. But yeah, I could be there.
Geoff Huston 31:30
So 1/5 of all the duplicates come from the same IP address. Jaw
dropping moment.
George Michaelson 31:36
Wait, I'm sorry. 45% of all the duplicates are rapid fire.
Geoff Huston 31:40
Yes,
George Michaelson 31:40
and 1/5 of the duplicates
Geoff Huston 31:43
of the rapid fire duplicates of the ones that are tap tap come
from the same IP address.
George Michaelson 31:49
I'm not using two people to ask this question for me. I'm asking
it twice. Wow,
Geoff Huston 31:54
what's your
problem, George? Sorry, what?
George Michaelson 31:57
Didn't hear you. What did you say?
Geoff Huston 31:58
Yeah, what? Sorry, what?
Sorry, what? I'm in a group. Sorry, what?
George Michaelson 32:02
You know, entering the room as a tragedy of the commons. I'm
sitting here, frankly, aghast, Geoff, because this is taking a
service that is notionally free, but actually has a publicly
distributed cost. It is not free to provide the DNS either as an
authority or as a public resolver. It's a cost, and it's a packet
cost in terms of congestion on the wire. And you've just said to
me there could be two or three or four or 10 duplications to get
an answer from a system that we already agree is overwhelmingly
reliable, even if you're using the UDP shout into the void
protocol,
Geoff Huston 32:41
60% of the duplicates are only single duplicates. If I say one or
two duplicates, so in other words, a total of three packets, I get
up to 80% So most of the time, it really is tap tap, and it's so
bizarre. Not tap tap tap.
George Michaelson 32:55
But
isn't that an abuse of the system?
Geoff Huston 32:57
Well, from my selfish perspective, as the resolver, I just want
the answer, dude, and I'm incredibly impatient. So someone else
has the cost to bear. Tough on them.
George Michaelson 33:06
Ah,
Geoff Huston 33:07
it's the selfish. It's the tragedy of the commons all over again.
And I'm sitting there going, "Wow, where is this?" Well, we know
where it is because we know who's asking.
George Michaelson 33:18
Oh dear, name and shame time. That comes up so often in these
measurements, doesn't it?
Geoff Huston 33:23
It does, but I'll be vague. I'll be general. The hot spots appear
to be where there is, if you will, network infrastructure that may
not be so crash hot. So oddly enough, the Indian subcontinent,
including Bangladesh, we see a lot of duplicates. Double the rate
coming out of the U.S. and Europe. That doesn't mean the U.S. and
Europe don't have it. They do, but the rate is much higher in our
servers that serve the Indian subcontinent, and oddly enough, in
China and Hong Kong in that area, where again the duplicate rate
is pretty much double what we see in other parts of the world.
George Michaelson 33:57
So China, we know, has a national strategic view of DNS in the
context of Internet, and it is extremely common for people to use
intermediary devices because they can satisfy national delivery
goals by buying legitimated systems that do the things the Chinese
state considers necessary for safe Internet. So, for China, my
spidey sense is saying this sounds like intermediaries.
Geoff Huston 34:25
It's hard to say, George, because in our experiment we don't get
to see what happened to start this off. We know an ad got placed
in a browser. We know that, but we can't trace the DNS from that
machine all the way through to the one machine we can see the
machine that is asking our server the question. So between the
user and that final sort of hopping off point, it's all black to
us. We can't tell. But it does strike me as amazing that almost
one half of the DNS are ghost cars. They're nothing. They're not
real. They're just duplicates, and most of those duplicates occur
almost back to back with the original query. And so the answer
set, the answer set, is going out to all of them, duplicate and
original. [George: Yeah], it's not as if I don't answer the
duplicate. I've already answered one.
George Michaelson 35:17
No,
Geoff Huston 35:17
we answer everything.
George Michaelson 35:18
I'm kind of wishing there was the facility in the packet of DNS
under this condition of behavior that allowed a flag to be put in
the question, saying this is the second time I've asked. Because
the thing you don't know is the origination order of these two
packets. You know the arrival order. You can't actually know which
of the two was.
Geoff Huston 35:38
You can't tell, and I can't tell. But if it's the same IP address
and all that's varied is the source port, does it matter which was
the original?
George Michaelson 35:46
No, not really. I mean, it's kind of irrelevant.
Geoff Huston 35:49
Exactly. So, in some ways, the most bizarre behavior, where it's
almost as if the switch is duplicating packets. Now, that used to
happen, and I remember seeing it in about 1990, that certain
switches from a certain really cheap vendor decided that when life
was tough, instead of really switching one packet, they'd turn one
packet into two for free.
George Michaelson 36:11
Because the world always gets better when you put more stuff on
the network, right?
Geoff Huston 36:15
More packets equals more joy, you know. But those days,
thankfully,
George Michaelson 36:21
every packet's an adventure, Geoff.
Geoff Huston 36:23
Even when it gets cloned, but you know they were cloned packets.
These are point of difference. It's it's a DNS source port that's
actually changed.
George Michaelson 36:32
Wow, Geoff, this is really fascinating. I think you might need to
move to other forms of experimentation, like doing packet trace on
the client side resolver to see if you can actually see two packet
queries being initiated out of some system.
Geoff Huston 36:47
Oh, at some point we resort back to understanding which resolvers
are doing this and understanding who runs them, and actually
dropping them a note saying, "Hi, we've noticed something rather
bizarre about the DNS. Yeah, you are running. Who's vendor's code?
Are you using? What version is it? How have you customized it?"
George Michaelson 37:06
So, dear listener, if you happen to know the reason that this
behavior is now ubiquitous in the global Internet, Geoff would love
to know.
Geoff Huston 37:15
Oh,
I would love to know. So much of the Internet happens without us
looking, and even if you were running a system that did it, do you
know? Nobody knows.
George Michaelson 37:24
Oh God, I wouldn't have a clue.
Geoff Huston 37:26
No, no
one looks down at this detail. Quite frankly, even if you run an
authoritative name server trying to find these rapid duplicates,
you've actually really got to look hard. And so you kind of
wonder, George, what else is out there of the Internet? Complete
swap. The only difference is no one's bothered looking for it.
[George: Yeah] you know it's out there, just waiting to be
uncovered.
George Michaelson 37:49
Folks, there's a world of measurement out there. We should all be
looking at this stuff. That's really fascinating, Geoff. You've
written this one up on your blog.
Geoff Huston 37:56
It's
coming up in the next couple of days. I'm just in the final
process of crossing the eyes, dotting the Ts, or whatever you do
with I's and T s.
George Michaelson 38:04
By the time this one goes to air, I'm sure you'll have it online,
so we'll include it in the web page with it. That was really
great, Geoff. Thanks.
Geoff Huston 38:11
Thanks, George. Cheers. Till next time.
George Michaelson 38:14
If you've got a story or research to share here on Ping, why not
get in contact by email to ping at APNIC. net or via the APNIC
social media channels. Also, remember the measurement at APNIC.
net mailing list on Orbit is there to discuss and share relevant
collaborative opportunities, grants and funding opportunities,
jobs or graduate placings, or to seek feedback from the community
on your own measurement projects, be sure to check out the APNIC
website for all your resource and community needs. Until next
time.