Geoff Huston 0:00
Well, we kind of wanted to do the same thing in networks, because
wouldn't it be good if we could take the service, the content, and
pre-distribute it closer to consumers who wanted it, and then when
we say it's here, it's everywhere, and do we need to even wait and
pre-provision? Can we use the secret of things like the DNS
itself, where the first answer gets copied by the DNS
infrastructure in computers close to who's asking, close to but
not the same, so all your other customers in that network who ask
for the same domain name shortly after the first simply get a copy
of what the first answer was. What's the IP address of google.com
Oh, that's boring. I've already done that 20 times today. Have a
cache copy.
George Michaelson 1:01
You're listening to Ping, a podcast by APNIC, discussing all
things related to measuring the Internet. I'm your host, George
Michaelson. This time I'm talking again to Geoff Huston from APNIC
Labs, in his regular monthly spot on Ping. In several recent
episodes, Geoff has asserted the routing function in the Internet
has moved from IP level packet forwarding models into a process
driven in the name to address lookup function. Who you are and
where you might want to go is no longer just about your IP address
and what you might think is the other end across a network, it's
now about where you are and where intermediary services can best
service your request from, but how does this really work in
practice? What kinds of processes determine where you should be
served from, and who decides? IP level routing hasn't gone away.
There's still a very important routing method called Anycast for
selecting the closest server to be used, and this is deployed at
scale, but it's increasingly being augmented by name-based DNS-
based methods to determine where to go. Geoff, welcome back to
Ping.
Geoff Huston 2:17
Hi, George, how are you today?
George Michaelson 2:19
I'm good, Geoff, but I'm confused because you have spent so many
of your recent ping recordings talking about how names are
replacing routing as the steering artifact for the global
Internet, and I keep meaning to ask, how does this work,
Geoff Huston 2:36
right? So let's get into that, because I actually think this topic
is totally fascinating, and I actually think it's revolutionary in
the change that it's making to the concept of computer networks,
you know. Computer networks were, I suppose, a lot like telephone
networks. I'm here, you're there, and the job of the telephone
network is to actually carry my voice to you and carry your voice
to me. [George: Yeah], I can't teleport you, I can't change you, I
can't duplicate you. The network is there to bridge the two of us
together without either of us getting out of our chairs. Great.
Now, when we built computer networks, the first thing we did is
build it in the image of the thing we knew. Didn't have any other
model? [George: Yeah], so we kind of built them as a way of having
conversations between computers. It's over there, mine is over
here. You know, I tap on my keyboard, little packets get scurrying
over to your computer. You tap on your packet, your keyboard,
little packets go scurrying back.
George Michaelson 3:39
It's kind of inherently, me to you, you to me. It's not typically
a form of communication that's inviting the idea of me to 16 other
people or me to two people at once. It's building it point to
point, you to me, me to you,
Geoff Huston 3:55
right. And, oddly enough, you know, there were these other models,
and we have their paradigms to some extent, you know, radio and
television is a form of broadcast where everyone within a certain
radius can listen to a single point source, and the source is
available to anyone within that served geography, and multicast is
slightly more selective, it says I wish to simultaneously get a
bunch of people to listen, and I'd actually like to agree that
they're listeners, but you know, we have computer analogies,
computer networking analogies of those other models as well.
Although fascinatingly, they were really hard, and they never took
off in the mass market. [George: Yeah]
Geoff Huston 4:40
you know, part of the problem with multicast is that it was, to be
perfectly frank, a really fragile technology, and the only folk
who tried to use it were actually folk who wanted to have a
digital version of radio broadcast, and the idea that the audience
would all line up at the same time to listen. Listen to you seemed
a bit far-fetched in today's world. They don't do it like that,
and so it never really caught on.
George Michaelson 5:06
Well, I've just watched the launch of Artemis Two on the start of
its journey, mankind reclaiming outer space one rocket at a time,
and because of the nature of an event that is happening at a set
time, there are still things - sporting functions, opening of
parliament, royal marriages - where people want to line up at the
gate, but if you're talking about the vast majority of digital
content that we used to see as a collective engagement, it's no
longer available at 9pm and only at 9pm it's you watch it when you
want it, so that model of let's get everyone together in time and
say "go", it kind of narrowed down,
Geoff Huston 5:51
it's gone,
George Michaelson 5:52
it's gone.
Geoff Huston 5:52
I didn't see the launch, I'll catch it up later on streaming,
George Michaelson 5:56
ha,
Geoff Huston 5:57
me and billions of others, no doubt, so yes, that paradigm of
simultaneous audience, I think, doesn't exist anymore. I think the
problem was, in so far as the original version of computer
networking was actually small transactions, they weren't large
ones. It was more about replacing the office fax with electronic
mail. It was web pages that were realistically short, sweet, and
simple.
George Michaelson 6:26
Yeah,
Geoff Huston 6:26
images rather than movies.
George Michaelson 6:28
If we can make the event of getting it really quick, you can get
on with reading it at the speed of human eyeballs, and me, the
network, I don't have to sit there, I can deal with someone else's
problems,
Geoff Huston 6:40
right? So the task of the network was simply to carry the
consumer, the user, the client to the server, where whatever it is
they wanted was, and carry the answers back. And by and large,
this served us for a couple of decades, but as our computing
technology got better, our displays got better, and we were
digitizing almost everything. What actually happened, of course,
was that computing was getting more interesting. You could do
bigger things.
George Michaelson 7:13
When I transited from Britain to Australia, this was around the
time that Ken Thompson, one of the two people who built the Unix
operating system we all live and depend on nowadays. He was one of
the privileged people who had a server at home with a 600 megabyte
disk drive that he had dedicated to recordings of music, because
he was exploring psychoacoustic effects under compression of
audio, and I remember having conversations with people saying you
couldn't afford to send music like that down a network, if you had
a modem, it would take a month to get even an hour's worth of
music out of it, but the thing is, he was researching something
that, as you've just said, has come into the fore as computers
have got faster and as networks have got faster, we can now do
things that we simply couldn't imagine when this form of
networking was being invented.
Geoff Huston 8:07
So, yes, computers have got faster and cheaper, storage has got
cheaper, 600 megabytes. I'm sorry, a significant problem today is
600 petabytes, a major problem is 60 terabytes, but under that,
get into the store and buy a bit more, either spinning rust or or
your solid state storage, and you're done. So, in some ways, we're
living in a world of abundance, but one thing hasn't changed, and
oddly enough, that's the speed of light over distance, and when
you think about computer protocol performance, distance is a pain,
because the way we actually make networks work is a feedback loop.
Oh, go a bit faster, there's room in the network, says your
feedback, and it's this delay between desire and reality. If you
stretch it out, it's like living in a world where everything
happens a second later, you have to slow yourself down, because
you do something, you have to wait for a response. And the issue
was that, although we had heaps of bandwidth, heaps of computing,
heaps of storage, we still had this problem that distance was
killing us,
George Michaelson 9:19
right, because the effect of the speed of light means that the
feedback loop you to me to either tell me speed up or slow down,
that message necessarily has the delay of the speed of light
between you and me. The further apart we are, the more delay there
is between you saying stop and me actually stopping, which means
my responsiveness to changes in the network, whether it's full or
empty, congested or noisy, is necessarily bound in that delay.
Geoff Huston 9:50
So it's really hard, oddly enough, even though we have high-speed
networks, very high-speed networks, it's really hard for a user, a
client. To actually exercise that, so if I want to stream a video
from I don't know, I live in Australia, so stream a video from the
United Kingdom, from France or Germany, from the other side of the
world, even though every individual component might be megabits of
capacity and there might be available megabits of capacity, I
wouldn't get it, and this was kind of the issue that was
confronting originally folk like Microsoft and Apple. Their
problem was that they released operating system updates every
couple of months and said, "Right, there's a new update for
Windows available, come to Seattle with your network and pull it
down from our servers. Now, nothing wrong with that, other than a
few 100 million people wanting to do so simultaneously.
George Michaelson 10:49
Yeah. Oh my gosh, the Windows releases midnight in America. I'm
going to wait till one minute past midnight and fetch it.
Geoff Huston 10:57
You and 100 million of your best mates, and so things melted and
didn't take us long to figure out. Well, if computing and storage
is so cheap, let's replicate the content in advance. It's like in
the old days, a book publisher would send in advance copies of
Harry Potter and whatever it was to bookstores, and then announce
the opening day, and the stores had already been pre-provisioned
with the book.
George Michaelson 11:23
Yes, there were a few lawsuits where the stores were selling pre-
release copies of the book under the counter to journalists, and
people were breaching the magic wall of
Geoff Huston 11:31
The Secret Covenant, or whatever it was. Well, we kind of wanted
to do the same thing in networks, because wouldn't it be good if
we could take the service, the content, and pre-distribute it
closer to consumers who wanted it, and then when we say it's here,
it's everywhere, and do we need to even wait and pre-provision?
Can we use the secret of things like the DNS itself, where the
first answer gets copied by the DNS infrastructure in computers
close to who's asking close to, but not the same, so all your
other customers in that network who ask for the same domain name
shortly after the first simply get a copy of what the first answer
was, what's the IP address of google.com Oh, that's boring. I've
already done that 20 times today. Have a cache copy.
George Michaelson 12:28
Cache, the magic word emerges. Caching.
Geoff Huston 12:32
Well, I'm actually thinking it's distribution. Caching is what we
call it in the DNS, but in the server world we call it content
distribution. That's what you're trying to do is to take a big
thing, be it a video or any other kind of large volume server or
server's content, and as you stream it outwards, you keep a copy
close to, but not right inside the end user's environment in their
ISP, for example, so now when anyone else asks for the same video,
and I bet you they will, I'll be able to serve it remotely and not
touch the network. Brilliant, all of a sudden I've got rid of the
network, it's gone.
George Michaelson 13:14
Well, you needed a pipe, you could push the stuff down to get it
to all those places you're choosing to put it, so "a" network has
to exist for you to push stuff to the places you want to cache
copies of it, but "a" network isn't the same as "the" network, is
it?
Geoff Huston 13:36
Well, let me make your brain explode, then, because let's Geoff's
content distribution system rig up a bunch of virtual servers on
borrowed hardware or leased hardware in a few 1000 places all
around the world, close to populations, and they will serve
Geoff's content locally over the Internet protocol. Fine, how do I
feed those front ends? Well, I have the Geoff content factory
somewhere else, doesn't really matter where, and I want to feed it
in an advance. Oh, that means I could trickle feed it. I don't
need to be fast, it's not on demand, I'm pre-provisioning. What
protocol should I use to do that pre-provisioning? Apple Talk
doesn't matter,
George Michaelson 14:28
private,
Geoff Huston 14:28
you don't see it,
George Michaelson 14:29
it's private,
Geoff Huston 14:30
it's all private, it's private, and whatever protocol I choose to
use is kind of my business between my master server and my front-
end, you know, retail points of presence, my content points of
presence. So, in some ways, all I'm trying to do is deliver this
same content all over the world through the last hop, but how I do
so internally is kind of my problem, and no one else's, and so
kind of doesn't matter, and so. Google internally uses BBR or IBM
SNA. No one cares. It's not our problem, it's their problem.
George Michaelson 15:09
Yeah, from talking with the engineers who manage the FreeBSD
distribution framework, I think they are existing in the margins
where they have to try and reduce cost here, and they've
constructed almost, I believe, a two-tier model. They have
regional nodes that they push to, which means the center only has
to arrange for maybe five copies to be pushed out, and then local
nodes do a pre-fetching, a pre-provisioning fetch from a regional
node. So they've kind of invented two levels of behavior, and
they've got less complexity, but in the end, nobody in the real
world has to care, because all they see is fetch-free BSD from a
local mirror.
Geoff Huston 15:52
You could be talking about Akamai, or Fastly, or Google, or any of
the others. They all use the same techniques. It's really common,
and it works. That's a good thing. It works. So, okay,
distribution is kind of a solved problem, but let's go to the
other side of this. You want to go to Geoff's favorite videos from
Geoff's Content Distribution Network, and I have 1000 points of
presence now. If you make a dud choice and pick a server that is
not near you, but on the other side of the world. No one's better
off, are they? [George: No] you're going to have a bad experience,
because that was a bad choice.
George Michaelson 16:32
I feel like this has been the longest preamble in the history of
podcasting, Geoff, because we've just arrived at the entry point.
How, how does it work?
Geoff Huston 16:43
Well, I was wanting to kind of motivate it, and
George Michaelson 16:45
you got there
Geoff Huston 16:46
everyone to understand the nature of the problem and kind of why
it's not straightforward, and there are a number of techniques, if
you will, to actually answer this, and the first approach, which
oddly enough was very controversial, was actually to put the same
computer platform everywhere that you're serving, so this computer
platform uses address 192.168 dot 0.1 and I put another one in
Paris, 192.168 dot 0.1 and I put another one in New York, one in
Sydney, one in Singapore, one in Beijing. I use the same address
everywhere, and I go to the routing system and go, Hi, I'm over
here, and here, and here, and here.
George Michaelson 17:33
Now you've been at some pains when we talk BGP and the behavior of
BGP selection of a prefix to route to, you've talked about how, if
you announce a big address and then announce some more specific
wins, that becomes the thing you look at, and separately, you've
talked about the loop detection mechanism becoming a cost measure
that lets you determine which is the lowest cost path to take, and
I'm looking at this situation where you're announcing exactly the
same address in BGP in lots of places, and I'm asking myself, what
mechanism makes me pick one near me and you pick one near you, and
some Parisian pick one in Paris.
Geoff Huston 18:18
Well, you said it yourself, BGP picks the route with the smallest
metric. Now, in BGP, the metric is the number of networks I need
to traverse to get to where it's being announced. So, if this
conversation is you and I sitting in, say, America, and the server
is being announced in, say, Paris. It is likely there might be a
local network, a transatlantic transit network, and another local
network. It's likely the cost to get to the Parisian server is
three, [George: right] But I've also got my server in America. It
may well be that it's very, very close to you, an adjacent network
cost one
George Michaelson 19:05
right,
Geoff Huston 19:06
BGP picks the lowest cost, and so if you go to 192.168 dot 0.1 and
you're in America, you will go to the server that the routing
system BGP thinks is closest to you,
George Michaelson 19:20
so in a sense it's almost like a decision-free outcome. I don't
have to pick where to go. Rooting optimizes the decision facing me
and the guy next door, and simultaneously some Parisian guy
sitting in a flat. We all get optimized in rooting to the closest
point,
Geoff Huston 19:40
right! And you go, geez, that sounds weird. Let me point out a few
major uses of this anycast approach, where you put the same
address everywhere, and I point to the root of the DNS. The root
of the DNS is served by 13 unique V4 addresses. And 13 unique V6
addresses,
George Michaelson 20:02
right.
Geoff Huston 20:03
Each notional server has a V address and a V6 address. So, there
are 13 root servers. There are 13 machines that serve the root.
No, there are 1500 of them. Well, how does that work? Each root
server actually has a large number of servers, servers all over
the world, all over the Internet, all listening on the same IP
address, the same, so there are 13 separate little what we call
anycast clouds in v4 another 13 in V6 and the routing system
automatically takes you to the closest server. No one does
anything. BGP looks after all of this, and it's actually not a bad
approach. You've kind of outsourced your problem to BGP, knock
yourself out, kiddies, you know. So not only does it keep the root
of the DNS evenly loaded across the world, and stops the root
servers from melting through overuse or any individual server. The
technique also works for efficiency. You just get taken to the
closest one. Who uses that technique? Cloudflare. So, Cloudflare
have a very small number of IP addresses, but if you look hard at
their network, you actually find the same IP address pops up
everywhere, because they're anycasting and therefore direct you to
the closest content on Cloudflare's network is certainly left to
BGP, easy.
George Michaelson 21:36
So nothing comes for free, Geoff, and I'm going to imagine there
are potential pitfalls in this, I mean, I'm sitting here thinking,
if one of these boxes breaks and I need to log into it, I can't
use this magic anycast address to log in. There's got to be at
least another address on it that is unique to each, each box.
Geoff Huston 21:53
Oh, you need a, you need a service address that's unique, but let
me tell you what the problem is. The network, in topology sense,
is not long and stringy, is short and fat. The average AS length
on the Internet is a little over four networks, four, so the
average AS metric is four, which is really coarse, and if you
attach to a big spanning network, so if Geoff's ISP spans all of
America and Europe with one AS [George: ah right] Then, in terms
of BGP metrics, every location in Europe and America is the same
distance. How do you know what's closest,
George Michaelson 22:36
right?
Geoff Huston 22:36
And the answer is BGP doesn't, and so that will give you poor
outcomes if you are dealing with networks that span large
distances, and guess what, there are networks that span large
distances, and at that way anycast kind of makes, unfortunately,
suboptimal decisions, and we see this.
George Michaelson 22:59
right, you could wind up being in America with a resource in
France looking equal cost to a resource quite close to you
logistically on the continent, and you could incur that time delay
of a sub oceanic fiber optic link, because BGP can't tell you you
should look somewhere else,
Geoff Huston 23:18
right? And the right research folk, one of their early experiments
with Atlas and mapping was to actually ask their Atlas nodes,
which instance of A.root-servers.net do you go to when you ask the
root of the DNS, which instance of B, and so they were trying to
actually isolate how well does this root server any cast system
actually work, and obviously when you go looking for anomalies,
you find the world is full of them, and of course, there were
Atlas nodes in Europe who thought via BGP that the closest anycast
instance of one of these servers was in places over the other side
of the Atlantic in America, and so on, and it wasn't rare. It's
common, because BGP is pretty coarse, you know. Short fat networks
don't really do a good job with anycast. Okay,
George Michaelson 24:12
and if I front up on a big network and say, hey, could you make
your network a bit more granular? Could you grind it up a bit to
divide it into the chunks that would make my life easier. The
inevitable point is, yeah, I could do that,
Geoff Huston 24:25
but you've got to pay me. You've got to pay me large amounts of
money, and the real answer is, I don't think you're going to,
because I'm not hosting your content, and I'm not hosting your
customers, I'm just the transit dude in the middle, and so the
economics don't really make this any better, and trying to make a
longer, thinner network out of a short, fat network is kind of
anti-gravity. No one wants it, so it's not going to get any
better. Oddly enough, it's going to get worse as we move on. So,
in some ways, anycast has its limits, and if you really want to
start shaving things down. If you really think there should be a
difference between Frankfurt and Paris, if there should be a
difference between Paris and Lyon, there should be a difference
between Paris and what's a town very, very close to Paris, Paris,
Orleans, a few kilometers away. If you really want to make that
difference, or a distinction between Tokyo and Asaka BGP isn't
going to do it, it's not
George Michaelson 25:23
right. So, BGP was almost the decision-free mechanism where you,
the content provider, and me, the client, have nothing we had to
do to find what was considered the air quotes best choice, and
what we're kind of getting to is, well, you actually are going to
have to make a decision now, and there's that magic moment. Who's
making the decision? How do they do it?
Geoff Huston 25:46
So, if I can't use anycast and get a free ride with routing, let's
try a different approach. So, again, I have my content
distribution network, and I give each of these points of presence
a different IP address, 192 168 1,2,3,4,5 so let's have 1000 of
them. You asked for Geoff's content.com Now I've got to go. Oh, I
could give you in the DNS 1000 different answers. Which one is the
answer that's good for you? And you put your little hand up,
going, well, I didn't really want all 1000 answers, I can't cope
with that, not my problem, but don't give me the one that's
furthest away, please, just don't do that, Anchorage, Alaska is
only good if you're living in Anchorage, Alaska,
George Michaelson 26:36
So I trigger a decision having to be made because I say I want to
get somewhere, and you, because I'm getting somewhere by asking
about, tell me the address of a name, your systems, or someone on
your behalf is getting that call that's saying, I, George, I want
to do this, and the problem is, How can I know based on you asking
where the best place for you is? What do I know to help decide.
Geoff Huston 27:03
So, let's make a sweeping assumption, and it's probably not that
bad a sweeping assumption. You and your DNS agent, your recursive
resolver, the thing that actually asks your questions for you is
normally provided by your ISP. Let's assume that that's close to
you, so now I have a third party, your resolver, and I have it
asking me, the authoritative server for my content distribution
network, a question about content, content that I'm serving, so
now I've got a simple problem, or simpler, an IP address is asking
me about a server. I have 1000 possible IP address answers for
that name. I just need to pick the one that's closest. Okay, we
have geolocation databases in its crudest sense. I can map to
countries. Let's put a country code on your resolver that asks the
question. Fair enough. You live in Japan. Let's shuffle through my
list, find an answer where the server is located in Japan. That's
the answer. So now you're getting steered to something that I
think is good.
George Michaelson 28:13
It definitely sounds like an improvement against large flat
networks that potentially span interoceanic networks and different
economies, but I'm not feeling like you've kind of got perfection
yet, Geoff, because some economies are very long and thin. Chile
is a massive, long, thin stripe down the side of South America,
and if you happen to be right down the bottom at Ushuaia, you are
not close to Santiago, right up in the north, Japan is a very
long, thin island. It's quite a lot bigger than people think,
because of the Mercator projection. And if you're right up the
north, you are not close to right down the south. So you fix the
problem by keeping it within Japan. I'm giving you a thumbs up. It
still hasn't got close to me.
Geoff Huston 29:00
Well, before you start throwing rocks, let me observe that this is
the model used by Akamai, one of the major content players on the
planet. They do this by DNS. It's not even a clever trick, it's an
obvious thing. When you serve content through Akamai, when you say
take this content and serve Akamai, they say to you, look, instead
of serving or giving an IP address of your server in the DNS
CName, it alias it to us, and so they say, if I want to serve
example.com they might say, in your DNS master zone, use the CName
for this service, example.www.example.com.www.example.com and C
name it to dub dub dub.example.com dot edgekey.net Who owns
edgekey.net Akamai. Oh, now I've got dub dub dub.example.com dot
edgekey.net Where does that. That map to another CName. Now the
trick is that that second mapping depends on who's asking, because
I have maybe have 4000 possible answers that identify an edge
location by name,
George Michaelson 30:17
right?
Geoff Huston 30:17
And I do a magic thing that says www.example.com dot edgekey.net
CNames to the server in Tokyo dot akamai edge.net What's its IP
address? Well, that's constant, that's easy. The server in Tokyo
dot akamyedge.net you know, is 23.1 95 dot 84.2 41 or whatever,
George Michaelson 30:39
right?
Geoff Huston 30:39
And so now with just a little bit of mucking around with Cnames,
I've got a decent way of doing this, as long as I can triangulate
and say, "A, you live close to your DNS resolver, so if I optimize
my answer against who's asking, I can steer you to one of the
servers that serve the content you're after that's closest to you,
George Michaelson 31:04
so having kind of kibitz and said yeah, don't like this, this
don't think it's perfect, I want to back off and observe, I'm sure
I'm using Akamai services all the time, proof by example, I'm not
seeing massive amounts of "this kind of sucks", so as a consumer
of services, I think we might be in 80-20 land. Geoff, this works
mostly if I'm not doing stupid tricks with my DNS, and if I'm
mostly using a DNS server close logistically to where I live and
work. It basically seems to work.
Geoff Huston 31:37
Many folk might bank with the Bank of Akamai, they might, it
works, and the beauty of this is also that you can't see those
translation names when your browser then connects to this IP
address of an Akamai server close to you. You say, Hi server, I
want to connect to www.my bank, and I want to create a TLS
connection, and the Akamai server goes got a problem. I am
configured to be an alias for that banking name. Here is a key,
here is a key exchange, here is a TLS connection. Let's do this,
so the entire thing is seamless to the user, and I've now managed
to by doing a simple sort of trick with the DNS, steer you to
something is really close to you. Presto, it's fast, it's cheap.
Did I mention it's fast? It's really fast.
George Michaelson 32:31
Yeah,
Geoff Huston 32:31
because it's local, so kind of works. Yes. So now, George, let's
start throwing rocks.
George Michaelson 32:38
Yeah, so I want to watch content from the BBC, and the BBC says,
'Sorry, chum, your DNS is in Australia. So I ring up a mate with
the DNS server in the UK, and I let him let me ask questions
through him, and the BBC go, 'Not a problem, buddy. I see your
questions coming from the UK. Here is Doctor Who'. It feels like I
can drive through geo-fenced IPR rules if I can make my questions
come from the UK
Geoff Huston 33:10
more than 10 years ago with Netflix. If you changed by simply
going into the config part of your computer and changed who your
DNS resolver was to your country of choice, Netflix would open up
the library of that country because it assumed you were your DNS.
George Michaelson 33:27
Yeah,
Geoff Huston 33:28
and yes, it's it is, was, and probably still is abusable from that
respect.
George Michaelson 33:34
Yeah,
Geoff Huston 33:34
and the other problem that's just as bad are these open DNS
resolvers,
George Michaelson 33:40
yeah,
Geoff Huston 33:40
like 1.1.1.1 and 8.8.8.8 Well, the thing is, which actual DNS
engine did your query go to? You don't know, but when Akamai get
the query coming from one of these open DNS resolvers, say Google,
it's not clear that the address being used by that open DNS
resolver, when it asks Akamai is close to you. Oh, you're coming
from Paris. Oh, you're coming from Moscow. Oh, you're coming from
Madrid. You know, it's really, really hard to nail it down.
[George: Yeah], and so under some pressure, Akamai and a few other
content folk convinced the big content distribution networks to
add a security leak, a nightmare. When you ask the authoritative
server, why don't you attach the client address? Oh, can't do
that. Don't want to do that. Why don't you attach the client's
subnet? Oh, that's okay. That's not very privacy leaking rubbish.
We'll use that, so we actually intrude into the DNS between the
recursive and the authoritative an item of data that actually
points to the original user,
George Michaelson 34:51
right? So the person asking the question is the primary vehicle
that can at least target in a lot of cases down to an economy.
Economy, but we've changed the protocol to intrude a flag going
out asking the question, this geezer in this suburb in Ginza is
asking about this TV show, and the belief I have is, oh, well, you
just identified me to Ginza, there's a million neighbors, and the
belief you have is nobody's lying here. He really wants the server
in Tokyo, not the server in Hiroshima, because he said Ginza, so
it kind of sounds better, Geoff.
Geoff Huston 35:30
It kind of sounds better, because in essence, you're giving a few
more clues of information. Okay, how good is a subnet? It's a
privacy protection. Meh, not very good, you're leaking information
about you that maybe you didn't want. The plus side of this, who
controls the mapping from name to number? You see, in BGP Anycast,
the service provider had absolutely no control,
George Michaelson 35:57
right?
Geoff Huston 35:57
BGP, but here I am, my DNS servers control that mapping, I can
twiddle and play if the server in Asaka is flat on its back,
howling at the moon, and the only available capacity is somewhere
else, Tokyo, whatever, I can actually intrude into this, and with
a short life answer, I can tell you somewhere else,
George Michaelson 36:21
short life, that would be because these kinds of answers wind up
having a cache behavior, and if I've cached in your system, you
have to go over here for the lifetime of the timer, I'm going over
there.
Geoff Huston 36:34
The thing that makes the DNS work so well, caching kind of is your
enemy if you're trying to do some degree of dynamic load balancing
or optimization, because once you give an answer in the DNS to
that resolver, everyone else who queries for that name is going to
get the same cached answer,
George Michaelson 36:53
right,
Geoff Huston 36:53
and if you want to recompute, then the only thing is give it a
short lifetime in the cache, because otherwise the first answer is
sticky for an unknown number of subsequent people, and so to do
this, you lose cache.
George Michaelson 37:07
So, these are two different kinds of rocks being thrown here. One
is it's a bit of an anti-pattern for efficiency in the system.
Okay, we have to deal with that, and the other is it's a big
privacy concern. You don't really get to keep your privacy when
you ask these questions anymore.
Geoff Huston 37:24
Yeah, tough. So, what do you do? Well, we've just discussed, I
think, almost the transition of networking and technology. We
originally started with a smart network and dumb devices, it was
called telephony, and we started to put more and more smarts in
the edge devices and strip out the middle, so using BGP kind of
relies on magic, and BGP devices need to do very, very little.
Using the DNS lifts it up one level, so now you've got the DNS
doing this distance mapping and trying to get you to the optimal
point. Why stop there? You could do what Netflix do. There is an
Amazon server, or 20, or 1000 where all the initial transactions
for Netflix content get passed to on the West Coast of America.
Oops, but that's okay, because then there's an application
problem, because you divide your content into manageable chunks,
little bits, and you have the ability to go well, user, you're
making contact with me, and you want this content. Now, it's not a
DNS question, it's not even an IP routing question. Now it is a
content streaming question in the application level,
George Michaelson 38:43
and it's between me, the user who is paying you Netflix for this
service. I'm giving you money, and so this anonymity and privacy
thing. Well, I was never going to have any privacy in you, because
I was always going to pay you to give me the things that I tell
you I want to get,
Geoff Huston 39:05
and you want me to know it's you, because you're paid me, and you
want that service, so you know
George Michaelson 39:10
I do.
Geoff Huston 39:11
This is now consensual disclosure.
George Michaelson 39:14
So, there's a third thing kind of buried underneath here, which
isn't really material for the technique we're talking about, but I
think it's part of the story. You also, when you send me things in
Netflix, you send it fast enough that I can build up quite a big
buffer in time. I get content from you that spans probably two or
three minutes in me, so that even if you have to stutter a bit
sending stuff to me, I don't see that watching this movie, and so
that chunking technique, where you're saying, "Oh, that server's
dying, I need to move you somewhere else". In your clients, you
know, most of them have a bit of a buffer here, where if they have
to flick open a link somewhere else and start dealing with it, it
isn't actually going to affect their streaming.
Geoff Huston 39:59
Everything is a compromise, isn't it? You can do, and for a long
time we did do high watermark, low watermark, trying to get the
playback buffer in the device that you know the customer is using,
the television, if you will, trying to keep that buffer full, and
when it got to a low mortar mark, you just slam data down the
line, and then went quiescent until the playback got to the low
watermark threshold, we call that the Netflix spike, and it's a
bit like hitting the network with a sledgehammer once every few
minutes, and it's just as painful. Modern thinking says don't
stress the network like that, bad things happen, you just drive
the network into bad places, try and even your load, [George:
right] Try not to do stress techniques, but do the gentle feed.
Yes, it makes it a little bit more difficult to change where your
stuff is being served from when you need to, but at the same time
the network is going to behave a whole lot better. The good news
about this approach, because now it's at the application level,
and I have control over who is delivering you elements of the
service. Is I can start playing with what if questions.
George Michaelson 41:10
Yeah,
Geoff Huston 41:10
I've been serving you from Asaka for quite some time now. I wonder
if Tokyo could deliver you better service. We're up to chunk 1023
Let's deliver 1024 from Tokyo, and tell me, because it's my
software at the other end. Tell me, how that was for you.
George Michaelson 41:27
Yeah,
Geoff Huston 41:27
and I can then dither and try and get really fine-grained accuracy
on what is the best service at this point in time,
George Michaelson 41:38
and so this mechanistic behavior, it might be using
intermediaries, it might be using locally best service anycast in
all kinds of weird ways, but we have moved back into application
layer decisions between you, the sender, and me, the recipient.
It's an instrumented dialog, but you get to understand it without
intermediaries having to make inferences. There's ways that you
can control the behavior, and I can control the behavior in
software.
Geoff Huston 42:06
So, as we've moved content up into video, as the content stuff
gets bigger, and we're now talking about, you know, hundreds of
megabytes or gigabytes, then the whole idea of make a decision
once on what's best and stick with it, which is Akamai's decision
doesn't really work very well, and so you have to constantly use
some kind of adaptive feedback to check if that original decision
could be bettered if there's another way of doing it to improve
the response, which is where Netflix and YouTube and a number of
other streaming platforms have gone to, because it becomes an
application issue that can best actually address this problem.
George Michaelson 42:43
So, let me just run history on the quick track. We had its point
to point, and you could even have operators plugging between patch
boards to connect you to me. We had its packets, and we're going
to give them addresses, and suddenly it's about knowing how to get
there, which means routing. We had routing can be made to look
like one thing is anywhere, and you have to use a routing trick to
pick which is the best one. [Geoff: Yes], we then say that isn't
good enough. Can we go back to something that's about names? Who
am I, and what do I want to get? And we start using mechanisms in
names, and you've now moved up to well, it is names, but it's also
got a lot of what are you actually doing with me? Let's measure
this, let's optimize this kind of sounds like it's name-based,
it's a bit of the routing story still in it, but it's also
smarter.
Geoff Huston 43:37
Well, it's kind of pushing the problem further and further up into
where there I think there is intelligence and information to make
it so the DNS problem is I have control of that translation, I can
select a server for you, but then we're both stuck in that
selection, we can't change it easily. When I start to dither at
the application level, I have a huge amount of flexibility, and I
can actually figure out was my original decision good or bad, and
if it's bad, are there alternatives that can offer a better
service? Don't forget, too, that there's a lot of remedies. I can
reduce the quality of the signal to reduce the volume, or
preferably I can stay at the same playback quality, but move to a
server that has a better chance of keeping that quality up for
your experience, and so that flexibility sort of plays into hands
of the video streamers going, I'm really interested in user
experience and the quality of that, I think I could do a better
job. What does this mean, I suppose, for interoperability and
standards and all that kind of stuff. The answer is no, not really
anymore.
George Michaelson 44:45
No, this is moved out of the domain of let's do it all the same
way into I've got special secret sauce.
Geoff Huston 44:53
There's Geoff's secret sauce, there's George secret sauce, and a
playback from Geoff's server might not necessarily. Necessarily
work with George's receiver, it's not that kind of
interoperability anymore.
George Michaelson 45:06
Yeah,
Geoff Huston 45:06
each of these application level content factories and distribution
systems work in their own ecosystem.
George Michaelson 45:14
Yeah, we're building vertical markets and vertically walled
gardens. You want to do this behavior, it's going to be good in
me, and my special sauce makes it better in me than from the other
chaps. Stay with me,
Geoff Huston 45:27
right? And it kind of goes well, but I'm a minnow in this world,
I'm just a starter. Oh God, I'm sorry. You're back to banging the
rocks together, dude. You know, you're doing a web-based playback
buffer because you're not big enough to build your own ecosystem,
[George: yeah] And it kind of does play into a centrality
argument, yes. But on the other hand, it also says you can do a
better job, but you, you need to be able to reach deep inside the
applications and the systems that are making this happen and
customizing it for your needs. [George: Yeah] and that's sort of
where we stand in this, that you know, it's a game for a small
number of high-tuned approaches to this problem, delivering
experiences I think that are amazing, but limited to a few large
players to actually optimize that.
George Michaelson 46:16
Yeah, so we've kind of got where we needed to be, name-based
service delivery efficient and local, but along the way we also
bumped into some realities in the modern world. Nothing is free,
and this behavior has emerged in a way that's tended to
concentrate market power a bit. On the other hand, as consumers,
we're making decisions all the time in this environment, Geoff.
Maybe this is something that we have to be prepared to do, eyes
wide open.
Geoff Huston 46:46
As consumers, we make choices. Yes, and this is all driven by
those choices. I want better quality video, I want quality that's
better than what I can even get from local broadcast, and so on
and so forth. And that's what drives this, and I suppose the
observation these days is, if you look at the proportion of
traffic volume, and even the proportion of consumer money, we are
a video streaming network streaming down the last mile, that's
what the money says, everything else is a marginal, a marginal
behavior with a few percentage points of revenue on the side,
George Michaelson 47:23
yeah,
Geoff Huston 47:23
and of course the big volumes drive the industry as a whole. This
is what we invest in, so very capable, very fast last mile
networks, g 6g fiber, you know, you name it, that's what customers
want, and a densely driven local data center network where the
content is replicated densely at a level of cities, not ISPs, not
countries, at a level of cities of population dense points. That's
what we're building at the moment. And you go, well, how does this
work out in the future? I don't know. AI data centers are so big
and expensive, we're back to a couple per country, if you're
lucky, you know, the power systems, and so on.
George Michaelson 48:04
Yeah,
Geoff Huston 48:05
I suppose the real answer is we really don't know what we're
building.
George Michaelson 48:08
No,
Geoff Huston 48:08
but we're spending a lot of money building it anyway.
George Michaelson 48:11
Yeah,
Geoff Huston 48:11
and that, I think, is another topic. The AI, you know, is this a
boom and about to go bust, or is there something behind it? But we
can talk about that.
George Michaelson 48:21
That's a story for another day, Geoff.
Geoff Huston 48:23
Another day, George.
George Michaelson 48:24
But this has been absolutely fascinating. Thanks, Geoff.
Geoff Huston 48:28
Thanks, George. Thanks.
George Michaelson 48:31
If you've got a story or research to share here on Ping, why not
get in contact by email to ping@apnic.net or via the APNIC social
media channels. Also, remember the measurement@apnic.net mailing
list on Orbit is there to discuss and share relevant collaborative
opportunities, grants and funding opportunities, jobs, and
graduate placings, or to seek feedback from the community on your
own measurement projects, be sure to check out the APNIC website
for all your resource and community needs. Until next time.
Well, we kind of wanted to do the same thing in networks, because
wouldn't it be good if we could take the service, the content, and
pre-distribute it closer to consumers who wanted it, and then when
we say it's here, it's everywhere, and do we need to even wait and
pre-provision? Can we use the secret of things like the DNS
itself, where the first answer gets copied by the DNS
infrastructure in computers close to who's asking, close to but
not the same, so all your other customers in that network who ask
for the same domain name shortly after the first simply get a copy
of what the first answer was. What's the IP address of google.com
Oh, that's boring. I've already done that 20 times today. Have a
cache copy.
George Michaelson 1:01
You're listening to Ping, a podcast by APNIC, discussing all
things related to measuring the Internet. I'm your host, George
Michaelson. This time I'm talking again to Geoff Huston from APNIC
Labs, in his regular monthly spot on Ping. In several recent
episodes, Geoff has asserted the routing function in the Internet
has moved from IP level packet forwarding models into a process
driven in the name to address lookup function. Who you are and
where you might want to go is no longer just about your IP address
and what you might think is the other end across a network, it's
now about where you are and where intermediary services can best
service your request from, but how does this really work in
practice? What kinds of processes determine where you should be
served from, and who decides? IP level routing hasn't gone away.
There's still a very important routing method called Anycast for
selecting the closest server to be used, and this is deployed at
scale, but it's increasingly being augmented by name-based DNS-
based methods to determine where to go. Geoff, welcome back to
Ping.
Geoff Huston 2:17
Hi, George, how are you today?
George Michaelson 2:19
I'm good, Geoff, but I'm confused because you have spent so many
of your recent ping recordings talking about how names are
replacing routing as the steering artifact for the global
Internet, and I keep meaning to ask, how does this work,
Geoff Huston 2:36
right? So let's get into that, because I actually think this topic
is totally fascinating, and I actually think it's revolutionary in
the change that it's making to the concept of computer networks,
you know. Computer networks were, I suppose, a lot like telephone
networks. I'm here, you're there, and the job of the telephone
network is to actually carry my voice to you and carry your voice
to me. [George: Yeah], I can't teleport you, I can't change you, I
can't duplicate you. The network is there to bridge the two of us
together without either of us getting out of our chairs. Great.
Now, when we built computer networks, the first thing we did is
build it in the image of the thing we knew. Didn't have any other
model? [George: Yeah], so we kind of built them as a way of having
conversations between computers. It's over there, mine is over
here. You know, I tap on my keyboard, little packets get scurrying
over to your computer. You tap on your packet, your keyboard,
little packets go scurrying back.
George Michaelson 3:39
It's kind of inherently, me to you, you to me. It's not typically
a form of communication that's inviting the idea of me to 16 other
people or me to two people at once. It's building it point to
point, you to me, me to you,
Geoff Huston 3:55
right. And, oddly enough, you know, there were these other models,
and we have their paradigms to some extent, you know, radio and
television is a form of broadcast where everyone within a certain
radius can listen to a single point source, and the source is
available to anyone within that served geography, and multicast is
slightly more selective, it says I wish to simultaneously get a
bunch of people to listen, and I'd actually like to agree that
they're listeners, but you know, we have computer analogies,
computer networking analogies of those other models as well.
Although fascinatingly, they were really hard, and they never took
off in the mass market. [George: Yeah]
Geoff Huston 4:40
you know, part of the problem with multicast is that it was, to be
perfectly frank, a really fragile technology, and the only folk
who tried to use it were actually folk who wanted to have a
digital version of radio broadcast, and the idea that the audience
would all line up at the same time to listen. Listen to you seemed
a bit far-fetched in today's world. They don't do it like that,
and so it never really caught on.
George Michaelson 5:06
Well, I've just watched the launch of Artemis Two on the start of
its journey, mankind reclaiming outer space one rocket at a time,
and because of the nature of an event that is happening at a set
time, there are still things - sporting functions, opening of
parliament, royal marriages - where people want to line up at the
gate, but if you're talking about the vast majority of digital
content that we used to see as a collective engagement, it's no
longer available at 9pm and only at 9pm it's you watch it when you
want it, so that model of let's get everyone together in time and
say "go", it kind of narrowed down,
Geoff Huston 5:51
it's gone,
George Michaelson 5:52
it's gone.
Geoff Huston 5:52
I didn't see the launch, I'll catch it up later on streaming,
George Michaelson 5:56
ha,
Geoff Huston 5:57
me and billions of others, no doubt, so yes, that paradigm of
simultaneous audience, I think, doesn't exist anymore. I think the
problem was, in so far as the original version of computer
networking was actually small transactions, they weren't large
ones. It was more about replacing the office fax with electronic
mail. It was web pages that were realistically short, sweet, and
simple.
George Michaelson 6:26
Yeah,
Geoff Huston 6:26
images rather than movies.
George Michaelson 6:28
If we can make the event of getting it really quick, you can get
on with reading it at the speed of human eyeballs, and me, the
network, I don't have to sit there, I can deal with someone else's
problems,
Geoff Huston 6:40
right? So the task of the network was simply to carry the
consumer, the user, the client to the server, where whatever it is
they wanted was, and carry the answers back. And by and large,
this served us for a couple of decades, but as our computing
technology got better, our displays got better, and we were
digitizing almost everything. What actually happened, of course,
was that computing was getting more interesting. You could do
bigger things.
George Michaelson 7:13
When I transited from Britain to Australia, this was around the
time that Ken Thompson, one of the two people who built the Unix
operating system we all live and depend on nowadays. He was one of
the privileged people who had a server at home with a 600 megabyte
disk drive that he had dedicated to recordings of music, because
he was exploring psychoacoustic effects under compression of
audio, and I remember having conversations with people saying you
couldn't afford to send music like that down a network, if you had
a modem, it would take a month to get even an hour's worth of
music out of it, but the thing is, he was researching something
that, as you've just said, has come into the fore as computers
have got faster and as networks have got faster, we can now do
things that we simply couldn't imagine when this form of
networking was being invented.
Geoff Huston 8:07
So, yes, computers have got faster and cheaper, storage has got
cheaper, 600 megabytes. I'm sorry, a significant problem today is
600 petabytes, a major problem is 60 terabytes, but under that,
get into the store and buy a bit more, either spinning rust or or
your solid state storage, and you're done. So, in some ways, we're
living in a world of abundance, but one thing hasn't changed, and
oddly enough, that's the speed of light over distance, and when
you think about computer protocol performance, distance is a pain,
because the way we actually make networks work is a feedback loop.
Oh, go a bit faster, there's room in the network, says your
feedback, and it's this delay between desire and reality. If you
stretch it out, it's like living in a world where everything
happens a second later, you have to slow yourself down, because
you do something, you have to wait for a response. And the issue
was that, although we had heaps of bandwidth, heaps of computing,
heaps of storage, we still had this problem that distance was
killing us,
George Michaelson 9:19
right, because the effect of the speed of light means that the
feedback loop you to me to either tell me speed up or slow down,
that message necessarily has the delay of the speed of light
between you and me. The further apart we are, the more delay there
is between you saying stop and me actually stopping, which means
my responsiveness to changes in the network, whether it's full or
empty, congested or noisy, is necessarily bound in that delay.
Geoff Huston 9:50
So it's really hard, oddly enough, even though we have high-speed
networks, very high-speed networks, it's really hard for a user, a
client. To actually exercise that, so if I want to stream a video
from I don't know, I live in Australia, so stream a video from the
United Kingdom, from France or Germany, from the other side of the
world, even though every individual component might be megabits of
capacity and there might be available megabits of capacity, I
wouldn't get it, and this was kind of the issue that was
confronting originally folk like Microsoft and Apple. Their
problem was that they released operating system updates every
couple of months and said, "Right, there's a new update for
Windows available, come to Seattle with your network and pull it
down from our servers. Now, nothing wrong with that, other than a
few 100 million people wanting to do so simultaneously.
George Michaelson 10:49
Yeah. Oh my gosh, the Windows releases midnight in America. I'm
going to wait till one minute past midnight and fetch it.
Geoff Huston 10:57
You and 100 million of your best mates, and so things melted and
didn't take us long to figure out. Well, if computing and storage
is so cheap, let's replicate the content in advance. It's like in
the old days, a book publisher would send in advance copies of
Harry Potter and whatever it was to bookstores, and then announce
the opening day, and the stores had already been pre-provisioned
with the book.
George Michaelson 11:23
Yes, there were a few lawsuits where the stores were selling pre-
release copies of the book under the counter to journalists, and
people were breaching the magic wall of
Geoff Huston 11:31
The Secret Covenant, or whatever it was. Well, we kind of wanted
to do the same thing in networks, because wouldn't it be good if
we could take the service, the content, and pre-distribute it
closer to consumers who wanted it, and then when we say it's here,
it's everywhere, and do we need to even wait and pre-provision?
Can we use the secret of things like the DNS itself, where the
first answer gets copied by the DNS infrastructure in computers
close to who's asking close to, but not the same, so all your
other customers in that network who ask for the same domain name
shortly after the first simply get a copy of what the first answer
was, what's the IP address of google.com Oh, that's boring. I've
already done that 20 times today. Have a cache copy.
George Michaelson 12:28
Cache, the magic word emerges. Caching.
Geoff Huston 12:32
Well, I'm actually thinking it's distribution. Caching is what we
call it in the DNS, but in the server world we call it content
distribution. That's what you're trying to do is to take a big
thing, be it a video or any other kind of large volume server or
server's content, and as you stream it outwards, you keep a copy
close to, but not right inside the end user's environment in their
ISP, for example, so now when anyone else asks for the same video,
and I bet you they will, I'll be able to serve it remotely and not
touch the network. Brilliant, all of a sudden I've got rid of the
network, it's gone.
George Michaelson 13:14
Well, you needed a pipe, you could push the stuff down to get it
to all those places you're choosing to put it, so "a" network has
to exist for you to push stuff to the places you want to cache
copies of it, but "a" network isn't the same as "the" network, is
it?
Geoff Huston 13:36
Well, let me make your brain explode, then, because let's Geoff's
content distribution system rig up a bunch of virtual servers on
borrowed hardware or leased hardware in a few 1000 places all
around the world, close to populations, and they will serve
Geoff's content locally over the Internet protocol. Fine, how do I
feed those front ends? Well, I have the Geoff content factory
somewhere else, doesn't really matter where, and I want to feed it
in an advance. Oh, that means I could trickle feed it. I don't
need to be fast, it's not on demand, I'm pre-provisioning. What
protocol should I use to do that pre-provisioning? Apple Talk
doesn't matter,
George Michaelson 14:28
private,
Geoff Huston 14:28
you don't see it,
George Michaelson 14:29
it's private,
Geoff Huston 14:30
it's all private, it's private, and whatever protocol I choose to
use is kind of my business between my master server and my front-
end, you know, retail points of presence, my content points of
presence. So, in some ways, all I'm trying to do is deliver this
same content all over the world through the last hop, but how I do
so internally is kind of my problem, and no one else's, and so
kind of doesn't matter, and so. Google internally uses BBR or IBM
SNA. No one cares. It's not our problem, it's their problem.
George Michaelson 15:09
Yeah, from talking with the engineers who manage the FreeBSD
distribution framework, I think they are existing in the margins
where they have to try and reduce cost here, and they've
constructed almost, I believe, a two-tier model. They have
regional nodes that they push to, which means the center only has
to arrange for maybe five copies to be pushed out, and then local
nodes do a pre-fetching, a pre-provisioning fetch from a regional
node. So they've kind of invented two levels of behavior, and
they've got less complexity, but in the end, nobody in the real
world has to care, because all they see is fetch-free BSD from a
local mirror.
Geoff Huston 15:52
You could be talking about Akamai, or Fastly, or Google, or any of
the others. They all use the same techniques. It's really common,
and it works. That's a good thing. It works. So, okay,
distribution is kind of a solved problem, but let's go to the
other side of this. You want to go to Geoff's favorite videos from
Geoff's Content Distribution Network, and I have 1000 points of
presence now. If you make a dud choice and pick a server that is
not near you, but on the other side of the world. No one's better
off, are they? [George: No] you're going to have a bad experience,
because that was a bad choice.
George Michaelson 16:32
I feel like this has been the longest preamble in the history of
podcasting, Geoff, because we've just arrived at the entry point.
How, how does it work?
Geoff Huston 16:43
Well, I was wanting to kind of motivate it, and
George Michaelson 16:45
you got there
Geoff Huston 16:46
everyone to understand the nature of the problem and kind of why
it's not straightforward, and there are a number of techniques, if
you will, to actually answer this, and the first approach, which
oddly enough was very controversial, was actually to put the same
computer platform everywhere that you're serving, so this computer
platform uses address 192.168 dot 0.1 and I put another one in
Paris, 192.168 dot 0.1 and I put another one in New York, one in
Sydney, one in Singapore, one in Beijing. I use the same address
everywhere, and I go to the routing system and go, Hi, I'm over
here, and here, and here, and here.
George Michaelson 17:33
Now you've been at some pains when we talk BGP and the behavior of
BGP selection of a prefix to route to, you've talked about how, if
you announce a big address and then announce some more specific
wins, that becomes the thing you look at, and separately, you've
talked about the loop detection mechanism becoming a cost measure
that lets you determine which is the lowest cost path to take, and
I'm looking at this situation where you're announcing exactly the
same address in BGP in lots of places, and I'm asking myself, what
mechanism makes me pick one near me and you pick one near you, and
some Parisian pick one in Paris.
Geoff Huston 18:18
Well, you said it yourself, BGP picks the route with the smallest
metric. Now, in BGP, the metric is the number of networks I need
to traverse to get to where it's being announced. So, if this
conversation is you and I sitting in, say, America, and the server
is being announced in, say, Paris. It is likely there might be a
local network, a transatlantic transit network, and another local
network. It's likely the cost to get to the Parisian server is
three, [George: right] But I've also got my server in America. It
may well be that it's very, very close to you, an adjacent network
cost one
George Michaelson 19:05
right,
Geoff Huston 19:06
BGP picks the lowest cost, and so if you go to 192.168 dot 0.1 and
you're in America, you will go to the server that the routing
system BGP thinks is closest to you,
George Michaelson 19:20
so in a sense it's almost like a decision-free outcome. I don't
have to pick where to go. Rooting optimizes the decision facing me
and the guy next door, and simultaneously some Parisian guy
sitting in a flat. We all get optimized in rooting to the closest
point,
Geoff Huston 19:40
right! And you go, geez, that sounds weird. Let me point out a few
major uses of this anycast approach, where you put the same
address everywhere, and I point to the root of the DNS. The root
of the DNS is served by 13 unique V4 addresses. And 13 unique V6
addresses,
George Michaelson 20:02
right.
Geoff Huston 20:03
Each notional server has a V address and a V6 address. So, there
are 13 root servers. There are 13 machines that serve the root.
No, there are 1500 of them. Well, how does that work? Each root
server actually has a large number of servers, servers all over
the world, all over the Internet, all listening on the same IP
address, the same, so there are 13 separate little what we call
anycast clouds in v4 another 13 in V6 and the routing system
automatically takes you to the closest server. No one does
anything. BGP looks after all of this, and it's actually not a bad
approach. You've kind of outsourced your problem to BGP, knock
yourself out, kiddies, you know. So not only does it keep the root
of the DNS evenly loaded across the world, and stops the root
servers from melting through overuse or any individual server. The
technique also works for efficiency. You just get taken to the
closest one. Who uses that technique? Cloudflare. So, Cloudflare
have a very small number of IP addresses, but if you look hard at
their network, you actually find the same IP address pops up
everywhere, because they're anycasting and therefore direct you to
the closest content on Cloudflare's network is certainly left to
BGP, easy.
George Michaelson 21:36
So nothing comes for free, Geoff, and I'm going to imagine there
are potential pitfalls in this, I mean, I'm sitting here thinking,
if one of these boxes breaks and I need to log into it, I can't
use this magic anycast address to log in. There's got to be at
least another address on it that is unique to each, each box.
Geoff Huston 21:53
Oh, you need a, you need a service address that's unique, but let
me tell you what the problem is. The network, in topology sense,
is not long and stringy, is short and fat. The average AS length
on the Internet is a little over four networks, four, so the
average AS metric is four, which is really coarse, and if you
attach to a big spanning network, so if Geoff's ISP spans all of
America and Europe with one AS [George: ah right] Then, in terms
of BGP metrics, every location in Europe and America is the same
distance. How do you know what's closest,
George Michaelson 22:36
right?
Geoff Huston 22:36
And the answer is BGP doesn't, and so that will give you poor
outcomes if you are dealing with networks that span large
distances, and guess what, there are networks that span large
distances, and at that way anycast kind of makes, unfortunately,
suboptimal decisions, and we see this.
George Michaelson 22:59
right, you could wind up being in America with a resource in
France looking equal cost to a resource quite close to you
logistically on the continent, and you could incur that time delay
of a sub oceanic fiber optic link, because BGP can't tell you you
should look somewhere else,
Geoff Huston 23:18
right? And the right research folk, one of their early experiments
with Atlas and mapping was to actually ask their Atlas nodes,
which instance of A.root-servers.net do you go to when you ask the
root of the DNS, which instance of B, and so they were trying to
actually isolate how well does this root server any cast system
actually work, and obviously when you go looking for anomalies,
you find the world is full of them, and of course, there were
Atlas nodes in Europe who thought via BGP that the closest anycast
instance of one of these servers was in places over the other side
of the Atlantic in America, and so on, and it wasn't rare. It's
common, because BGP is pretty coarse, you know. Short fat networks
don't really do a good job with anycast. Okay,
George Michaelson 24:12
and if I front up on a big network and say, hey, could you make
your network a bit more granular? Could you grind it up a bit to
divide it into the chunks that would make my life easier. The
inevitable point is, yeah, I could do that,
Geoff Huston 24:25
but you've got to pay me. You've got to pay me large amounts of
money, and the real answer is, I don't think you're going to,
because I'm not hosting your content, and I'm not hosting your
customers, I'm just the transit dude in the middle, and so the
economics don't really make this any better, and trying to make a
longer, thinner network out of a short, fat network is kind of
anti-gravity. No one wants it, so it's not going to get any
better. Oddly enough, it's going to get worse as we move on. So,
in some ways, anycast has its limits, and if you really want to
start shaving things down. If you really think there should be a
difference between Frankfurt and Paris, if there should be a
difference between Paris and Lyon, there should be a difference
between Paris and what's a town very, very close to Paris, Paris,
Orleans, a few kilometers away. If you really want to make that
difference, or a distinction between Tokyo and Asaka BGP isn't
going to do it, it's not
George Michaelson 25:23
right. So, BGP was almost the decision-free mechanism where you,
the content provider, and me, the client, have nothing we had to
do to find what was considered the air quotes best choice, and
what we're kind of getting to is, well, you actually are going to
have to make a decision now, and there's that magic moment. Who's
making the decision? How do they do it?
Geoff Huston 25:46
So, if I can't use anycast and get a free ride with routing, let's
try a different approach. So, again, I have my content
distribution network, and I give each of these points of presence
a different IP address, 192 168 1,2,3,4,5 so let's have 1000 of
them. You asked for Geoff's content.com Now I've got to go. Oh, I
could give you in the DNS 1000 different answers. Which one is the
answer that's good for you? And you put your little hand up,
going, well, I didn't really want all 1000 answers, I can't cope
with that, not my problem, but don't give me the one that's
furthest away, please, just don't do that, Anchorage, Alaska is
only good if you're living in Anchorage, Alaska,
George Michaelson 26:36
So I trigger a decision having to be made because I say I want to
get somewhere, and you, because I'm getting somewhere by asking
about, tell me the address of a name, your systems, or someone on
your behalf is getting that call that's saying, I, George, I want
to do this, and the problem is, How can I know based on you asking
where the best place for you is? What do I know to help decide.
Geoff Huston 27:03
So, let's make a sweeping assumption, and it's probably not that
bad a sweeping assumption. You and your DNS agent, your recursive
resolver, the thing that actually asks your questions for you is
normally provided by your ISP. Let's assume that that's close to
you, so now I have a third party, your resolver, and I have it
asking me, the authoritative server for my content distribution
network, a question about content, content that I'm serving, so
now I've got a simple problem, or simpler, an IP address is asking
me about a server. I have 1000 possible IP address answers for
that name. I just need to pick the one that's closest. Okay, we
have geolocation databases in its crudest sense. I can map to
countries. Let's put a country code on your resolver that asks the
question. Fair enough. You live in Japan. Let's shuffle through my
list, find an answer where the server is located in Japan. That's
the answer. So now you're getting steered to something that I
think is good.
George Michaelson 28:13
It definitely sounds like an improvement against large flat
networks that potentially span interoceanic networks and different
economies, but I'm not feeling like you've kind of got perfection
yet, Geoff, because some economies are very long and thin. Chile
is a massive, long, thin stripe down the side of South America,
and if you happen to be right down the bottom at Ushuaia, you are
not close to Santiago, right up in the north, Japan is a very
long, thin island. It's quite a lot bigger than people think,
because of the Mercator projection. And if you're right up the
north, you are not close to right down the south. So you fix the
problem by keeping it within Japan. I'm giving you a thumbs up. It
still hasn't got close to me.
Geoff Huston 29:00
Well, before you start throwing rocks, let me observe that this is
the model used by Akamai, one of the major content players on the
planet. They do this by DNS. It's not even a clever trick, it's an
obvious thing. When you serve content through Akamai, when you say
take this content and serve Akamai, they say to you, look, instead
of serving or giving an IP address of your server in the DNS
CName, it alias it to us, and so they say, if I want to serve
example.com they might say, in your DNS master zone, use the CName
for this service, example.www.example.com.www.example.com and C
name it to dub dub dub.example.com dot edgekey.net Who owns
edgekey.net Akamai. Oh, now I've got dub dub dub.example.com dot
edgekey.net Where does that. That map to another CName. Now the
trick is that that second mapping depends on who's asking, because
I have maybe have 4000 possible answers that identify an edge
location by name,
George Michaelson 30:17
right?
Geoff Huston 30:17
And I do a magic thing that says www.example.com dot edgekey.net
CNames to the server in Tokyo dot akamai edge.net What's its IP
address? Well, that's constant, that's easy. The server in Tokyo
dot akamyedge.net you know, is 23.1 95 dot 84.2 41 or whatever,
George Michaelson 30:39
right?
Geoff Huston 30:39
And so now with just a little bit of mucking around with Cnames,
I've got a decent way of doing this, as long as I can triangulate
and say, "A, you live close to your DNS resolver, so if I optimize
my answer against who's asking, I can steer you to one of the
servers that serve the content you're after that's closest to you,
George Michaelson 31:04
so having kind of kibitz and said yeah, don't like this, this
don't think it's perfect, I want to back off and observe, I'm sure
I'm using Akamai services all the time, proof by example, I'm not
seeing massive amounts of "this kind of sucks", so as a consumer
of services, I think we might be in 80-20 land. Geoff, this works
mostly if I'm not doing stupid tricks with my DNS, and if I'm
mostly using a DNS server close logistically to where I live and
work. It basically seems to work.
Geoff Huston 31:37
Many folk might bank with the Bank of Akamai, they might, it
works, and the beauty of this is also that you can't see those
translation names when your browser then connects to this IP
address of an Akamai server close to you. You say, Hi server, I
want to connect to www.my bank, and I want to create a TLS
connection, and the Akamai server goes got a problem. I am
configured to be an alias for that banking name. Here is a key,
here is a key exchange, here is a TLS connection. Let's do this,
so the entire thing is seamless to the user, and I've now managed
to by doing a simple sort of trick with the DNS, steer you to
something is really close to you. Presto, it's fast, it's cheap.
Did I mention it's fast? It's really fast.
George Michaelson 32:31
Yeah,
Geoff Huston 32:31
because it's local, so kind of works. Yes. So now, George, let's
start throwing rocks.
George Michaelson 32:38
Yeah, so I want to watch content from the BBC, and the BBC says,
'Sorry, chum, your DNS is in Australia. So I ring up a mate with
the DNS server in the UK, and I let him let me ask questions
through him, and the BBC go, 'Not a problem, buddy. I see your
questions coming from the UK. Here is Doctor Who'. It feels like I
can drive through geo-fenced IPR rules if I can make my questions
come from the UK
Geoff Huston 33:10
more than 10 years ago with Netflix. If you changed by simply
going into the config part of your computer and changed who your
DNS resolver was to your country of choice, Netflix would open up
the library of that country because it assumed you were your DNS.
George Michaelson 33:27
Yeah,
Geoff Huston 33:28
and yes, it's it is, was, and probably still is abusable from that
respect.
George Michaelson 33:34
Yeah,
Geoff Huston 33:34
and the other problem that's just as bad are these open DNS
resolvers,
George Michaelson 33:40
yeah,
Geoff Huston 33:40
like 1.1.1.1 and 8.8.8.8 Well, the thing is, which actual DNS
engine did your query go to? You don't know, but when Akamai get
the query coming from one of these open DNS resolvers, say Google,
it's not clear that the address being used by that open DNS
resolver, when it asks Akamai is close to you. Oh, you're coming
from Paris. Oh, you're coming from Moscow. Oh, you're coming from
Madrid. You know, it's really, really hard to nail it down.
[George: Yeah], and so under some pressure, Akamai and a few other
content folk convinced the big content distribution networks to
add a security leak, a nightmare. When you ask the authoritative
server, why don't you attach the client address? Oh, can't do
that. Don't want to do that. Why don't you attach the client's
subnet? Oh, that's okay. That's not very privacy leaking rubbish.
We'll use that, so we actually intrude into the DNS between the
recursive and the authoritative an item of data that actually
points to the original user,
George Michaelson 34:51
right? So the person asking the question is the primary vehicle
that can at least target in a lot of cases down to an economy.
Economy, but we've changed the protocol to intrude a flag going
out asking the question, this geezer in this suburb in Ginza is
asking about this TV show, and the belief I have is, oh, well, you
just identified me to Ginza, there's a million neighbors, and the
belief you have is nobody's lying here. He really wants the server
in Tokyo, not the server in Hiroshima, because he said Ginza, so
it kind of sounds better, Geoff.
Geoff Huston 35:30
It kind of sounds better, because in essence, you're giving a few
more clues of information. Okay, how good is a subnet? It's a
privacy protection. Meh, not very good, you're leaking information
about you that maybe you didn't want. The plus side of this, who
controls the mapping from name to number? You see, in BGP Anycast,
the service provider had absolutely no control,
George Michaelson 35:57
right?
Geoff Huston 35:57
BGP, but here I am, my DNS servers control that mapping, I can
twiddle and play if the server in Asaka is flat on its back,
howling at the moon, and the only available capacity is somewhere
else, Tokyo, whatever, I can actually intrude into this, and with
a short life answer, I can tell you somewhere else,
George Michaelson 36:21
short life, that would be because these kinds of answers wind up
having a cache behavior, and if I've cached in your system, you
have to go over here for the lifetime of the timer, I'm going over
there.
Geoff Huston 36:34
The thing that makes the DNS work so well, caching kind of is your
enemy if you're trying to do some degree of dynamic load balancing
or optimization, because once you give an answer in the DNS to
that resolver, everyone else who queries for that name is going to
get the same cached answer,
George Michaelson 36:53
right,
Geoff Huston 36:53
and if you want to recompute, then the only thing is give it a
short lifetime in the cache, because otherwise the first answer is
sticky for an unknown number of subsequent people, and so to do
this, you lose cache.
George Michaelson 37:07
So, these are two different kinds of rocks being thrown here. One
is it's a bit of an anti-pattern for efficiency in the system.
Okay, we have to deal with that, and the other is it's a big
privacy concern. You don't really get to keep your privacy when
you ask these questions anymore.
Geoff Huston 37:24
Yeah, tough. So, what do you do? Well, we've just discussed, I
think, almost the transition of networking and technology. We
originally started with a smart network and dumb devices, it was
called telephony, and we started to put more and more smarts in
the edge devices and strip out the middle, so using BGP kind of
relies on magic, and BGP devices need to do very, very little.
Using the DNS lifts it up one level, so now you've got the DNS
doing this distance mapping and trying to get you to the optimal
point. Why stop there? You could do what Netflix do. There is an
Amazon server, or 20, or 1000 where all the initial transactions
for Netflix content get passed to on the West Coast of America.
Oops, but that's okay, because then there's an application
problem, because you divide your content into manageable chunks,
little bits, and you have the ability to go well, user, you're
making contact with me, and you want this content. Now, it's not a
DNS question, it's not even an IP routing question. Now it is a
content streaming question in the application level,
George Michaelson 38:43
and it's between me, the user who is paying you Netflix for this
service. I'm giving you money, and so this anonymity and privacy
thing. Well, I was never going to have any privacy in you, because
I was always going to pay you to give me the things that I tell
you I want to get,
Geoff Huston 39:05
and you want me to know it's you, because you're paid me, and you
want that service, so you know
George Michaelson 39:10
I do.
Geoff Huston 39:11
This is now consensual disclosure.
George Michaelson 39:14
So, there's a third thing kind of buried underneath here, which
isn't really material for the technique we're talking about, but I
think it's part of the story. You also, when you send me things in
Netflix, you send it fast enough that I can build up quite a big
buffer in time. I get content from you that spans probably two or
three minutes in me, so that even if you have to stutter a bit
sending stuff to me, I don't see that watching this movie, and so
that chunking technique, where you're saying, "Oh, that server's
dying, I need to move you somewhere else". In your clients, you
know, most of them have a bit of a buffer here, where if they have
to flick open a link somewhere else and start dealing with it, it
isn't actually going to affect their streaming.
Geoff Huston 39:59
Everything is a compromise, isn't it? You can do, and for a long
time we did do high watermark, low watermark, trying to get the
playback buffer in the device that you know the customer is using,
the television, if you will, trying to keep that buffer full, and
when it got to a low mortar mark, you just slam data down the
line, and then went quiescent until the playback got to the low
watermark threshold, we call that the Netflix spike, and it's a
bit like hitting the network with a sledgehammer once every few
minutes, and it's just as painful. Modern thinking says don't
stress the network like that, bad things happen, you just drive
the network into bad places, try and even your load, [George:
right] Try not to do stress techniques, but do the gentle feed.
Yes, it makes it a little bit more difficult to change where your
stuff is being served from when you need to, but at the same time
the network is going to behave a whole lot better. The good news
about this approach, because now it's at the application level,
and I have control over who is delivering you elements of the
service. Is I can start playing with what if questions.
George Michaelson 41:10
Yeah,
Geoff Huston 41:10
I've been serving you from Asaka for quite some time now. I wonder
if Tokyo could deliver you better service. We're up to chunk 1023
Let's deliver 1024 from Tokyo, and tell me, because it's my
software at the other end. Tell me, how that was for you.
George Michaelson 41:27
Yeah,
Geoff Huston 41:27
and I can then dither and try and get really fine-grained accuracy
on what is the best service at this point in time,
George Michaelson 41:38
and so this mechanistic behavior, it might be using
intermediaries, it might be using locally best service anycast in
all kinds of weird ways, but we have moved back into application
layer decisions between you, the sender, and me, the recipient.
It's an instrumented dialog, but you get to understand it without
intermediaries having to make inferences. There's ways that you
can control the behavior, and I can control the behavior in
software.
Geoff Huston 42:06
So, as we've moved content up into video, as the content stuff
gets bigger, and we're now talking about, you know, hundreds of
megabytes or gigabytes, then the whole idea of make a decision
once on what's best and stick with it, which is Akamai's decision
doesn't really work very well, and so you have to constantly use
some kind of adaptive feedback to check if that original decision
could be bettered if there's another way of doing it to improve
the response, which is where Netflix and YouTube and a number of
other streaming platforms have gone to, because it becomes an
application issue that can best actually address this problem.
George Michaelson 42:43
So, let me just run history on the quick track. We had its point
to point, and you could even have operators plugging between patch
boards to connect you to me. We had its packets, and we're going
to give them addresses, and suddenly it's about knowing how to get
there, which means routing. We had routing can be made to look
like one thing is anywhere, and you have to use a routing trick to
pick which is the best one. [Geoff: Yes], we then say that isn't
good enough. Can we go back to something that's about names? Who
am I, and what do I want to get? And we start using mechanisms in
names, and you've now moved up to well, it is names, but it's also
got a lot of what are you actually doing with me? Let's measure
this, let's optimize this kind of sounds like it's name-based,
it's a bit of the routing story still in it, but it's also
smarter.
Geoff Huston 43:37
Well, it's kind of pushing the problem further and further up into
where there I think there is intelligence and information to make
it so the DNS problem is I have control of that translation, I can
select a server for you, but then we're both stuck in that
selection, we can't change it easily. When I start to dither at
the application level, I have a huge amount of flexibility, and I
can actually figure out was my original decision good or bad, and
if it's bad, are there alternatives that can offer a better
service? Don't forget, too, that there's a lot of remedies. I can
reduce the quality of the signal to reduce the volume, or
preferably I can stay at the same playback quality, but move to a
server that has a better chance of keeping that quality up for
your experience, and so that flexibility sort of plays into hands
of the video streamers going, I'm really interested in user
experience and the quality of that, I think I could do a better
job. What does this mean, I suppose, for interoperability and
standards and all that kind of stuff. The answer is no, not really
anymore.
George Michaelson 44:45
No, this is moved out of the domain of let's do it all the same
way into I've got special secret sauce.
Geoff Huston 44:53
There's Geoff's secret sauce, there's George secret sauce, and a
playback from Geoff's server might not necessarily. Necessarily
work with George's receiver, it's not that kind of
interoperability anymore.
George Michaelson 45:06
Yeah,
Geoff Huston 45:06
each of these application level content factories and distribution
systems work in their own ecosystem.
George Michaelson 45:14
Yeah, we're building vertical markets and vertically walled
gardens. You want to do this behavior, it's going to be good in
me, and my special sauce makes it better in me than from the other
chaps. Stay with me,
Geoff Huston 45:27
right? And it kind of goes well, but I'm a minnow in this world,
I'm just a starter. Oh God, I'm sorry. You're back to banging the
rocks together, dude. You know, you're doing a web-based playback
buffer because you're not big enough to build your own ecosystem,
[George: yeah] And it kind of does play into a centrality
argument, yes. But on the other hand, it also says you can do a
better job, but you, you need to be able to reach deep inside the
applications and the systems that are making this happen and
customizing it for your needs. [George: Yeah] and that's sort of
where we stand in this, that you know, it's a game for a small
number of high-tuned approaches to this problem, delivering
experiences I think that are amazing, but limited to a few large
players to actually optimize that.
George Michaelson 46:16
Yeah, so we've kind of got where we needed to be, name-based
service delivery efficient and local, but along the way we also
bumped into some realities in the modern world. Nothing is free,
and this behavior has emerged in a way that's tended to
concentrate market power a bit. On the other hand, as consumers,
we're making decisions all the time in this environment, Geoff.
Maybe this is something that we have to be prepared to do, eyes
wide open.
Geoff Huston 46:46
As consumers, we make choices. Yes, and this is all driven by
those choices. I want better quality video, I want quality that's
better than what I can even get from local broadcast, and so on
and so forth. And that's what drives this, and I suppose the
observation these days is, if you look at the proportion of
traffic volume, and even the proportion of consumer money, we are
a video streaming network streaming down the last mile, that's
what the money says, everything else is a marginal, a marginal
behavior with a few percentage points of revenue on the side,
George Michaelson 47:23
yeah,
Geoff Huston 47:23
and of course the big volumes drive the industry as a whole. This
is what we invest in, so very capable, very fast last mile
networks, g 6g fiber, you know, you name it, that's what customers
want, and a densely driven local data center network where the
content is replicated densely at a level of cities, not ISPs, not
countries, at a level of cities of population dense points. That's
what we're building at the moment. And you go, well, how does this
work out in the future? I don't know. AI data centers are so big
and expensive, we're back to a couple per country, if you're
lucky, you know, the power systems, and so on.
George Michaelson 48:04
Yeah,
Geoff Huston 48:05
I suppose the real answer is we really don't know what we're
building.
George Michaelson 48:08
No,
Geoff Huston 48:08
but we're spending a lot of money building it anyway.
George Michaelson 48:11
Yeah,
Geoff Huston 48:11
and that, I think, is another topic. The AI, you know, is this a
boom and about to go bust, or is there something behind it? But we
can talk about that.
George Michaelson 48:21
That's a story for another day, Geoff.
Geoff Huston 48:23
Another day, George.
George Michaelson 48:24
But this has been absolutely fascinating. Thanks, Geoff.
Geoff Huston 48:28
Thanks, George. Thanks.
George Michaelson 48:31
If you've got a story or research to share here on Ping, why not
get in contact by email to ping@apnic.net or via the APNIC social
media channels. Also, remember the measurement@apnic.net mailing
list on Orbit is there to discuss and share relevant collaborative
opportunities, grants and funding opportunities, jobs, and
graduate placings, or to seek feedback from the community on your
own measurement projects, be sure to check out the APNIC website
for all your resource and community needs. Until next time.