Geoff Huston 0:00
A lot of routing protocols at the time that didn't use TCP had
this short minded or amnesiac counter every 90 seconds, every
period, it said you could be losing your marbles. Let me refresh
your memory. Here's everything I know, and those broadcast storms
of just refreshing stuff you already were told was a great fail
safe in an unreliable transport, but ultimately was incredibly
inefficient making this TCP say, I've told you, I'm never going to
tell you again, unless it changes. That was good. That cut down
the amount of routing traffic like crazy. The next thing is
looping. BGP had a very clever way of detecting looping, because
if A talks to B, talks to C, talks to D, talks to A, how do you
know it's a loop? It's kind of hard sometimes, if you think about
it, but what BGP did is it attached, almost like a snail trail, a
line of thread through the maze. Every time a network, its prefix
went through a network, it added the number of that network to its
trail. So let's say it's network A, the path is A. When it gets
sent to B, it goes well, the path is BA. When that set sent to C,
the path is CBA. In other words, the networks that have seen this
advertisement sign it and attach it, that signature, that their
number, their identifier, to that update. So if you have loop it,
you're seeing yourself. And thought, seen this. This is not real.
I've already seen it. I'm going to reject up front path vector was
the big difference between BGP and rip the Routing Information
Protocol, which did not have that and consequently, the only way
it figured out a loop was to actually count to infinity.
George Michaelson 1:57
You're listening to ping. A podcast by APNIC discussing all things
related to measuring the Internet. I'm your host, George
Michaelson, this time, I'm talking to Geoff Huston from APNIC labs
again in his regular monthly spot on ping. for over two decades,
Geoff has been publishing the CIDR report, a daily web update on
the state of BGP worldwide with information on routing table size
and where it comes from, along with information about anomalous
announcements started by Tony Bates at Cisco in the mid 90s, then
carried on by Phil Smith, who was also at Cisco at the time, and
then Geoff. It's a remarkable continuous measurement of the state
of routing. The CIDR report grew out of concerns with the
consequences of worldwide growth in the size of the routing table,
and especially a trend towards de-aggregation, when as holders
deliberately announce more specific prefixes from Their blocks to
optimize routing behaviors in a process called traffic
engineering. But what is CIDR or classless Internet domain
routing? What problem was it designed to solve and in a modern
Internet, Is this the problem we think it is? Has the CIDR report
been overtaken by other events in routing? Geoff, welcome back to
ping. What shall we talk about today?
Geoff Huston 3:23
Hi George. Look, I think it's time we talked about routing. It's
been a continuous Internet subject for the last eon or two.
George Michaelson 3:31
Oh, it's been a while since we did that, at least two episodes.
Geoff Huston 3:35
Oh, this time I actually want to kick into a report that I've been
producing for at least 20 odd years, and has been produced for
more than 30, I think about a sort of a summary of the state of
the routing system. It's called the CIDR report,
George Michaelson 3:50
CIDR
Geoff Huston 3:51
But what I'd like to do first is to Yes, now it's not about
something you drink, truly. I'd like to sort of talk about the
routing problem, as it was, right from the word go, and how the
CIDR report sort of snuck into all this
George Michaelson 4:07
Okey dokey, roll the clock back 20 years. 30 years? Is that
enough?
Geoff Huston 4:12
Oh, keep going. We're talking 1980s and I suppose one of the
things that was totally different about the desire of computer
networks, as distinct from, say, telephone networks and similar
ones of their day, is that computer networks were meant to be self
learning, that if you had one device, it wasn't a network. If you
had two, they were meant to be able to discover each other,
[George: right] And if you had a collection of these things, some
of them connected over local area networks, big thick Ethernet
wires, and some connected over thinner wires. Y'all meant to
discover each other.
George Michaelson 4:54
That was very much driven in the experience of pre existing
structural networks that were built with much more like a god
designer who said, Thou shalt connect this to that, and thou shalt
fan it out in these ways, to these remote entry vehicles and these
concentrators and will gather data from these locations and send
it along these links. It's like somebody laid out a map of a
network designed around devices and a computer. And the 80s people
were saying, we don't want to have to do that.
Geoff Huston 5:23
Well, that's right. And I think in some ways, the 80s saw the
beginning of Ethernet as a local area network. And the thing about
Ethernet, which was, I make it sounds like nothing today, but at
the time, you just plugged your computer into that common wire,
that bus cable, originally with a thick yellow cable you got at
your drill and your very expensive transceiver, drilled into the
coaxial cable on the little black mark, tapped it, connected your
computer, and you're able to talk to everyone else, other than the
need for a drill and an expensive transceiver. There wasn't a
controller of these Ethernet there was no one in charge. You just
connected your network, and while we sort of tried to push them
further out into the wider area, we wanted to take that plug and
play convenience with us. We wanted to use networks that did not
have a controller, that were actually an amalgam of almost self
organized networks that became off a self organizing beast. And
that was the desire of the routing protocols at the time [George:
right] to exchange the role of the network controller, the Big Fat
Controller, if you're a fan of Thomas the Tank Engine, to exchange
that role into one that was realistically one which was brokered
by the protocol itself, [George: right] So that was the magic
desire. And there are a number of attempts at this by various
vendors in various ways. The one that a lot of us had experience
with was actually a vendor network called DECnet.
George Michaelson 6:59
Yeah, used much as you did Geoff. It was a regular part of my
life. The thing
Geoff Huston 7:04
about DECnet was It was organized into two sort of tiers. There
was an area where machines inside an area had a detailed
conversation about their connections. You're on the same Ethernet
as me. I'm connected by a dedicated wire to you, etc. And so
inside an area, all the computers needed to be interconnected one
way or another. Didn't need to be a mesh, but any computer in a
network in an area needed to get to any other computer in the same
area without leaving that area right, just sort of like a local
city or something. And then DECnet had a second upper level
protocol, almost like an inter area protocol, where there were
area borders, and the area borders spoke via dedicated connections
to other area borders. And they said, Hi, I can reach all of area
one. How about you? I can reach all of area two. And so to send a
packet between a computer in area one to one in area two, you
directed your packet to your local area border router, who passed
it over to the right area, who then sent it on to the right
computer in that area.
George Michaelson 8:20
So Geoff the way you're describing this, it's like there's a bit
from Box A and a bit from Box B, because if you're inside the
area, you're making it sound as if it used the technique to
discover all the things within that area, the way the protocol
worked, you'd eventually discover every printer, every other host,
every other computer. But if you want to go to another area, there
was a higher structural layer that was having to do some
management and passing right,
Geoff Huston 8:45
right. Well, things almost as a an atlas, a mapping system. If
you're in a city, it's nice to know every street, but if you're
trying to get a packet to a different city, you really don't care.
You want to get it to that city and let the destination city sort
itself out. You didn't want to start the network and the protocol
interactions with all the detail of remote location, so you
literally did divide and conquer. Divide into areas, and at the
inter area level, you only looked after the connections between
these area border routers, and within an area, you looked after
the detail,
George Michaelson 9:21
you've already introduced the ideas of something local and a
boundary with a border and a border router, and the concept of an
area. And to me, these are four terms that are very interesting to
establish inside this model this early on in people's thinking.
Geoff Huston 9:37
I think DECnet was a model that actually we could have borrowed
and pushed into the Internet, almost subconsciously, I think so
many of us had cut our teeth from that particular vendor network,
and this concept of areas Radia Perlman had a lot to do with it
when she was working in digital was actually really seductively,
right? It sort of pushed detail into areas where you need a
detail. But when you're simply trying to move around a country or
around a region, you didn't need to do that. You just simply
talked about areas. And there were some pretty massive DECnet
networks, the NASA Science Internet, the high energy physics
network. They really did span most of the Northern Henry one way
or another. They weren't fantastically big networks, you know,
kilobits per second, but they did have a lot, a lot of members,
even digital's own corporate network,
George Michaelson 10:29
even this semi automatic mechanism of things being discovered and
things forming areas and communicating with each other, worked at
the scale of the whole of North America, the whole of Britain,
places of that kind of scale could quite comfortably build one of
these things.
Geoff Huston 10:44
So, yeah, digital built a network with 100,000 computers in it as
their corporate network. It was possible to scale it up. And so
when we started to build this Internet beyond the limited Research
Project Agency, the ARPAnet exercise, into the bigger field, that
routing concept we carried with us, and the Internet itself, I
think it really took off when the National Science Foundation
transformed the Internet by supporting a thing called the NSF net
across the United States to connect their super computers. And,
you know, all of a sudden, instead of a few 100 computers, which
was the ARPANET, the NSF net was talking in terms of 5000 10,000
20,000 network, there 20,000 computers inside a few months. Not
only that, but they wanted to interconnect to the NASA network, to
the high energy physics network. They wanted to interconnect to
the sort of quickly dying ARPANET, the old defense network, old
defense sponsored network. So the problem was complex. It wasn't
simple, and they needed to use much the same trick. But what they
did, I think, was quite insightful. It was that area and time when
divide and conquer made sense, that there was no single corporate
and what they very quickly realized was that the networking,
routing protocol you used inside a network an area didn't need to
be the same one as you used outside that area.
George Michaelson 12:14
That is quite a radical idea when you compare it to other
behaviors extant in the Internet in these days, the idea that
things didn't have to be the same, in some ways, ran counter to a
culture that said, we're going to standardize the format of email
addresses across the entire surface of this network. We're going
to make telnet and ftp pretty much use the same kind of logical
structure. And here you are saying divide and conquer allowed
people to stand up and say, let's actually think about doing it
differently inside an area and differently between areas. That's
interesting. An interesting difference
Geoff Huston 12:50
actually, like a difference between a commercial service offering
and the so called, you know, the research people at the time, the
commercial thought, I solve all your problems and so DECnet really
was, I think, two different routing protocols smashed into one
because, you know, DECnet needed to do everything, whereas in the
research world, you know, the academic research world, it really
was. Look, we're not going to tell you what to do. Here are some
things, but if you don't like them, go run your own. And there
were, in the 80s, a lot of university campuses, I know I lived in
one where driving your own and building your own local area
network protocol was a sport, a competitive hobby. I worked at the
ANU, home of ANU net at the University of Monash in Victoria,
Monnet and on and on and on.
George Michaelson 13:38
And UQ net, which is where I was and csironet, which is where I
was before that. And you could build a surprisingly large amount
of interconnected state using incredibly simple protocols, like
RIP, the protocol we use to construct the network at UQ, right,
Geoff Huston 13:53
right. So there was this sort of thing of, roll your own local
network, roll your own network within sort of your network, and
then a case of, how do you hook them together? And that's kind of
well, in some ways, it's quite easy, if you think that a routing
protocol is simply to tell you where addresses are, address
prefixes. So each network has a bundle of IP addresses, in this
case, and hopefully they're contiguous. Well, at the time they
were and they came in three sacks, the Class C sack, and that had
256 addresses, the Class B sack, which had 65,000 and for the
privileged few, there was the class A sack with 17 million
addresses.
George Michaelson 14:39
Yeah, but this kind of anti Goldilocks moment, because the class C
sack that anyone could get easily was really too small, and if you
were big right on the Class B sack was running a bit, and the
class A sack, well, you couldn't get it.
Geoff Huston 14:52
We will get there. But I want to talk about this for a second,
because what the routing protocol needed to do was in the exterior
sense, to say to all the other networks, or at least to the ones
that were immediately adjacent, as networks, not as computers, as
networks. Hey, I can reach these following Class C addresses
bundles and a Class B bundle. Here have the bundles that I can
reach. You've got a packet address. Any of those bundles, send
them to me. I'm your network. You tell me what you can get to
George Michaelson 15:28
I'm your network, which is a magic moment in information sharing,
isn't it, because it takes a thing that could be an enumerated
list. Here's 255 things, and it says, Nah I'm just going to refer
to it as a common bit and how many there are, and that's all I
have to tell you.
Geoff Huston 15:30
And so this idea of bundling up and talking about reachability by
network sacks or prefixes appeared, and the original
implementations of this was the EGP, interestingly, the exterior
gateway protocol, and there were a few sort of things around that,
but it was all relatively simple. But oddly enough, the problem
they were solving was not simple, because already we had nine
regional networks, a National Science backbone network, a NASA
network, a high energy physics network, a network in the United
Kingdom, a network, I think in Germany, very quickly, there are a
bunch of these coming up and getting from one place to another was
a non trivial problem, and you couldn't just send the details of
every connection to every part of the net. There was no bandwidth,
no capability. It would work. So we needed a better form of
exterior network routing. And you know, the answer was pretty
simple, as it turned out, it came out in 1989 when IBM's Jakov
Rechter, IBM was the contractor to NSF, and Yaakov was the
principal scientist over there at IBM, and Cisco's early employee,
Kurt Lockheed, came up with this protocol called the Border
Gateway routing protocol. Now it was simple because it relied on
something that I think was originally in a paper in 1956 the
Bellman Ford, distance vector protocol. I'll tell you everything I
know when you hear what I say, add one to every metric about what
I know. I can see this address prefix in one hop so you go, very
good, Geoff, I see it. Therefore in two hops. I see this one in
five, because people have told me and told me, and you say, very
good. I'll make it six.
George Michaelson 17:29
So the essential quality here is that we've stayed in the space of
we don't want to have to have a big boss that tells us what to
say. We all shout out into the world the stuff about us. But
instead of shouting the entire list of what we got, we now come to
consider, well, I consider myself an area. I shout out efficiently
the prefixes the block of things I have. And you've just said, we
added a twist to that, that when you do it, you use this technique
from the 50s called Bellman Ford, shouting it out is cost of one,
and the guy who hears it shouts it out again, but adds one to make
it a cost of two. And if you were the third in the chain and you
heard it, you'd shout it out, and you'd add a cost to make it
three, right?
Geoff Huston 18:09
right And if you had two connections, people shouting at you, and
one of them says, I see this address prefix with a cost of one,
and the other one says, I too see it, but my cost is eight.
Obviously you'd pick the cost of one, you know, so you naturally
prefer the shorter cost.
George Michaelson 18:26
So we have one unit, one unit of cost and its distance. And if
everyone obeys the rules, you can rationally determine the lowest,
the best cost.
Geoff Huston 18:36
Right? Simple distance. You and I could connect over anything we
wanted, everything was cost one if it moved from one network to
another. Wasn't very sophisticated, but I want my big fat wire to
be more, more preferable than your tiny little wire, and it's kind
of No, no. BGP doesn't do that. The distance between a network is
one stop trying to be clever.
George Michaelson 18:59
There's no dimension for that.
Geoff Huston 19:01
However, there were a few things that were clever. The first thing
they did, actually, I think it was, was one of the big fundamental
things, is that the transport that BGP used was actually TCP,
which means, when I tell you, and you send me an acknowledgement,
you know it, I know you know it because you acknowledge the
packet. I never need to tell you again, only if it changes.
George Michaelson 19:27
So BGP didn't have to construct its own way of saying. Did you
hear that? Are you sure about that? Can I believe that because TCP
told you they got it.
Geoff Huston 19:37
A lot of routing protocols at the time that didn't use TCP had
this short minded or amnesiac counter every 90 seconds, every
period it said you could be losing your marbles. Let me refresh
your memory. Here's everything I know and those broadcast storms
of just refreshing stuff you already were told was a great fail
safe in an unreliable transfer. Bought, but ultimately was
incredibly inefficient making this TCP said, I've told you, I'm
never going to tell you again, unless it changes. That was good.
That cut down the amount of routing traffic like crazy. The next
thing is looping. BGP had a very clever way of detecting looping,
because if A talks to B, talks to C, talks to D, talks to a how do
you know it's a loop. It's kind of hard sometimes, if you think
about it, but what BGP did is it attached, almost like a snail
trail, a line of thread through the maze. Every time a network,
its prefix went through a network, it added the number of that
work to its trail. So let's say it's network A, the path is A.
When it gets sent to B, it goes well, the path is BA. When that
set sent to C, the path is CBA. In other words, the networks that
have seen this advertisement sign it and attach it. That signature
that their number, their identifier, to that update. So if you
ever loop it, you're seeing yourself and thought, seen this. This
is not real. I've already seen it. I'm going to reject up front
path. It was the big difference between BGP and rip the Routing
Information Protocol, which did not have that. And consequently,
the only way it figured out a loop was to actually count to
infinity.
George Michaelson 21:24
That takes a long time.
Geoff Huston 21:26
Well, they made it 16, but it's still not very good. So we've
talked about the separation between inside your network and
between networks. We've talked about TCP, we've talked about path,
vector and transport. But the other thing I think we've just
alluded to, but it is very, very important. Is no one's in
control. There are no permissions. It is a network of peers.
George Michaelson 21:47
Well, we could say there's a kind of moment here that we're not
going to call control of routing. But if you said, put a number, B
puts its number, those numbers do have to be kept distinct. And so
there's a function outside of routing, handing out the numbers to
make sure you only hand them out once, and you don't make two
people think they both claim to be the same unique identity. But
that aside, nothing in routing determines certain parts are built.
Geoff Huston 22:15
We had a registry run by Stanford research that handed out address
prefixes, those bags of addresses, and it handed out what we
called autonomous system numbers, these identifiers for networks.
And there are very few qualifications to those. You sent a fax
because it was the time of faxes over to these folk, and they fax
back saying your autonomous system number is this number, and if
you ask for address prefix, here's the address prefix, Class A, B
or C, and off we all went. And this kind of worked through
relatively well. The early days of the expansion of the Internet,
when the NSF started work, there were a few 100 such networks. And
over the ensuing years, you know, 90 91 92 then NSF heyday, it
really did swing up to 15,000 component networks, which is
amazingly fast.
George Michaelson 23:11
And this mechanism of simple BGP, a path mechanism to detect
loops, assignment of unique numbers, this was essentially working,
and meant that prior state you've described, the several networks
in America, the discrete networks in Britain and Germany, the
emerging network in Australia, they were able to come into a union
of exchanging information with each other. No central controller.
It was just fundamentally working.
Geoff Huston 23:37
You just needed to connect to someone who was connected. It was
this was that you just added yourself to the edge, and that edge
based system was phenomenally successful, brilliantly successful.
By 1994 we were faced, and had been facing, for some years, some
predictions, which were truly awesome, because all of a sudden
this was the consumption problem. We were going to consume all the
available addresses. And let me qualify that we weren't really
there are 4 billion addresses in IPV4, but there are only only 16
bits of the class B bags, only 65,000 and by the mid 90s, in fact,
early 90s, we were consuming them at a great rate. By 1994 we had
about two years to go when we'd run out no more middle size
networks. We had lots of C's, lots of little networks, but, you
know, that was its own problem. We had an astonishing number of
those, but the routers weren't big enough to route them all, and
that was just going to drown us in noise. We had a small number of
the big networks, but there were only 25 .. 127 of them. And so at
some point we're going to run out if we ever started using them
heavily. So how do we stop this? Well, the answer was, dispense
with the pre packaged bag sizes, get rid of the Goldilocks
problem. And. Say, look, here's a bunch of addresses that all
share a common prefix, and here's the size of that prefix. So you
could have networks with 256 hosts, or you could move it one bit
to the left and have 512, hosts, or 1024, hosts. You could
customize the size of your network. That's a great idea. It scales
brilliantly. But this routing protocol, BGP only thought about
class A, class B, class C,
George Michaelson 25:30
so that magic of how the Goldilocks bags was designed meant you
only had to look at top bits, a small number of bits at the front
of the address to work out the size of the bag. And you've said,
if we're running out of bags, we're going to have to find a way to
split things into different sizes, but you're no longer going to
be able to use just the front bits to work out how big is the bag,
right?
Geoff Huston 25:56
right. Every sort of routing object is now a common prefix value
and the length of that prefix value. So a Class C network in the
old speak was actually the first 24 bits of an address plus the
size 24 [George: right]. But that also allowed me to talk about,
say, a bunch of four Class Cs as a bunch of 22 bits of common
prefix followed by the number 22 which is the length of that
prefix.
George Michaelson 26:30
Yeah, it's a kind of slightly odd moment for people, because what
you're measuring is the length of the thing that is set. And the
natural way a lot of people think about this is it couldn't be a
count of things that are set but a count of things that are free.
But the fact is, we made a decision. It's the count of how long is
set.
Geoff Huston 26:49
right. So you're not out of the woods yet. You defined a new
address plan, which will get us out of the problem of running out
of class B addresses really neatly. We're not going to run out
anymore, because we haven't got that slop of trying to fit
everything into a class B, we can customize. But if the routing
protocol doesn't recognize that notation, that way of describing
addresses, you're no better off, literally no better off. And so
what we had to do was to change the routing protocols for RIP, we
had the inventive name rip v2 which included
George Michaelson 27:24
good name, Don't knock it, Geoff, that's a very good name.
Geoff Huston 27:28
It included prefix sizes. And for BGP, we went from BGP three to
the imaginatively named BGP four,
George Michaelson 27:36
good name, Geoff, Don't knock it.
Geoff Huston 27:39
It included classless inter domain routing, and this is where the
acronym CIDR
George Michaelson 27:45
comes from, C, I, D, R, CIDR.
Geoff Huston 27:48
CIDR. Now there was an IETF meeting, an Internet engineering task
force meeting, in March 1994 and the whole thing about let's all
run BGP four was given a thorough airing, and at the time, Cisco,
who had been instrumental in pushing this out, and was one of the
major router vendors of the day, deployed it and pushed it out
with their customers. And effects were dramatic and immediate. In
early March 94 we had just kicked 20,000 routes in the routing
table, class, A's, B's and C's a magic number, because one of the
networks, the US military network, Milnet, could only take 20,000
routes, and it just couldn't handle any more connectivity. Oops.
In the next sort of eight weeks after that meeting, that number of
routing entries dropped from 2000 down to 17,500 because we're
able to summarize more effectively, like a whole bunch of little
C's
George Michaelson 28:48
It had an immediate benefit that it was able to allow you to put
forward more efficiently statements analogous to the old bag
sizes, but now representing bigger chunks.
Geoff Huston 29:00
Er Yes and by doing more accurate and single chunks for larger and
larger addresses, what we were having because everyone was making
up for the shortfall in Bs with lots and lots of Cs, we were
having exponential growth in the routing table. And it wasn't that
there was an exponential number of new networks as such. There
really was a growth in the number of C networks, because the
medium sized networks were expressed as eight CS and 10 CS and so
on. And we were trying to relate that by making that a single
route object again. And interestingly, in the ensuing years, you
know, 94 95 96 the exponential growth in the routing table
disappeared and became largely linear again. And we all thought,
well, that's pretty cool, but we kind of didn't quite think it
through, because
George Michaelson 29:54
simplify, yes, simplifying assumptions, where you don't think
about all the data. Dimensions of problem and risk. And after
you've been running it for a while, probably you now depend on it.
You learn that it carries the seeds of its own destruction.
Geoff Huston 30:10
You know, when I said the distance, the metric between two
adjacent networks is always one, and the issue is I could connect
to you with a big, multi mega bit connection, and that had a
distance of one I could connect to you with a dial up modem with
the speed of, you know, I don't know, a few bits per fortnight,
and that would too, have a metric of one. How do I say to the
world, please use the link over there. Please don't use this tiny
link. It's only there for emergencies. How do I perform what we
currently call traffic engineering and load up a network whose
links aren't all the same, they're not uniform quality and size?
Somehow I want to express that, and BGP says, Well, don't look at
me, dude. And the reason why BGP can't do it is that, if I'm
trying to express a metric that's common to 20,000 independent
different networks, what's the metric? What's the unit? What does
it mean?
George Michaelson 31:15
Yeah, there's no controller. There's no one in charge. There's no
one to say from Monday, if you write this magic value in this is
what it means
Geoff Huston 31:24
that link is of cost five. This link is of cost 10. No one knows
that, and we kind of gave up on that, but we had another trick.
And this is again, clever in a, I suppose, useful way, but almost
perverse at the same time. So I have a bunch of Class Cs. Let's
say I have 16 of them, 16, that's a prefix of length, 20 -16 Class
Cs so I have a bunch of networks, but I can also advertise as well
as the 20 a couple of slash 24s drawn from inside that first
network. What do you mean? As it says, I can take more specific
prefixes and also advertise them. You know? Why would you want to
do that? Geoff, ah, because of a quirk in the way BGP actually
works. Because not only does BGP try to select the route with the
minimum metric, the minimum number of networks to traverse it
also, and absolutely prefers the most specific.
George Michaelson 32:32
So if you talk about something two ways, and one of them is, oh,
just casually in passing, there's a million things I know about
hosts in this block, but you also specifically say, here's some
magic info about a small section of it. You're meant to prefer the
magic info if you're walking along and that's the one you're
heading to, whatever was said about that smaller chunk that's more
important to you.
Geoff Huston 32:57
That's more important. So if I want to bias traffic being into me.
So if you need to buy us the information about how to reach me,
because I've got multiple paths, because, you know, networking is
getting serious, I can do just selectively advertising my more
specific prefixes that I think are going to attract traffic down
links which are either cheaper or better capacity, I can engineer
my traffic to match my policies as to how I connect to the rest of
the Internet.
George Michaelson 33:28
But Geoff, I'm starting to think you said earlier that the network
had been growing as an exponential curve and increase in the
amount of things being announced, and when this CIDR trick was
announced, it magically became both smaller and a bit more linear,
and you've now said but we found a way to do traffic engineering
if you announce lots and lots and lots and lots more than you were
announcing.
Geoff Huston 33:53
Yeah. Shame about the temper growth of BGP, because we got into
this wholesale of announcing more specifics, and yet again, the
race was on that we're actually finding that the routing table
wasn't just a bunch of addresses that you can reach, but a large
amount of that table size and a large amount of the routing
protocol was actually processing refinements to that more
specifics
George Michaelson 34:22
specific things that are specific to you because you are
attempting to manage engineering aspects of how you need to be
reached. So the whole of the world has to see this so that you can
manage something that's probably quite close to you, but that's
the only mechanism we have.
Geoff Huston 34:39
That's the only mechanism we have. So it didn't take long or half
of the routing table, [George half] be populated by more specifics
aggregates in the other half of the routing table, half
George Michaelson 34:53
of all the information we shared was this attempt to engineer
management of link utilization. Wow.
Geoff Huston 34:59
Because it was easier to make, if you will, lots of inadequate
links than it was to do a small number of monster links, because
they cost a fortune. And so it was quite common practice to have
richer connectivity, more connections, but then try and balance
your traffic across them using BGP in creative ways. So what
happens if everyone does it? Well, we all have to carry routers
with larger and larger routing tables. That's interesting. How
long have I got to look up an address inside a routing table?
Well, you've got the time it takes for one packet because you need
to get ready for the next packet to make a look up decision. So as
our links increase in speed, we've got less time to do a lookup,
and as the number of entries in that table increases, we've got
more things we need to look through to find a match. And we're
kind of struggling at the point because routing is getting more
expensive per unit
George Michaelson 36:02
right? So money has entered the room, money expressed as the
capitalized cost of making a machine that you can buy now and
foreseeably cope with this increase in the rate of speed and the
increase in the amount of information. This is not a trivial
decision. Geoff, this is quite a consequential one.
Geoff Huston 36:24
Well, it's it's a classic tragedy of the commons. I can optimize
and engineer my position in the inter domain routing space by
slicing and dicing my advertised networks into a fine grained
storm of prefixes and selectively advertise it. And indeed, some
folk made a business of optimize advertisements to actually
engineer across multiple links, constantly changing those
advertise more specifics to balance traffic. So I'm optimizing
myself, but the cost is everybody. And I mean, everybody else has
to carry those additional routes, so no one else benefits just me,
and that's the tragedy of the commons in a sort of a simple
restatement, collectively, income is bad, individually happy, but
only at the cost of everyone else.
George Michaelson 37:21
So this becomes a problem that's quite important for people to
understand the shape of the problem. How big is this problem? How
big is this problem growing? How does it impact me individually,
making my capital acquisitions? How do I plan this? That's a
classic situation where measurement comes to the fore, isn't it?
Geoff Huston 37:41
Well, it is, but I'm going to phrase it somewhat differently. How
can you get folk to stop it? I don't mention no one's in charge.
So it's kind of well, you know, hit me. There's no forcing
functions. Nobody's in control. So trying to put a lid on that
behavior, try to say, look, that is really anti social. You know,
I can appreciate you've got a problem to solve, but can you
appreciate the rest of us have to carry your load? Try and
exercise some constraint.
George Michaelson 38:13
So it's measurement. It's measurement to a different outcome. It's
kind of like the shame list of people who've left empty bottles in
the fridge at uni. It's name and shame, isn't it?
Geoff Huston 38:24
You've got it name and shame. Try and find the worst abusers of
this practice and say, just please how much you're doing to the
routing system, how bad your individual contribution is. And look,
I hope you, or maybe your people around you chastise you and
moderate what you're doing, because everyone is suffering from
you. And oddly enough, that is the CIDR report.
George Michaelson 38:48
Blimey,
Geoff Huston 38:49
Naming and Shaming,
George Michaelson 38:50
It's social engineering.
Geoff Huston 38:52
Oh yes, very much. So, very much. So originally started by Tony
Bates when he was in Cisco, I believe it was after he'd gone
through InternetMCI, and taken over by Philip Smith, also at
Cisco. And then I ended up picking it up when I was, I was in
Telstra, I think, at the time, and carrying it through. And it
was, it was a report that simply said, This is how big the routing
table is, and this is the daily size for the last seven days. Ooh,
look it rowing. And by the way, if we got rid of all those more
specifics, if we aggregated like crazy, here's the size it could
have been. It's got the same information content, the same
reachability content, but here's how much less of a load it could
be. Please appreciate this. But the report actually went one
further. Not only did it list those networks which were really,
really abusing this, it also listed what each of them could do in
terms of withdrawing more specifics and even adding one or two
aggregates to help that which would have the. Same traffic
engineering outcomes would preserve all of their properties of
incoming traffic, but drastically reduce their footprint in the
routing system.
George Michaelson 40:09
So it's both a name and shame, and here's how to get yourself out
of the mess. It's a classic the way out of the hole is to stop
digging, but adding with it, here's a ladder to lift you out of
the hole.
Geoff Huston 40:22
Yes, here's a way to actually, if you invest some effort, you can
achieve the same outcomes. And this is how. So, you know, please
consider this is a way through. So how good was it? Broadly
enough, there are actually a few folk who studied the CIDR report
as an exercise either in chastisement, Han Nusbacher and Barry
green reported to NANOG, because they manually followed up each
week. CIDR report got in touch with the big 10 and said, Hey, what
are you doing? Can you do better? Here's what the CIDR report says
about you, and they reported some limited levels of success,
[George: right] But there was actually a student at MIT in 2011
Stephen Woodrow, who got his PhD on the CIDR report, yay.
Stephen, who actually did analyze in some depth how effective was
naming and shaming.
George Michaelson 41:14
Geoff, you're talking about the CIDR report in the past tense,
but you haven't actually stopped publishing. Have you?
Geoff Huston 41:20
still coming out every day, still coming out. But this report was
in 2011 there's some time ago. And Stephen's report was actually
interesting. He said, Look, in the first few years it was
effective, it really did have some traction with the routing
community. And folk did exercise constraint after they had been
their attention to be drawn to it, and, you know, lifted their
game. But even in 2011 he said, Look, that's declined a lot, and
by 2011 it really didn't make much difference. And if you look at
the stats,
George Michaelson 41:54
but still we go on
Geoff Huston 41:56
you look at the stats that more specific good news is it hasn't
got worse. The bad news is it hasn't gotten any better. And as we
get to 1 million and just exceeded routing table entries in v4
half a million do not add reachability, do not add information,
they simply refine the traffic.
George Michaelson 42:16
And they could be done more effectively and be a bit more
efficient,
Geoff Huston 42:21
you'd have smaller, cheaper routers. And what about V6? Well, the
answer was, V6 did start with the best of intentions, and in 2004
2005 when, admittedly, the routing table was tiny, only about 20%
of the routes were more specific, and there was this feeling even
by 2011 at the time of Steven's report that we could do better,
and we're doing better, because in 2011 around 20% of the V6
routes were more specific. It was pretty stable, but
George Michaelson 42:53
good place to be.
Geoff Huston 42:54
When was V6? day, June the sixth 2012 or something? Lets be
serious about V6? Well, what does getting serious really mean? It
means using it for real traffic. It means using it for traffic
engineering. It means Yes, you guessed it, advertising more
specifics to actually bias the flow of traffic down your wires in
V6. And not only did we achieve
George Michaelson 43:18
But surely people don't disaggregate to the same extent, please.
Geoff, tell me they're slightly more effective in six.
Geoff Huston 43:25
Oh God, no, it's up to 60% worse. Now I don't know if these six
were more serious about V6. I don't know what the story is, but we
actually were doing worse. That was a peak. By the way. We have
got better. We're now down at 58% and that's pretty stable over
the last few years. It's really no it's worse. It really isn't
much better.
George Michaelson 43:46
This is quite an interesting situation, because we've done a
historical walk through, getting to the point where we understand
why people do what they do, and an emerging tragedy of the
commons. And we have three people, Tony Hain Phil Smith and you
who carry forward this report the CIDR report that someone in 2011
has said, Yeah, it did work, but it's not so effective now in
terms of its impact on things. But you haven't stopped publishing
the report, and we still know the shape of the hole and the depth
of the hole that we're digging ourselves into that interest me.
Geoff, does this still matter?
Geoff Huston 44:25
Well, the beauty of the report, of course, is just a program. It
doesn't need human intervention. So publishing it is without cost,
drawing people's attention to it is much harder, and getting folk
to act on it is much harder. But I suspect there's something else
at play as well. You see, at the time, 90s, early 2000s we had a
network which was used to connect people to service, people to
content, whatever was there was over there, and you were here, and
the network was the way you got to there. And. What it assumed,
interestingly, was that storage, computation, processing,
everything that makes a service is expensive and difficult, and
comms is cheap, because if I can get your packets over there, you
can access the service. We're all happy. You know, Microsoft were
busy doing updates of their windows operating system from a
massive barn of servers located in Seattle, in Washington and in
North West America.
George Michaelson 45:29
So routing was absolutely vital to deliver service models that at
that time, you did it by using a routing fabric to get there.
Geoff Huston 45:38
Routing really mattered. But as it turned out, the underlying
assumption, computation is expensive, storage is expensive, you
know, mounting services is expensive, communications is cheap. Is
the exact opposite of what we see today. Fiber optics has, you
know, completely changed things around to some extent, but what's
really changed, really changed is Moore's law. Reputation is dirt
cheap. Storage is dirt cheap. Why didn't I just replicate this
service 4000 times around the entire planet and position my
content beside every single major eyeball network says Netflix,
says YouTube, says Akamai. And so now we've actually found it
easier not to route. And this is this whole argument.
George Michaelson 46:27
Oh, that is a strong statement. Geoff, no routing necessary to get
the things you want to see and do.
Geoff Huston 46:35
Well, if I can get my content really close to you all I've got to
go through is the last mile transit, the last sorry, the last mile
access. There's no routing. It's just switching. Because cost
equals distance squared. If I can deliver that service very, very
close to you, it's really cheap because I don't have to go through
long, skinny pipes around the world. It's fast, so I can not only
take on the old broadcast television networks, but I can up the
ante by a factor of 10, faster, cheaper, better by simply locating
that everywhere. And so now about 90% of most eyeball networks
have their traffic being passed from a data center down the road
to an eyeball the other side of the road, and the total traffic,
the total distance phone, is tiny. There's no transit. If there's
no transit, there's no routing. And so in some ways, the
Internet's routing system is a curious historical artifact that a
few poor people have to rely on, but everyone else is over it. The
vast pool of money has moved onward into data centers and
localized delivery, and the CIDR report has accurately documented
that in ways that I think we didn't envisage at the time. So we've
just moved on from routing, and that's why the CIDR report is no
longer fundamentally an important report. It's just a historical
artifact,
George Michaelson 48:01
interesting, fascinating, but not actually driving the engine room
of what makes the Internet special.
Geoff Huston 48:08
We thought routing was everything. We were wrong. Interesting.
George Michaelson 48:11
So is there somewhere people can go to read the CIDR report on the
web?
Geoff Huston 48:17
www.cidr-report.org, and there's even a V6 version as well on
that page. Yes, it's all there.
Thanks, Geoff,
thank you, and thank you listeners.
George Michaelson 48:29
If you've got a story or research to share here on ping, why not
get in contact by email to ping@apnic.net or via the APNIC social
media channels, also remember the measurement@apnic.net mailing
list on orbit is there to discuss and share relevant collaborative
opportunities, grants and funding opportunities, jobs and graduate
placings, or to seek feedback from the community on your own
measurement projects. Be sure to check out the APNIC website for
all your resource and community needs until next time you.
A lot of routing protocols at the time that didn't use TCP had
this short minded or amnesiac counter every 90 seconds, every
period, it said you could be losing your marbles. Let me refresh
your memory. Here's everything I know, and those broadcast storms
of just refreshing stuff you already were told was a great fail
safe in an unreliable transport, but ultimately was incredibly
inefficient making this TCP say, I've told you, I'm never going to
tell you again, unless it changes. That was good. That cut down
the amount of routing traffic like crazy. The next thing is
looping. BGP had a very clever way of detecting looping, because
if A talks to B, talks to C, talks to D, talks to A, how do you
know it's a loop? It's kind of hard sometimes, if you think about
it, but what BGP did is it attached, almost like a snail trail, a
line of thread through the maze. Every time a network, its prefix
went through a network, it added the number of that network to its
trail. So let's say it's network A, the path is A. When it gets
sent to B, it goes well, the path is BA. When that set sent to C,
the path is CBA. In other words, the networks that have seen this
advertisement sign it and attach it, that signature, that their
number, their identifier, to that update. So if you have loop it,
you're seeing yourself. And thought, seen this. This is not real.
I've already seen it. I'm going to reject up front path vector was
the big difference between BGP and rip the Routing Information
Protocol, which did not have that and consequently, the only way
it figured out a loop was to actually count to infinity.
George Michaelson 1:57
You're listening to ping. A podcast by APNIC discussing all things
related to measuring the Internet. I'm your host, George
Michaelson, this time, I'm talking to Geoff Huston from APNIC labs
again in his regular monthly spot on ping. for over two decades,
Geoff has been publishing the CIDR report, a daily web update on
the state of BGP worldwide with information on routing table size
and where it comes from, along with information about anomalous
announcements started by Tony Bates at Cisco in the mid 90s, then
carried on by Phil Smith, who was also at Cisco at the time, and
then Geoff. It's a remarkable continuous measurement of the state
of routing. The CIDR report grew out of concerns with the
consequences of worldwide growth in the size of the routing table,
and especially a trend towards de-aggregation, when as holders
deliberately announce more specific prefixes from Their blocks to
optimize routing behaviors in a process called traffic
engineering. But what is CIDR or classless Internet domain
routing? What problem was it designed to solve and in a modern
Internet, Is this the problem we think it is? Has the CIDR report
been overtaken by other events in routing? Geoff, welcome back to
ping. What shall we talk about today?
Geoff Huston 3:23
Hi George. Look, I think it's time we talked about routing. It's
been a continuous Internet subject for the last eon or two.
George Michaelson 3:31
Oh, it's been a while since we did that, at least two episodes.
Geoff Huston 3:35
Oh, this time I actually want to kick into a report that I've been
producing for at least 20 odd years, and has been produced for
more than 30, I think about a sort of a summary of the state of
the routing system. It's called the CIDR report,
George Michaelson 3:50
CIDR
Geoff Huston 3:51
But what I'd like to do first is to Yes, now it's not about
something you drink, truly. I'd like to sort of talk about the
routing problem, as it was, right from the word go, and how the
CIDR report sort of snuck into all this
George Michaelson 4:07
Okey dokey, roll the clock back 20 years. 30 years? Is that
enough?
Geoff Huston 4:12
Oh, keep going. We're talking 1980s and I suppose one of the
things that was totally different about the desire of computer
networks, as distinct from, say, telephone networks and similar
ones of their day, is that computer networks were meant to be self
learning, that if you had one device, it wasn't a network. If you
had two, they were meant to be able to discover each other,
[George: right] And if you had a collection of these things, some
of them connected over local area networks, big thick Ethernet
wires, and some connected over thinner wires. Y'all meant to
discover each other.
George Michaelson 4:54
That was very much driven in the experience of pre existing
structural networks that were built with much more like a god
designer who said, Thou shalt connect this to that, and thou shalt
fan it out in these ways, to these remote entry vehicles and these
concentrators and will gather data from these locations and send
it along these links. It's like somebody laid out a map of a
network designed around devices and a computer. And the 80s people
were saying, we don't want to have to do that.
Geoff Huston 5:23
Well, that's right. And I think in some ways, the 80s saw the
beginning of Ethernet as a local area network. And the thing about
Ethernet, which was, I make it sounds like nothing today, but at
the time, you just plugged your computer into that common wire,
that bus cable, originally with a thick yellow cable you got at
your drill and your very expensive transceiver, drilled into the
coaxial cable on the little black mark, tapped it, connected your
computer, and you're able to talk to everyone else, other than the
need for a drill and an expensive transceiver. There wasn't a
controller of these Ethernet there was no one in charge. You just
connected your network, and while we sort of tried to push them
further out into the wider area, we wanted to take that plug and
play convenience with us. We wanted to use networks that did not
have a controller, that were actually an amalgam of almost self
organized networks that became off a self organizing beast. And
that was the desire of the routing protocols at the time [George:
right] to exchange the role of the network controller, the Big Fat
Controller, if you're a fan of Thomas the Tank Engine, to exchange
that role into one that was realistically one which was brokered
by the protocol itself, [George: right] So that was the magic
desire. And there are a number of attempts at this by various
vendors in various ways. The one that a lot of us had experience
with was actually a vendor network called DECnet.
George Michaelson 6:59
Yeah, used much as you did Geoff. It was a regular part of my
life. The thing
Geoff Huston 7:04
about DECnet was It was organized into two sort of tiers. There
was an area where machines inside an area had a detailed
conversation about their connections. You're on the same Ethernet
as me. I'm connected by a dedicated wire to you, etc. And so
inside an area, all the computers needed to be interconnected one
way or another. Didn't need to be a mesh, but any computer in a
network in an area needed to get to any other computer in the same
area without leaving that area right, just sort of like a local
city or something. And then DECnet had a second upper level
protocol, almost like an inter area protocol, where there were
area borders, and the area borders spoke via dedicated connections
to other area borders. And they said, Hi, I can reach all of area
one. How about you? I can reach all of area two. And so to send a
packet between a computer in area one to one in area two, you
directed your packet to your local area border router, who passed
it over to the right area, who then sent it on to the right
computer in that area.
George Michaelson 8:20
So Geoff the way you're describing this, it's like there's a bit
from Box A and a bit from Box B, because if you're inside the
area, you're making it sound as if it used the technique to
discover all the things within that area, the way the protocol
worked, you'd eventually discover every printer, every other host,
every other computer. But if you want to go to another area, there
was a higher structural layer that was having to do some
management and passing right,
Geoff Huston 8:45
right. Well, things almost as a an atlas, a mapping system. If
you're in a city, it's nice to know every street, but if you're
trying to get a packet to a different city, you really don't care.
You want to get it to that city and let the destination city sort
itself out. You didn't want to start the network and the protocol
interactions with all the detail of remote location, so you
literally did divide and conquer. Divide into areas, and at the
inter area level, you only looked after the connections between
these area border routers, and within an area, you looked after
the detail,
George Michaelson 9:21
you've already introduced the ideas of something local and a
boundary with a border and a border router, and the concept of an
area. And to me, these are four terms that are very interesting to
establish inside this model this early on in people's thinking.
Geoff Huston 9:37
I think DECnet was a model that actually we could have borrowed
and pushed into the Internet, almost subconsciously, I think so
many of us had cut our teeth from that particular vendor network,
and this concept of areas Radia Perlman had a lot to do with it
when she was working in digital was actually really seductively,
right? It sort of pushed detail into areas where you need a
detail. But when you're simply trying to move around a country or
around a region, you didn't need to do that. You just simply
talked about areas. And there were some pretty massive DECnet
networks, the NASA Science Internet, the high energy physics
network. They really did span most of the Northern Henry one way
or another. They weren't fantastically big networks, you know,
kilobits per second, but they did have a lot, a lot of members,
even digital's own corporate network,
George Michaelson 10:29
even this semi automatic mechanism of things being discovered and
things forming areas and communicating with each other, worked at
the scale of the whole of North America, the whole of Britain,
places of that kind of scale could quite comfortably build one of
these things.
Geoff Huston 10:44
So, yeah, digital built a network with 100,000 computers in it as
their corporate network. It was possible to scale it up. And so
when we started to build this Internet beyond the limited Research
Project Agency, the ARPAnet exercise, into the bigger field, that
routing concept we carried with us, and the Internet itself, I
think it really took off when the National Science Foundation
transformed the Internet by supporting a thing called the NSF net
across the United States to connect their super computers. And,
you know, all of a sudden, instead of a few 100 computers, which
was the ARPANET, the NSF net was talking in terms of 5000 10,000
20,000 network, there 20,000 computers inside a few months. Not
only that, but they wanted to interconnect to the NASA network, to
the high energy physics network. They wanted to interconnect to
the sort of quickly dying ARPANET, the old defense network, old
defense sponsored network. So the problem was complex. It wasn't
simple, and they needed to use much the same trick. But what they
did, I think, was quite insightful. It was that area and time when
divide and conquer made sense, that there was no single corporate
and what they very quickly realized was that the networking,
routing protocol you used inside a network an area didn't need to
be the same one as you used outside that area.
George Michaelson 12:14
That is quite a radical idea when you compare it to other
behaviors extant in the Internet in these days, the idea that
things didn't have to be the same, in some ways, ran counter to a
culture that said, we're going to standardize the format of email
addresses across the entire surface of this network. We're going
to make telnet and ftp pretty much use the same kind of logical
structure. And here you are saying divide and conquer allowed
people to stand up and say, let's actually think about doing it
differently inside an area and differently between areas. That's
interesting. An interesting difference
Geoff Huston 12:50
actually, like a difference between a commercial service offering
and the so called, you know, the research people at the time, the
commercial thought, I solve all your problems and so DECnet really
was, I think, two different routing protocols smashed into one
because, you know, DECnet needed to do everything, whereas in the
research world, you know, the academic research world, it really
was. Look, we're not going to tell you what to do. Here are some
things, but if you don't like them, go run your own. And there
were, in the 80s, a lot of university campuses, I know I lived in
one where driving your own and building your own local area
network protocol was a sport, a competitive hobby. I worked at the
ANU, home of ANU net at the University of Monash in Victoria,
Monnet and on and on and on.
George Michaelson 13:38
And UQ net, which is where I was and csironet, which is where I
was before that. And you could build a surprisingly large amount
of interconnected state using incredibly simple protocols, like
RIP, the protocol we use to construct the network at UQ, right,
Geoff Huston 13:53
right. So there was this sort of thing of, roll your own local
network, roll your own network within sort of your network, and
then a case of, how do you hook them together? And that's kind of
well, in some ways, it's quite easy, if you think that a routing
protocol is simply to tell you where addresses are, address
prefixes. So each network has a bundle of IP addresses, in this
case, and hopefully they're contiguous. Well, at the time they
were and they came in three sacks, the Class C sack, and that had
256 addresses, the Class B sack, which had 65,000 and for the
privileged few, there was the class A sack with 17 million
addresses.
George Michaelson 14:39
Yeah, but this kind of anti Goldilocks moment, because the class C
sack that anyone could get easily was really too small, and if you
were big right on the Class B sack was running a bit, and the
class A sack, well, you couldn't get it.
Geoff Huston 14:52
We will get there. But I want to talk about this for a second,
because what the routing protocol needed to do was in the exterior
sense, to say to all the other networks, or at least to the ones
that were immediately adjacent, as networks, not as computers, as
networks. Hey, I can reach these following Class C addresses
bundles and a Class B bundle. Here have the bundles that I can
reach. You've got a packet address. Any of those bundles, send
them to me. I'm your network. You tell me what you can get to
George Michaelson 15:28
I'm your network, which is a magic moment in information sharing,
isn't it, because it takes a thing that could be an enumerated
list. Here's 255 things, and it says, Nah I'm just going to refer
to it as a common bit and how many there are, and that's all I
have to tell you.
Geoff Huston 15:30
And so this idea of bundling up and talking about reachability by
network sacks or prefixes appeared, and the original
implementations of this was the EGP, interestingly, the exterior
gateway protocol, and there were a few sort of things around that,
but it was all relatively simple. But oddly enough, the problem
they were solving was not simple, because already we had nine
regional networks, a National Science backbone network, a NASA
network, a high energy physics network, a network in the United
Kingdom, a network, I think in Germany, very quickly, there are a
bunch of these coming up and getting from one place to another was
a non trivial problem, and you couldn't just send the details of
every connection to every part of the net. There was no bandwidth,
no capability. It would work. So we needed a better form of
exterior network routing. And you know, the answer was pretty
simple, as it turned out, it came out in 1989 when IBM's Jakov
Rechter, IBM was the contractor to NSF, and Yaakov was the
principal scientist over there at IBM, and Cisco's early employee,
Kurt Lockheed, came up with this protocol called the Border
Gateway routing protocol. Now it was simple because it relied on
something that I think was originally in a paper in 1956 the
Bellman Ford, distance vector protocol. I'll tell you everything I
know when you hear what I say, add one to every metric about what
I know. I can see this address prefix in one hop so you go, very
good, Geoff, I see it. Therefore in two hops. I see this one in
five, because people have told me and told me, and you say, very
good. I'll make it six.
George Michaelson 17:29
So the essential quality here is that we've stayed in the space of
we don't want to have to have a big boss that tells us what to
say. We all shout out into the world the stuff about us. But
instead of shouting the entire list of what we got, we now come to
consider, well, I consider myself an area. I shout out efficiently
the prefixes the block of things I have. And you've just said, we
added a twist to that, that when you do it, you use this technique
from the 50s called Bellman Ford, shouting it out is cost of one,
and the guy who hears it shouts it out again, but adds one to make
it a cost of two. And if you were the third in the chain and you
heard it, you'd shout it out, and you'd add a cost to make it
three, right?
Geoff Huston 18:09
right And if you had two connections, people shouting at you, and
one of them says, I see this address prefix with a cost of one,
and the other one says, I too see it, but my cost is eight.
Obviously you'd pick the cost of one, you know, so you naturally
prefer the shorter cost.
George Michaelson 18:26
So we have one unit, one unit of cost and its distance. And if
everyone obeys the rules, you can rationally determine the lowest,
the best cost.
Geoff Huston 18:36
Right? Simple distance. You and I could connect over anything we
wanted, everything was cost one if it moved from one network to
another. Wasn't very sophisticated, but I want my big fat wire to
be more, more preferable than your tiny little wire, and it's kind
of No, no. BGP doesn't do that. The distance between a network is
one stop trying to be clever.
George Michaelson 18:59
There's no dimension for that.
Geoff Huston 19:01
However, there were a few things that were clever. The first thing
they did, actually, I think it was, was one of the big fundamental
things, is that the transport that BGP used was actually TCP,
which means, when I tell you, and you send me an acknowledgement,
you know it, I know you know it because you acknowledge the
packet. I never need to tell you again, only if it changes.
George Michaelson 19:27
So BGP didn't have to construct its own way of saying. Did you
hear that? Are you sure about that? Can I believe that because TCP
told you they got it.
Geoff Huston 19:37
A lot of routing protocols at the time that didn't use TCP had
this short minded or amnesiac counter every 90 seconds, every
period it said you could be losing your marbles. Let me refresh
your memory. Here's everything I know and those broadcast storms
of just refreshing stuff you already were told was a great fail
safe in an unreliable transfer. Bought, but ultimately was
incredibly inefficient making this TCP said, I've told you, I'm
never going to tell you again, unless it changes. That was good.
That cut down the amount of routing traffic like crazy. The next
thing is looping. BGP had a very clever way of detecting looping,
because if A talks to B, talks to C, talks to D, talks to a how do
you know it's a loop. It's kind of hard sometimes, if you think
about it, but what BGP did is it attached, almost like a snail
trail, a line of thread through the maze. Every time a network,
its prefix went through a network, it added the number of that
work to its trail. So let's say it's network A, the path is A.
When it gets sent to B, it goes well, the path is BA. When that
set sent to C, the path is CBA. In other words, the networks that
have seen this advertisement sign it and attach it. That signature
that their number, their identifier, to that update. So if you
ever loop it, you're seeing yourself and thought, seen this. This
is not real. I've already seen it. I'm going to reject up front
path. It was the big difference between BGP and rip the Routing
Information Protocol, which did not have that. And consequently,
the only way it figured out a loop was to actually count to
infinity.
George Michaelson 21:24
That takes a long time.
Geoff Huston 21:26
Well, they made it 16, but it's still not very good. So we've
talked about the separation between inside your network and
between networks. We've talked about TCP, we've talked about path,
vector and transport. But the other thing I think we've just
alluded to, but it is very, very important. Is no one's in
control. There are no permissions. It is a network of peers.
George Michaelson 21:47
Well, we could say there's a kind of moment here that we're not
going to call control of routing. But if you said, put a number, B
puts its number, those numbers do have to be kept distinct. And so
there's a function outside of routing, handing out the numbers to
make sure you only hand them out once, and you don't make two
people think they both claim to be the same unique identity. But
that aside, nothing in routing determines certain parts are built.
Geoff Huston 22:15
We had a registry run by Stanford research that handed out address
prefixes, those bags of addresses, and it handed out what we
called autonomous system numbers, these identifiers for networks.
And there are very few qualifications to those. You sent a fax
because it was the time of faxes over to these folk, and they fax
back saying your autonomous system number is this number, and if
you ask for address prefix, here's the address prefix, Class A, B
or C, and off we all went. And this kind of worked through
relatively well. The early days of the expansion of the Internet,
when the NSF started work, there were a few 100 such networks. And
over the ensuing years, you know, 90 91 92 then NSF heyday, it
really did swing up to 15,000 component networks, which is
amazingly fast.
George Michaelson 23:11
And this mechanism of simple BGP, a path mechanism to detect
loops, assignment of unique numbers, this was essentially working,
and meant that prior state you've described, the several networks
in America, the discrete networks in Britain and Germany, the
emerging network in Australia, they were able to come into a union
of exchanging information with each other. No central controller.
It was just fundamentally working.
Geoff Huston 23:37
You just needed to connect to someone who was connected. It was
this was that you just added yourself to the edge, and that edge
based system was phenomenally successful, brilliantly successful.
By 1994 we were faced, and had been facing, for some years, some
predictions, which were truly awesome, because all of a sudden
this was the consumption problem. We were going to consume all the
available addresses. And let me qualify that we weren't really
there are 4 billion addresses in IPV4, but there are only only 16
bits of the class B bags, only 65,000 and by the mid 90s, in fact,
early 90s, we were consuming them at a great rate. By 1994 we had
about two years to go when we'd run out no more middle size
networks. We had lots of C's, lots of little networks, but, you
know, that was its own problem. We had an astonishing number of
those, but the routers weren't big enough to route them all, and
that was just going to drown us in noise. We had a small number of
the big networks, but there were only 25 .. 127 of them. And so at
some point we're going to run out if we ever started using them
heavily. So how do we stop this? Well, the answer was, dispense
with the pre packaged bag sizes, get rid of the Goldilocks
problem. And. Say, look, here's a bunch of addresses that all
share a common prefix, and here's the size of that prefix. So you
could have networks with 256 hosts, or you could move it one bit
to the left and have 512, hosts, or 1024, hosts. You could
customize the size of your network. That's a great idea. It scales
brilliantly. But this routing protocol, BGP only thought about
class A, class B, class C,
George Michaelson 25:30
so that magic of how the Goldilocks bags was designed meant you
only had to look at top bits, a small number of bits at the front
of the address to work out the size of the bag. And you've said,
if we're running out of bags, we're going to have to find a way to
split things into different sizes, but you're no longer going to
be able to use just the front bits to work out how big is the bag,
right?
Geoff Huston 25:56
right. Every sort of routing object is now a common prefix value
and the length of that prefix value. So a Class C network in the
old speak was actually the first 24 bits of an address plus the
size 24 [George: right]. But that also allowed me to talk about,
say, a bunch of four Class Cs as a bunch of 22 bits of common
prefix followed by the number 22 which is the length of that
prefix.
George Michaelson 26:30
Yeah, it's a kind of slightly odd moment for people, because what
you're measuring is the length of the thing that is set. And the
natural way a lot of people think about this is it couldn't be a
count of things that are set but a count of things that are free.
But the fact is, we made a decision. It's the count of how long is
set.
Geoff Huston 26:49
right. So you're not out of the woods yet. You defined a new
address plan, which will get us out of the problem of running out
of class B addresses really neatly. We're not going to run out
anymore, because we haven't got that slop of trying to fit
everything into a class B, we can customize. But if the routing
protocol doesn't recognize that notation, that way of describing
addresses, you're no better off, literally no better off. And so
what we had to do was to change the routing protocols for RIP, we
had the inventive name rip v2 which included
George Michaelson 27:24
good name, Don't knock it, Geoff, that's a very good name.
Geoff Huston 27:28
It included prefix sizes. And for BGP, we went from BGP three to
the imaginatively named BGP four,
George Michaelson 27:36
good name, Geoff, Don't knock it.
Geoff Huston 27:39
It included classless inter domain routing, and this is where the
acronym CIDR
George Michaelson 27:45
comes from, C, I, D, R, CIDR.
Geoff Huston 27:48
CIDR. Now there was an IETF meeting, an Internet engineering task
force meeting, in March 1994 and the whole thing about let's all
run BGP four was given a thorough airing, and at the time, Cisco,
who had been instrumental in pushing this out, and was one of the
major router vendors of the day, deployed it and pushed it out
with their customers. And effects were dramatic and immediate. In
early March 94 we had just kicked 20,000 routes in the routing
table, class, A's, B's and C's a magic number, because one of the
networks, the US military network, Milnet, could only take 20,000
routes, and it just couldn't handle any more connectivity. Oops.
In the next sort of eight weeks after that meeting, that number of
routing entries dropped from 2000 down to 17,500 because we're
able to summarize more effectively, like a whole bunch of little
C's
George Michaelson 28:48
It had an immediate benefit that it was able to allow you to put
forward more efficiently statements analogous to the old bag
sizes, but now representing bigger chunks.
Geoff Huston 29:00
Er Yes and by doing more accurate and single chunks for larger and
larger addresses, what we were having because everyone was making
up for the shortfall in Bs with lots and lots of Cs, we were
having exponential growth in the routing table. And it wasn't that
there was an exponential number of new networks as such. There
really was a growth in the number of C networks, because the
medium sized networks were expressed as eight CS and 10 CS and so
on. And we were trying to relate that by making that a single
route object again. And interestingly, in the ensuing years, you
know, 94 95 96 the exponential growth in the routing table
disappeared and became largely linear again. And we all thought,
well, that's pretty cool, but we kind of didn't quite think it
through, because
George Michaelson 29:54
simplify, yes, simplifying assumptions, where you don't think
about all the data. Dimensions of problem and risk. And after
you've been running it for a while, probably you now depend on it.
You learn that it carries the seeds of its own destruction.
Geoff Huston 30:10
You know, when I said the distance, the metric between two
adjacent networks is always one, and the issue is I could connect
to you with a big, multi mega bit connection, and that had a
distance of one I could connect to you with a dial up modem with
the speed of, you know, I don't know, a few bits per fortnight,
and that would too, have a metric of one. How do I say to the
world, please use the link over there. Please don't use this tiny
link. It's only there for emergencies. How do I perform what we
currently call traffic engineering and load up a network whose
links aren't all the same, they're not uniform quality and size?
Somehow I want to express that, and BGP says, Well, don't look at
me, dude. And the reason why BGP can't do it is that, if I'm
trying to express a metric that's common to 20,000 independent
different networks, what's the metric? What's the unit? What does
it mean?
George Michaelson 31:15
Yeah, there's no controller. There's no one in charge. There's no
one to say from Monday, if you write this magic value in this is
what it means
Geoff Huston 31:24
that link is of cost five. This link is of cost 10. No one knows
that, and we kind of gave up on that, but we had another trick.
And this is again, clever in a, I suppose, useful way, but almost
perverse at the same time. So I have a bunch of Class Cs. Let's
say I have 16 of them, 16, that's a prefix of length, 20 -16 Class
Cs so I have a bunch of networks, but I can also advertise as well
as the 20 a couple of slash 24s drawn from inside that first
network. What do you mean? As it says, I can take more specific
prefixes and also advertise them. You know? Why would you want to
do that? Geoff, ah, because of a quirk in the way BGP actually
works. Because not only does BGP try to select the route with the
minimum metric, the minimum number of networks to traverse it
also, and absolutely prefers the most specific.
George Michaelson 32:32
So if you talk about something two ways, and one of them is, oh,
just casually in passing, there's a million things I know about
hosts in this block, but you also specifically say, here's some
magic info about a small section of it. You're meant to prefer the
magic info if you're walking along and that's the one you're
heading to, whatever was said about that smaller chunk that's more
important to you.
Geoff Huston 32:57
That's more important. So if I want to bias traffic being into me.
So if you need to buy us the information about how to reach me,
because I've got multiple paths, because, you know, networking is
getting serious, I can do just selectively advertising my more
specific prefixes that I think are going to attract traffic down
links which are either cheaper or better capacity, I can engineer
my traffic to match my policies as to how I connect to the rest of
the Internet.
George Michaelson 33:28
But Geoff, I'm starting to think you said earlier that the network
had been growing as an exponential curve and increase in the
amount of things being announced, and when this CIDR trick was
announced, it magically became both smaller and a bit more linear,
and you've now said but we found a way to do traffic engineering
if you announce lots and lots and lots and lots more than you were
announcing.
Geoff Huston 33:53
Yeah. Shame about the temper growth of BGP, because we got into
this wholesale of announcing more specifics, and yet again, the
race was on that we're actually finding that the routing table
wasn't just a bunch of addresses that you can reach, but a large
amount of that table size and a large amount of the routing
protocol was actually processing refinements to that more
specifics
George Michaelson 34:22
specific things that are specific to you because you are
attempting to manage engineering aspects of how you need to be
reached. So the whole of the world has to see this so that you can
manage something that's probably quite close to you, but that's
the only mechanism we have.
Geoff Huston 34:39
That's the only mechanism we have. So it didn't take long or half
of the routing table, [George half] be populated by more specifics
aggregates in the other half of the routing table, half
George Michaelson 34:53
of all the information we shared was this attempt to engineer
management of link utilization. Wow.
Geoff Huston 34:59
Because it was easier to make, if you will, lots of inadequate
links than it was to do a small number of monster links, because
they cost a fortune. And so it was quite common practice to have
richer connectivity, more connections, but then try and balance
your traffic across them using BGP in creative ways. So what
happens if everyone does it? Well, we all have to carry routers
with larger and larger routing tables. That's interesting. How
long have I got to look up an address inside a routing table?
Well, you've got the time it takes for one packet because you need
to get ready for the next packet to make a look up decision. So as
our links increase in speed, we've got less time to do a lookup,
and as the number of entries in that table increases, we've got
more things we need to look through to find a match. And we're
kind of struggling at the point because routing is getting more
expensive per unit
George Michaelson 36:02
right? So money has entered the room, money expressed as the
capitalized cost of making a machine that you can buy now and
foreseeably cope with this increase in the rate of speed and the
increase in the amount of information. This is not a trivial
decision. Geoff, this is quite a consequential one.
Geoff Huston 36:24
Well, it's it's a classic tragedy of the commons. I can optimize
and engineer my position in the inter domain routing space by
slicing and dicing my advertised networks into a fine grained
storm of prefixes and selectively advertise it. And indeed, some
folk made a business of optimize advertisements to actually
engineer across multiple links, constantly changing those
advertise more specifics to balance traffic. So I'm optimizing
myself, but the cost is everybody. And I mean, everybody else has
to carry those additional routes, so no one else benefits just me,
and that's the tragedy of the commons in a sort of a simple
restatement, collectively, income is bad, individually happy, but
only at the cost of everyone else.
George Michaelson 37:21
So this becomes a problem that's quite important for people to
understand the shape of the problem. How big is this problem? How
big is this problem growing? How does it impact me individually,
making my capital acquisitions? How do I plan this? That's a
classic situation where measurement comes to the fore, isn't it?
Geoff Huston 37:41
Well, it is, but I'm going to phrase it somewhat differently. How
can you get folk to stop it? I don't mention no one's in charge.
So it's kind of well, you know, hit me. There's no forcing
functions. Nobody's in control. So trying to put a lid on that
behavior, try to say, look, that is really anti social. You know,
I can appreciate you've got a problem to solve, but can you
appreciate the rest of us have to carry your load? Try and
exercise some constraint.
George Michaelson 38:13
So it's measurement. It's measurement to a different outcome. It's
kind of like the shame list of people who've left empty bottles in
the fridge at uni. It's name and shame, isn't it?
Geoff Huston 38:24
You've got it name and shame. Try and find the worst abusers of
this practice and say, just please how much you're doing to the
routing system, how bad your individual contribution is. And look,
I hope you, or maybe your people around you chastise you and
moderate what you're doing, because everyone is suffering from
you. And oddly enough, that is the CIDR report.
George Michaelson 38:48
Blimey,
Geoff Huston 38:49
Naming and Shaming,
George Michaelson 38:50
It's social engineering.
Geoff Huston 38:52
Oh yes, very much. So, very much. So originally started by Tony
Bates when he was in Cisco, I believe it was after he'd gone
through InternetMCI, and taken over by Philip Smith, also at
Cisco. And then I ended up picking it up when I was, I was in
Telstra, I think, at the time, and carrying it through. And it
was, it was a report that simply said, This is how big the routing
table is, and this is the daily size for the last seven days. Ooh,
look it rowing. And by the way, if we got rid of all those more
specifics, if we aggregated like crazy, here's the size it could
have been. It's got the same information content, the same
reachability content, but here's how much less of a load it could
be. Please appreciate this. But the report actually went one
further. Not only did it list those networks which were really,
really abusing this, it also listed what each of them could do in
terms of withdrawing more specifics and even adding one or two
aggregates to help that which would have the. Same traffic
engineering outcomes would preserve all of their properties of
incoming traffic, but drastically reduce their footprint in the
routing system.
George Michaelson 40:09
So it's both a name and shame, and here's how to get yourself out
of the mess. It's a classic the way out of the hole is to stop
digging, but adding with it, here's a ladder to lift you out of
the hole.
Geoff Huston 40:22
Yes, here's a way to actually, if you invest some effort, you can
achieve the same outcomes. And this is how. So, you know, please
consider this is a way through. So how good was it? Broadly
enough, there are actually a few folk who studied the CIDR report
as an exercise either in chastisement, Han Nusbacher and Barry
green reported to NANOG, because they manually followed up each
week. CIDR report got in touch with the big 10 and said, Hey, what
are you doing? Can you do better? Here's what the CIDR report says
about you, and they reported some limited levels of success,
[George: right] But there was actually a student at MIT in 2011
Stephen Woodrow, who got his PhD on the CIDR report, yay.
Stephen, who actually did analyze in some depth how effective was
naming and shaming.
George Michaelson 41:14
Geoff, you're talking about the CIDR report in the past tense,
but you haven't actually stopped publishing. Have you?
Geoff Huston 41:20
still coming out every day, still coming out. But this report was
in 2011 there's some time ago. And Stephen's report was actually
interesting. He said, Look, in the first few years it was
effective, it really did have some traction with the routing
community. And folk did exercise constraint after they had been
their attention to be drawn to it, and, you know, lifted their
game. But even in 2011 he said, Look, that's declined a lot, and
by 2011 it really didn't make much difference. And if you look at
the stats,
George Michaelson 41:54
but still we go on
Geoff Huston 41:56
you look at the stats that more specific good news is it hasn't
got worse. The bad news is it hasn't gotten any better. And as we
get to 1 million and just exceeded routing table entries in v4
half a million do not add reachability, do not add information,
they simply refine the traffic.
George Michaelson 42:16
And they could be done more effectively and be a bit more
efficient,
Geoff Huston 42:21
you'd have smaller, cheaper routers. And what about V6? Well, the
answer was, V6 did start with the best of intentions, and in 2004
2005 when, admittedly, the routing table was tiny, only about 20%
of the routes were more specific, and there was this feeling even
by 2011 at the time of Steven's report that we could do better,
and we're doing better, because in 2011 around 20% of the V6
routes were more specific. It was pretty stable, but
George Michaelson 42:53
good place to be.
Geoff Huston 42:54
When was V6? day, June the sixth 2012 or something? Lets be
serious about V6? Well, what does getting serious really mean? It
means using it for real traffic. It means using it for traffic
engineering. It means Yes, you guessed it, advertising more
specifics to actually bias the flow of traffic down your wires in
V6. And not only did we achieve
George Michaelson 43:18
But surely people don't disaggregate to the same extent, please.
Geoff, tell me they're slightly more effective in six.
Geoff Huston 43:25
Oh God, no, it's up to 60% worse. Now I don't know if these six
were more serious about V6. I don't know what the story is, but we
actually were doing worse. That was a peak. By the way. We have
got better. We're now down at 58% and that's pretty stable over
the last few years. It's really no it's worse. It really isn't
much better.
George Michaelson 43:46
This is quite an interesting situation, because we've done a
historical walk through, getting to the point where we understand
why people do what they do, and an emerging tragedy of the
commons. And we have three people, Tony Hain Phil Smith and you
who carry forward this report the CIDR report that someone in 2011
has said, Yeah, it did work, but it's not so effective now in
terms of its impact on things. But you haven't stopped publishing
the report, and we still know the shape of the hole and the depth
of the hole that we're digging ourselves into that interest me.
Geoff, does this still matter?
Geoff Huston 44:25
Well, the beauty of the report, of course, is just a program. It
doesn't need human intervention. So publishing it is without cost,
drawing people's attention to it is much harder, and getting folk
to act on it is much harder. But I suspect there's something else
at play as well. You see, at the time, 90s, early 2000s we had a
network which was used to connect people to service, people to
content, whatever was there was over there, and you were here, and
the network was the way you got to there. And. What it assumed,
interestingly, was that storage, computation, processing,
everything that makes a service is expensive and difficult, and
comms is cheap, because if I can get your packets over there, you
can access the service. We're all happy. You know, Microsoft were
busy doing updates of their windows operating system from a
massive barn of servers located in Seattle, in Washington and in
North West America.
George Michaelson 45:29
So routing was absolutely vital to deliver service models that at
that time, you did it by using a routing fabric to get there.
Geoff Huston 45:38
Routing really mattered. But as it turned out, the underlying
assumption, computation is expensive, storage is expensive, you
know, mounting services is expensive, communications is cheap. Is
the exact opposite of what we see today. Fiber optics has, you
know, completely changed things around to some extent, but what's
really changed, really changed is Moore's law. Reputation is dirt
cheap. Storage is dirt cheap. Why didn't I just replicate this
service 4000 times around the entire planet and position my
content beside every single major eyeball network says Netflix,
says YouTube, says Akamai. And so now we've actually found it
easier not to route. And this is this whole argument.
George Michaelson 46:27
Oh, that is a strong statement. Geoff, no routing necessary to get
the things you want to see and do.
Geoff Huston 46:35
Well, if I can get my content really close to you all I've got to
go through is the last mile transit, the last sorry, the last mile
access. There's no routing. It's just switching. Because cost
equals distance squared. If I can deliver that service very, very
close to you, it's really cheap because I don't have to go through
long, skinny pipes around the world. It's fast, so I can not only
take on the old broadcast television networks, but I can up the
ante by a factor of 10, faster, cheaper, better by simply locating
that everywhere. And so now about 90% of most eyeball networks
have their traffic being passed from a data center down the road
to an eyeball the other side of the road, and the total traffic,
the total distance phone, is tiny. There's no transit. If there's
no transit, there's no routing. And so in some ways, the
Internet's routing system is a curious historical artifact that a
few poor people have to rely on, but everyone else is over it. The
vast pool of money has moved onward into data centers and
localized delivery, and the CIDR report has accurately documented
that in ways that I think we didn't envisage at the time. So we've
just moved on from routing, and that's why the CIDR report is no
longer fundamentally an important report. It's just a historical
artifact,
George Michaelson 48:01
interesting, fascinating, but not actually driving the engine room
of what makes the Internet special.
Geoff Huston 48:08
We thought routing was everything. We were wrong. Interesting.
George Michaelson 48:11
So is there somewhere people can go to read the CIDR report on the
web?
Geoff Huston 48:17
www.cidr-report.org, and there's even a V6 version as well on
that page. Yes, it's all there.
Thanks, Geoff,
thank you, and thank you listeners.
George Michaelson 48:29
If you've got a story or research to share here on ping, why not
get in contact by email to ping@apnic.net or via the APNIC social
media channels, also remember the measurement@apnic.net mailing
list on orbit is there to discuss and share relevant collaborative
opportunities, grants and funding opportunities, jobs and graduate
placings, or to seek feedback from the community on your own
measurement projects. Be sure to check out the APNIC website for
all your resource and community needs until next time you.