WEBVTT
1
00:00:02.080 --> 00:00:06.400
Welcome to fedsock Forums, a podcast
of the Federal Society's Practice groups. I'm
2
00:00:06.440 --> 00:00:10.039
Ny kas Merrick, Vice President and
Director of Practice Groups at the Federal Society.
3
00:00:10.240 --> 00:00:14.640
For exclusive access to live recordings of
fedsock Forum programs, become a Federal
4
00:00:14.679 --> 00:00:20.679
Society member today at fedsoc dot org. Hello everyone, and welcome to this
5
00:00:20.719 --> 00:00:24.839
Federal Society virtual event. My name
is Emily Manning, and I'm an Associate
6
00:00:24.879 --> 00:00:29.000
director of Practice Groups with the Federal
Society. Today we're excited to host a
7
00:00:29.039 --> 00:00:35.240
discussion titled AI meets Copyright Understanding New
York Times v OpenAI. We're joined today
8
00:00:35.320 --> 00:00:40.320
by Charles Dwan V. Rosen,
Stephen M. Tepp, and our moderator
9
00:00:40.320 --> 00:00:44.960
today is John P. Moran of
Council at Holland and Knight. John has
10
00:00:45.000 --> 00:00:50.840
experienced litigating many patent, trademark and
trade secret cases and Federal District Court and
11
00:00:51.000 --> 00:00:55.280
argues appeals at the US Court of
Appeals for the Federal Circuit the US Court
12
00:00:55.280 --> 00:01:00.000
of Appeals for the Fourth Circuit.
He has prosecuted or directly supervised the prosecute
13
00:01:00.359 --> 00:01:06.680
of hundreds of patent applications and many
different technologies, including telecommunication systems and equipment
14
00:01:07.120 --> 00:01:15.480
robotics, artificial intelligence, imaging technology, nuclear reactor instrumentation, semiconductor devices and
15
00:01:15.560 --> 00:01:19.879
manufacturing processes and medical devices. If
you'd like to learn more about today's speakers,
16
00:01:19.920 --> 00:01:23.959
their full bios can be viewed on
our website FEDSOC dot org. After
17
00:01:25.000 --> 00:01:27.359
our speakers give their opening remarks,
we will turn to you the audience,
18
00:01:27.359 --> 00:01:30.480
for questions. If you have a
question, please enter it into the Q
19
00:01:30.560 --> 00:01:34.200
and A function at the bottom of
your zoom window, and we will do
20
00:01:34.239 --> 00:01:37.560
our best answer as many as we
can. Finally, I'll note that,
21
00:01:37.640 --> 00:01:41.079
as always, all expressions of opinion
today are those of our guest speakers,
22
00:01:41.319 --> 00:01:44.480
not the Federal Society. With that, thank you for joining us today,
23
00:01:44.519 --> 00:01:48.680
and John, the floor is yours. Thank you Emily for that introduction.
24
00:01:49.599 --> 00:01:53.840
Well, this afternoon's panel is Emily
indicated. We'll discuss the copyrighted issues raised
25
00:01:55.040 --> 00:02:00.480
in the complaint brought by The New
York Times against Microsoft and several Open AI
26
00:02:00.840 --> 00:02:06.400
entity. The complaint, which was
fought last December, has seven counts.
27
00:02:06.760 --> 00:02:12.960
Four of the counts are for copyright
infringement counts. There's a Digital Millennium Copyright
28
00:02:13.080 --> 00:02:20.439
Act count and a common unfair competition
by misappropriation of New York Times intellectual property.
29
00:02:21.080 --> 00:02:25.800
There's also a trademark dilution count,
which will not be addressed today.
30
00:02:27.199 --> 00:02:34.759
In support of the copyright infringement allegations, New York Times included in the complaint
31
00:02:34.919 --> 00:02:40.400
one hundred examples of outputs that are
nearly identicate and identical to the copyrighted New
32
00:02:40.479 --> 00:02:47.240
York Times content. Before we address
the copyright issues, a bit of technology
33
00:02:47.319 --> 00:02:53.960
vocabulary at a very high level may
help the discussion, so to begin with,
34
00:02:53.159 --> 00:03:00.639
the complaint focuses on the open aised
GPT models as it relates to today's
35
00:03:00.639 --> 00:03:09.719
discussion. The acronym stands for generative
pre trained transformer. The generative, which
36
00:03:09.759 --> 00:03:15.479
we will probably discussed today, indicates
that the model takes inputs from users and
37
00:03:15.599 --> 00:03:22.240
generates an output such as one hundred
examples of New York Times content and the
38
00:03:22.280 --> 00:03:28.479
complaint. The P for pre trained
indicates that the model is pre trained on
39
00:03:28.599 --> 00:03:32.719
the volume of data such as the
New York Times content, and lastly,
40
00:03:34.159 --> 00:03:39.800
the T for transformer very generally relates
to a way of processing data to provide
41
00:03:39.960 --> 00:03:46.639
context for the data, such as
a sequence of words. One note that
42
00:03:46.759 --> 00:03:52.479
the complaint uses the term embedded.
It appears that it uses that term in
43
00:03:52.560 --> 00:03:57.240
the non artificial intelligence sense, that
is, the sense to the stone is
44
00:03:57.240 --> 00:04:05.840
embedded in concrete, other than the
process of embedding converting words into corresponding numerical
45
00:04:05.960 --> 00:04:13.599
values. Those topics were probably those
concepts will probably discussed during the panel's discussion
46
00:04:13.680 --> 00:04:19.759
today. So as emily indicated to
discuss the New York Times allegations of copyright
47
00:04:19.800 --> 00:04:27.680
infringement, we have three experts.
Zev Rosen is an assistant professor, School
48
00:04:27.680 --> 00:04:33.319
of Law at Southern Illinois University.
It was the Abraham Chemistine Scholar in Residence
49
00:04:33.399 --> 00:04:40.600
at the United States Copyright Office and
is particularly renowned for his expertise on copyright
50
00:04:40.720 --> 00:04:47.199
law, especially in its historical development. Charles One is an assistant professor at
51
00:04:47.240 --> 00:04:54.079
the American University College of Law.
It was previously a post doctoral fellow at
52
00:04:54.160 --> 00:05:00.439
Cornarial Tech. Working with professional James
Grimmleman. Charles focuses his research on the
53
00:05:00.439 --> 00:05:08.079
public effects of technology policy, intellectual
property law, and primarily patents and copyrights.
54
00:05:09.720 --> 00:05:15.079
Steve Tapp is a president and CEO
of Sentinel Worldwide, which is an
55
00:05:15.079 --> 00:05:20.959
intellectual property consultancy. He's also a
lecturer in law at the George Washington University
56
00:05:21.800 --> 00:05:29.240
School of law. More recently,
Steve has co founded rights Cliques, which
57
00:05:29.279 --> 00:05:33.560
is a suite of software tools for
independent creators to register, manage, and
58
00:05:33.639 --> 00:05:40.839
enforce their copyrights. As Emily indicated, we'll start with opening statements, will
59
00:05:40.879 --> 00:05:45.560
begin with Zeve's, then Charles and
followed by Steve. So with that,
60
00:05:46.560 --> 00:05:50.199
I'll turn it over to z.
All right, everyone, good afternoon or
61
00:05:50.279 --> 00:05:57.600
morning, depending where you are there
are. This causes really fascinating because we're
62
00:05:57.639 --> 00:06:04.800
finally getting at a host of important
issues regarding AI and the Internet, and
63
00:06:05.000 --> 00:06:11.240
what we call traditional producer is a
cultural works we of New York Times or
64
00:06:11.319 --> 00:06:15.920
many other that I think will probably
be honest cot tales. There's a couple
65
00:06:15.920 --> 00:06:19.319
of legal issues I want to highlight, and I'd love some questions to follow
66
00:06:19.399 --> 00:06:26.199
this once my colleagues over opening statements. The initial issue here is really direct
67
00:06:26.279 --> 00:06:30.560
infringement, which is to say,
by taking all of these New York Times
68
00:06:30.560 --> 00:06:36.959
stories, they are effectively copying them
or creating derivative works of them. A
69
00:06:38.000 --> 00:06:42.600
derivative work is a work which recasts, adapts, or transforms and existing work.
70
00:06:45.000 --> 00:06:48.399
The scope of derivative works right is
frankly not terribly well understood because it
71
00:06:48.639 --> 00:06:54.399
usually dovetails of copying and is typically
going to be one or the other.
72
00:06:55.319 --> 00:07:01.639
But that is a big part of
what is being claimed here. The resolution
73
00:07:01.759 --> 00:07:06.279
of it is really, I think
going to be a question of a is
74
00:07:06.319 --> 00:07:12.079
one of those activities happening. I
tend to think it is, and honestly,
75
00:07:12.480 --> 00:07:15.480
I'm not quite positive how much even
opening I'm sure, I'm sure we'll
76
00:07:15.519 --> 00:07:19.879
disputed, but I think that's probably
an easier case. But at least some
77
00:07:20.000 --> 00:07:24.759
copying is occurring, although there's details
of that that I'm sure some of my
78
00:07:24.800 --> 00:07:28.680
colleagues will discuss. But then fair
use is going to be a major issue
79
00:07:28.759 --> 00:07:35.160
there, particularly whether or not it's
transformative. What transformative means is really a
80
00:07:35.199 --> 00:07:41.839
hard question, but too recent.
The sort of formative case is the Acoff
81
00:07:42.160 --> 00:07:47.399
versus Campbell case about two live crews
Pretty woman, But I really think this
82
00:07:47.519 --> 00:07:53.800
case is going to be on fair
use, a conflict of the Google versus
83
00:07:53.800 --> 00:08:03.279
Oracle case, which found that the
Java program called Api eyes that their reimplementation
84
00:08:03.480 --> 00:08:09.160
or copying depending on your perspective,
I suppose into the droid operating system was
85
00:08:09.199 --> 00:08:13.920
fair use and transformative, and the
flip side Warhol versus Goldsmith holding that the
86
00:08:15.639 --> 00:08:20.680
Andy Warhol transformation into of Lyn Goldsmith's
photo of prints into an Andy Warhol silk
87
00:08:20.680 --> 00:08:28.439
screen of prints and then is licensing
to contiinast was not fair use and not
88
00:08:28.519 --> 00:08:33.720
transformative, So that's going to be
a bit of way. I suspect there
89
00:08:33.799 --> 00:08:39.279
was a motion to dismiss which is
pending, and that's all we have right
90
00:08:39.320 --> 00:08:41.799
now. We don't have, and
so fair use I don't think is ripe
91
00:08:41.879 --> 00:08:45.080
yet. I think you have to
wait for at leat summary judgment on that.
92
00:08:45.320 --> 00:08:50.559
But I kind of suspect if it
doesn't get dismissed and we have to
93
00:08:50.600 --> 00:08:56.679
get to that, we'll see on
that. You have to show direct infringement
94
00:08:56.720 --> 00:09:00.360
to get to any of the other
stuff except some of them more except for
95
00:09:00.879 --> 00:09:07.559
the last one for contributory infringement,
which is also alleged a lot of you
96
00:09:07.600 --> 00:09:09.799
us have to do, particularly with
Microsoft, which is not open AI,
97
00:09:11.759 --> 00:09:16.080
but as the complaint notes, is
kind of an alter ego, and we
98
00:09:16.159 --> 00:09:20.519
have all sorts of rights and open
AI, and we also own a lot
99
00:09:20.519 --> 00:09:26.600
of the IP. Contributory infringement is
defendant have knowledge of direction infringement and defendant
100
00:09:26.679 --> 00:09:37.519
materially contributing to that infringement from loaded
terms there knowledge knowledge you know question.
101
00:09:37.759 --> 00:09:41.600
We've had a conflict that happens a
lot in the context of the Internet and
102
00:09:41.639 --> 00:09:48.879
takedowns that whether knowing requires sort of
red flag or general knowledge. Where's a
103
00:09:48.919 --> 00:09:56.240
case of all the coximmunications and it
was very adequately flagging for users who repeat
104
00:09:56.279 --> 00:10:03.919
infringers. The details of that case
law are somehow twenty six years after the
105
00:10:03.960 --> 00:10:13.120
law was passed, still in flux
to a degree and material contribution. I
106
00:10:13.159 --> 00:10:16.240
think that that one is easier to
show here. I think it's going to
107
00:10:16.440 --> 00:10:18.440
hinge on knowledge, but I've been
wrong before. If that's the other key
108
00:10:18.480 --> 00:10:26.879
part vicarious liability. The case VET
I always teach when I talk about vicarious
109
00:10:26.919 --> 00:10:31.600
liability the case called Funavisia versus Cherry
Orchard, and it's a case VET held
110
00:10:31.600 --> 00:10:35.080
a out of fleet. I mean
it's contributory as well, but a case
111
00:10:35.080 --> 00:10:41.320
it held out of flea market that
was basically a hotbed of pirated media was
112
00:10:41.440 --> 00:10:45.759
engaging with vicarious liability because they had
a right to control the infringing activity.
113
00:10:45.799 --> 00:10:50.799
The infringing activity being the resale of
pirated discs installed at the flea market that
114
00:10:50.840 --> 00:10:54.960
were rented out and arriving a financial
commercial benefit from it, which was once
115
00:10:54.960 --> 00:11:00.519
again they were getting paid the fees
or restalls. Both of other I think
116
00:11:00.919 --> 00:11:07.919
are going to come out quite substantially. But other sort of pseudo IP issue
117
00:11:07.960 --> 00:11:11.440
I want to flag, if it's
really interesting, claims that New York Times
118
00:11:11.440 --> 00:11:18.919
owns Wirecutter, which is product recommendations, and they claim it is misappropriation because
119
00:11:18.960 --> 00:11:26.720
they are taking Wirecutter reviews and are
not getting affiliate payments. The large open
120
00:11:26.720 --> 00:11:31.720
AI are claiming this is preemptive by
a copyright lack which has at all legal
121
00:11:31.799 --> 00:11:37.840
or equitable rights that are within the
general scope of copyright are preempted. And
122
00:11:37.919 --> 00:11:41.639
on the other hand you have this
case is International News Service versus Associated Press,
123
00:11:41.679 --> 00:11:48.679
Supreme Court nineteen eighteen, holding the
misappropriation of hot news. In other
124
00:11:48.720 --> 00:11:52.600
words, news which was fresh out
of battle line the World War One was
125
00:11:52.799 --> 00:12:00.159
a misappropriation and a violation of common
law rights that was not preempted. The
126
00:12:00.200 --> 00:12:05.080
scope of VINS versus the AP nowadays
has been questioned for some cases have limited
127
00:12:05.080 --> 00:12:09.279
it, but it's not dead,
and so I think it's be interesting to
128
00:12:09.279 --> 00:12:13.600
see how that plays out. I
think, well, I'm looking forward to
129
00:12:13.759 --> 00:12:16.000
more of your questions. I think
we'll turn over Charles from more first head
130
00:12:16.039 --> 00:12:24.120
thoughts. All right, uh,
thanks thanks to me for that, you
131
00:12:24.159 --> 00:12:28.320
know, really excellent introduction to kind
of what the major issues are in this
132
00:12:28.440 --> 00:12:35.279
case and what the kind of key
doctrines of copyright and other law are that
133
00:12:35.360 --> 00:12:37.759
are at play. As you can
see, this case has a lot of
134
00:12:37.799 --> 00:12:39.879
things going on, and so you
know what I'll try to do is I'll
135
00:12:39.919 --> 00:12:43.399
try to number one, just go
through kind of what's been going on in
136
00:12:43.440 --> 00:12:46.600
the case specifically, so I'll talk
about particularly the motions to the SPISS that
137
00:12:46.639 --> 00:12:52.399
have been filed, and then take
a couple of guesses. As you mentioned,
138
00:12:52.679 --> 00:12:54.840
we don't know what the fair use
defense is going to look like,
139
00:12:54.879 --> 00:12:58.279
but I'll try to talk a little
bit about, you know what i'd expect,
140
00:12:58.159 --> 00:13:03.519
particularly, try to guess kind of
what I would imagine open AI would
141
00:13:03.559 --> 00:13:09.120
try to argue, and then maybe
leave with a couple of broader thoughts about
142
00:13:09.200 --> 00:13:15.639
kind of where this fits into the
larger debate over copyright and AI. So
143
00:13:15.960 --> 00:13:20.000
as far as the procedure of the
lawsuit goes, we started with a complaint
144
00:13:20.039 --> 00:13:24.919
in the lawsuit filed in December twenty
twenty three, and now we have on
145
00:13:24.960 --> 00:13:28.759
the table two motions to dismiss,
one from open Ai that was filed late
146
00:13:28.799 --> 00:13:35.159
February and one from Microsoft that was
filed about a week ago. These motions
147
00:13:35.159 --> 00:13:39.159
to dismiss don't go to the entirety
of the case, which is somewhat interesting.
148
00:13:39.200 --> 00:13:43.639
They only go to what the what
Microsoft and open Ai describe as sort
149
00:13:43.639 --> 00:13:48.080
of ancillary issues, and so the
ones that they talk about, in particular
150
00:13:48.200 --> 00:13:52.039
the contributory liability issue, they say
that there's a lack of knowledge, as
151
00:13:52.159 --> 00:13:54.879
Via mentioned, so they say that
that one should be dismissed. The common
152
00:13:54.960 --> 00:14:00.759
law misappropriation issue that's you mentioned this
hot news I in S is AP case
153
00:14:01.000 --> 00:14:03.639
kind of theory. They say that
that's preempted by the Copyright Act, that
154
00:14:03.720 --> 00:14:11.440
copyright protection being a federal law overrides
what the states, what state protection is
155
00:14:11.480 --> 00:14:13.559
given there, and they point to
know a number of cases that show that
156
00:14:13.600 --> 00:14:18.519
the I S versus AP doctrine is
kind of begrudgingly accepted at this point.
157
00:14:18.600 --> 00:14:22.919
That's kind of their argument for that. John mentioned also that there was this
158
00:14:22.000 --> 00:14:28.399
Digital Millennium Copyright Act count in the
complaint, and that one is actually somewhat
159
00:14:28.399 --> 00:14:33.320
interesting. So I'll just I'll go
into that just a little bit. Some
160
00:14:33.399 --> 00:14:37.399
of you may be familiar with the
Digital Millennium Copyright Act in terms of its
161
00:14:37.440 --> 00:14:43.279
anti circumvention provisions, the rules that
say that you're not allowed to kind of
162
00:14:43.320 --> 00:14:46.519
break digital rights management. This case
actually deals with a different part of the
163
00:14:46.600 --> 00:14:52.919
DMCA, Section twelve oh two,
which relates to copyright management information, so
164
00:14:52.039 --> 00:15:00.639
basically the inclusion of metadata like authors
or titles or copyright notices inside files,
165
00:15:00.679 --> 00:15:07.840
typically digital files. And so The
Times argues that in training these generative AI
166
00:15:09.039 --> 00:15:15.279
systems, Open AI and Microsoft removed
that information and as a result, violated
167
00:15:15.320 --> 00:15:18.320
section twelve h two. The difficulty
that they're going to face improving that,
168
00:15:18.559 --> 00:15:22.320
as Microsoft and open AI point out, is that in order to show a
169
00:15:22.399 --> 00:15:26.679
violation of the section, you have
to show what's called a double cienter requirement.
170
00:15:28.240 --> 00:15:31.559
Number one that's open AI knew that
it was removing and number two that
171
00:15:31.639 --> 00:15:35.919
it knew that the result of removal, or at least should have known that
172
00:15:35.960 --> 00:15:41.399
the result of removal would be further
infringement. And so open and Microsoft argue
173
00:15:41.480 --> 00:15:45.559
that number one, there's no efforden
is that this information was actually removed during
174
00:15:45.600 --> 00:15:48.559
the training process, but number two
that they won't be able to satisfy the
175
00:15:50.000 --> 00:15:54.279
knowledge requirements. They also raise a
time bar question. They say that there's
176
00:15:54.320 --> 00:16:02.440
a three year period look back period
that limits the extent that the copyright allegations
177
00:16:02.480 --> 00:16:06.080
can go back. That's actually a
case that's being considered by the Supreme Court
178
00:16:06.159 --> 00:16:10.399
right now. It was just argued
a couple of weeks ago, and so
179
00:16:10.799 --> 00:16:14.840
that's just another argument that they bring
up. But this again doesn't get to
180
00:16:14.919 --> 00:16:19.480
the substitutive questions that you mentioned,
the questions of whether or not there actually
181
00:16:19.720 --> 00:16:26.200
is copyright infringement in the training of
these systems. Using the New York Times
182
00:16:26.200 --> 00:16:33.360
and other articles, there are three
points that the Times identifies as where the
183
00:16:33.399 --> 00:16:37.639
infringement could occur. Number One,
they say that the collection of the articles
184
00:16:37.720 --> 00:16:41.960
to make the training data use to
train these AI systems, that was an
185
00:16:42.000 --> 00:16:45.720
infringement because you were making a lot
of copies in order to collect them.
186
00:16:45.120 --> 00:16:51.919
Second, they allege that the model
itself all of the data parameters I think
187
00:16:51.960 --> 00:16:59.480
one point seven trillion numbers that make
up the the GPT systems that's somewhere embedded
188
00:16:59.519 --> 00:17:03.879
in there is all of the information
necessary to replicate a lot of the articles,
189
00:17:03.879 --> 00:17:07.480
and therefore the model itself is a
potential infringement. And third, they
190
00:17:07.519 --> 00:17:12.119
say that when you use the model
in such a way that it generates that
191
00:17:12.559 --> 00:17:17.200
generates infringing content, that that use
is sort of a public performance. It
192
00:17:17.240 --> 00:17:23.160
allows you to get the information out
in order to show that these in order
193
00:17:23.200 --> 00:17:26.680
to show infringement, what the Times
would have to show is number one,
194
00:17:26.680 --> 00:17:30.799
that this is copyrightable subject matter.
There are some interesting questions there because you
195
00:17:30.839 --> 00:17:33.920
know, a lot of the information
that's being drawn is factual, and so
196
00:17:33.079 --> 00:17:37.799
maybe there'll be that issue that comes
out. Generally factual information is not considered
197
00:17:37.839 --> 00:17:41.000
copyrightable, but you know, like
I said, that's probably not going to
198
00:17:41.039 --> 00:17:45.440
be the lead argument. There also
are questions of what exactly counts as an
199
00:17:45.480 --> 00:17:48.839
infringing act. You know, is
something that's internally inside the model that nobody
200
00:17:48.880 --> 00:17:52.680
could actually see or understand, is
that an infringement. There's actually sort of
201
00:17:52.680 --> 00:17:56.119
an interesting question about that. But
again, the large issue that open ai
202
00:17:56.200 --> 00:18:00.720
and Microsoft intend to raise, in
fact, they say that they are they're
203
00:18:00.759 --> 00:18:06.839
actually very excited they say in their
in their motions of dismiss to litigate this
204
00:18:06.920 --> 00:18:11.799
issue is the fair use question,
assuming that it is an act of copying
205
00:18:11.880 --> 00:18:14.839
or it is a derivative work to
do all of those things I just mentioned,
206
00:18:15.400 --> 00:18:18.759
does the fair use doctrine permit it? And so courts have used the
207
00:18:18.799 --> 00:18:22.240
fair use doctrine a variety of situations, Google versus Oracle and the software context.
208
00:18:22.759 --> 00:18:27.480
On the Warhol case that was an
artistic use. On the Campical case
209
00:18:27.480 --> 00:18:30.240
that was parody. Oh, fairiuse
is sort of this jack of all trades
210
00:18:30.279 --> 00:18:36.200
doctrine ends up being used in all
sorts of places. In that line,
211
00:18:36.559 --> 00:18:38.759
there are a number of cases that
will help open AI quite a bit,
212
00:18:40.039 --> 00:18:44.920
although possibly to a limited extent.
There was a case over Google Images where
213
00:18:44.960 --> 00:18:48.039
Google had collected a bunch of images
and was displaying them using an image search
214
00:18:48.079 --> 00:18:53.720
engine. Courts said that because Google
had really downsampled them and they didn't really
215
00:18:53.759 --> 00:18:59.839
serve as replacement, that database of
images was fair use. There's a case
216
00:18:59.839 --> 00:19:04.160
called Eye Paradigms in which a company
made a plagiarism detection program and there was
217
00:19:04.160 --> 00:19:08.160
a question of whether or not the
inputs to the plagiarism detection program, which
218
00:19:08.200 --> 00:19:12.559
were basically the essays that were being
detected, whether or not that was an
219
00:19:12.599 --> 00:19:15.680
infringe And again the court said,
you know, this is sort of a
220
00:19:15.720 --> 00:19:22.039
new tool, the actual articles aren't
retrievable. Similarly, with the Google Books
221
00:19:22.119 --> 00:19:25.440
case, it was alleged that Google
scanning of a bunch of books to create
222
00:19:25.480 --> 00:19:30.880
the Google Book search engine was copyright
infringe and again a court said, I
223
00:19:30.880 --> 00:19:34.880
think the second Circuit said that this
was fair use on the grounds that the
224
00:19:36.039 --> 00:19:40.759
mass scanning of books to provide a
service that didn't really replicate the value of
225
00:19:40.759 --> 00:19:48.480
the books themselves was allowable and as
a result not a copyright infringement. Courts
226
00:19:48.559 --> 00:19:53.279
usually applied this four factor test in
which they look at the the nature of
227
00:19:53.319 --> 00:20:00.240
the copyrighted work, the nature of
the use, the purpose in character of
228
00:20:00.279 --> 00:20:03.920
the use, the amount that was
used. And finally, and probably this
229
00:20:03.039 --> 00:20:07.160
is the most important factor by many
measures, the economic impact on the market
230
00:20:07.200 --> 00:20:11.359
for the original copyrighted work. And
I think that that's going to be the
231
00:20:11.359 --> 00:20:15.640
most interesting one to follow in this
case, because, on the one hand,
232
00:20:15.079 --> 00:20:22.440
these are incredibly valuable systems. Right
Artificial intelligence has huge potential in terms
233
00:20:22.599 --> 00:20:30.559
of business uses, commercial uses,
uses for individual consumers. It can be
234
00:20:30.680 --> 00:20:37.480
used as the platform for many other
technologies. Is that part of the market
235
00:20:37.839 --> 00:20:42.200
that inheres in the copyright that The
New York Times has in all of its
236
00:20:42.279 --> 00:20:45.240
articles, or that a novelist has
in all the novels that are used in
237
00:20:45.279 --> 00:20:48.359
training. Right, the novelists or
the New York Times, they would say,
238
00:20:48.480 --> 00:20:52.200
yes, the point of our articles
it is to provide information, and
239
00:20:52.279 --> 00:20:57.079
that information is being used to provide
a valuable service through things like SHATGBT,
240
00:20:57.400 --> 00:21:03.039
and as a result, there should
be some cut of that open AI of
241
00:21:03.079 --> 00:21:04.960
course, would argue the other way
around. They would say, look,
242
00:21:06.000 --> 00:21:10.279
this is a completely different service.
It doesn't serve as a replacement for the
243
00:21:10.319 --> 00:21:15.359
original articles. It's transformative in the
way that Svia mentioned, and transformative something
244
00:21:15.400 --> 00:21:18.960
that the courts have really looked at, so that I think is going to
245
00:21:18.960 --> 00:21:22.400
be sort of the parameters of debate. We do have these cases about mass
246
00:21:22.440 --> 00:21:26.519
text and data mining which are different
from this case, but you know,
247
00:21:26.039 --> 00:21:32.319
provide some basis for understanding where fair
use goes and that transformativelopment and what effects
248
00:21:33.119 --> 00:21:38.160
the availability of these systems has on
those copyrighted works. I think is going
249
00:21:38.200 --> 00:21:47.279
to be really important and something really
to watch as this case progresses. All
250
00:21:47.359 --> 00:21:52.200
right, I think that makes it
my turn. So let me begin by
251
00:21:52.200 --> 00:21:56.640
saying thank you to the Federalist Society
for inviting me today, and Emily for
252
00:21:56.759 --> 00:22:00.799
organizing this panel, and of course
the On for his kind introduction, and
253
00:22:00.839 --> 00:22:06.440
my fellow panelists for their opening remarks. Let me note that my remarks are
254
00:22:06.519 --> 00:22:10.640
my own and do not necessarily reflect
the views of any client or employer.
255
00:22:11.440 --> 00:22:15.759
I want to begin by putting this
case on others like it into a broader
256
00:22:15.799 --> 00:22:19.440
perspective. Those of us who've been
working in copyright law and policy over the
257
00:22:19.440 --> 00:22:25.039
past thirty or so years, and
I'm afraid my gray hair gives that away,
258
00:22:26.480 --> 00:22:30.079
have seen history repeat itself over and
over. First, a new technology
259
00:22:30.079 --> 00:22:34.920
comes along and makes it easier than
ever to copyright, to obtain copyrighted works.
260
00:22:36.480 --> 00:22:41.759
Now, in the ideal scenario that
development is mutually beneficial, creators can
261
00:22:41.799 --> 00:22:47.680
reach new audiences expanded audiences more easily, and the widespread availability of creative works
262
00:22:47.759 --> 00:22:53.160
drives demand for the technology. Everybody
wins. But in practice, the operators
263
00:22:53.200 --> 00:22:59.160
of the technology in the past have
often made choices that allocate to themselves the
264
00:22:59.200 --> 00:23:03.039
lions share of the income. In
some cases, those choices have included willfully
265
00:23:03.119 --> 00:23:10.319
tolerating infringement on platforms, knowing full
well that creators, especially independent creators,
266
00:23:10.880 --> 00:23:15.680
like both the means and tools to
achieve meaningful vindication of their rights. So
267
00:23:15.839 --> 00:23:22.359
here we are with generative AI systems
built on large language models. It's deja
268
00:23:22.400 --> 00:23:26.680
vu all over again. Such systems
require massive volumes of works in order to
269
00:23:26.680 --> 00:23:33.240
be capable of producing the commercially valuable
outputs the designers seek to market. That
270
00:23:33.319 --> 00:23:37.400
fact is not in dispute, but
how those works are obtained is a commercial
271
00:23:37.480 --> 00:23:41.759
choice. There is nothing in the
nature of the technology that requires those works
272
00:23:41.799 --> 00:23:48.440
to be scraped without notice, without
authorization, or without compensation. Yet that's
273
00:23:48.480 --> 00:23:53.680
precisely what's happened. In the public
policy sphere, people willing to defend those
274
00:23:53.720 --> 00:23:59.880
decisions often try to create a false
dichotomy between the massive, unauthorized scraping and
275
00:24:00.039 --> 00:24:04.079
the existence of generative AI. The
reality is that there are companies that have
276
00:24:04.160 --> 00:24:11.279
built generative AI systems unlicensed materials.
Those that choose to do otherwise are not
277
00:24:11.359 --> 00:24:15.920
engaged in a crusade for the betterment
of humanity. They are commercial enterprises trying
278
00:24:15.920 --> 00:24:19.960
to avoid paying for critical inputs.
As the chairwoman of the Federal Trade Commission
279
00:24:21.000 --> 00:24:26.200
recently said very plainly, firms cannot
use claims of innovations as an excuse for
280
00:24:26.319 --> 00:24:32.279
law breaking. So let me turn
to the particular legal issues in this case,
281
00:24:32.359 --> 00:24:37.400
and I'm going to focus on the
direct copyright infringement issues. One would
282
00:24:37.440 --> 00:24:41.680
think that in a circumstance when computers
were and our programmed to crawl the Internet
283
00:24:41.720 --> 00:24:47.480
and copy literally billions of works,
the largest copying effort in history, that
284
00:24:47.559 --> 00:24:51.440
it would be beyond serious contention that
the reproduction right of those works has been
285
00:24:51.480 --> 00:24:56.599
implicated, And yet open AI and
others are trying to put that exact matter
286
00:24:56.640 --> 00:25:00.960
in dispute. Anyone who is even
a passing understanding of how computers operate knows
287
00:25:00.960 --> 00:25:04.920
that computers must make copies in order
to process what has been input into them,
288
00:25:06.960 --> 00:25:10.000
and the New York Times evidence shows
that by inputting a certain set of
289
00:25:10.000 --> 00:25:14.960
prompts, substantially, if not strikingly, similar copies of their original works will
290
00:25:15.000 --> 00:25:18.759
be output by open AI, So
it seems self evident that copies of the
291
00:25:18.799 --> 00:25:22.000
original works must be in the computer
memory in order for that to happen.
292
00:25:23.599 --> 00:25:29.279
Still, open AI argues to the
contrary, if the internal operation of the
293
00:25:29.279 --> 00:25:33.119
system were transparent to the public,
we would have real insight into the facts
294
00:25:33.119 --> 00:25:37.480
of how it operates. But despite
the name, it was given Open AI,
295
00:25:37.559 --> 00:25:41.240
like other generative AI systems, is
in fact locked up tight. Perhaps
296
00:25:41.279 --> 00:25:45.839
some of this will come out in
discovery, but in any event, I
297
00:25:45.880 --> 00:25:51.240
am deeply skeptical that there's any serious
argument other than that the unauthorized scraping of
298
00:25:51.279 --> 00:25:56.720
copyrighted works does implicate the reproduction right, which means, as my fellow panelists
299
00:25:56.720 --> 00:26:00.839
have already said, the real action
in this case be in the fair use
300
00:26:00.960 --> 00:26:04.920
argument. As has already been noted, the fair use assessment in the context
301
00:26:04.920 --> 00:26:11.240
of the first factor is likely to
involve consideration of whether opening eyes copying constitutes
302
00:26:11.279 --> 00:26:17.359
a transformative use. While the term
transformative has been part of copyright jurisprudence for
303
00:26:17.359 --> 00:26:21.799
a very long time, it was
given special significance by the Supreme Court decision
304
00:26:21.880 --> 00:26:26.400
in the two Life Crew case Campbell
Vis's A Cuff Rose to give the proper
305
00:26:26.400 --> 00:26:30.480
caption in nineteen ninety four. Since
that time, lower courts has struggled to
306
00:26:30.519 --> 00:26:36.480
apply this term, sometimes resulting in
extreme results, such as when Google's forbatim
307
00:26:36.519 --> 00:26:40.920
copying of tens of millions of books
was held to be highly transformative by the
308
00:26:40.920 --> 00:26:45.599
Second Circuit. Fortunately, the Supreme
Court had occasion to revisit this doctrine in
309
00:26:45.640 --> 00:26:52.000
twenty twenty two in Warhol Foundation versus
Goldsmith and articulated a much more reasonable and
310
00:26:52.039 --> 00:26:56.839
workable framework. So I think the
lower court fair use decisions that pre date
311
00:26:56.920 --> 00:27:03.599
Goldsmith are now of questionable applicability.
The first factor begins with a contrast between
312
00:27:03.680 --> 00:27:07.519
nonprofit use, which is favored,
and commercial use, which is disfavored.
313
00:27:07.200 --> 00:27:11.400
Prior to Goldsmith, some courts were
finding that a transformative use not only negated
314
00:27:11.400 --> 00:27:17.880
the commerciality but essentially overtook all the
other fair use factors as well. But
315
00:27:17.960 --> 00:27:21.240
a Goldsmith, the Supreme Court was
much more measured, holding the weight of
316
00:27:21.240 --> 00:27:26.319
commerciality against fair use can be lessened
by the degree to which the use is
317
00:27:26.319 --> 00:27:30.200
transformative, that is, has a
further purpose or different character. It's a
318
00:27:30.200 --> 00:27:37.119
sliding scale, not a Boolean analysis. So what constitutes transformative use? Some
319
00:27:37.200 --> 00:27:41.480
courts had gone so far in finding
any new element of the use to be
320
00:27:41.519 --> 00:27:45.960
transformative that many commentators wondered what if
anything, was left of the statutory right
321
00:27:47.200 --> 00:27:52.640
to authorize the creation of derivative works. As they mentioned, the Goldsmith Court
322
00:27:52.680 --> 00:27:56.720
wrote, to make transformative use of
an original must go beyond that required to
323
00:27:56.759 --> 00:28:03.039
qualify as a derivative use that has
a distinct purposes justified because it furthers the
324
00:28:03.039 --> 00:28:06.880
goal of copyright, namely to promote
the progress of science and the arts,
325
00:28:06.920 --> 00:28:14.519
without diminishing the incentive to create.
Quoting author's gildmersus Google, the court continued,
326
00:28:14.960 --> 00:28:18.799
the more the appropriator is using the
copied material for new transformative purposes,
327
00:28:19.079 --> 00:28:23.000
the less likely it is that the
appropriation will serve as a substitute for the
328
00:28:23.039 --> 00:28:29.200
original or its plausible derivatives. So
we see that substitution is a key part
329
00:28:29.200 --> 00:28:33.079
of the analysis, one that has
the power to disqualify use from transformative status.
330
00:28:33.880 --> 00:28:40.279
And the Times has shown that OpenAI
is capable of preproducing substantially similar copies
331
00:28:40.319 --> 00:28:45.000
of its works. That's powerful evidence
on the substitution effect that AI will have
332
00:28:45.039 --> 00:28:52.640
to contend with when this case moves
forward. The Supreme Court in Goldsmith summarized
333
00:28:52.640 --> 00:28:56.480
its first factor analysis with the passage
quote. The first factor considers whether the
334
00:28:56.599 --> 00:29:00.519
use of a copyrighted work has a
further purpose or different character, which is
335
00:29:00.559 --> 00:29:04.519
a matter of degree, and the
degree of difference must be balanced against the
336
00:29:04.559 --> 00:29:10.079
commercial nature of the use quote quote. Those who are sympathetic to the AI
337
00:29:10.200 --> 00:29:12.359
side will point to the creation of
new works by the use of the resulting
338
00:29:12.400 --> 00:29:18.680
system. But this misses a key
step. Generative AI systems do not create
339
00:29:18.720 --> 00:29:22.720
images or text or anything on their
own initiative. Human interaction is required in
340
00:29:22.759 --> 00:29:27.119
the form of prompts, and without
that human interaction, the resulting product would
341
00:29:27.160 --> 00:29:33.319
have no protectable creative expression. So
it isn't the AI system that generates new
342
00:29:33.319 --> 00:29:37.839
works. It is the human users. And it's been good law for a
343
00:29:37.960 --> 00:29:41.599
very long time that the copyer of
works may not stand in the shoes of
344
00:29:41.640 --> 00:29:47.319
the users of those copies to justify
the initial copying. That, of course,
345
00:29:47.359 --> 00:29:51.640
is Michigan Document Services versus Princeton University
Press. We can talk about that
346
00:29:51.680 --> 00:29:55.759
if it comes up in the Q
and A. As for the remaining fair
347
00:29:55.880 --> 00:29:59.319
use factors, it seems to me
they all favor the copyright owner in varying
348
00:29:59.359 --> 00:30:03.519
degrees. The works were largely creative, although perhaps some news articles marginally less
349
00:30:03.559 --> 00:30:08.119
so still I think it cleared.
The second factor favors the Times. The
350
00:30:08.279 --> 00:30:14.119
entirety of every work was copied,
so the third factor is strongly against fair
351
00:30:14.240 --> 00:30:17.920
use, and the fourth factor,
harmed to the current or potential market,
352
00:30:18.039 --> 00:30:22.799
very strongly favors the Times. Again, substitution is probably the strongest evidence of
353
00:30:22.839 --> 00:30:26.079
harm, followed closely by the existence
of a licensing market for the same use,
354
00:30:26.799 --> 00:30:30.799
both of which we have here.
So if open AI is going to
355
00:30:30.799 --> 00:30:33.920
prevail on fair use, it seems
to me they will need a very strong
356
00:30:33.960 --> 00:30:37.559
finding on the first factor in order
to overcome not only the commercial nature of
357
00:30:37.599 --> 00:30:41.200
the use, but all the other
factors as well. On the law,
358
00:30:41.240 --> 00:30:45.759
this seems unlikely. I can only
imagine a finding of fair use frankly if
359
00:30:45.759 --> 00:30:49.920
the court is taken in by the
cool factor arguments. But it would be
360
00:30:49.960 --> 00:30:55.960
a travesty for creators works to lose
essentially all protection in the aihe and it
361
00:30:56.000 --> 00:31:00.000
would certainly harm the incentive to create, which is what the Goldsmith Court kept
362
00:31:00.079 --> 00:31:07.680
certainly in mind. Thanks, we
have a couple of questions before I go
363
00:31:07.799 --> 00:31:11.079
there. I think you addressed this, Steve, and that is the substitution,
364
00:31:11.319 --> 00:31:18.839
the competitiveness in the market, and
in particular that particular fact which the
365
00:31:19.160 --> 00:31:23.640
complaint I believe goes to some links
to try to create a competitive market.
366
00:31:23.920 --> 00:31:30.599
And how does that get them out
of or diminish the Google Books copying case.
367
00:31:33.599 --> 00:31:37.720
Well, I'm not sure the degree
to which Google Books is still good
368
00:31:37.799 --> 00:31:42.559
law after Goldsmith. Not the Goldsmith
overturned that decision in any way, but
369
00:31:42.799 --> 00:31:51.440
the analysis was so different that it's
hard to make that case any longer.
370
00:31:51.480 --> 00:31:57.880
That the Google books really is a
useful president when you look at this at
371
00:31:57.920 --> 00:32:04.400
ai generative systems in the most general
way and kind of departing for a moment
372
00:32:04.440 --> 00:32:07.680
from the particulars of the law,
but in a matter of social policy,
373
00:32:08.359 --> 00:32:17.920
we have systems where people's work product
was copied without their knowledge, permission,
374
00:32:19.640 --> 00:32:25.880
or compensation, and it's being used
to design and implement technology that will then
375
00:32:25.960 --> 00:32:30.839
put many of those people out of
business entirely or substantially reduce their income by
376
00:32:30.880 --> 00:32:37.200
competing directly with them. There's also
questions about, you know, well,
377
00:32:37.200 --> 00:32:40.839
why did you copy this work versus
that work? Does any individual work have
378
00:32:42.279 --> 00:32:49.559
particular value? Works were chosen because
they are valuable for what they are,
379
00:32:50.559 --> 00:32:54.599
and so there is value there.
And I think at the highest level of
380
00:32:55.079 --> 00:33:02.519
public policy analysis, value was taken
and value has been created in companies that
381
00:33:02.559 --> 00:33:12.160
are already valuing themselves at over trillion
dollars, and there's got to be a
382
00:33:12.200 --> 00:33:19.960
way to compensate the people whose works
were taken to build those systems. It's
383
00:33:20.000 --> 00:33:23.319
beyond this panel talk about exactly how
that works, going back to the particulars
384
00:33:23.359 --> 00:33:28.839
of the copyright law, the fact
that this goes back to get to the
385
00:33:28.880 --> 00:33:34.160
reproduction right. It seems self evident
to me that copies were made, and
386
00:33:34.200 --> 00:33:39.319
I've just made the fair use analysis
that I think is appropriate. So I'm
387
00:33:39.640 --> 00:33:45.039
I wouldn't be surprised to see the
times prevail in this case. But time
388
00:33:45.079 --> 00:33:52.160
will tell. So we have a
question that goes to ask quickly on I'm
389
00:33:52.200 --> 00:33:55.960
sorry, go ahead, Yes,
I just wanted to say I think it's
390
00:33:55.960 --> 00:34:00.960
slight different tack on Google books.
I think it's very much of a piece
391
00:34:01.000 --> 00:34:04.720
of Google image Search, and I
think it has to be read as fact
392
00:34:04.759 --> 00:34:07.480
specific. I mean, whether or
not it's good law after Warhol, I
393
00:34:07.519 --> 00:34:14.079
don't know, but I really think
that the focus on stipid view and directing
394
00:34:14.119 --> 00:34:19.079
people to booksellers and gootle books,
and the focus on directing people to websites
395
00:34:19.079 --> 00:34:23.039
and Google image search with with key
too decisions in terms of a market effect
396
00:34:23.079 --> 00:34:27.679
and transformation. If the use was
not a substitute in any way, at
397
00:34:27.760 --> 00:34:31.320
least in the courts view, but
rather a channeling use of sorts and so
398
00:34:31.440 --> 00:34:37.039
I do think it has to be
a difference there where there's no channeling whatsoever
399
00:34:37.199 --> 00:34:40.679
here. But like Steve said,
we'll see, well so if I can
400
00:34:40.800 --> 00:34:45.199
just kind of jump in a very
quickly. So one of the arguments that
401
00:34:45.480 --> 00:34:49.719
is sort of previewed in the in
the motions of the Smiths is this question
402
00:34:49.840 --> 00:34:58.440
of to what extent the actual substitutional
use is actually relevant. And so open
403
00:34:58.480 --> 00:35:02.239
Aye and Microsoft both point exhibit jay
of the complaint, which is where you
404
00:35:02.280 --> 00:35:07.320
get sort of the original duplication or
memorization as they call it, in the
405
00:35:07.320 --> 00:35:14.079
industry of articles. And as those
examples show, in order to get a
406
00:35:14.519 --> 00:35:19.559
in order to prompt chatterp to produce
a New York Times article, you first
407
00:35:19.559 --> 00:35:22.280
have to give it the first half
the article, essentially. And so their
408
00:35:22.400 --> 00:35:24.960
argument then is that at that point, if you already have the first half
409
00:35:24.960 --> 00:35:28.679
of the article, it's not much
of a substitution to say that you get
410
00:35:28.679 --> 00:35:30.599
the second half, especially when the
second half is available in other places on
411
00:35:30.639 --> 00:35:35.960
the Internet. And so that's an
interesting argument. I haven't heard that before,
412
00:35:36.000 --> 00:35:40.199
and it's sort of unique to this
sort of situation, but I think
413
00:35:40.199 --> 00:35:44.360
that what it does mean is that
it's going to be a little more complicated,
414
00:35:44.360 --> 00:35:49.360
I think to prove the substitution element
compared to some of the other cases.
415
00:35:49.400 --> 00:35:52.199
That doesn't mean that, you know, I think that's it's a slamm
416
00:35:52.199 --> 00:35:53.960
don win for them. It certainly
is the case that you can get a
417
00:35:54.039 --> 00:35:59.880
large chunk of copyright material. It's
just that it's not the normal way that
418
00:36:00.039 --> 00:36:05.280
one would expect to retrieve it.
There's a lot of focus on that this
419
00:36:05.440 --> 00:36:12.079
is a bug rather they want it
would legally shouldn't be significant, but has
420
00:36:12.119 --> 00:36:15.840
come up before, so who knows. Yeah, my understanding is that there
421
00:36:15.880 --> 00:36:22.079
also is an amount of research on
just how to stop at a computer level
422
00:36:22.360 --> 00:36:23.599
of the sort of thing from happening. You know, is there a way
423
00:36:23.599 --> 00:36:28.400
that they can just put filters at
the end, or can they use some
424
00:36:28.440 --> 00:36:31.840
sort of special training processes. That
is sort of an interesting research question that
425
00:36:31.880 --> 00:36:35.400
I know is open. I think
that the difficulty is that a lot of
426
00:36:35.440 --> 00:36:39.719
times articles are available online and slightly
different variations, like somebody might admit a
427
00:36:39.719 --> 00:36:46.039
paragraph, and that's why it's hard
for it's hard to simply detect identical copies.
428
00:36:46.800 --> 00:36:49.719
But it's sort of interesting that,
you know, there may be a
429
00:36:49.800 --> 00:36:52.920
technical solution in addition to the legal
approach, and there's a question of which
430
00:36:52.960 --> 00:36:55.440
one you think should lead. Right. On the one hand, maybe we
431
00:36:55.480 --> 00:36:59.719
want to say that the law encourages
people to develop these sorts of technologies that
432
00:37:00.000 --> 00:37:02.840
on void duplication. On the other
hand, maybe we say, maybe we're
433
00:37:02.840 --> 00:37:07.239
worried that by imposing the hammer of
the law too early, we prevent people
434
00:37:07.280 --> 00:37:13.360
from coming up with these sorts of
technologies that can satisfy better middle grounds,
435
00:37:13.400 --> 00:37:15.760
and we end up, you know, one of the one of the remedies
436
00:37:15.800 --> 00:37:20.960
that's asked for is destruction of chat
GPT, and so there's you know,
437
00:37:21.000 --> 00:37:23.840
there is a question of whether or
not it would be premature to be applying
438
00:37:23.880 --> 00:37:29.480
this before we really know what the
sort of opportunities the technology ultimately look like.
439
00:37:30.599 --> 00:37:34.119
So that's what I was referencing earlier
when I said that some people want
440
00:37:34.159 --> 00:37:37.800
to look at this as copyright versus
AI, and you can't have both.
441
00:37:38.880 --> 00:37:45.639
Had had open Ai chosen to license
the materially used, that no one would
442
00:37:45.679 --> 00:37:51.199
be asking to destroy chat GPT,
nor is anyone asking to destroy AI systems
443
00:37:51.239 --> 00:37:57.239
that were built on licensed material.
So putting that aside for a moment,
444
00:37:57.880 --> 00:38:02.559
the substitute effect had several layers to
it. First and foremost, the fact
445
00:38:02.559 --> 00:38:09.840
that you can get that material out
of chat GPT is strong evidence that the
446
00:38:10.119 --> 00:38:15.000
copyrightable material is somewhere in the chat
GPT system, the open AI system,
447
00:38:15.320 --> 00:38:22.280
and therefore the reproduction right is implicated. Then on the question of whether the
448
00:38:22.320 --> 00:38:29.679
output is infringing and the substitution effect. There we kind of have two or
449
00:38:29.719 --> 00:38:37.159
three different layers in traditional copyright analysis. Obviously, substantial similarity is an access
450
00:38:37.519 --> 00:38:46.960
are the two problems to proving copying
and infringing reproduction. So the fact that
451
00:38:47.000 --> 00:38:53.920
it may be difficult or unusual to
get chat GPT to put that out doesn't
452
00:38:53.960 --> 00:38:58.599
mean it's less infringing. It just
means it may take a little more work,
453
00:38:58.639 --> 00:39:05.400
but it's still infringing. More generally, more, there's a question about
454
00:39:05.920 --> 00:39:15.239
the degree to which these generative AI
systems will overall all diminish demand for the
455
00:39:15.280 --> 00:39:21.400
copyrighted works produced by the people who
produce the works that were used to build
456
00:39:21.440 --> 00:39:29.039
the system, and that admittedly departs
somewhat from a direct infringement analysis to more
457
00:39:29.079 --> 00:39:34.840
public policy question of is it right
that the people whose works were taken without
458
00:39:34.920 --> 00:39:38.800
notice, without permission, and without
compensation should be told, sorry, you're
459
00:39:38.840 --> 00:39:45.199
out of lock, and these new
systems that we built with all the things
460
00:39:45.239 --> 00:39:53.719
we took from you are now going
to put you out of business. We
461
00:39:53.760 --> 00:40:00.320
have a question from the audience concerning
generally to the panelists together, and the
462
00:40:00.400 --> 00:40:07.360
specific question is relevance of sega versus
accolade and the concept of intermediate cop copying.
463
00:40:08.360 --> 00:40:14.559
Can any of you guys speak to
that? Yeah, so this is
464
00:40:14.719 --> 00:40:19.800
also I was mentioning that there is
this sort of question of whether or not
465
00:40:20.239 --> 00:40:27.000
sort of like purely internal use would
count as an infringement. And there are
466
00:40:27.000 --> 00:40:31.639
a couple of cases that say that
at least that's an interesting question. So
467
00:40:30.400 --> 00:40:37.639
so the say atholate case and the
reverse engineering cases they raise that question.
468
00:40:37.639 --> 00:40:42.760
I think you probably remember this is
the cartoon network case, right, which
469
00:40:42.800 --> 00:40:46.639
talks about transitory copying. So you
have a number of you have a number
470
00:40:46.639 --> 00:40:52.480
of cases that support the sort of
idea that maybe like internal copying or copying
471
00:40:52.519 --> 00:41:00.719
that is precedent to some other public
use might not receive the same analogy assists.
472
00:41:00.480 --> 00:41:04.960
And so that's obviously going to be
relevant here because you know to the
473
00:41:05.039 --> 00:41:07.920
extent that the allegation is that the
copying is inside the model. The model
474
00:41:07.960 --> 00:41:14.239
is not something anybody can see,
and so that might fall within those doctrines.
475
00:41:14.239 --> 00:41:17.159
Similarly, the training process, again
that's not something anybody actually sees,
476
00:41:17.719 --> 00:41:23.639
and as a result, those aren't
those potentially aren't necessarily the sorts of things
477
00:41:23.639 --> 00:41:30.159
that are of concern to copyright law. In view of those cases, it's
478
00:41:30.199 --> 00:41:32.119
really kind of you know, what's
going on the output end that has the
479
00:41:32.119 --> 00:41:36.679
commercial impact. That's where the fair
use question is going to really come up.
480
00:41:36.760 --> 00:41:38.760
That's where the market effect is going
to come up. That's where transformativeness
481
00:41:38.840 --> 00:41:44.079
is going to come out. Yeah, that that that's that that's at least
482
00:41:44.079 --> 00:41:50.199
part of kind of where where those
cases fit in. Go ahead, see
483
00:41:50.239 --> 00:41:52.960
I will say, I mean my
view of both Accolae case and Cartoon Network
484
00:41:53.039 --> 00:41:59.800
that they boiled down to if you're
doing some internal copying for what otherwise illegally
485
00:42:00.199 --> 00:42:02.880
and non copying use. Now it
will say where Cartoon Network with plenty of
486
00:42:02.920 --> 00:42:08.840
copying, but it was all licensed
and and in Siga xacol it it's interoperability
487
00:42:08.840 --> 00:42:14.760
copying and just that you could reverse
engineer the interface for the game console.
488
00:42:15.320 --> 00:42:17.800
I think those cases basically boiled down
to, well, yeah, if you're
489
00:42:17.840 --> 00:42:22.119
if you're if not actually infringing in
any way. It's just a way of
490
00:42:22.119 --> 00:42:27.239
getting to that conclusion. And with
the other words, totally legal intermediate copying
491
00:42:27.280 --> 00:42:29.480
is okay. But I don't think
that's the case here. The whole problem
492
00:42:29.559 --> 00:42:34.639
if it chat GPT is spitting out
lots of copyrighted content and it's copyright content
493
00:42:34.679 --> 00:42:37.119
in and out, I think that's
the problem. Well, I mean that
494
00:42:37.199 --> 00:42:38.280
that's sort of the question of the
case, right, if it turns out
495
00:42:38.280 --> 00:42:40.800
that that is fair use. Right, if it turns out that, you
496
00:42:40.800 --> 00:42:44.519
know, they accept the argument that
this is a bug and you know,
497
00:42:44.880 --> 00:42:47.079
you know deminimus spitting out of exact
copy. You know, I don't know
498
00:42:47.199 --> 00:42:51.280
exactly how the court would rule,
but they say that that's okay, then
499
00:42:51.320 --> 00:42:54.000
the precedent copying should be okay as
well. We don't want we don't we
500
00:42:54.000 --> 00:42:58.960
don't want basically copyright law to be
focused on trivialities in a sense. So
501
00:42:59.280 --> 00:43:01.840
you know, kind of the question
is it's sort of a standardfall together situation
502
00:43:02.079 --> 00:43:08.440
where if the outputs end up being
problematic not not protected under fair use,
503
00:43:08.800 --> 00:43:13.480
then the whole thing probably wouldn't be
fair use either. But if it is
504
00:43:13.639 --> 00:43:17.440
then chances are the earlier copying is
okay as well in view of those presents.
505
00:43:19.400 --> 00:43:22.559
I think that's a little bit too
narrow in one respect, Sega versus
506
00:43:22.559 --> 00:43:30.199
Accolade was predicated on the fact the
ruling was predicated on the fact that the
507
00:43:30.239 --> 00:43:36.000
resulting products that Accolade wanted to create
were not competing with what they copied in
508
00:43:36.079 --> 00:43:42.599
any way, not direct substitutes or
even trying to create competing products as alternatives,
509
00:43:43.159 --> 00:43:45.880
but rather games rather as oppose to
the operating system, which is what
510
00:43:45.920 --> 00:43:53.079
they copied, games that would simply
work on that operating system. Whether generative
511
00:43:53.119 --> 00:44:01.719
AI systems are outputting substantially similar material
to what was copy or even just generally
512
00:44:01.800 --> 00:44:12.000
competing with creators, it's much more
of a market effect than we had in
513
00:44:12.039 --> 00:44:17.320
any way in Sega Versus Accolade,
and I have to on the Cartoon Network
514
00:44:17.320 --> 00:44:22.400
case. I don't really see much
of an analogy there to any part of
515
00:44:22.440 --> 00:44:30.039
this case that dealt with you know, it was all three aspects of the
516
00:44:30.079 --> 00:44:39.039
Cartoon Network case were and remain outlying
outlier decisions on both the temporary copy issue
517
00:44:39.360 --> 00:44:45.840
and the public performance issue, as
well as the authorizing the copy or the
518
00:44:50.000 --> 00:44:53.400
I can't think of the word anyway, none of those if this is not
519
00:44:53.480 --> 00:44:58.800
about public performance, this is not
about temporary copies. Those copies are in
520
00:44:58.840 --> 00:45:00.960
there for a long period of time. I think that you actually doualize that
521
00:45:00.960 --> 00:45:04.920
as a public performance. Interestingly,
I'm not exactly sure why, but I
522
00:45:04.920 --> 00:45:08.039
think that was actually the place.
Well, yes, but it's not the
523
00:45:08.039 --> 00:45:13.800
same sort of issue where oh,
there's only one copy versus ten thousand copies,
524
00:45:14.360 --> 00:45:20.559
And ironically the ten thousand copies is
not a public performance, said Cartoon
525
00:45:20.559 --> 00:45:24.199
Network. In the event, of
all the holdings of Cartoon Network, that
526
00:45:24.280 --> 00:45:30.239
element is surely the weakest. After
the Supreme Court decision in Areo, which,
527
00:45:30.280 --> 00:45:37.079
while not explicitly overruling Cartoon Network,
articulated very clearly an analysis at odds
528
00:45:37.119 --> 00:45:42.480
with the Second Circuits analysis of public
performance in Cartoon Network. I don't see
529
00:45:42.480 --> 00:45:52.880
any way that that's good law.
Do any of you have a view of
530
00:45:53.320 --> 00:46:00.679
the likelihood of a licensing solution to
this or do you think it's going to
531
00:46:00.760 --> 00:46:06.960
resolve on by the court Supreme Court? Probably I think the heavily licensed I
532
00:46:07.000 --> 00:46:09.119
think that I mean not saying I
don't want to say inevitably, you never
533
00:46:09.199 --> 00:46:13.719
know this could be litigated for a
dozen years. But they'll note in the
534
00:46:13.840 --> 00:46:20.000
complaint that New York Times is when
was heavily used sources for open AI,
535
00:46:20.960 --> 00:46:27.559
but other sources, notably Associated Press
or others mentioned in the complaint, had
536
00:46:29.239 --> 00:46:31.639
been paid for licensing. So I
think it's there, But I think,
537
00:46:34.719 --> 00:46:37.280
well, I mean, the old
line is that litigation is negotiation by other
538
00:46:37.360 --> 00:46:40.039
means. I'm not sure it's always
true, but I think it is true
539
00:46:40.039 --> 00:46:45.360
here. Yeah, I mean I
think that you know, there there are,
540
00:46:45.679 --> 00:46:49.320
as Steve mentioned, there are a
number of companies that are out there
541
00:46:49.360 --> 00:46:52.159
developing generative systems that are that are
that are based on that are that are
542
00:46:52.199 --> 00:46:58.320
based on licensing. The difficulty,
of course, which is which I'm sure
543
00:46:58.320 --> 00:47:00.400
opening I would be happy point to
on believe that they point to this in
544
00:47:00.519 --> 00:47:06.599
some of their papers, is that
a lot of the reason why these systems
545
00:47:06.679 --> 00:47:12.239
work particularly well is just the breadth
of content, the fact that they have
546
00:47:13.840 --> 00:47:20.000
large quantities, tremendously large quantities of
data to work off of, which would
547
00:47:20.039 --> 00:47:22.559
be substantially harder to obtain by a
licensing. And so, you know,
548
00:47:22.599 --> 00:47:29.719
while I think it's correct that this
isn't a question of just AI versus licensing
549
00:47:29.760 --> 00:47:32.639
that you that you can have both
at the same time. You aren't changing
550
00:47:32.639 --> 00:47:37.320
the parameters of what the technological development
looks like, right, and that could
551
00:47:37.400 --> 00:47:39.519
be a good thing. I'm not
necessarily saying that it's a bad thing if
552
00:47:39.519 --> 00:47:46.119
computer scientists are forced to work with
a different set of legal parameters. There
553
00:47:46.159 --> 00:47:52.440
could be value in specifically working with
licensed copies. But one thing to keep
554
00:47:52.480 --> 00:47:57.679
in mind is that there is also
a pretty good chunk of poly domain information
555
00:47:57.719 --> 00:48:02.000
that's available that had been the basis
for training of artificial intelligence for many years,
556
00:48:02.239 --> 00:48:05.960
and one of the reasons that companies
have moved away from that is that
557
00:48:05.960 --> 00:48:10.639
it produced very strange systems. A
lot of that material is hundreds of years
558
00:48:10.639 --> 00:48:16.679
old material that's sufficiently out of copyright, or material that is very strange to
559
00:48:16.800 --> 00:48:23.360
use as training data. The en
run email data base was a common one
560
00:48:23.400 --> 00:48:28.280
that was used, and as people
have pointed out, that ended up producing
561
00:48:28.400 --> 00:48:34.800
systems that incorporated all sorts of very
strange biases and so's you know. I
562
00:48:34.840 --> 00:48:37.400
think that it's on the one hand, yeah, you can't have a world
563
00:48:37.480 --> 00:48:43.480
in which licensing is prominent. You
do end up with a different technological environment,
564
00:48:43.639 --> 00:48:47.679
and I think it's it's an interesting
question of what that environment looks like
565
00:48:47.880 --> 00:48:51.440
and how you really feel about it, what it does to the true actually
566
00:48:51.440 --> 00:48:54.280
of the technology. Well, Charles, I think you just made a great
567
00:48:54.320 --> 00:48:59.159
case for the value of the copyrighted
works that were copying in order to build
568
00:48:59.480 --> 00:49:05.320
generative AI systems. Does it take
a little longer? Might it cost more
569
00:49:05.920 --> 00:49:10.199
to compensate creators for using their works
to build your system? Well? Yes,
570
00:49:10.280 --> 00:49:15.960
of course, the same way that
it would take a lot longer if
571
00:49:16.119 --> 00:49:21.800
the AI companies didn't have Nvidia chips
because they didn't want to pay for them,
572
00:49:22.519 --> 00:49:25.960
their processing capacity would be greatly diminished. But they need them and they
573
00:49:27.000 --> 00:49:30.920
want them, so they pay for
them. And I would suggest that the
574
00:49:30.039 --> 00:49:38.360
social aspect of this of telling creators
you're the one input into generative AI that
575
00:49:38.440 --> 00:49:43.079
no one has to pay for.
The coders they get paid, developers,
576
00:49:43.840 --> 00:49:49.559
the chip makers they get paid.
The power utilities that provide all that electricity
577
00:49:49.599 --> 00:49:54.760
to the data farms they get paid. But the creators, sorry, you're
578
00:49:54.840 --> 00:50:00.039
out of lock. That I think
there's an it looked at me in in
579
00:50:00.119 --> 00:50:05.559
that context. It's a pretty difficult
argument to sustain, but I think that
580
00:50:05.760 --> 00:50:12.960
generally correct this idea that if there's
some value out there, then possibly we
581
00:50:13.000 --> 00:50:15.199
want to make sure that there's compensation
with the value. The difficulty here,
582
00:50:15.239 --> 00:50:19.880
of course, is that copyright doesn't
cover every possible value of a work.
583
00:50:20.400 --> 00:50:24.159
Right, you have factual databases which
simply have no copyright protection because facts are
584
00:50:24.159 --> 00:50:29.880
not copyrightable, and so there is
sort of this larger question of, you
585
00:50:29.920 --> 00:50:31.960
know, how do you deal with
the problem value. Similarly, a lot
586
00:50:32.000 --> 00:50:37.039
of these systems use private data,
right, they use individual personal data.
587
00:50:37.079 --> 00:50:40.039
That's data that's not copyright protectable at
all, right, and that is value.
588
00:50:40.119 --> 00:50:43.599
So you know, one way you
could solve as by trying to create
589
00:50:43.599 --> 00:50:46.400
some sort of property right in private
information, but that actually causes a lot
590
00:50:46.440 --> 00:50:52.119
of other problems, as you know, a fair amount of scholarship that's out
591
00:50:52.159 --> 00:50:57.679
there. So I think that you
want to separate out the questions of how
592
00:50:57.679 --> 00:51:04.280
do you deal with the value propose
of information from the copyright questions, and
593
00:51:04.880 --> 00:51:07.840
if you're talking about just sort of
the value propositions, maybe once you're really
594
00:51:07.840 --> 00:51:10.320
asking for some sort of regulatory agency, which you know, we can debate
595
00:51:10.360 --> 00:51:15.840
that, but then you also have
to have the second question of is copyright
596
00:51:15.880 --> 00:51:22.880
the correct vehicle for doing that given
that copyright is historically limited to certain things
597
00:51:22.960 --> 00:51:27.159
which are a little bit different from
what exactly is going on with artificial intelligence
598
00:51:27.159 --> 00:51:32.239
systems. The fact versus the idea
versus expression, the fact versus expression dichotomies,
599
00:51:34.119 --> 00:51:37.239
and of course the fair use doctrine, these all play into kind of
600
00:51:37.239 --> 00:51:43.039
what the boundaries of copyright law are
that don't make value exactly aligned with the
601
00:51:43.119 --> 00:51:46.559
with the rights provide unfederal law.
I do agree there are elements of these
602
00:51:46.599 --> 00:51:52.800
issues that go beyond copyright law,
and you're seeing all sorts of proposals in
603
00:51:52.800 --> 00:52:01.239
those veins, from privacy protections to
state and federal legion on issues of publicity
604
00:52:01.320 --> 00:52:07.760
rights, narrowing it to the copyright
issues, And because that is the focus
605
00:52:07.800 --> 00:52:14.760
of this panel, I do think
that of course collections of data can be
606
00:52:14.840 --> 00:52:19.199
protectable as compilations by their selection,
arrangement coordination, as I'm sure you know.
607
00:52:20.639 --> 00:52:24.320
And then many of the works that
were taken were taken precisely because they
608
00:52:24.360 --> 00:52:36.079
are valuable expression of contemporary American society, and we don't want people aren't going
609
00:52:36.119 --> 00:52:40.079
to pay for AI systems that speak
in old English or forsooth or what have
610
00:52:40.199 --> 00:52:46.679
you. So when I speak of
the value. I'm speaking both generally beyond
611
00:52:46.719 --> 00:52:52.400
copyright, but also in the context
of the fourth factor, where there's harm
612
00:52:52.480 --> 00:53:00.960
to the current and potential market or
for the works that were copied. I
613
00:53:00.000 --> 00:53:06.679
have a question from the audience's question
of possible to the panel, I'm tired
614
00:53:07.119 --> 00:53:12.800
is tirty? Is is there a
place in this discussion for Sony safe harbor?
615
00:53:15.239 --> 00:53:21.840
Yeah? For Sony you said,
right, yes, yeah, So
616
00:53:21.960 --> 00:53:24.920
I actually pulled along from Sony.
Everyone reads Sony for this first half,
617
00:53:24.960 --> 00:53:31.280
which is capable of substance of you
know, subject of commercially significant non infringing
618
00:53:31.440 --> 00:53:37.880
uses. Of course, in Sony
the court held that time shifting was okay,
619
00:53:37.159 --> 00:53:40.920
but librarying was probably not. And
everyone forgets to the second half.
620
00:53:44.320 --> 00:53:46.840
There's room for it, but I
don't think it gets you there in this
621
00:53:47.000 --> 00:53:53.199
case. I think that clearly open
Aye had well. Versually is an interesting
622
00:53:53.280 --> 00:53:57.599
question. It's a capable of non
infringing uses, I mean through it depends
623
00:53:57.639 --> 00:54:04.679
how you define it. What what
point is non fringing? And I mean
624
00:54:05.079 --> 00:54:08.000
actually a lot of the most of
dismiss was on sexual limitations, but a
625
00:54:08.039 --> 00:54:13.519
hepstink discovery rule was going to bring
in the ingestion of a lot of that
626
00:54:13.599 --> 00:54:16.280
as well. Potentially of course,
then you know we might be waiting to
627
00:54:16.280 --> 00:54:21.719
find out what win our shop versus
Neely holds, or if anything is held
628
00:54:21.760 --> 00:54:27.320
in that case. This is a
three year sexual limitations case. Yes,
629
00:54:27.639 --> 00:54:31.920
so on the Sony case. This
is for those of you who weren't familiar
630
00:54:31.960 --> 00:54:37.639
with it. This is Sony Betamax
issue from the early nineteen eighties when the
631
00:54:37.840 --> 00:54:45.639
movie studios sued claiming infringement by virtue
of early video cassette recording of broadcast television.
632
00:54:45.280 --> 00:54:50.320
And as we said there was there
were two findings. One was that
633
00:54:50.840 --> 00:54:58.800
recording a show that was broadcast over
free broadcast television strictly for the purpose of
634
00:54:58.880 --> 00:55:05.440
watching it later at a more convenient
time is a fair use, but keeping
635
00:55:05.480 --> 00:55:10.199
libraries of those is not. And
then second, in terms of the contributory
636
00:55:10.280 --> 00:55:19.239
liability of the manufacturer of the Betamax
machine, Sony in that case, that
637
00:55:20.079 --> 00:55:27.039
the existence of commercially significant, substantial
noninfringing uses that the court would not impute
638
00:55:27.159 --> 00:55:30.920
knowledge of the infringement, and knowledge
being one of the two problems of contributory
639
00:55:30.960 --> 00:55:37.239
liability. The Supreme Court later in
Rockster made very clear that there is no
640
00:55:37.400 --> 00:55:40.880
safe harbor. This question was framed
as a safe harbor, and Sony that
641
00:55:42.000 --> 00:55:45.440
is not the correct framing of the
law. There is not a safe harbor.
642
00:55:47.000 --> 00:55:52.079
If there are commercially significant, non
infringing uses, then Sony is still
643
00:55:52.119 --> 00:55:57.519
good law that the court will not
impute knowledge for that reason. That doesn't
644
00:55:57.559 --> 00:56:00.719
mean the court might not impute knowledge
for another reason, in which it did
645
00:56:00.760 --> 00:56:06.679
in Brockster, and that reason being
in that case that the defendant had induced
646
00:56:06.840 --> 00:56:12.320
the infringements by the users of the
Groxter system. So all of that,
647
00:56:12.719 --> 00:56:15.800
of course is in the context of
contributory liability, which is a doctrine of
648
00:56:15.840 --> 00:56:22.159
secondary or third party liability in US
law. But there are direct infringement issues
649
00:56:22.239 --> 00:56:27.440
here that I think need to be
resolved first, and that's where the fair
650
00:56:27.559 --> 00:56:31.480
use comes in. The fair use
argument comes in which Sony has nothing to
651
00:56:31.519 --> 00:56:37.639
say about in this context. Yeah, so I I know we're pretty close
652
00:56:37.639 --> 00:56:39.119
to out of time. But if
I'm just say a couple of things about
653
00:56:39.159 --> 00:56:44.400
Sony, so Sony is simultaneously irrelevant
and also highly relevant to this case.
654
00:56:45.239 --> 00:56:49.239
So the reason that's not terribly relevant
is that Sony dealt with this contributory infringement
655
00:56:49.280 --> 00:56:52.880
issue and here the allegation is direct
infringement, right, And so as a
656
00:56:52.880 --> 00:56:58.199
result, by doctrine it doesn't help
us other than for the time shifting and
657
00:56:58.239 --> 00:57:01.280
the question of what counts as fair
use. But the way it is relevant
658
00:57:01.360 --> 00:57:06.920
is that Sony is within a line
of cases that deal with the intersection between
659
00:57:06.920 --> 00:57:09.760
copyright law and technology. Right.
So you have a new technology, the
660
00:57:09.840 --> 00:57:15.199
VCR, which doesn't itself say much
about any particular copyright work, but is
661
00:57:15.239 --> 00:57:21.960
a vehicle by which infringements can occur. And similarly, you had cases about
662
00:57:22.039 --> 00:57:24.079
xerox machines, photographs. You can
go all the way back to the printing
663
00:57:24.119 --> 00:57:28.519
press if you really want, and
you recognize that you have these general purpose
664
00:57:28.519 --> 00:57:35.840
technologies that can enable copyright infringement,
that can enable easier copying. How do
665
00:57:35.880 --> 00:57:39.519
you deal with that sort of balance? And really the answer comes down to
666
00:57:39.639 --> 00:57:45.239
what exactly are the boundaries of the
copyright? Right. If we say that
667
00:57:45.280 --> 00:57:49.119
copyright involves any sort of copying at
all, then of course all of these
668
00:57:49.119 --> 00:57:52.239
things are infringement, and you wouldn't
be able to have printing process, photographs,
669
00:57:52.320 --> 00:57:55.320
xerox machines, any of these sorts
of things. But of course there
670
00:57:55.360 --> 00:57:58.880
is a fair use doctrine. That
fair use doctrine exists for a reason.
671
00:57:58.920 --> 00:58:04.440
It's to set a bound around what
the copyright right protects and to allow for
672
00:58:04.480 --> 00:58:07.079
technological innovation beyond that space. And
so in a sense, what this case
673
00:58:07.119 --> 00:58:12.599
really is all about is what exactly
is that line. Where does the line
674
00:58:12.800 --> 00:58:19.719
fall such that an act that colloquially
we would call copying falls outside of copying
675
00:58:19.760 --> 00:58:22.760
and falls within the space of permissible
innovation. That's a line that's been pretty
676
00:58:22.760 --> 00:58:27.800
important throughout history and a lot of
different technologies. And so in that sense,
677
00:58:27.880 --> 00:58:31.039
this case is not terribly new.
It's simply an iteration of that same
678
00:58:31.159 --> 00:58:36.840
problem that has come up many many
times for copyright law and for a lot
679
00:58:36.880 --> 00:58:42.760
of other areas of law, how
it intersects with new technologies. So that's
680
00:58:42.880 --> 00:58:46.199
the false dichotomy of we wouldn't have
the printing pressor's xerox machines if we didn't
681
00:58:46.199 --> 00:58:51.400
have fair use, because there's no
such thing as licensing. That's of course
682
00:58:51.480 --> 00:58:57.039
not the case, Nor do I
believe that there's some sort of penumber around
683
00:58:57.079 --> 00:59:06.519
Sony or other copyright andechnological innovation cases. And what I do agree with is
684
00:59:06.559 --> 00:59:13.000
that copyright has always been a law
that has reacted to a variety of factors,
685
00:59:13.039 --> 00:59:17.320
developments in the marketplace, consumer preferences, and technological evolution. Absolutely,
686
00:59:19.599 --> 00:59:23.159
and the court looks at some of
the particulars and fair use has a key
687
00:59:23.280 --> 00:59:29.440
role in that, of course.
So we have decisions like Rockster, as
688
00:59:29.440 --> 00:59:32.920
I mentioned, Ario, another one
I mentioned, and many others where the
689
00:59:34.000 --> 00:59:37.400
court said, no, this is
a model based on infringement and we're not
690
00:59:37.440 --> 00:59:42.079
going to permit it. We have
others where the court, like in Sony,
691
00:59:42.199 --> 00:59:46.000
said this is inherently meant to be
an innocent business that could be used
692
00:59:46.039 --> 00:59:50.760
for something afarious, but we don't. We're not going to hold the manufacturer
693
00:59:50.840 --> 00:59:55.360
or the device secondarily liable just because
it could be used in a bad way.
694
00:59:57.519 --> 01:00:00.760
In the case before us, these
generatives AI systems that made a conscious
695
01:00:00.800 --> 01:00:07.519
decision to copy copyrighted works by the
billions. I think it's a lot more
696
01:00:07.679 --> 01:00:17.800
like Ario and Rockster than it is
like the innocent and business models Aily.
697
01:00:17.880 --> 01:00:22.800
I know we're up against the hour, and I have just one question for
698
01:00:22.880 --> 01:00:25.360
the panel as a whole, and
the question is what if we remove the
699
01:00:25.400 --> 01:00:30.880
technology from this discussion and just put
it into a business sense that for some
700
01:00:30.920 --> 01:00:36.320
reason the company starts a business and
happens to go into the New York Times
701
01:00:36.400 --> 01:00:40.239
archives and carts it all off into
their warehouse, and then they adhire a
702
01:00:40.360 --> 01:00:45.239
whole bunch of employees for people who
come into the front counter and ask them
703
01:00:45.320 --> 01:00:50.039
questions, and they go back and
skirt around the New York Times information and
704
01:00:50.039 --> 01:00:55.599
give them an answer or or an
asset that's directly competing with what the New
705
01:00:55.719 --> 01:01:01.000
York Times uses is archives for Is
is that a valid analogy? Is it,
706
01:01:01.079 --> 01:01:06.920
you know, technologically and misplaced or
does that make a difference in the
707
01:01:06.960 --> 01:01:12.239
discussion. If you remove the technology
from what's actually happening with the copyright,
708
01:01:12.280 --> 01:01:17.039
it works, John, I think
I think you're a little description. It's
709
01:01:17.119 --> 01:01:22.159
kind of close to a library,
which is okay, but you do get
710
01:01:22.199 --> 01:01:24.559
into some But if you tweak it
in two ways. One, if you
711
01:01:24.599 --> 01:01:28.199
say, what if they give copies
of Times that are clothes I cart it
712
01:01:28.239 --> 01:01:32.320
off instead of just giving them summaries, or even if like a copy cut
713
01:01:32.360 --> 01:01:37.639
and paste out sentences instead of a
ransom note style answer, I think you're
714
01:01:37.679 --> 01:01:42.760
getting closer to the fat current here. And then the flip side is what
715
01:01:42.920 --> 01:01:49.599
if they grab the New York Times
from you know that from how offer presses
716
01:01:50.320 --> 01:01:54.280
in the East Coast? Why are
summaries of a story so West coast and
717
01:01:54.320 --> 01:01:59.159
print from there? And of course
that's ions versus AP, which was held
718
01:01:59.159 --> 01:02:04.760
to be young awful. So I
don't think you can make it technology independent
719
01:02:04.800 --> 01:02:08.559
because all the technology dictates some of
illegal questions. I mean, to make
720
01:02:08.599 --> 01:02:14.199
things even worse. They do talk
about hallucinations, the problem in which sometimes
721
01:02:14.239 --> 01:02:16.400
you will ask for a New York
Times article from chatchipg, I'll give you
722
01:02:16.400 --> 01:02:20.639
something completely made up that has nothing
to do with that. So the analogy
723
01:02:20.679 --> 01:02:23.519
really is a library reference desk that
gives you, every once in a while
724
01:02:23.639 --> 01:02:27.480
an exact copy of an article that
I found, and every once in a
725
01:02:27.519 --> 01:02:32.199
while gives you totally useless factual information. It's a little hard to draw analogies
726
01:02:32.239 --> 01:02:37.920
to that point, I think,
but it does, and I think it
727
01:02:37.000 --> 01:02:40.039
kind of shows why it's a little
difficult to work with analogies, and sometimes
728
01:02:40.079 --> 01:02:44.039
it really is worth just figuring out
exactly what's going on with the technology.
729
01:02:44.079 --> 01:02:51.159
Yet it's at its proper levels.
I'll just add that what you described was
730
01:02:51.199 --> 01:02:54.880
a lot like a library. If
it's a nonprofit organization, you didn't stipulate
731
01:02:54.920 --> 01:03:01.159
which way that organizations operating nonprofit ormercial. But the point I want to make
732
01:03:01.280 --> 01:03:12.679
is Congress has taken pains to enact
specific statutory exemptions for libraries to provide library
733
01:03:12.679 --> 01:03:19.719
patrons with a certain degree of service
without going so far that it implicates the
734
01:03:20.000 --> 01:03:24.599
incentive to create in the first place. And libraries are not entitled to just
735
01:03:25.440 --> 01:03:30.119
take copies of everything that they want
to have in their collections without paying for
736
01:03:30.159 --> 01:03:36.119
them, without permission and so on. And in fact, in the online
737
01:03:36.159 --> 01:03:39.559
context, some of these issues are
being litigated right now. There's an entity
738
01:03:40.239 --> 01:03:46.039
that calls itself the Internet Archive that
has created a website where it just lets
739
01:03:46.079 --> 01:03:51.760
people take copies of copyrighted works,
and they want to call themselves the library,
740
01:03:52.519 --> 01:03:57.719
but they don't qualify under any of
the library exceptions, and thus far
741
01:03:57.760 --> 01:04:01.320
in the pending litigation they've been quite
unsuccessful. It's on appeals, so we'll
742
01:04:01.320 --> 01:04:03.800
see where it goes. But I
don't I wouldn't put a lot of money
743
01:04:03.800 --> 01:04:13.559
on them. Over to you,
Emily all right on behalf of the Federal
744
01:04:13.599 --> 01:04:16.239
Society. Thank you all for joining
us for this great discussion today. Thank
745
01:04:16.280 --> 01:04:19.880
you also to our audience for joining
us. We greatly appreciate your participation.
746
01:04:20.440 --> 01:04:25.599
Check out our website fedsoc dot org
or follow us on all major social media
747
01:04:25.639 --> 01:04:29.880
platforms at fedsoc to stay up to
date with announcements and upcoming webinars. Thank
748
01:04:29.920 --> 01:04:33.320
you once more for tuning in,
and we are adjourned. Thank you for
749
01:04:33.400 --> 01:04:38.639
listening to this episode of FEDSOC Forums, a podcast of the Federal Societies Practice
750
01:04:38.639 --> 01:04:42.920
Groups. For more information about the
Federal Society, the practice groups, and
751
01:04:42.960 --> 01:04:46.199
to become a Federal Society member,
please visit our website at fedsoc dot or org
1
00:00:02.080 --> 00:00:06.400
Welcome to fedsock Forums, a podcast
of the Federal Society's Practice groups. I'm
2
00:00:06.440 --> 00:00:10.039
Ny kas Merrick, Vice President and
Director of Practice Groups at the Federal Society.
3
00:00:10.240 --> 00:00:14.640
For exclusive access to live recordings of
fedsock Forum programs, become a Federal
4
00:00:14.679 --> 00:00:20.679
Society member today at fedsoc dot org. Hello everyone, and welcome to this
5
00:00:20.719 --> 00:00:24.839
Federal Society virtual event. My name
is Emily Manning, and I'm an Associate
6
00:00:24.879 --> 00:00:29.000
director of Practice Groups with the Federal
Society. Today we're excited to host a
7
00:00:29.039 --> 00:00:35.240
discussion titled AI meets Copyright Understanding New
York Times v OpenAI. We're joined today
8
00:00:35.320 --> 00:00:40.320
by Charles Dwan V. Rosen,
Stephen M. Tepp, and our moderator
9
00:00:40.320 --> 00:00:44.960
today is John P. Moran of
Council at Holland and Knight. John has
10
00:00:45.000 --> 00:00:50.840
experienced litigating many patent, trademark and
trade secret cases and Federal District Court and
11
00:00:51.000 --> 00:00:55.280
argues appeals at the US Court of
Appeals for the Federal Circuit the US Court
12
00:00:55.280 --> 00:01:00.000
of Appeals for the Fourth Circuit.
He has prosecuted or directly supervised the prosecute
13
00:01:00.359 --> 00:01:06.680
of hundreds of patent applications and many
different technologies, including telecommunication systems and equipment
14
00:01:07.120 --> 00:01:15.480
robotics, artificial intelligence, imaging technology, nuclear reactor instrumentation, semiconductor devices and
15
00:01:15.560 --> 00:01:19.879
manufacturing processes and medical devices. If
you'd like to learn more about today's speakers,
16
00:01:19.920 --> 00:01:23.959
their full bios can be viewed on
our website FEDSOC dot org. After
17
00:01:25.000 --> 00:01:27.359
our speakers give their opening remarks,
we will turn to you the audience,
18
00:01:27.359 --> 00:01:30.480
for questions. If you have a
question, please enter it into the Q
19
00:01:30.560 --> 00:01:34.200
and A function at the bottom of
your zoom window, and we will do
20
00:01:34.239 --> 00:01:37.560
our best answer as many as we
can. Finally, I'll note that,
21
00:01:37.640 --> 00:01:41.079
as always, all expressions of opinion
today are those of our guest speakers,
22
00:01:41.319 --> 00:01:44.480
not the Federal Society. With that, thank you for joining us today,
23
00:01:44.519 --> 00:01:48.680
and John, the floor is yours. Thank you Emily for that introduction.
24
00:01:49.599 --> 00:01:53.840
Well, this afternoon's panel is Emily
indicated. We'll discuss the copyrighted issues raised
25
00:01:55.040 --> 00:02:00.480
in the complaint brought by The New
York Times against Microsoft and several Open AI
26
00:02:00.840 --> 00:02:06.400
entity. The complaint, which was
fought last December, has seven counts.
27
00:02:06.760 --> 00:02:12.960
Four of the counts are for copyright
infringement counts. There's a Digital Millennium Copyright
28
00:02:13.080 --> 00:02:20.439
Act count and a common unfair competition
by misappropriation of New York Times intellectual property.
29
00:02:21.080 --> 00:02:25.800
There's also a trademark dilution count,
which will not be addressed today.
30
00:02:27.199 --> 00:02:34.759
In support of the copyright infringement allegations, New York Times included in the complaint
31
00:02:34.919 --> 00:02:40.400
one hundred examples of outputs that are
nearly identicate and identical to the copyrighted New
32
00:02:40.479 --> 00:02:47.240
York Times content. Before we address
the copyright issues, a bit of technology
33
00:02:47.319 --> 00:02:53.960
vocabulary at a very high level may
help the discussion, so to begin with,
34
00:02:53.159 --> 00:03:00.639
the complaint focuses on the open aised
GPT models as it relates to today's
35
00:03:00.639 --> 00:03:09.719
discussion. The acronym stands for generative
pre trained transformer. The generative, which
36
00:03:09.759 --> 00:03:15.479
we will probably discussed today, indicates
that the model takes inputs from users and
37
00:03:15.599 --> 00:03:22.240
generates an output such as one hundred
examples of New York Times content and the
38
00:03:22.280 --> 00:03:28.479
complaint. The P for pre trained
indicates that the model is pre trained on
39
00:03:28.599 --> 00:03:32.719
the volume of data such as the
New York Times content, and lastly,
40
00:03:34.159 --> 00:03:39.800
the T for transformer very generally relates
to a way of processing data to provide
41
00:03:39.960 --> 00:03:46.639
context for the data, such as
a sequence of words. One note that
42
00:03:46.759 --> 00:03:52.479
the complaint uses the term embedded.
It appears that it uses that term in
43
00:03:52.560 --> 00:03:57.240
the non artificial intelligence sense, that
is, the sense to the stone is
44
00:03:57.240 --> 00:04:05.840
embedded in concrete, other than the
process of embedding converting words into corresponding numerical
45
00:04:05.960 --> 00:04:13.599
values. Those topics were probably those
concepts will probably discussed during the panel's discussion
46
00:04:13.680 --> 00:04:19.759
today. So as emily indicated to
discuss the New York Times allegations of copyright
47
00:04:19.800 --> 00:04:27.680
infringement, we have three experts.
Zev Rosen is an assistant professor, School
48
00:04:27.680 --> 00:04:33.319
of Law at Southern Illinois University.
It was the Abraham Chemistine Scholar in Residence
49
00:04:33.399 --> 00:04:40.600
at the United States Copyright Office and
is particularly renowned for his expertise on copyright
50
00:04:40.720 --> 00:04:47.199
law, especially in its historical development. Charles One is an assistant professor at
51
00:04:47.240 --> 00:04:54.079
the American University College of Law.
It was previously a post doctoral fellow at
52
00:04:54.160 --> 00:05:00.439
Cornarial Tech. Working with professional James
Grimmleman. Charles focuses his research on the
53
00:05:00.439 --> 00:05:08.079
public effects of technology policy, intellectual
property law, and primarily patents and copyrights.
54
00:05:09.720 --> 00:05:15.079
Steve Tapp is a president and CEO
of Sentinel Worldwide, which is an
55
00:05:15.079 --> 00:05:20.959
intellectual property consultancy. He's also a
lecturer in law at the George Washington University
56
00:05:21.800 --> 00:05:29.240
School of law. More recently,
Steve has co founded rights Cliques, which
57
00:05:29.279 --> 00:05:33.560
is a suite of software tools for
independent creators to register, manage, and
58
00:05:33.639 --> 00:05:40.839
enforce their copyrights. As Emily indicated, we'll start with opening statements, will
59
00:05:40.879 --> 00:05:45.560
begin with Zeve's, then Charles and
followed by Steve. So with that,
60
00:05:46.560 --> 00:05:50.199
I'll turn it over to z.
All right, everyone, good afternoon or
61
00:05:50.279 --> 00:05:57.600
morning, depending where you are there
are. This causes really fascinating because we're
62
00:05:57.639 --> 00:06:04.800
finally getting at a host of important
issues regarding AI and the Internet, and
63
00:06:05.000 --> 00:06:11.240
what we call traditional producer is a
cultural works we of New York Times or
64
00:06:11.319 --> 00:06:15.920
many other that I think will probably
be honest cot tales. There's a couple
65
00:06:15.920 --> 00:06:19.319
of legal issues I want to highlight, and I'd love some questions to follow
66
00:06:19.399 --> 00:06:26.199
this once my colleagues over opening statements. The initial issue here is really direct
67
00:06:26.279 --> 00:06:30.560
infringement, which is to say,
by taking all of these New York Times
68
00:06:30.560 --> 00:06:36.959
stories, they are effectively copying them
or creating derivative works of them. A
69
00:06:38.000 --> 00:06:42.600
derivative work is a work which recasts, adapts, or transforms and existing work.
70
00:06:45.000 --> 00:06:48.399
The scope of derivative works right is
frankly not terribly well understood because it
71
00:06:48.639 --> 00:06:54.399
usually dovetails of copying and is typically
going to be one or the other.
72
00:06:55.319 --> 00:07:01.639
But that is a big part of
what is being claimed here. The resolution
73
00:07:01.759 --> 00:07:06.279
of it is really, I think
going to be a question of a is
74
00:07:06.319 --> 00:07:12.079
one of those activities happening. I
tend to think it is, and honestly,
75
00:07:12.480 --> 00:07:15.480
I'm not quite positive how much even
opening I'm sure, I'm sure we'll
76
00:07:15.519 --> 00:07:19.879
disputed, but I think that's probably
an easier case. But at least some
77
00:07:20.000 --> 00:07:24.759
copying is occurring, although there's details
of that that I'm sure some of my
78
00:07:24.800 --> 00:07:28.680
colleagues will discuss. But then fair
use is going to be a major issue
79
00:07:28.759 --> 00:07:35.160
there, particularly whether or not it's
transformative. What transformative means is really a
80
00:07:35.199 --> 00:07:41.839
hard question, but too recent.
The sort of formative case is the Acoff
81
00:07:42.160 --> 00:07:47.399
versus Campbell case about two live crews
Pretty woman, But I really think this
82
00:07:47.519 --> 00:07:53.800
case is going to be on fair
use, a conflict of the Google versus
83
00:07:53.800 --> 00:08:03.279
Oracle case, which found that the
Java program called Api eyes that their reimplementation
84
00:08:03.480 --> 00:08:09.160
or copying depending on your perspective,
I suppose into the droid operating system was
85
00:08:09.199 --> 00:08:13.920
fair use and transformative, and the
flip side Warhol versus Goldsmith holding that the
86
00:08:15.639 --> 00:08:20.680
Andy Warhol transformation into of Lyn Goldsmith's
photo of prints into an Andy Warhol silk
87
00:08:20.680 --> 00:08:28.439
screen of prints and then is licensing
to contiinast was not fair use and not
88
00:08:28.519 --> 00:08:33.720
transformative, So that's going to be
a bit of way. I suspect there
89
00:08:33.799 --> 00:08:39.279
was a motion to dismiss which is
pending, and that's all we have right
90
00:08:39.320 --> 00:08:41.799
now. We don't have, and
so fair use I don't think is ripe
91
00:08:41.879 --> 00:08:45.080
yet. I think you have to
wait for at leat summary judgment on that.
92
00:08:45.320 --> 00:08:50.559
But I kind of suspect if it
doesn't get dismissed and we have to
93
00:08:50.600 --> 00:08:56.679
get to that, we'll see on
that. You have to show direct infringement
94
00:08:56.720 --> 00:09:00.360
to get to any of the other
stuff except some of them more except for
95
00:09:00.879 --> 00:09:07.559
the last one for contributory infringement,
which is also alleged a lot of you
96
00:09:07.600 --> 00:09:09.799
us have to do, particularly with
Microsoft, which is not open AI,
97
00:09:11.759 --> 00:09:16.080
but as the complaint notes, is
kind of an alter ego, and we
98
00:09:16.159 --> 00:09:20.519
have all sorts of rights and open
AI, and we also own a lot
99
00:09:20.519 --> 00:09:26.600
of the IP. Contributory infringement is
defendant have knowledge of direction infringement and defendant
100
00:09:26.679 --> 00:09:37.519
materially contributing to that infringement from loaded
terms there knowledge knowledge you know question.
101
00:09:37.759 --> 00:09:41.600
We've had a conflict that happens a
lot in the context of the Internet and
102
00:09:41.639 --> 00:09:48.879
takedowns that whether knowing requires sort of
red flag or general knowledge. Where's a
103
00:09:48.919 --> 00:09:56.240
case of all the coximmunications and it
was very adequately flagging for users who repeat
104
00:09:56.279 --> 00:10:03.919
infringers. The details of that case
law are somehow twenty six years after the
105
00:10:03.960 --> 00:10:13.120
law was passed, still in flux
to a degree and material contribution. I
106
00:10:13.159 --> 00:10:16.240
think that that one is easier to
show here. I think it's going to
107
00:10:16.440 --> 00:10:18.440
hinge on knowledge, but I've been
wrong before. If that's the other key
108
00:10:18.480 --> 00:10:26.879
part vicarious liability. The case VET
I always teach when I talk about vicarious
109
00:10:26.919 --> 00:10:31.600
liability the case called Funavisia versus Cherry
Orchard, and it's a case VET held
110
00:10:31.600 --> 00:10:35.080
a out of fleet. I mean
it's contributory as well, but a case
111
00:10:35.080 --> 00:10:41.320
it held out of flea market that
was basically a hotbed of pirated media was
112
00:10:41.440 --> 00:10:45.759
engaging with vicarious liability because they had
a right to control the infringing activity.
113
00:10:45.799 --> 00:10:50.799
The infringing activity being the resale of
pirated discs installed at the flea market that
114
00:10:50.840 --> 00:10:54.960
were rented out and arriving a financial
commercial benefit from it, which was once
115
00:10:54.960 --> 00:11:00.519
again they were getting paid the fees
or restalls. Both of other I think
116
00:11:00.919 --> 00:11:07.919
are going to come out quite substantially. But other sort of pseudo IP issue
117
00:11:07.960 --> 00:11:11.440
I want to flag, if it's
really interesting, claims that New York Times
118
00:11:11.440 --> 00:11:18.919
owns Wirecutter, which is product recommendations, and they claim it is misappropriation because
119
00:11:18.960 --> 00:11:26.720
they are taking Wirecutter reviews and are
not getting affiliate payments. The large open
120
00:11:26.720 --> 00:11:31.720
AI are claiming this is preemptive by
a copyright lack which has at all legal
121
00:11:31.799 --> 00:11:37.840
or equitable rights that are within the
general scope of copyright are preempted. And
122
00:11:37.919 --> 00:11:41.639
on the other hand you have this
case is International News Service versus Associated Press,
123
00:11:41.679 --> 00:11:48.679
Supreme Court nineteen eighteen, holding the
misappropriation of hot news. In other
124
00:11:48.720 --> 00:11:52.600
words, news which was fresh out
of battle line the World War One was
125
00:11:52.799 --> 00:12:00.159
a misappropriation and a violation of common
law rights that was not preempted. The
126
00:12:00.200 --> 00:12:05.080
scope of VINS versus the AP nowadays
has been questioned for some cases have limited
127
00:12:05.080 --> 00:12:09.279
it, but it's not dead,
and so I think it's be interesting to
128
00:12:09.279 --> 00:12:13.600
see how that plays out. I
think, well, I'm looking forward to
129
00:12:13.759 --> 00:12:16.000
more of your questions. I think
we'll turn over Charles from more first head
130
00:12:16.039 --> 00:12:24.120
thoughts. All right, uh,
thanks thanks to me for that, you
131
00:12:24.159 --> 00:12:28.320
know, really excellent introduction to kind
of what the major issues are in this
132
00:12:28.440 --> 00:12:35.279
case and what the kind of key
doctrines of copyright and other law are that
133
00:12:35.360 --> 00:12:37.759
are at play. As you can
see, this case has a lot of
134
00:12:37.799 --> 00:12:39.879
things going on, and so you
know what I'll try to do is I'll
135
00:12:39.919 --> 00:12:43.399
try to number one, just go
through kind of what's been going on in
136
00:12:43.440 --> 00:12:46.600
the case specifically, so I'll talk
about particularly the motions to the SPISS that
137
00:12:46.639 --> 00:12:52.399
have been filed, and then take
a couple of guesses. As you mentioned,
138
00:12:52.679 --> 00:12:54.840
we don't know what the fair use
defense is going to look like,
139
00:12:54.879 --> 00:12:58.279
but I'll try to talk a little
bit about, you know what i'd expect,
140
00:12:58.159 --> 00:13:03.519
particularly, try to guess kind of
what I would imagine open AI would
141
00:13:03.559 --> 00:13:09.120
try to argue, and then maybe
leave with a couple of broader thoughts about
142
00:13:09.200 --> 00:13:15.639
kind of where this fits into the
larger debate over copyright and AI. So
143
00:13:15.960 --> 00:13:20.000
as far as the procedure of the
lawsuit goes, we started with a complaint
144
00:13:20.039 --> 00:13:24.919
in the lawsuit filed in December twenty
twenty three, and now we have on
145
00:13:24.960 --> 00:13:28.759
the table two motions to dismiss,
one from open Ai that was filed late
146
00:13:28.799 --> 00:13:35.159
February and one from Microsoft that was
filed about a week ago. These motions
147
00:13:35.159 --> 00:13:39.159
to dismiss don't go to the entirety
of the case, which is somewhat interesting.
148
00:13:39.200 --> 00:13:43.639
They only go to what the what
Microsoft and open Ai describe as sort
149
00:13:43.639 --> 00:13:48.080
of ancillary issues, and so the
ones that they talk about, in particular
150
00:13:48.200 --> 00:13:52.039
the contributory liability issue, they say
that there's a lack of knowledge, as
151
00:13:52.159 --> 00:13:54.879
Via mentioned, so they say that
that one should be dismissed. The common
152
00:13:54.960 --> 00:14:00.759
law misappropriation issue that's you mentioned this
hot news I in S is AP case
153
00:14:01.000 --> 00:14:03.639
kind of theory. They say that
that's preempted by the Copyright Act, that
154
00:14:03.720 --> 00:14:11.440
copyright protection being a federal law overrides
what the states, what state protection is
155
00:14:11.480 --> 00:14:13.559
given there, and they point to
know a number of cases that show that
156
00:14:13.600 --> 00:14:18.519
the I S versus AP doctrine is
kind of begrudgingly accepted at this point.
157
00:14:18.600 --> 00:14:22.919
That's kind of their argument for that. John mentioned also that there was this
158
00:14:22.000 --> 00:14:28.399
Digital Millennium Copyright Act count in the
complaint, and that one is actually somewhat
159
00:14:28.399 --> 00:14:33.320
interesting. So I'll just I'll go
into that just a little bit. Some
160
00:14:33.399 --> 00:14:37.399
of you may be familiar with the
Digital Millennium Copyright Act in terms of its
161
00:14:37.440 --> 00:14:43.279
anti circumvention provisions, the rules that
say that you're not allowed to kind of
162
00:14:43.320 --> 00:14:46.519
break digital rights management. This case
actually deals with a different part of the
163
00:14:46.600 --> 00:14:52.919
DMCA, Section twelve oh two,
which relates to copyright management information, so
164
00:14:52.039 --> 00:15:00.639
basically the inclusion of metadata like authors
or titles or copyright notices inside files,
165
00:15:00.679 --> 00:15:07.840
typically digital files. And so The
Times argues that in training these generative AI
166
00:15:09.039 --> 00:15:15.279
systems, Open AI and Microsoft removed
that information and as a result, violated
167
00:15:15.320 --> 00:15:18.320
section twelve h two. The difficulty
that they're going to face improving that,
168
00:15:18.559 --> 00:15:22.320
as Microsoft and open AI point out, is that in order to show a
169
00:15:22.399 --> 00:15:26.679
violation of the section, you have
to show what's called a double cienter requirement.
170
00:15:28.240 --> 00:15:31.559
Number one that's open AI knew that
it was removing and number two that
171
00:15:31.639 --> 00:15:35.919
it knew that the result of removal, or at least should have known that
172
00:15:35.960 --> 00:15:41.399
the result of removal would be further
infringement. And so open and Microsoft argue
173
00:15:41.480 --> 00:15:45.559
that number one, there's no efforden
is that this information was actually removed during
174
00:15:45.600 --> 00:15:48.559
the training process, but number two
that they won't be able to satisfy the
175
00:15:50.000 --> 00:15:54.279
knowledge requirements. They also raise a
time bar question. They say that there's
176
00:15:54.320 --> 00:16:02.440
a three year period look back period
that limits the extent that the copyright allegations
177
00:16:02.480 --> 00:16:06.080
can go back. That's actually a
case that's being considered by the Supreme Court
178
00:16:06.159 --> 00:16:10.399
right now. It was just argued
a couple of weeks ago, and so
179
00:16:10.799 --> 00:16:14.840
that's just another argument that they bring
up. But this again doesn't get to
180
00:16:14.919 --> 00:16:19.480
the substitutive questions that you mentioned,
the questions of whether or not there actually
181
00:16:19.720 --> 00:16:26.200
is copyright infringement in the training of
these systems. Using the New York Times
182
00:16:26.200 --> 00:16:33.360
and other articles, there are three
points that the Times identifies as where the
183
00:16:33.399 --> 00:16:37.639
infringement could occur. Number One,
they say that the collection of the articles
184
00:16:37.720 --> 00:16:41.960
to make the training data use to
train these AI systems, that was an
185
00:16:42.000 --> 00:16:45.720
infringement because you were making a lot
of copies in order to collect them.
186
00:16:45.120 --> 00:16:51.919
Second, they allege that the model
itself all of the data parameters I think
187
00:16:51.960 --> 00:16:59.480
one point seven trillion numbers that make
up the the GPT systems that's somewhere embedded
188
00:16:59.519 --> 00:17:03.879
in there is all of the information
necessary to replicate a lot of the articles,
189
00:17:03.879 --> 00:17:07.480
and therefore the model itself is a
potential infringement. And third, they
190
00:17:07.519 --> 00:17:12.119
say that when you use the model
in such a way that it generates that
191
00:17:12.559 --> 00:17:17.200
generates infringing content, that that use
is sort of a public performance. It
192
00:17:17.240 --> 00:17:23.160
allows you to get the information out
in order to show that these in order
193
00:17:23.200 --> 00:17:26.680
to show infringement, what the Times
would have to show is number one,
194
00:17:26.680 --> 00:17:30.799
that this is copyrightable subject matter.
There are some interesting questions there because you
195
00:17:30.839 --> 00:17:33.920
know, a lot of the information
that's being drawn is factual, and so
196
00:17:33.079 --> 00:17:37.799
maybe there'll be that issue that comes
out. Generally factual information is not considered
197
00:17:37.839 --> 00:17:41.000
copyrightable, but you know, like
I said, that's probably not going to
198
00:17:41.039 --> 00:17:45.440
be the lead argument. There also
are questions of what exactly counts as an
199
00:17:45.480 --> 00:17:48.839
infringing act. You know, is
something that's internally inside the model that nobody
200
00:17:48.880 --> 00:17:52.680
could actually see or understand, is
that an infringement. There's actually sort of
201
00:17:52.680 --> 00:17:56.119
an interesting question about that. But
again, the large issue that open ai
202
00:17:56.200 --> 00:18:00.720
and Microsoft intend to raise, in
fact, they say that they are they're
203
00:18:00.759 --> 00:18:06.839
actually very excited they say in their
in their motions of dismiss to litigate this
204
00:18:06.920 --> 00:18:11.799
issue is the fair use question,
assuming that it is an act of copying
205
00:18:11.880 --> 00:18:14.839
or it is a derivative work to
do all of those things I just mentioned,
206
00:18:15.400 --> 00:18:18.759
does the fair use doctrine permit it? And so courts have used the
207
00:18:18.799 --> 00:18:22.240
fair use doctrine a variety of situations, Google versus Oracle and the software context.
208
00:18:22.759 --> 00:18:27.480
On the Warhol case that was an
artistic use. On the Campical case
209
00:18:27.480 --> 00:18:30.240
that was parody. Oh, fairiuse
is sort of this jack of all trades
210
00:18:30.279 --> 00:18:36.200
doctrine ends up being used in all
sorts of places. In that line,
211
00:18:36.559 --> 00:18:38.759
there are a number of cases that
will help open AI quite a bit,
212
00:18:40.039 --> 00:18:44.920
although possibly to a limited extent.
There was a case over Google Images where
213
00:18:44.960 --> 00:18:48.039
Google had collected a bunch of images
and was displaying them using an image search
214
00:18:48.079 --> 00:18:53.720
engine. Courts said that because Google
had really downsampled them and they didn't really
215
00:18:53.759 --> 00:18:59.839
serve as replacement, that database of
images was fair use. There's a case
216
00:18:59.839 --> 00:19:04.160
called Eye Paradigms in which a company
made a plagiarism detection program and there was
217
00:19:04.160 --> 00:19:08.160
a question of whether or not the
inputs to the plagiarism detection program, which
218
00:19:08.200 --> 00:19:12.559
were basically the essays that were being
detected, whether or not that was an
219
00:19:12.599 --> 00:19:15.680
infringe And again the court said,
you know, this is sort of a
220
00:19:15.720 --> 00:19:22.039
new tool, the actual articles aren't
retrievable. Similarly, with the Google Books
221
00:19:22.119 --> 00:19:25.440
case, it was alleged that Google
scanning of a bunch of books to create
222
00:19:25.480 --> 00:19:30.880
the Google Book search engine was copyright
infringe and again a court said, I
223
00:19:30.880 --> 00:19:34.880
think the second Circuit said that this
was fair use on the grounds that the
224
00:19:36.039 --> 00:19:40.759
mass scanning of books to provide a
service that didn't really replicate the value of
225
00:19:40.759 --> 00:19:48.480
the books themselves was allowable and as
a result not a copyright infringement. Courts
226
00:19:48.559 --> 00:19:53.279
usually applied this four factor test in
which they look at the the nature of
227
00:19:53.319 --> 00:20:00.240
the copyrighted work, the nature of
the use, the purpose in character of
228
00:20:00.279 --> 00:20:03.920
the use, the amount that was
used. And finally, and probably this
229
00:20:03.039 --> 00:20:07.160
is the most important factor by many
measures, the economic impact on the market
230
00:20:07.200 --> 00:20:11.359
for the original copyrighted work. And
I think that that's going to be the
231
00:20:11.359 --> 00:20:15.640
most interesting one to follow in this
case, because, on the one hand,
232
00:20:15.079 --> 00:20:22.440
these are incredibly valuable systems. Right
Artificial intelligence has huge potential in terms
233
00:20:22.599 --> 00:20:30.559
of business uses, commercial uses,
uses for individual consumers. It can be
234
00:20:30.680 --> 00:20:37.480
used as the platform for many other
technologies. Is that part of the market
235
00:20:37.839 --> 00:20:42.200
that inheres in the copyright that The
New York Times has in all of its
236
00:20:42.279 --> 00:20:45.240
articles, or that a novelist has
in all the novels that are used in
237
00:20:45.279 --> 00:20:48.359
training. Right, the novelists or
the New York Times, they would say,
238
00:20:48.480 --> 00:20:52.200
yes, the point of our articles
it is to provide information, and
239
00:20:52.279 --> 00:20:57.079
that information is being used to provide
a valuable service through things like SHATGBT,
240
00:20:57.400 --> 00:21:03.039
and as a result, there should
be some cut of that open AI of
241
00:21:03.079 --> 00:21:04.960
course, would argue the other way
around. They would say, look,
242
00:21:06.000 --> 00:21:10.279
this is a completely different service.
It doesn't serve as a replacement for the
243
00:21:10.319 --> 00:21:15.359
original articles. It's transformative in the
way that Svia mentioned, and transformative something
244
00:21:15.400 --> 00:21:18.960
that the courts have really looked at, so that I think is going to
245
00:21:18.960 --> 00:21:22.400
be sort of the parameters of debate. We do have these cases about mass
246
00:21:22.440 --> 00:21:26.519
text and data mining which are different
from this case, but you know,
247
00:21:26.039 --> 00:21:32.319
provide some basis for understanding where fair
use goes and that transformativelopment and what effects
248
00:21:33.119 --> 00:21:38.160
the availability of these systems has on
those copyrighted works. I think is going
249
00:21:38.200 --> 00:21:47.279
to be really important and something really
to watch as this case progresses. All
250
00:21:47.359 --> 00:21:52.200
right, I think that makes it
my turn. So let me begin by
251
00:21:52.200 --> 00:21:56.640
saying thank you to the Federalist Society
for inviting me today, and Emily for
252
00:21:56.759 --> 00:22:00.799
organizing this panel, and of course
the On for his kind introduction, and
253
00:22:00.839 --> 00:22:06.440
my fellow panelists for their opening remarks. Let me note that my remarks are
254
00:22:06.519 --> 00:22:10.640
my own and do not necessarily reflect
the views of any client or employer.
255
00:22:11.440 --> 00:22:15.759
I want to begin by putting this
case on others like it into a broader
256
00:22:15.799 --> 00:22:19.440
perspective. Those of us who've been
working in copyright law and policy over the
257
00:22:19.440 --> 00:22:25.039
past thirty or so years, and
I'm afraid my gray hair gives that away,
258
00:22:26.480 --> 00:22:30.079
have seen history repeat itself over and
over. First, a new technology
259
00:22:30.079 --> 00:22:34.920
comes along and makes it easier than
ever to copyright, to obtain copyrighted works.
260
00:22:36.480 --> 00:22:41.759
Now, in the ideal scenario that
development is mutually beneficial, creators can
261
00:22:41.799 --> 00:22:47.680
reach new audiences expanded audiences more easily, and the widespread availability of creative works
262
00:22:47.759 --> 00:22:53.160
drives demand for the technology. Everybody
wins. But in practice, the operators
263
00:22:53.200 --> 00:22:59.160
of the technology in the past have
often made choices that allocate to themselves the
264
00:22:59.200 --> 00:23:03.039
lions share of the income. In
some cases, those choices have included willfully
265
00:23:03.119 --> 00:23:10.319
tolerating infringement on platforms, knowing full
well that creators, especially independent creators,
266
00:23:10.880 --> 00:23:15.680
like both the means and tools to
achieve meaningful vindication of their rights. So
267
00:23:15.839 --> 00:23:22.359
here we are with generative AI systems
built on large language models. It's deja
268
00:23:22.400 --> 00:23:26.680
vu all over again. Such systems
require massive volumes of works in order to
269
00:23:26.680 --> 00:23:33.240
be capable of producing the commercially valuable
outputs the designers seek to market. That
270
00:23:33.319 --> 00:23:37.400
fact is not in dispute, but
how those works are obtained is a commercial
271
00:23:37.480 --> 00:23:41.759
choice. There is nothing in the
nature of the technology that requires those works
272
00:23:41.799 --> 00:23:48.440
to be scraped without notice, without
authorization, or without compensation. Yet that's
273
00:23:48.480 --> 00:23:53.680
precisely what's happened. In the public
policy sphere, people willing to defend those
274
00:23:53.720 --> 00:23:59.880
decisions often try to create a false
dichotomy between the massive, unauthorized scraping and
275
00:24:00.039 --> 00:24:04.079
the existence of generative AI. The
reality is that there are companies that have
276
00:24:04.160 --> 00:24:11.279
built generative AI systems unlicensed materials.
Those that choose to do otherwise are not
277
00:24:11.359 --> 00:24:15.920
engaged in a crusade for the betterment
of humanity. They are commercial enterprises trying
278
00:24:15.920 --> 00:24:19.960
to avoid paying for critical inputs.
As the chairwoman of the Federal Trade Commission
279
00:24:21.000 --> 00:24:26.200
recently said very plainly, firms cannot
use claims of innovations as an excuse for
280
00:24:26.319 --> 00:24:32.279
law breaking. So let me turn
to the particular legal issues in this case,
281
00:24:32.359 --> 00:24:37.400
and I'm going to focus on the
direct copyright infringement issues. One would
282
00:24:37.440 --> 00:24:41.680
think that in a circumstance when computers
were and our programmed to crawl the Internet
283
00:24:41.720 --> 00:24:47.480
and copy literally billions of works,
the largest copying effort in history, that
284
00:24:47.559 --> 00:24:51.440
it would be beyond serious contention that
the reproduction right of those works has been
285
00:24:51.480 --> 00:24:56.599
implicated, And yet open AI and
others are trying to put that exact matter
286
00:24:56.640 --> 00:25:00.960
in dispute. Anyone who is even
a passing understanding of how computers operate knows
287
00:25:00.960 --> 00:25:04.920
that computers must make copies in order
to process what has been input into them,
288
00:25:06.960 --> 00:25:10.000
and the New York Times evidence shows
that by inputting a certain set of
289
00:25:10.000 --> 00:25:14.960
prompts, substantially, if not strikingly, similar copies of their original works will
290
00:25:15.000 --> 00:25:18.759
be output by open AI, So
it seems self evident that copies of the
291
00:25:18.799 --> 00:25:22.000
original works must be in the computer
memory in order for that to happen.
292
00:25:23.599 --> 00:25:29.279
Still, open AI argues to the
contrary, if the internal operation of the
293
00:25:29.279 --> 00:25:33.119
system were transparent to the public,
we would have real insight into the facts
294
00:25:33.119 --> 00:25:37.480
of how it operates. But despite
the name, it was given Open AI,
295
00:25:37.559 --> 00:25:41.240
like other generative AI systems, is
in fact locked up tight. Perhaps
296
00:25:41.279 --> 00:25:45.839
some of this will come out in
discovery, but in any event, I
297
00:25:45.880 --> 00:25:51.240
am deeply skeptical that there's any serious
argument other than that the unauthorized scraping of
298
00:25:51.279 --> 00:25:56.720
copyrighted works does implicate the reproduction right, which means, as my fellow panelists
299
00:25:56.720 --> 00:26:00.839
have already said, the real action
in this case be in the fair use
300
00:26:00.960 --> 00:26:04.920
argument. As has already been noted, the fair use assessment in the context
301
00:26:04.920 --> 00:26:11.240
of the first factor is likely to
involve consideration of whether opening eyes copying constitutes
302
00:26:11.279 --> 00:26:17.359
a transformative use. While the term
transformative has been part of copyright jurisprudence for
303
00:26:17.359 --> 00:26:21.799
a very long time, it was
given special significance by the Supreme Court decision
304
00:26:21.880 --> 00:26:26.400
in the two Life Crew case Campbell
Vis's A Cuff Rose to give the proper
305
00:26:26.400 --> 00:26:30.480
caption in nineteen ninety four. Since
that time, lower courts has struggled to
306
00:26:30.519 --> 00:26:36.480
apply this term, sometimes resulting in
extreme results, such as when Google's forbatim
307
00:26:36.519 --> 00:26:40.920
copying of tens of millions of books
was held to be highly transformative by the
308
00:26:40.920 --> 00:26:45.599
Second Circuit. Fortunately, the Supreme
Court had occasion to revisit this doctrine in
309
00:26:45.640 --> 00:26:52.000
twenty twenty two in Warhol Foundation versus
Goldsmith and articulated a much more reasonable and
310
00:26:52.039 --> 00:26:56.839
workable framework. So I think the
lower court fair use decisions that pre date
311
00:26:56.920 --> 00:27:03.599
Goldsmith are now of questionable applicability.
The first factor begins with a contrast between
312
00:27:03.680 --> 00:27:07.519
nonprofit use, which is favored,
and commercial use, which is disfavored.
313
00:27:07.200 --> 00:27:11.400
Prior to Goldsmith, some courts were
finding that a transformative use not only negated
314
00:27:11.400 --> 00:27:17.880
the commerciality but essentially overtook all the
other fair use factors as well. But
315
00:27:17.960 --> 00:27:21.240
a Goldsmith, the Supreme Court was
much more measured, holding the weight of
316
00:27:21.240 --> 00:27:26.319
commerciality against fair use can be lessened
by the degree to which the use is
317
00:27:26.319 --> 00:27:30.200
transformative, that is, has a
further purpose or different character. It's a
318
00:27:30.200 --> 00:27:37.119
sliding scale, not a Boolean analysis. So what constitutes transformative use? Some
319
00:27:37.200 --> 00:27:41.480
courts had gone so far in finding
any new element of the use to be
320
00:27:41.519 --> 00:27:45.960
transformative that many commentators wondered what if
anything, was left of the statutory right
321
00:27:47.200 --> 00:27:52.640
to authorize the creation of derivative works. As they mentioned, the Goldsmith Court
322
00:27:52.680 --> 00:27:56.720
wrote, to make transformative use of
an original must go beyond that required to
323
00:27:56.759 --> 00:28:03.039
qualify as a derivative use that has
a distinct purposes justified because it furthers the
324
00:28:03.039 --> 00:28:06.880
goal of copyright, namely to promote
the progress of science and the arts,
325
00:28:06.920 --> 00:28:14.519
without diminishing the incentive to create.
Quoting author's gildmersus Google, the court continued,
326
00:28:14.960 --> 00:28:18.799
the more the appropriator is using the
copied material for new transformative purposes,
327
00:28:19.079 --> 00:28:23.000
the less likely it is that the
appropriation will serve as a substitute for the
328
00:28:23.039 --> 00:28:29.200
original or its plausible derivatives. So
we see that substitution is a key part
329
00:28:29.200 --> 00:28:33.079
of the analysis, one that has
the power to disqualify use from transformative status.
330
00:28:33.880 --> 00:28:40.279
And the Times has shown that OpenAI
is capable of preproducing substantially similar copies
331
00:28:40.319 --> 00:28:45.000
of its works. That's powerful evidence
on the substitution effect that AI will have
332
00:28:45.039 --> 00:28:52.640
to contend with when this case moves
forward. The Supreme Court in Goldsmith summarized
333
00:28:52.640 --> 00:28:56.480
its first factor analysis with the passage
quote. The first factor considers whether the
334
00:28:56.599 --> 00:29:00.519
use of a copyrighted work has a
further purpose or different character, which is
335
00:29:00.559 --> 00:29:04.519
a matter of degree, and the
degree of difference must be balanced against the
336
00:29:04.559 --> 00:29:10.079
commercial nature of the use quote quote. Those who are sympathetic to the AI
337
00:29:10.200 --> 00:29:12.359
side will point to the creation of
new works by the use of the resulting
338
00:29:12.400 --> 00:29:18.680
system. But this misses a key
step. Generative AI systems do not create
339
00:29:18.720 --> 00:29:22.720
images or text or anything on their
own initiative. Human interaction is required in
340
00:29:22.759 --> 00:29:27.119
the form of prompts, and without
that human interaction, the resulting product would
341
00:29:27.160 --> 00:29:33.319
have no protectable creative expression. So
it isn't the AI system that generates new
342
00:29:33.319 --> 00:29:37.839
works. It is the human users. And it's been good law for a
343
00:29:37.960 --> 00:29:41.599
very long time that the copyer of
works may not stand in the shoes of
344
00:29:41.640 --> 00:29:47.319
the users of those copies to justify
the initial copying. That, of course,
345
00:29:47.359 --> 00:29:51.640
is Michigan Document Services versus Princeton University
Press. We can talk about that
346
00:29:51.680 --> 00:29:55.759
if it comes up in the Q
and A. As for the remaining fair
347
00:29:55.880 --> 00:29:59.319
use factors, it seems to me
they all favor the copyright owner in varying
348
00:29:59.359 --> 00:30:03.519
degrees. The works were largely creative, although perhaps some news articles marginally less
349
00:30:03.559 --> 00:30:08.119
so still I think it cleared.
The second factor favors the Times. The
350
00:30:08.279 --> 00:30:14.119
entirety of every work was copied,
so the third factor is strongly against fair
351
00:30:14.240 --> 00:30:17.920
use, and the fourth factor,
harmed to the current or potential market,
352
00:30:18.039 --> 00:30:22.799
very strongly favors the Times. Again, substitution is probably the strongest evidence of
353
00:30:22.839 --> 00:30:26.079
harm, followed closely by the existence
of a licensing market for the same use,
354
00:30:26.799 --> 00:30:30.799
both of which we have here.
So if open AI is going to
355
00:30:30.799 --> 00:30:33.920
prevail on fair use, it seems
to me they will need a very strong
356
00:30:33.960 --> 00:30:37.559
finding on the first factor in order
to overcome not only the commercial nature of
357
00:30:37.599 --> 00:30:41.200
the use, but all the other
factors as well. On the law,
358
00:30:41.240 --> 00:30:45.759
this seems unlikely. I can only
imagine a finding of fair use frankly if
359
00:30:45.759 --> 00:30:49.920
the court is taken in by the
cool factor arguments. But it would be
360
00:30:49.960 --> 00:30:55.960
a travesty for creators works to lose
essentially all protection in the aihe and it
361
00:30:56.000 --> 00:31:00.000
would certainly harm the incentive to create, which is what the Goldsmith Court kept
362
00:31:00.079 --> 00:31:07.680
certainly in mind. Thanks, we
have a couple of questions before I go
363
00:31:07.799 --> 00:31:11.079
there. I think you addressed this, Steve, and that is the substitution,
364
00:31:11.319 --> 00:31:18.839
the competitiveness in the market, and
in particular that particular fact which the
365
00:31:19.160 --> 00:31:23.640
complaint I believe goes to some links
to try to create a competitive market.
366
00:31:23.920 --> 00:31:30.599
And how does that get them out
of or diminish the Google Books copying case.
367
00:31:33.599 --> 00:31:37.720
Well, I'm not sure the degree
to which Google Books is still good
368
00:31:37.799 --> 00:31:42.559
law after Goldsmith. Not the Goldsmith
overturned that decision in any way, but
369
00:31:42.799 --> 00:31:51.440
the analysis was so different that it's
hard to make that case any longer.
370
00:31:51.480 --> 00:31:57.880
That the Google books really is a
useful president when you look at this at
371
00:31:57.920 --> 00:32:04.400
ai generative systems in the most general
way and kind of departing for a moment
372
00:32:04.440 --> 00:32:07.680
from the particulars of the law,
but in a matter of social policy,
373
00:32:08.359 --> 00:32:17.920
we have systems where people's work product
was copied without their knowledge, permission,
374
00:32:19.640 --> 00:32:25.880
or compensation, and it's being used
to design and implement technology that will then
375
00:32:25.960 --> 00:32:30.839
put many of those people out of
business entirely or substantially reduce their income by
376
00:32:30.880 --> 00:32:37.200
competing directly with them. There's also
questions about, you know, well,
377
00:32:37.200 --> 00:32:40.839
why did you copy this work versus
that work? Does any individual work have
378
00:32:42.279 --> 00:32:49.559
particular value? Works were chosen because
they are valuable for what they are,
379
00:32:50.559 --> 00:32:54.599
and so there is value there.
And I think at the highest level of
380
00:32:55.079 --> 00:33:02.519
public policy analysis, value was taken
and value has been created in companies that
381
00:33:02.559 --> 00:33:12.160
are already valuing themselves at over trillion
dollars, and there's got to be a
382
00:33:12.200 --> 00:33:19.960
way to compensate the people whose works
were taken to build those systems. It's
383
00:33:20.000 --> 00:33:23.319
beyond this panel talk about exactly how
that works, going back to the particulars
384
00:33:23.359 --> 00:33:28.839
of the copyright law, the fact
that this goes back to get to the
385
00:33:28.880 --> 00:33:34.160
reproduction right. It seems self evident
to me that copies were made, and
386
00:33:34.200 --> 00:33:39.319
I've just made the fair use analysis
that I think is appropriate. So I'm
387
00:33:39.640 --> 00:33:45.039
I wouldn't be surprised to see the
times prevail in this case. But time
388
00:33:45.079 --> 00:33:52.160
will tell. So we have a
question that goes to ask quickly on I'm
389
00:33:52.200 --> 00:33:55.960
sorry, go ahead, Yes,
I just wanted to say I think it's
390
00:33:55.960 --> 00:34:00.960
slight different tack on Google books.
I think it's very much of a piece
391
00:34:01.000 --> 00:34:04.720
of Google image Search, and I
think it has to be read as fact
392
00:34:04.759 --> 00:34:07.480
specific. I mean, whether or
not it's good law after Warhol, I
393
00:34:07.519 --> 00:34:14.079
don't know, but I really think
that the focus on stipid view and directing
394
00:34:14.119 --> 00:34:19.079
people to booksellers and gootle books,
and the focus on directing people to websites
395
00:34:19.079 --> 00:34:23.039
and Google image search with with key
too decisions in terms of a market effect
396
00:34:23.079 --> 00:34:27.679
and transformation. If the use was
not a substitute in any way, at
397
00:34:27.760 --> 00:34:31.320
least in the courts view, but
rather a channeling use of sorts and so
398
00:34:31.440 --> 00:34:37.039
I do think it has to be
a difference there where there's no channeling whatsoever
399
00:34:37.199 --> 00:34:40.679
here. But like Steve said,
we'll see, well so if I can
400
00:34:40.800 --> 00:34:45.199
just kind of jump in a very
quickly. So one of the arguments that
401
00:34:45.480 --> 00:34:49.719
is sort of previewed in the in
the motions of the Smiths is this question
402
00:34:49.840 --> 00:34:58.440
of to what extent the actual substitutional
use is actually relevant. And so open
403
00:34:58.480 --> 00:35:02.239
Aye and Microsoft both point exhibit jay
of the complaint, which is where you
404
00:35:02.280 --> 00:35:07.320
get sort of the original duplication or
memorization as they call it, in the
405
00:35:07.320 --> 00:35:14.079
industry of articles. And as those
examples show, in order to get a
406
00:35:14.519 --> 00:35:19.559
in order to prompt chatterp to produce
a New York Times article, you first
407
00:35:19.559 --> 00:35:22.280
have to give it the first half
the article, essentially. And so their
408
00:35:22.400 --> 00:35:24.960
argument then is that at that point, if you already have the first half
409
00:35:24.960 --> 00:35:28.679
of the article, it's not much
of a substitution to say that you get
410
00:35:28.679 --> 00:35:30.599
the second half, especially when the
second half is available in other places on
411
00:35:30.639 --> 00:35:35.960
the Internet. And so that's an
interesting argument. I haven't heard that before,
412
00:35:36.000 --> 00:35:40.199
and it's sort of unique to this
sort of situation, but I think
413
00:35:40.199 --> 00:35:44.360
that what it does mean is that
it's going to be a little more complicated,
414
00:35:44.360 --> 00:35:49.360
I think to prove the substitution element
compared to some of the other cases.
415
00:35:49.400 --> 00:35:52.199
That doesn't mean that, you know, I think that's it's a slamm
416
00:35:52.199 --> 00:35:53.960
don win for them. It certainly
is the case that you can get a
417
00:35:54.039 --> 00:35:59.880
large chunk of copyright material. It's
just that it's not the normal way that
418
00:36:00.039 --> 00:36:05.280
one would expect to retrieve it.
There's a lot of focus on that this
419
00:36:05.440 --> 00:36:12.079
is a bug rather they want it
would legally shouldn't be significant, but has
420
00:36:12.119 --> 00:36:15.840
come up before, so who knows. Yeah, my understanding is that there
421
00:36:15.880 --> 00:36:22.079
also is an amount of research on
just how to stop at a computer level
422
00:36:22.360 --> 00:36:23.599
of the sort of thing from happening. You know, is there a way
423
00:36:23.599 --> 00:36:28.400
that they can just put filters at
the end, or can they use some
424
00:36:28.440 --> 00:36:31.840
sort of special training processes. That
is sort of an interesting research question that
425
00:36:31.880 --> 00:36:35.400
I know is open. I think
that the difficulty is that a lot of
426
00:36:35.440 --> 00:36:39.719
times articles are available online and slightly
different variations, like somebody might admit a
427
00:36:39.719 --> 00:36:46.039
paragraph, and that's why it's hard
for it's hard to simply detect identical copies.
428
00:36:46.800 --> 00:36:49.719
But it's sort of interesting that,
you know, there may be a
429
00:36:49.800 --> 00:36:52.920
technical solution in addition to the legal
approach, and there's a question of which
430
00:36:52.960 --> 00:36:55.440
one you think should lead. Right. On the one hand, maybe we
431
00:36:55.480 --> 00:36:59.719
want to say that the law encourages
people to develop these sorts of technologies that
432
00:37:00.000 --> 00:37:02.840
on void duplication. On the other
hand, maybe we say, maybe we're
433
00:37:02.840 --> 00:37:07.239
worried that by imposing the hammer of
the law too early, we prevent people
434
00:37:07.280 --> 00:37:13.360
from coming up with these sorts of
technologies that can satisfy better middle grounds,
435
00:37:13.400 --> 00:37:15.760
and we end up, you know, one of the one of the remedies
436
00:37:15.800 --> 00:37:20.960
that's asked for is destruction of chat
GPT, and so there's you know,
437
00:37:21.000 --> 00:37:23.840
there is a question of whether or
not it would be premature to be applying
438
00:37:23.880 --> 00:37:29.480
this before we really know what the
sort of opportunities the technology ultimately look like.
439
00:37:30.599 --> 00:37:34.119
So that's what I was referencing earlier
when I said that some people want
440
00:37:34.159 --> 00:37:37.800
to look at this as copyright versus
AI, and you can't have both.
441
00:37:38.880 --> 00:37:45.639
Had had open Ai chosen to license
the materially used, that no one would
442
00:37:45.679 --> 00:37:51.199
be asking to destroy chat GPT,
nor is anyone asking to destroy AI systems
443
00:37:51.239 --> 00:37:57.239
that were built on licensed material.
So putting that aside for a moment,
444
00:37:57.880 --> 00:38:02.559
the substitute effect had several layers to
it. First and foremost, the fact
445
00:38:02.559 --> 00:38:09.840
that you can get that material out
of chat GPT is strong evidence that the
446
00:38:10.119 --> 00:38:15.000
copyrightable material is somewhere in the chat
GPT system, the open AI system,
447
00:38:15.320 --> 00:38:22.280
and therefore the reproduction right is implicated. Then on the question of whether the
448
00:38:22.320 --> 00:38:29.679
output is infringing and the substitution effect. There we kind of have two or
449
00:38:29.719 --> 00:38:37.159
three different layers in traditional copyright analysis. Obviously, substantial similarity is an access
450
00:38:37.519 --> 00:38:46.960
are the two problems to proving copying
and infringing reproduction. So the fact that
451
00:38:47.000 --> 00:38:53.920
it may be difficult or unusual to
get chat GPT to put that out doesn't
452
00:38:53.960 --> 00:38:58.599
mean it's less infringing. It just
means it may take a little more work,
453
00:38:58.639 --> 00:39:05.400
but it's still infringing. More generally, more, there's a question about
454
00:39:05.920 --> 00:39:15.239
the degree to which these generative AI
systems will overall all diminish demand for the
455
00:39:15.280 --> 00:39:21.400
copyrighted works produced by the people who
produce the works that were used to build
456
00:39:21.440 --> 00:39:29.039
the system, and that admittedly departs
somewhat from a direct infringement analysis to more
457
00:39:29.079 --> 00:39:34.840
public policy question of is it right
that the people whose works were taken without
458
00:39:34.920 --> 00:39:38.800
notice, without permission, and without
compensation should be told, sorry, you're
459
00:39:38.840 --> 00:39:45.199
out of lock, and these new
systems that we built with all the things
460
00:39:45.239 --> 00:39:53.719
we took from you are now going
to put you out of business. We
461
00:39:53.760 --> 00:40:00.320
have a question from the audience concerning
generally to the panelists together, and the
462
00:40:00.400 --> 00:40:07.360
specific question is relevance of sega versus
accolade and the concept of intermediate cop copying.
463
00:40:08.360 --> 00:40:14.559
Can any of you guys speak to
that? Yeah, so this is
464
00:40:14.719 --> 00:40:19.800
also I was mentioning that there is
this sort of question of whether or not
465
00:40:20.239 --> 00:40:27.000
sort of like purely internal use would
count as an infringement. And there are
466
00:40:27.000 --> 00:40:31.639
a couple of cases that say that
at least that's an interesting question. So
467
00:40:30.400 --> 00:40:37.639
so the say atholate case and the
reverse engineering cases they raise that question.
468
00:40:37.639 --> 00:40:42.760
I think you probably remember this is
the cartoon network case, right, which
469
00:40:42.800 --> 00:40:46.639
talks about transitory copying. So you
have a number of you have a number
470
00:40:46.639 --> 00:40:52.480
of cases that support the sort of
idea that maybe like internal copying or copying
471
00:40:52.519 --> 00:41:00.719
that is precedent to some other public
use might not receive the same analogy assists.
472
00:41:00.480 --> 00:41:04.960
And so that's obviously going to be
relevant here because you know to the
473
00:41:05.039 --> 00:41:07.920
extent that the allegation is that the
copying is inside the model. The model
474
00:41:07.960 --> 00:41:14.239
is not something anybody can see,
and so that might fall within those doctrines.
475
00:41:14.239 --> 00:41:17.159
Similarly, the training process, again
that's not something anybody actually sees,
476
00:41:17.719 --> 00:41:23.639
and as a result, those aren't
those potentially aren't necessarily the sorts of things
477
00:41:23.639 --> 00:41:30.159
that are of concern to copyright law. In view of those cases, it's
478
00:41:30.199 --> 00:41:32.119
really kind of you know, what's
going on the output end that has the
479
00:41:32.119 --> 00:41:36.679
commercial impact. That's where the fair
use question is going to really come up.
480
00:41:36.760 --> 00:41:38.760
That's where the market effect is going
to come up. That's where transformativeness
481
00:41:38.840 --> 00:41:44.079
is going to come out. Yeah, that that that's that that's at least
482
00:41:44.079 --> 00:41:50.199
part of kind of where where those
cases fit in. Go ahead, see
483
00:41:50.239 --> 00:41:52.960
I will say, I mean my
view of both Accolae case and Cartoon Network
484
00:41:53.039 --> 00:41:59.800
that they boiled down to if you're
doing some internal copying for what otherwise illegally
485
00:42:00.199 --> 00:42:02.880
and non copying use. Now it
will say where Cartoon Network with plenty of
486
00:42:02.920 --> 00:42:08.840
copying, but it was all licensed
and and in Siga xacol it it's interoperability
487
00:42:08.840 --> 00:42:14.760
copying and just that you could reverse
engineer the interface for the game console.
488
00:42:15.320 --> 00:42:17.800
I think those cases basically boiled down
to, well, yeah, if you're
489
00:42:17.840 --> 00:42:22.119
if you're if not actually infringing in
any way. It's just a way of
490
00:42:22.119 --> 00:42:27.239
getting to that conclusion. And with
the other words, totally legal intermediate copying
491
00:42:27.280 --> 00:42:29.480
is okay. But I don't think
that's the case here. The whole problem
492
00:42:29.559 --> 00:42:34.639
if it chat GPT is spitting out
lots of copyrighted content and it's copyright content
493
00:42:34.679 --> 00:42:37.119
in and out, I think that's
the problem. Well, I mean that
494
00:42:37.199 --> 00:42:38.280
that's sort of the question of the
case, right, if it turns out
495
00:42:38.280 --> 00:42:40.800
that that is fair use. Right, if it turns out that, you
496
00:42:40.800 --> 00:42:44.519
know, they accept the argument that
this is a bug and you know,
497
00:42:44.880 --> 00:42:47.079
you know deminimus spitting out of exact
copy. You know, I don't know
498
00:42:47.199 --> 00:42:51.280
exactly how the court would rule,
but they say that that's okay, then
499
00:42:51.320 --> 00:42:54.000
the precedent copying should be okay as
well. We don't want we don't we
500
00:42:54.000 --> 00:42:58.960
don't want basically copyright law to be
focused on trivialities in a sense. So
501
00:42:59.280 --> 00:43:01.840
you know, kind of the question
is it's sort of a standardfall together situation
502
00:43:02.079 --> 00:43:08.440
where if the outputs end up being
problematic not not protected under fair use,
503
00:43:08.800 --> 00:43:13.480
then the whole thing probably wouldn't be
fair use either. But if it is
504
00:43:13.639 --> 00:43:17.440
then chances are the earlier copying is
okay as well in view of those presents.
505
00:43:19.400 --> 00:43:22.559
I think that's a little bit too
narrow in one respect, Sega versus
506
00:43:22.559 --> 00:43:30.199
Accolade was predicated on the fact the
ruling was predicated on the fact that the
507
00:43:30.239 --> 00:43:36.000
resulting products that Accolade wanted to create
were not competing with what they copied in
508
00:43:36.079 --> 00:43:42.599
any way, not direct substitutes or
even trying to create competing products as alternatives,
509
00:43:43.159 --> 00:43:45.880
but rather games rather as oppose to
the operating system, which is what
510
00:43:45.920 --> 00:43:53.079
they copied, games that would simply
work on that operating system. Whether generative
511
00:43:53.119 --> 00:44:01.719
AI systems are outputting substantially similar material
to what was copy or even just generally
512
00:44:01.800 --> 00:44:12.000
competing with creators, it's much more
of a market effect than we had in
513
00:44:12.039 --> 00:44:17.320
any way in Sega Versus Accolade,
and I have to on the Cartoon Network
514
00:44:17.320 --> 00:44:22.400
case. I don't really see much
of an analogy there to any part of
515
00:44:22.440 --> 00:44:30.039
this case that dealt with you know, it was all three aspects of the
516
00:44:30.079 --> 00:44:39.039
Cartoon Network case were and remain outlying
outlier decisions on both the temporary copy issue
517
00:44:39.360 --> 00:44:45.840
and the public performance issue, as
well as the authorizing the copy or the
518
00:44:50.000 --> 00:44:53.400
I can't think of the word anyway, none of those if this is not
519
00:44:53.480 --> 00:44:58.800
about public performance, this is not
about temporary copies. Those copies are in
520
00:44:58.840 --> 00:45:00.960
there for a long period of time. I think that you actually doualize that
521
00:45:00.960 --> 00:45:04.920
as a public performance. Interestingly,
I'm not exactly sure why, but I
522
00:45:04.920 --> 00:45:08.039
think that was actually the place.
Well, yes, but it's not the
523
00:45:08.039 --> 00:45:13.800
same sort of issue where oh,
there's only one copy versus ten thousand copies,
524
00:45:14.360 --> 00:45:20.559
And ironically the ten thousand copies is
not a public performance, said Cartoon
525
00:45:20.559 --> 00:45:24.199
Network. In the event, of
all the holdings of Cartoon Network, that
526
00:45:24.280 --> 00:45:30.239
element is surely the weakest. After
the Supreme Court decision in Areo, which,
527
00:45:30.280 --> 00:45:37.079
while not explicitly overruling Cartoon Network,
articulated very clearly an analysis at odds
528
00:45:37.119 --> 00:45:42.480
with the Second Circuits analysis of public
performance in Cartoon Network. I don't see
529
00:45:42.480 --> 00:45:52.880
any way that that's good law.
Do any of you have a view of
530
00:45:53.320 --> 00:46:00.679
the likelihood of a licensing solution to
this or do you think it's going to
531
00:46:00.760 --> 00:46:06.960
resolve on by the court Supreme Court? Probably I think the heavily licensed I
532
00:46:07.000 --> 00:46:09.119
think that I mean not saying I
don't want to say inevitably, you never
533
00:46:09.199 --> 00:46:13.719
know this could be litigated for a
dozen years. But they'll note in the
534
00:46:13.840 --> 00:46:20.000
complaint that New York Times is when
was heavily used sources for open AI,
535
00:46:20.960 --> 00:46:27.559
but other sources, notably Associated Press
or others mentioned in the complaint, had
536
00:46:29.239 --> 00:46:31.639
been paid for licensing. So I
think it's there, But I think,
537
00:46:34.719 --> 00:46:37.280
well, I mean, the old
line is that litigation is negotiation by other
538
00:46:37.360 --> 00:46:40.039
means. I'm not sure it's always
true, but I think it is true
539
00:46:40.039 --> 00:46:45.360
here. Yeah, I mean I
think that you know, there there are,
540
00:46:45.679 --> 00:46:49.320
as Steve mentioned, there are a
number of companies that are out there
541
00:46:49.360 --> 00:46:52.159
developing generative systems that are that are
that are based on that are that are
542
00:46:52.199 --> 00:46:58.320
based on licensing. The difficulty,
of course, which is which I'm sure
543
00:46:58.320 --> 00:47:00.400
opening I would be happy point to
on believe that they point to this in
544
00:47:00.519 --> 00:47:06.599
some of their papers, is that
a lot of the reason why these systems
545
00:47:06.679 --> 00:47:12.239
work particularly well is just the breadth
of content, the fact that they have
546
00:47:13.840 --> 00:47:20.000
large quantities, tremendously large quantities of
data to work off of, which would
547
00:47:20.039 --> 00:47:22.559
be substantially harder to obtain by a
licensing. And so, you know,
548
00:47:22.599 --> 00:47:29.719
while I think it's correct that this
isn't a question of just AI versus licensing
549
00:47:29.760 --> 00:47:32.639
that you that you can have both
at the same time. You aren't changing
550
00:47:32.639 --> 00:47:37.320
the parameters of what the technological development
looks like, right, and that could
551
00:47:37.400 --> 00:47:39.519
be a good thing. I'm not
necessarily saying that it's a bad thing if
552
00:47:39.519 --> 00:47:46.119
computer scientists are forced to work with
a different set of legal parameters. There
553
00:47:46.159 --> 00:47:52.440
could be value in specifically working with
licensed copies. But one thing to keep
554
00:47:52.480 --> 00:47:57.679
in mind is that there is also
a pretty good chunk of poly domain information
555
00:47:57.719 --> 00:48:02.000
that's available that had been the basis
for training of artificial intelligence for many years,
556
00:48:02.239 --> 00:48:05.960
and one of the reasons that companies
have moved away from that is that
557
00:48:05.960 --> 00:48:10.639
it produced very strange systems. A
lot of that material is hundreds of years
558
00:48:10.639 --> 00:48:16.679
old material that's sufficiently out of copyright, or material that is very strange to
559
00:48:16.800 --> 00:48:23.360
use as training data. The en
run email data base was a common one
560
00:48:23.400 --> 00:48:28.280
that was used, and as people
have pointed out, that ended up producing
561
00:48:28.400 --> 00:48:34.800
systems that incorporated all sorts of very
strange biases and so's you know. I
562
00:48:34.840 --> 00:48:37.400
think that it's on the one hand, yeah, you can't have a world
563
00:48:37.480 --> 00:48:43.480
in which licensing is prominent. You
do end up with a different technological environment,
564
00:48:43.639 --> 00:48:47.679
and I think it's it's an interesting
question of what that environment looks like
565
00:48:47.880 --> 00:48:51.440
and how you really feel about it, what it does to the true actually
566
00:48:51.440 --> 00:48:54.280
of the technology. Well, Charles, I think you just made a great
567
00:48:54.320 --> 00:48:59.159
case for the value of the copyrighted
works that were copying in order to build
568
00:48:59.480 --> 00:49:05.320
generative AI systems. Does it take
a little longer? Might it cost more
569
00:49:05.920 --> 00:49:10.199
to compensate creators for using their works
to build your system? Well? Yes,
570
00:49:10.280 --> 00:49:15.960
of course, the same way that
it would take a lot longer if
571
00:49:16.119 --> 00:49:21.800
the AI companies didn't have Nvidia chips
because they didn't want to pay for them,
572
00:49:22.519 --> 00:49:25.960
their processing capacity would be greatly diminished. But they need them and they
573
00:49:27.000 --> 00:49:30.920
want them, so they pay for
them. And I would suggest that the
574
00:49:30.039 --> 00:49:38.360
social aspect of this of telling creators
you're the one input into generative AI that
575
00:49:38.440 --> 00:49:43.079
no one has to pay for.
The coders they get paid, developers,
576
00:49:43.840 --> 00:49:49.559
the chip makers they get paid.
The power utilities that provide all that electricity
577
00:49:49.599 --> 00:49:54.760
to the data farms they get paid. But the creators, sorry, you're
578
00:49:54.840 --> 00:50:00.039
out of lock. That I think
there's an it looked at me in in
579
00:50:00.119 --> 00:50:05.559
that context. It's a pretty difficult
argument to sustain, but I think that
580
00:50:05.760 --> 00:50:12.960
generally correct this idea that if there's
some value out there, then possibly we
581
00:50:13.000 --> 00:50:15.199
want to make sure that there's compensation
with the value. The difficulty here,
582
00:50:15.239 --> 00:50:19.880
of course, is that copyright doesn't
cover every possible value of a work.
583
00:50:20.400 --> 00:50:24.159
Right, you have factual databases which
simply have no copyright protection because facts are
584
00:50:24.159 --> 00:50:29.880
not copyrightable, and so there is
sort of this larger question of, you
585
00:50:29.920 --> 00:50:31.960
know, how do you deal with
the problem value. Similarly, a lot
586
00:50:32.000 --> 00:50:37.039
of these systems use private data,
right, they use individual personal data.
587
00:50:37.079 --> 00:50:40.039
That's data that's not copyright protectable at
all, right, and that is value.
588
00:50:40.119 --> 00:50:43.599
So you know, one way you
could solve as by trying to create
589
00:50:43.599 --> 00:50:46.400
some sort of property right in private
information, but that actually causes a lot
590
00:50:46.440 --> 00:50:52.119
of other problems, as you know, a fair amount of scholarship that's out
591
00:50:52.159 --> 00:50:57.679
there. So I think that you
want to separate out the questions of how
592
00:50:57.679 --> 00:51:04.280
do you deal with the value propose
of information from the copyright questions, and
593
00:51:04.880 --> 00:51:07.840
if you're talking about just sort of
the value propositions, maybe once you're really
594
00:51:07.840 --> 00:51:10.320
asking for some sort of regulatory agency, which you know, we can debate
595
00:51:10.360 --> 00:51:15.840
that, but then you also have
to have the second question of is copyright
596
00:51:15.880 --> 00:51:22.880
the correct vehicle for doing that given
that copyright is historically limited to certain things
597
00:51:22.960 --> 00:51:27.159
which are a little bit different from
what exactly is going on with artificial intelligence
598
00:51:27.159 --> 00:51:32.239
systems. The fact versus the idea
versus expression, the fact versus expression dichotomies,
599
00:51:34.119 --> 00:51:37.239
and of course the fair use doctrine, these all play into kind of
600
00:51:37.239 --> 00:51:43.039
what the boundaries of copyright law are
that don't make value exactly aligned with the
601
00:51:43.119 --> 00:51:46.559
with the rights provide unfederal law.
I do agree there are elements of these
602
00:51:46.599 --> 00:51:52.800
issues that go beyond copyright law,
and you're seeing all sorts of proposals in
603
00:51:52.800 --> 00:52:01.239
those veins, from privacy protections to
state and federal legion on issues of publicity
604
00:52:01.320 --> 00:52:07.760
rights, narrowing it to the copyright
issues, And because that is the focus
605
00:52:07.800 --> 00:52:14.760
of this panel, I do think
that of course collections of data can be
606
00:52:14.840 --> 00:52:19.199
protectable as compilations by their selection,
arrangement coordination, as I'm sure you know.
607
00:52:20.639 --> 00:52:24.320
And then many of the works that
were taken were taken precisely because they
608
00:52:24.360 --> 00:52:36.079
are valuable expression of contemporary American society, and we don't want people aren't going
609
00:52:36.119 --> 00:52:40.079
to pay for AI systems that speak
in old English or forsooth or what have
610
00:52:40.199 --> 00:52:46.679
you. So when I speak of
the value. I'm speaking both generally beyond
611
00:52:46.719 --> 00:52:52.400
copyright, but also in the context
of the fourth factor, where there's harm
612
00:52:52.480 --> 00:53:00.960
to the current and potential market or
for the works that were copied. I
613
00:53:00.000 --> 00:53:06.679
have a question from the audience's question
of possible to the panel, I'm tired
614
00:53:07.119 --> 00:53:12.800
is tirty? Is is there a
place in this discussion for Sony safe harbor?
615
00:53:15.239 --> 00:53:21.840
Yeah? For Sony you said,
right, yes, yeah, So
616
00:53:21.960 --> 00:53:24.920
I actually pulled along from Sony.
Everyone reads Sony for this first half,
617
00:53:24.960 --> 00:53:31.280
which is capable of substance of you
know, subject of commercially significant non infringing
618
00:53:31.440 --> 00:53:37.880
uses. Of course, in Sony
the court held that time shifting was okay,
619
00:53:37.159 --> 00:53:40.920
but librarying was probably not. And
everyone forgets to the second half.
620
00:53:44.320 --> 00:53:46.840
There's room for it, but I
don't think it gets you there in this
621
00:53:47.000 --> 00:53:53.199
case. I think that clearly open
Aye had well. Versually is an interesting
622
00:53:53.280 --> 00:53:57.599
question. It's a capable of non
infringing uses, I mean through it depends
623
00:53:57.639 --> 00:54:04.679
how you define it. What what
point is non fringing? And I mean
624
00:54:05.079 --> 00:54:08.000
actually a lot of the most of
dismiss was on sexual limitations, but a
625
00:54:08.039 --> 00:54:13.519
hepstink discovery rule was going to bring
in the ingestion of a lot of that
626
00:54:13.599 --> 00:54:16.280
as well. Potentially of course,
then you know we might be waiting to
627
00:54:16.280 --> 00:54:21.719
find out what win our shop versus
Neely holds, or if anything is held
628
00:54:21.760 --> 00:54:27.320
in that case. This is a
three year sexual limitations case. Yes,
629
00:54:27.639 --> 00:54:31.920
so on the Sony case. This
is for those of you who weren't familiar
630
00:54:31.960 --> 00:54:37.639
with it. This is Sony Betamax
issue from the early nineteen eighties when the
631
00:54:37.840 --> 00:54:45.639
movie studios sued claiming infringement by virtue
of early video cassette recording of broadcast television.
632
00:54:45.280 --> 00:54:50.320
And as we said there was there
were two findings. One was that
633
00:54:50.840 --> 00:54:58.800
recording a show that was broadcast over
free broadcast television strictly for the purpose of
634
00:54:58.880 --> 00:55:05.440
watching it later at a more convenient
time is a fair use, but keeping
635
00:55:05.480 --> 00:55:10.199
libraries of those is not. And
then second, in terms of the contributory
636
00:55:10.280 --> 00:55:19.239
liability of the manufacturer of the Betamax
machine, Sony in that case, that
637
00:55:20.079 --> 00:55:27.039
the existence of commercially significant, substantial
noninfringing uses that the court would not impute
638
00:55:27.159 --> 00:55:30.920
knowledge of the infringement, and knowledge
being one of the two problems of contributory
639
00:55:30.960 --> 00:55:37.239
liability. The Supreme Court later in
Rockster made very clear that there is no
640
00:55:37.400 --> 00:55:40.880
safe harbor. This question was framed
as a safe harbor, and Sony that
641
00:55:42.000 --> 00:55:45.440
is not the correct framing of the
law. There is not a safe harbor.
642
00:55:47.000 --> 00:55:52.079
If there are commercially significant, non
infringing uses, then Sony is still
643
00:55:52.119 --> 00:55:57.519
good law that the court will not
impute knowledge for that reason. That doesn't
644
00:55:57.559 --> 00:56:00.719
mean the court might not impute knowledge
for another reason, in which it did
645
00:56:00.760 --> 00:56:06.679
in Brockster, and that reason being
in that case that the defendant had induced
646
00:56:06.840 --> 00:56:12.320
the infringements by the users of the
Groxter system. So all of that,
647
00:56:12.719 --> 00:56:15.800
of course is in the context of
contributory liability, which is a doctrine of
648
00:56:15.840 --> 00:56:22.159
secondary or third party liability in US
law. But there are direct infringement issues
649
00:56:22.239 --> 00:56:27.440
here that I think need to be
resolved first, and that's where the fair
650
00:56:27.559 --> 00:56:31.480
use comes in. The fair use
argument comes in which Sony has nothing to
651
00:56:31.519 --> 00:56:37.639
say about in this context. Yeah, so I I know we're pretty close
652
00:56:37.639 --> 00:56:39.119
to out of time. But if
I'm just say a couple of things about
653
00:56:39.159 --> 00:56:44.400
Sony, so Sony is simultaneously irrelevant
and also highly relevant to this case.
654
00:56:45.239 --> 00:56:49.239
So the reason that's not terribly relevant
is that Sony dealt with this contributory infringement
655
00:56:49.280 --> 00:56:52.880
issue and here the allegation is direct
infringement, right, And so as a
656
00:56:52.880 --> 00:56:58.199
result, by doctrine it doesn't help
us other than for the time shifting and
657
00:56:58.239 --> 00:57:01.280
the question of what counts as fair
use. But the way it is relevant
658
00:57:01.360 --> 00:57:06.920
is that Sony is within a line
of cases that deal with the intersection between
659
00:57:06.920 --> 00:57:09.760
copyright law and technology. Right.
So you have a new technology, the
660
00:57:09.840 --> 00:57:15.199
VCR, which doesn't itself say much
about any particular copyright work, but is
661
00:57:15.239 --> 00:57:21.960
a vehicle by which infringements can occur. And similarly, you had cases about
662
00:57:22.039 --> 00:57:24.079
xerox machines, photographs. You can
go all the way back to the printing
663
00:57:24.119 --> 00:57:28.519
press if you really want, and
you recognize that you have these general purpose
664
00:57:28.519 --> 00:57:35.840
technologies that can enable copyright infringement,
that can enable easier copying. How do
665
00:57:35.880 --> 00:57:39.519
you deal with that sort of balance? And really the answer comes down to
666
00:57:39.639 --> 00:57:45.239
what exactly are the boundaries of the
copyright? Right. If we say that
667
00:57:45.280 --> 00:57:49.119
copyright involves any sort of copying at
all, then of course all of these
668
00:57:49.119 --> 00:57:52.239
things are infringement, and you wouldn't
be able to have printing process, photographs,
669
00:57:52.320 --> 00:57:55.320
xerox machines, any of these sorts
of things. But of course there
670
00:57:55.360 --> 00:57:58.880
is a fair use doctrine. That
fair use doctrine exists for a reason.
671
00:57:58.920 --> 00:58:04.440
It's to set a bound around what
the copyright right protects and to allow for
672
00:58:04.480 --> 00:58:07.079
technological innovation beyond that space. And
so in a sense, what this case
673
00:58:07.119 --> 00:58:12.599
really is all about is what exactly
is that line. Where does the line
674
00:58:12.800 --> 00:58:19.719
fall such that an act that colloquially
we would call copying falls outside of copying
675
00:58:19.760 --> 00:58:22.760
and falls within the space of permissible
innovation. That's a line that's been pretty
676
00:58:22.760 --> 00:58:27.800
important throughout history and a lot of
different technologies. And so in that sense,
677
00:58:27.880 --> 00:58:31.039
this case is not terribly new.
It's simply an iteration of that same
678
00:58:31.159 --> 00:58:36.840
problem that has come up many many
times for copyright law and for a lot
679
00:58:36.880 --> 00:58:42.760
of other areas of law, how
it intersects with new technologies. So that's
680
00:58:42.880 --> 00:58:46.199
the false dichotomy of we wouldn't have
the printing pressor's xerox machines if we didn't
681
00:58:46.199 --> 00:58:51.400
have fair use, because there's no
such thing as licensing. That's of course
682
00:58:51.480 --> 00:58:57.039
not the case, Nor do I
believe that there's some sort of penumber around
683
00:58:57.079 --> 00:59:06.519
Sony or other copyright andechnological innovation cases. And what I do agree with is
684
00:59:06.559 --> 00:59:13.000
that copyright has always been a law
that has reacted to a variety of factors,
685
00:59:13.039 --> 00:59:17.320
developments in the marketplace, consumer preferences, and technological evolution. Absolutely,
686
00:59:19.599 --> 00:59:23.159
and the court looks at some of
the particulars and fair use has a key
687
00:59:23.280 --> 00:59:29.440
role in that, of course.
So we have decisions like Rockster, as
688
00:59:29.440 --> 00:59:32.920
I mentioned, Ario, another one
I mentioned, and many others where the
689
00:59:34.000 --> 00:59:37.400
court said, no, this is
a model based on infringement and we're not
690
00:59:37.440 --> 00:59:42.079
going to permit it. We have
others where the court, like in Sony,
691
00:59:42.199 --> 00:59:46.000
said this is inherently meant to be
an innocent business that could be used
692
00:59:46.039 --> 00:59:50.760
for something afarious, but we don't. We're not going to hold the manufacturer
693
00:59:50.840 --> 00:59:55.360
or the device secondarily liable just because
it could be used in a bad way.
694
00:59:57.519 --> 01:00:00.760
In the case before us, these
generatives AI systems that made a conscious
695
01:00:00.800 --> 01:00:07.519
decision to copy copyrighted works by the
billions. I think it's a lot more
696
01:00:07.679 --> 01:00:17.800
like Ario and Rockster than it is
like the innocent and business models Aily.
697
01:00:17.880 --> 01:00:22.800
I know we're up against the hour, and I have just one question for
698
01:00:22.880 --> 01:00:25.360
the panel as a whole, and
the question is what if we remove the
699
01:00:25.400 --> 01:00:30.880
technology from this discussion and just put
it into a business sense that for some
700
01:00:30.920 --> 01:00:36.320
reason the company starts a business and
happens to go into the New York Times
701
01:00:36.400 --> 01:00:40.239
archives and carts it all off into
their warehouse, and then they adhire a
702
01:00:40.360 --> 01:00:45.239
whole bunch of employees for people who
come into the front counter and ask them
703
01:00:45.320 --> 01:00:50.039
questions, and they go back and
skirt around the New York Times information and
704
01:00:50.039 --> 01:00:55.599
give them an answer or or an
asset that's directly competing with what the New
705
01:00:55.719 --> 01:01:01.000
York Times uses is archives for Is
is that a valid analogy? Is it,
706
01:01:01.079 --> 01:01:06.920
you know, technologically and misplaced or
does that make a difference in the
707
01:01:06.960 --> 01:01:12.239
discussion. If you remove the technology
from what's actually happening with the copyright,
708
01:01:12.280 --> 01:01:17.039
it works, John, I think
I think you're a little description. It's
709
01:01:17.119 --> 01:01:22.159
kind of close to a library,
which is okay, but you do get
710
01:01:22.199 --> 01:01:24.559
into some But if you tweak it
in two ways. One, if you
711
01:01:24.599 --> 01:01:28.199
say, what if they give copies
of Times that are clothes I cart it
712
01:01:28.239 --> 01:01:32.320
off instead of just giving them summaries, or even if like a copy cut
713
01:01:32.360 --> 01:01:37.639
and paste out sentences instead of a
ransom note style answer, I think you're
714
01:01:37.679 --> 01:01:42.760
getting closer to the fat current here. And then the flip side is what
715
01:01:42.920 --> 01:01:49.599
if they grab the New York Times
from you know that from how offer presses
716
01:01:50.320 --> 01:01:54.280
in the East Coast? Why are
summaries of a story so West coast and
717
01:01:54.320 --> 01:01:59.159
print from there? And of course
that's ions versus AP, which was held
718
01:01:59.159 --> 01:02:04.760
to be young awful. So I
don't think you can make it technology independent
719
01:02:04.800 --> 01:02:08.559
because all the technology dictates some of
illegal questions. I mean, to make
720
01:02:08.599 --> 01:02:14.199
things even worse. They do talk
about hallucinations, the problem in which sometimes
721
01:02:14.239 --> 01:02:16.400
you will ask for a New York
Times article from chatchipg, I'll give you
722
01:02:16.400 --> 01:02:20.639
something completely made up that has nothing
to do with that. So the analogy
723
01:02:20.679 --> 01:02:23.519
really is a library reference desk that
gives you, every once in a while
724
01:02:23.639 --> 01:02:27.480
an exact copy of an article that
I found, and every once in a
725
01:02:27.519 --> 01:02:32.199
while gives you totally useless factual information. It's a little hard to draw analogies
726
01:02:32.239 --> 01:02:37.920
to that point, I think,
but it does, and I think it
727
01:02:37.000 --> 01:02:40.039
kind of shows why it's a little
difficult to work with analogies, and sometimes
728
01:02:40.079 --> 01:02:44.039
it really is worth just figuring out
exactly what's going on with the technology.
729
01:02:44.079 --> 01:02:51.159
Yet it's at its proper levels.
I'll just add that what you described was
730
01:02:51.199 --> 01:02:54.880
a lot like a library. If
it's a nonprofit organization, you didn't stipulate
731
01:02:54.920 --> 01:03:01.159
which way that organizations operating nonprofit ormercial. But the point I want to make
732
01:03:01.280 --> 01:03:12.679
is Congress has taken pains to enact
specific statutory exemptions for libraries to provide library
733
01:03:12.679 --> 01:03:19.719
patrons with a certain degree of service
without going so far that it implicates the
734
01:03:20.000 --> 01:03:24.599
incentive to create in the first place. And libraries are not entitled to just
735
01:03:25.440 --> 01:03:30.119
take copies of everything that they want
to have in their collections without paying for
736
01:03:30.159 --> 01:03:36.119
them, without permission and so on. And in fact, in the online
737
01:03:36.159 --> 01:03:39.559
context, some of these issues are
being litigated right now. There's an entity
738
01:03:40.239 --> 01:03:46.039
that calls itself the Internet Archive that
has created a website where it just lets
739
01:03:46.079 --> 01:03:51.760
people take copies of copyrighted works,
and they want to call themselves the library,
740
01:03:52.519 --> 01:03:57.719
but they don't qualify under any of
the library exceptions, and thus far
741
01:03:57.760 --> 01:04:01.320
in the pending litigation they've been quite
unsuccessful. It's on appeals, so we'll
742
01:04:01.320 --> 01:04:03.800
see where it goes. But I
don't I wouldn't put a lot of money
743
01:04:03.800 --> 01:04:13.559
on them. Over to you,
Emily all right on behalf of the Federal
744
01:04:13.599 --> 01:04:16.239
Society. Thank you all for joining
us for this great discussion today. Thank
745
01:04:16.280 --> 01:04:19.880
you also to our audience for joining
us. We greatly appreciate your participation.
746
01:04:20.440 --> 01:04:25.599
Check out our website fedsoc dot org
or follow us on all major social media
747
01:04:25.639 --> 01:04:29.880
platforms at fedsoc to stay up to
date with announcements and upcoming webinars. Thank
748
01:04:29.920 --> 01:04:33.320
you once more for tuning in,
and we are adjourned. Thank you for
749
01:04:33.400 --> 01:04:38.639
listening to this episode of FEDSOC Forums, a podcast of the Federal Societies Practice
750
01:04:38.639 --> 01:04:42.920
Groups. For more information about the
Federal Society, the practice groups, and
751
01:04:42.960 --> 01:04:46.199
to become a Federal Society member,
please visit our website at fedsoc dot or org