1
00:00:10,680 --> 00:00:14,893
Hello, this is Anthony Diana from Reed Smith and welcome to Tech Law Talks.
2
00:00:14,893 --> 00:00:20,157
Today we are continuing a podcast series on AI enabled e-Discovery.
3
00:00:20,157 --> 00:00:31,024
This podcast series will focus on practical and legal issues when considering using AI
enabled e-Disovery with a focus on actual use cases, not just the theoretical.
4
00:00:31,024 --> 00:00:37,750
Joining me today is Dera Nevin from FTI, a good friend of mine and a good friend of Reed
Smith for many, many years.
5
00:00:37,750 --> 00:00:39,071
So welcome, Dera.
6
00:00:39,705 --> 00:00:41,232
Thank so much, Anthony.
7
00:00:41,232 --> 00:00:43,370
Thank you for having me on this podcast.
8
00:00:43,705 --> 00:00:44,105
Excellent.
9
00:00:44,105 --> 00:00:47,398
So, Dera let's talk about AI-enabled discovery.
10
00:00:47,398 --> 00:00:52,162
We know this is the new thing, particularly the use of GenAI and AI agents and all of
that.
11
00:00:52,162 --> 00:00:54,754
And it's obviously a lot of change, right?
12
00:00:54,754 --> 00:00:56,375
A lot of change in the e-discovery field.
13
00:00:56,375 --> 00:00:59,588
It has been pretty static, I would say, for the past 10 years.
14
00:00:59,588 --> 00:01:00,749
And now there's a lot of change.
15
00:01:00,749 --> 00:01:06,644
So I think there's a lot of my clients are really interested in what's happening out
there.
16
00:01:06,644 --> 00:01:10,487
Like what's available, what's good, what's bad, what challenges we have.
17
00:01:10,487 --> 00:01:13,819
So today I really wanted to sort of focus high level.
18
00:01:13,867 --> 00:01:16,190
on what you're seeing, your opinions on things.
19
00:01:16,190 --> 00:01:20,114
So let's start with, know, GenAI and AI agents and the like.
20
00:01:20,114 --> 00:01:28,278
What do you see in terms of the marketplace in terms of where we are today versus where
we're gonna be in the next one or two years?
21
00:01:28,278 --> 00:01:30,549
Okay, well that question's not broad at all, right?
22
00:01:30,549 --> 00:01:37,615
And the great irony of course is that eDiscovery has been pioneering the use of AI for
well over a decade, right?
23
00:01:37,615 --> 00:01:43,257
And so what's happening now is a completely different category of technology coming into
the mix.
24
00:01:43,909 --> 00:01:44,327
and various flavors of it.
25
00:01:44,327 --> 00:01:53,236
So, you know, I'm not going to talk about the AI that we've been using for years, even
though I think some people are suddenly realizing, I have been using AI for years, right.
26
00:01:53,236 --> 00:01:54,837
In a Cal model or retire model.
27
00:01:54,837 --> 00:02:02,504
And we'll start to focus on, you know, the generative AI use cases and how that is
starting to be used in e-discovery.
28
00:02:02,504 --> 00:02:04,786
And it's, I think we're still.
29
00:02:05,220 --> 00:02:07,121
at the beginning of it, right?
30
00:02:07,121 --> 00:02:15,239
Even though in the past year it has become much more pervasive and the quality of the
technology baked into commonly available technology, right?
31
00:02:15,239 --> 00:02:21,946
Such as the relativity platforms or the reveal platforms or some of the cloud-based
platforms, the quality is better.
32
00:02:21,946 --> 00:02:26,070
The interesting thing is this is the worst it's ever going to be.
33
00:02:26,070 --> 00:02:28,412
It's only going to get better from here.
34
00:02:28,412 --> 00:02:31,594
So a lot of the use cases that we're seeing are
35
00:02:31,826 --> 00:02:43,309
early stage experimentation, but we are already starting to see trends and we are starting
to identify places where it can have a real meaningful impact, not only on workflows, but
36
00:02:43,309 --> 00:02:46,462
on outcomes, either cost outcomes or speed outcomes.
37
00:02:46,462 --> 00:02:50,436
So I think we should start to dig into how that stuff is starting to show up.
38
00:02:50,476 --> 00:02:53,787
Yeah, and I like to hear you because again, I think I'm looking at it again.
39
00:02:53,787 --> 00:02:56,008
We talk about e-discovery, right?
40
00:02:56,008 --> 00:03:03,128
EDRM, like it's, and we're getting questions about preservation, collection, review,
production, privilege.
41
00:03:03,128 --> 00:03:09,664
mean, there's so many different areas where it's in use, but obviously there's a lot of
risk because it's new, right?
42
00:03:09,664 --> 00:03:18,407
And so, and I think, look, I think just a level set for the most part, actually, at least
for the major ones, you have to be careful about going.
43
00:03:18,407 --> 00:03:26,780
you know, too far afield of the major players, but things like security, it's not using
public ChatGPT I think that's been pretty much resolved.
44
00:03:26,780 --> 00:03:36,666
I know you still have to do your due diligence, but in terms of other things, like where
of the EDRM model, where are you seeing or beyond that, but where are you seeing it?
45
00:03:36,666 --> 00:03:38,717
Where's the most opportunity today, right?
46
00:03:38,717 --> 00:03:40,694
If somebody used it today, what should they say?
47
00:03:40,694 --> 00:03:43,535
Okay, I'm going to start, let me start experimenting.
48
00:03:43,535 --> 00:03:46,882
What's the, what's the type of, you know, process or tools you think
49
00:03:46,882 --> 00:03:47,762
Yeah.
50
00:03:48,843 --> 00:03:54,456
You know, a place where I am seeing good results, right?
51
00:03:54,456 --> 00:03:57,498
And by good results, I'm not saying perfect results, right?
52
00:03:57,498 --> 00:04:04,391
But a place that I am most interested in using it is at the earlier stages of the whole
process.
53
00:04:04,551 --> 00:04:15,927
So starting to interpose generative AI at the earlier stage of discovery, including kind
of as an ECA type mechanism.
54
00:04:15,927 --> 00:04:21,982
So I know a lot of people just want to go straight to the end, review my documents for me,
right, at the very end of the process.
55
00:04:21,982 --> 00:04:30,189
But we've started to see some really good results when we get a lot of inbound
documentation that we don't have a good sense of what its contents might be.
56
00:04:30,189 --> 00:04:41,227
Using LLMs to summarize or provide us with inventories or lists of what's inside to give
us a sense of, is this something that should actually go into the discovery workflow?
57
00:04:41,227 --> 00:04:43,506
Should it be processed, right?
58
00:04:43,506 --> 00:04:45,387
And at that, you don't need to be perfect.
59
00:04:45,387 --> 00:04:54,892
You just need to be directionally accurate or even able to interrogate it or learn what
questions or what keywords to ask at a very early stage when you're starting to think
60
00:04:54,892 --> 00:04:57,243
about discovery planning.
61
00:04:57,244 --> 00:05:08,733
There are some great opportunities to reduce volumes, improve speed to understanding and
start to help it, you know, can start to help you understand the nuance of the case a bit
62
00:05:08,733 --> 00:05:09,591
earlier.
63
00:05:09,591 --> 00:05:11,318
And that's a tremendous advantage.
64
00:05:11,318 --> 00:05:11,518
Yeah.
65
00:05:11,518 --> 00:05:17,271
And I think, I, I, and I think this is going to be a sea change because I think, you know,
normally the way this works, right?
66
00:05:17,271 --> 00:05:22,494
Whether it's complaint and investigation and like, usually say, okay, we have to start the
review.
67
00:05:22,494 --> 00:05:27,517
think even though we have had, had ECA tools and like, I don't know how much they've
really been used.
68
00:05:27,517 --> 00:05:28,907
I completely agree.
69
00:05:28,907 --> 00:05:31,959
The speed to knowledge, particularly early on.
70
00:05:31,959 --> 00:05:41,250
And I've used it for a case where you get all your documents and then the first thing you
do is put in the AI tool and you just get, you know, timelines and.
71
00:05:41,250 --> 00:05:43,390
people and it surfaces, who do I need to talk to?
72
00:05:43,390 --> 00:05:45,272
And you can ask him and it depends on the tool.
73
00:05:45,272 --> 00:05:48,164
Not all retools slightly different, but I agree.
74
00:05:48,164 --> 00:05:54,448
think this is going to be best practice is the first thing you do is gather documents, put
in the AI tool.
75
00:05:54,448 --> 00:05:56,579
Cause it doesn't, it doesn't cost that much, right?
76
00:05:56,579 --> 00:05:59,651
It's you're not starting a huge review process, whatever.
77
00:05:59,651 --> 00:06:01,051
And I think it's also important.
78
00:06:01,051 --> 00:06:03,052
It's, it's not reviewers.
79
00:06:03,052 --> 00:06:04,113
This is partner level.
80
00:06:04,113 --> 00:06:05,494
Like once it's done.
81
00:06:05,854 --> 00:06:09,456
The partners that the client can really understand.
82
00:06:09,528 --> 00:06:11,828
the issues of the case or start understanding the issues of case.
83
00:06:11,828 --> 00:06:15,560
And I think that's massive opportunity for everybody.
84
00:06:15,560 --> 00:06:20,652
Massive opportunity, uh a particularly great use case, and this is a file that my
colleague did.
85
00:06:20,652 --> 00:06:25,614
So I can talk about what happened, but I can't talk about it with extreme granularity.
86
00:06:25,634 --> 00:06:31,737
But there was a very significant inbound production from opposing party, right?
87
00:06:31,737 --> 00:06:33,077
Very significant.
88
00:06:33,077 --> 00:06:37,939
I don't want to say it was a data dump, but it was voluminous because of the nature of the
case and the issues in the case.
89
00:06:37,939 --> 00:06:44,922
So we had a very senior associate, an extremely knowledgeable experienced practitioner.
90
00:06:45,022 --> 00:06:55,896
work with our data scientist to train an LLM in the issues of that case and in the law and
in the domain area of that client.
91
00:06:55,896 --> 00:06:57,227
So train an LLM.
92
00:06:57,227 --> 00:07:02,478
And then we had that LLM review a subset of the corpus, right?
93
00:07:02,478 --> 00:07:06,529
So 500 documents to identify what was responsive.
94
00:07:06,529 --> 00:07:11,571
And then we fed that into the traditional TAR model, right?
95
00:07:11,571 --> 00:07:12,441
And then we did an
96
00:07:12,441 --> 00:07:19,273
iterative process to be able to surface documents from the inbound collection.
97
00:07:19,633 --> 00:07:24,014
And the lawyer was so impressed with that process, right?
98
00:07:24,014 --> 00:07:35,238
The degree to which the documents were accurately summarized, correctly reflected the risk
and the issues, surfaced additional information, proposed additional inquiries, and then
99
00:07:35,238 --> 00:07:38,619
the way the TAR model performed as a result.
100
00:07:38,619 --> 00:07:39,751
And because
101
00:07:39,751 --> 00:07:42,011
that they went through four or five or six rounds, right?
102
00:07:42,011 --> 00:07:54,351
So, you know, reviewed about two or 3000 documents, but out of a very significant inbound
corpus to be able to surface what was really necessary to be reviewed to prepare for the
103
00:07:54,351 --> 00:07:55,471
next steps, right?
104
00:07:55,471 --> 00:08:00,767
The depositions and the next steps that they then said, well, you know what?
105
00:08:00,933 --> 00:08:05,317
We're going to do this against our own outbound productions, which was also voluminous.
106
00:08:05,317 --> 00:08:08,521
And while we have all done this, it's all various reviewers.
107
00:08:08,521 --> 00:08:11,504
We have various leads on top of that process.
108
00:08:11,504 --> 00:08:13,917
And there's a certain degree of normalization.
109
00:08:13,917 --> 00:08:24,127
But having that LLM process on top of it with a highly trained specialized LLM, I think is
going to be a game changer as well for true understanding and case development.
110
00:08:24,143 --> 00:08:25,554
And I totally agree.
111
00:08:25,554 --> 00:08:30,939
think, and again, whether it's, and I could certainly see it particularly for
investigations, right?
112
00:08:30,939 --> 00:08:39,468
Where you're rushing to get the production out and you're doing the review, but sometimes
you still don't really, like the people really don't know when you're producing it, say,
113
00:08:39,468 --> 00:08:45,402
okay, it's even, you could obviously use it for QC, but I think the idea of when you get
it, you just run the tool.
114
00:08:45,463 --> 00:08:48,315
So, because look, what I'm going to tell people, which I think is true,
115
00:08:48,315 --> 00:08:49,986
The SEC is going to do it.
116
00:08:49,986 --> 00:08:51,496
The plaintiffs, everyone's going to do it.
117
00:08:51,496 --> 00:08:56,208
Your production set, they're going to put it in LLM and nobody's going to review every
document anymore.
118
00:08:56,208 --> 00:08:57,809
They're just going to use the tool.
119
00:08:57,809 --> 00:09:00,090
You might as well see what they're going to find, right?
120
00:09:00,090 --> 00:09:05,672
Like it's like, you're going to see exactly what they're seeing, which is remarkable in
some ways.
121
00:09:05,672 --> 00:09:11,084
But that also is scary because they're going to have speed to knowledge and they're going
to really understand the case, right?
122
00:09:11,084 --> 00:09:12,695
Or the investigation pretty quickly.
123
00:09:12,695 --> 00:09:17,517
So, yeah, I do think it's going to accelerate both those things, both early case
assessment.
124
00:09:17,847 --> 00:09:21,119
which is helpful, you you know, much more.
125
00:09:21,119 --> 00:09:31,064
And then at the end, when, when things are actually produced, both sides are going to have
a much better understanding of what are the key issues, the good, the bad and the ugly,
126
00:09:31,064 --> 00:09:35,397
like you said, the depositions, it's going to be pretty, it's going to level the playing
field.
127
00:09:35,397 --> 00:09:39,289
Like we're all going to be looking at it's different tools, but it's going to be pretty
close, right?
128
00:09:39,289 --> 00:09:45,042
Where everyone's going to have a pretty same view of the strengths and weaknesses of
allegations and the like.
129
00:09:45,042 --> 00:09:47,040
So it's going to be fascinating.
130
00:09:47,040 --> 00:09:56,706
that's going to be great, I think, because then really it will require both sides to
really form a point of view of the validity of their case and the angles of attack and how
131
00:09:56,706 --> 00:10:03,310
it might be prosecuted earlier on, which will lead to potentially different litigation
strategies.
132
00:10:03,310 --> 00:10:08,022
And I think the interesting thing is I've sat back and reflected on this.
133
00:10:08,022 --> 00:10:11,916
When you're leading a very large team in litigation, you're always having that
134
00:10:11,916 --> 00:10:13,137
multiple perspectives.
135
00:10:13,137 --> 00:10:18,179
And that multiple perspectives on the corpus of data, the evidence is very important.
136
00:10:18,179 --> 00:10:26,443
But having a single view about how does this all tie together, that's what a lot of teams
have really struggled to do is get a uniform view.
137
00:10:26,443 --> 00:10:36,888
And having an LLM be a uniform view that you can then test your own biases, hypotheticals,
scenarios against, I think will just improve strategic decision making.
138
00:10:36,926 --> 00:10:38,446
Yeah, absolutely.
139
00:10:38,446 --> 00:10:44,886
And I do think the other thing that you sort of pointed out, which I do think, I think
this is one of the things I think we're all going to, the industry is going to have to
140
00:10:44,886 --> 00:10:48,326
figure out is like TAR has been very effective, right?
141
00:10:48,326 --> 00:10:49,866
It actually works.
142
00:10:50,566 --> 00:10:55,646
One of the challenges, which you sort of talked about is to train the model takes time.
143
00:10:55,646 --> 00:11:04,225
I do think, like you sort of said, that having GenAI help with that initial training set
and figuring it out will
144
00:11:04,225 --> 00:11:06,077
bring down those costs, right?
145
00:11:06,077 --> 00:11:14,243
And hopefully have like this, like the whole point of TAR was supposed to be a senior
associate, like you just said, would actually be the ones reviewing those documents.
146
00:11:14,243 --> 00:11:21,429
I'm not sure how much that actually happens, because oftentimes clients for cost reasons
were like, well, we'll have someone else do it.
147
00:11:21,429 --> 00:11:22,991
But I do think that'll also help.
148
00:11:22,991 --> 00:11:24,272
And it also speeds it up, right?
149
00:11:24,272 --> 00:11:25,853
They want to get started with the production.
150
00:11:25,853 --> 00:11:29,247
You can say, well, let's do GenAI One, we have the early case assessment.
151
00:11:29,247 --> 00:11:30,278
That'll help quite a bit.
152
00:11:30,278 --> 00:11:32,604
And then two, let's do some training.
153
00:11:32,604 --> 00:11:40,057
and then actually still use TAR for producing the documents and review in part because
it's accepted by the other side, right?
154
00:11:40,057 --> 00:11:41,408
It's court accepted.
155
00:11:41,408 --> 00:11:44,989
You don't have to worry about disclosure and all that kind of crap.
156
00:11:44,989 --> 00:11:48,171
You could just really say, look, we're using TAR.
157
00:11:48,171 --> 00:11:48,981
It's accepted.
158
00:11:48,981 --> 00:11:50,111
Everyone's accepted it.
159
00:11:50,111 --> 00:11:56,434
And you don't have to get into, used GenAI and there's going to be all kinds of issues
about disclosure and what prompt did you use?
160
00:11:56,434 --> 00:12:01,606
And like, we know that's going to be litigated, but you're avoiding all that and you're
still getting the benefit.
161
00:12:01,606 --> 00:12:03,781
some real benefit on the production side.
162
00:12:03,781 --> 00:12:04,422
Okay.
163
00:12:04,422 --> 00:12:08,599
And then anything else that you're seeing, are you seeing anything on the privilege side?
164
00:12:08,599 --> 00:12:13,295
I know it's early days and I know people want to, but what are you seeing on the privilege
side?
165
00:12:13,408 --> 00:12:22,456
I think there's still some room for improvement both in LLM output quality and evaluation
and just also in workflow.
166
00:12:22,456 --> 00:12:27,761
Privilege is one of those areas where people are rightly focused on accuracy.
167
00:12:27,761 --> 00:12:32,775
And a lot of people want to use the LLM to review the document and do an individualized
document assessment.
168
00:12:32,775 --> 00:12:35,787
And I've tended to avoid that.
169
00:12:35,787 --> 00:12:37,016
uh
170
00:12:37,016 --> 00:12:48,232
We may get there at some point, but I think where we've seen greater success and this
continues to be a work in progress is really triaging the documents LLM as a
171
00:12:48,232 --> 00:12:48,943
countermeasure.
172
00:12:48,943 --> 00:12:55,286
So, know, like when we're putting it through TAR, we're often using keywords or
supplementation to do an initial screening, and then you're putting eyes on a lot of that
173
00:12:55,286 --> 00:13:00,710
stuff, but using LLM as a secondary check to also evaluate the documents.
174
00:13:00,710 --> 00:13:01,690
if you're...
175
00:13:01,918 --> 00:13:09,172
traditional method in your LLM both agree that the document is likely to be privileged,
you put it into a human-based workflow and do the analysis.
176
00:13:09,172 --> 00:13:17,047
If they both agree it's not privileged, it's probably not, where there is a conflict
between the two that goes in through a screening mechanism.
177
00:13:17,047 --> 00:13:27,043
So we've started to experiment with that and the early indications is that that's gonna be
a very promising mechanism to streamline the workflow and triage documents more
178
00:13:27,043 --> 00:13:28,354
effectively.
179
00:13:28,398 --> 00:13:35,247
into different workflows, which itself can improve the overall speed of the privilege
process.
180
00:13:35,508 --> 00:13:37,151
We're also then using those.
181
00:13:37,151 --> 00:13:47,079
this, the other key point for that, which I think is key for our clients out there, from a
cost perspective, because it's usually when does the law firm have to look at something
182
00:13:47,079 --> 00:13:50,431
versus when can you use your contract attorneys?
183
00:13:50,431 --> 00:13:56,716
They're doing privilege, but if you have the contract attorney says privilege and the tool
says privilege, you do it.
184
00:13:56,716 --> 00:13:58,367
You could obviously say, you know what?
185
00:13:58,367 --> 00:14:00,048
I'm not going to have the law firm look at it.
186
00:14:00,048 --> 00:14:01,893
Again, that's a risk decision.
187
00:14:01,893 --> 00:14:09,098
You may not want to do it individually, maybe do a sample to make sure it's working, but I
see that as a huge cost save and it obviously helps with the risk as well.
188
00:14:09,098 --> 00:14:19,336
And that's what we're talking about, triaging and then where there's conflict, that's when
you have the associate at the law firm at whatever $100 an hour, $800 an hour, look at and
189
00:14:19,336 --> 00:14:20,948
say, okay, yeah, this is the right decision.
190
00:14:20,948 --> 00:14:28,633
I do think that's certainly going to be the next phase for privilege and then we'll see
how it develops over time, but I tend to agree with you there.
191
00:14:28,718 --> 00:14:29,138
that's right.
192
00:14:29,138 --> 00:14:38,195
So we're seeing some really good techniques, again, training that LLM to understand the
case, understand the context, giving it instructions about privilege in context of the
193
00:14:38,195 --> 00:14:42,378
determinations are being made with a specifically trained LLM.
194
00:14:42,378 --> 00:14:49,416
We're also starting to use it with certain types of document summaries to facilitate the
generation of the privilege logs.
195
00:14:49,416 --> 00:14:52,085
A lot of the stuff that is automated.
196
00:14:52,386 --> 00:14:53,757
Not everything can be automated.
197
00:14:53,757 --> 00:15:06,892
So supplementing some of that automation process with LLMs is also, again, improving
preliminary outputs and then improving the speed at which a final log entries can be
198
00:15:06,892 --> 00:15:07,925
drafted.
199
00:15:07,925 --> 00:15:09,519
So help with that process.
200
00:15:09,519 --> 00:15:15,419
you seen any use or good uses of AI agents yet in the discovery process?
201
00:15:15,670 --> 00:15:18,486
You know, I'm seeing a lot of experimentation.
202
00:15:18,486 --> 00:15:23,430
I haven't seen anything where I'm prepared to say that this is ready.
203
00:15:23,430 --> 00:15:25,459
A good news, ready.
204
00:15:25,459 --> 00:15:27,682
impression is it's everyone's looking at it.
205
00:15:27,682 --> 00:15:36,021
And obviously there seems to be, there should be good use cases for it, but I haven't seen
yet someone say, here's an AI agent that can do QC or whatever it is.
206
00:15:36,021 --> 00:15:37,207
I'm not gonna say, yeah.
207
00:15:37,207 --> 00:15:47,655
Yeah, so we are seeing, I wouldn't say like quite agentic behavior, but we're seeing LLMs
train to perform tasks that I think will be taken over by agents.
208
00:15:47,695 --> 00:15:51,538
Again, I think we're at the worst that this technology is.
209
00:15:51,538 --> 00:16:01,434
And I think even more than deploying the technology, what I've really been focused on is
defining the use cases where we are going to get a meaningful.
210
00:16:01,434 --> 00:16:06,887
an accurate result and that there's a way that we can check it and benchmark it against
our expectations.
211
00:16:06,887 --> 00:16:12,655
That's always what is if we start to have discomfort with the process, can we explain
what's going on?
212
00:16:12,655 --> 00:16:14,237
Can we document what's going on?
213
00:16:14,237 --> 00:16:26,005
And to not have it substitute foundational decision making where it really shouldn't be,
where it really is implementing instructions or providing output for further human review
214
00:16:26,005 --> 00:16:28,187
tend to be good use cases.
215
00:16:28,187 --> 00:16:32,408
where like a high degree of benefit can obtain, right?
216
00:16:32,408 --> 00:16:33,428
absolutely.
217
00:16:33,428 --> 00:16:41,708
Yeah, and I think we've learned a lot of lessons learned from when, you know, technology
sits a review came out or whatever, even though we got the courts, there was a lot of
218
00:16:41,708 --> 00:16:45,308
people had bad experiences because of all the challenges.
219
00:16:45,348 --> 00:16:47,508
And then people stopped adopting it.
220
00:16:47,508 --> 00:16:56,108
I do think I think to your point originally, I think we're going to see a large adoption
of TAR now just because people are focused on and realize, OK, you know what?
221
00:16:56,108 --> 00:16:59,372
We made a bad experience 10 years ago, but we now know.
222
00:16:59,372 --> 00:17:02,572
the good, bad, and the ugly about it and how best to use it.
223
00:17:03,452 --> 00:17:04,742
So, see what we can only hope.
224
00:17:04,742 --> 00:17:10,948
people adopt TAR but not realize that it's AI because they now associate AI with
generative AI, right?
225
00:17:10,948 --> 00:17:14,491
So now they see TAR as like automation, which is kind of funny.
226
00:17:14,491 --> 00:17:19,036
So I also wanted to talk about like one of my favorite use cases, right?
227
00:17:19,036 --> 00:17:20,947
And you know, this is actually very important.
228
00:17:20,947 --> 00:17:27,322
There's still a lot of cases that deal with historical records or handwritten records.
229
00:17:27,896 --> 00:17:31,407
And really handling handwritten records can be challenging, right?
230
00:17:31,407 --> 00:17:42,380
You know, are a lot of uh things that will OCR that documentation, but when you get into
cursive writing, the quality of the OCR can be really quite terrible, right?
231
00:17:42,380 --> 00:17:43,791
Really, really quite terrible.
232
00:17:43,791 --> 00:17:55,694
So we've been innovating with some processes to not only use a different mechanism of OCR
using AI enablement to improve the quality of the OCR.
233
00:17:55,780 --> 00:18:07,255
But to then run an LLM on top of that OCR to predict likely words and improve the overall
quality, it is remarkable.
234
00:18:07,255 --> 00:18:08,335
It is remarkable.
235
00:18:08,335 --> 00:18:17,840
I've had a lot of cases with handwritten notes, historical data, a lot of medical notes
where looking at the record, I can't make out the handwriting.
236
00:18:17,840 --> 00:18:23,882
And then when I see the AI, the LLM generated post OCR cleanup.
237
00:18:23,939 --> 00:18:25,921
I'm like, now I know what this is saying.
238
00:18:25,921 --> 00:18:28,084
That's obviously what it's saying.
239
00:18:28,084 --> 00:18:37,397
And then to have all of that available within indexes for searching is transformative in
cases where there's large volumes of documentation.
240
00:18:37,397 --> 00:18:38,506
uh
241
00:18:38,506 --> 00:18:49,441
I've heard also from from others, like particularly where you have a lot of PDFs or
designs like, you know, where that is something that type of data, the like, because
242
00:18:49,441 --> 00:19:00,826
obviously, AI can work on images that does a really good job using image technology, not
necessarily the LMS, but the image technology, which is fascinating to me.
243
00:19:00,826 --> 00:19:04,418
m So it'd be again, I think we're early days and everybody's going to figure this out.
244
00:19:04,418 --> 00:19:05,256
But I think
245
00:19:05,256 --> 00:19:11,796
As we all said, there are so many use cases right now that people can, I think the theme
is you can use it now, right?
246
00:19:11,796 --> 00:19:14,336
There's no reason not to start using it.
247
00:19:14,456 --> 00:19:19,136
And again, I think everything that we've talked about is relatively low risk, right?
248
00:19:19,136 --> 00:19:23,936
We're not getting into, which I keep hearing, the other side is never going to let me.
249
00:19:23,936 --> 00:19:27,896
Like most of the stuff we just talked about, you don't need the other side to agree to
anything, right?
250
00:19:27,896 --> 00:19:29,456
It's early case assessment.
251
00:19:29,456 --> 00:19:34,118
Again, you may have to get approval to run the AI tool on their production, but that's
a...
252
00:19:34,118 --> 00:19:37,270
security, confidentiality issue, you just deal with that.
253
00:19:37,270 --> 00:19:43,354
But it's not getting into what are my prompts, how am I using it, which I think is really
beneficial.
254
00:19:43,354 --> 00:19:48,779
I do think, and again, I think we're starting to see like this year is gonna be the use of
GenAI.
255
00:19:48,779 --> 00:19:51,181
think almost every case is gonna start using it.
256
00:19:51,181 --> 00:19:51,474
No.
257
00:19:51,474 --> 00:19:55,465
some cases where there's a huge amount of video audio finding stuff, right?
258
00:19:55,465 --> 00:20:04,167
Like, I mean, I do this, like when I make my own videos to entertain my family, you know,
I'll you know, say, Hey, here's a whole bunch of videos.
259
00:20:04,167 --> 00:20:08,359
Find me where somebody says X or find me where somebody says Y right?
260
00:20:08,359 --> 00:20:15,007
Like with audio and video technology, which increasingly is prevalent in discovery
corpuses.
261
00:20:15,007 --> 00:20:19,199
I mean, it's gonna make a difference to our ability to get through that stuff quickly.
262
00:20:19,287 --> 00:20:20,018
Okay.
263
00:20:20,018 --> 00:20:21,969
Well, Dera thank you so much.
264
00:20:21,969 --> 00:20:22,729
I think this was great.
265
00:20:22,729 --> 00:20:26,242
I think it'd be a really good summary of where we are and where we're heading.
266
00:20:26,242 --> 00:20:30,745
And like I said, I think the theme is if you're not using GenAI, start using it.
267
00:20:30,745 --> 00:20:33,549
So thank you very much and we'll talk soon.
268
00:20:33,549 --> 00:20:36,080
And everyone else, please, I hope to hear from you.
269
00:20:36,080 --> 00:20:38,062
And obviously this podcast is a part of a series.
270
00:20:38,062 --> 00:20:40,014
So hope you join us next time.
271
00:20:40,014 --> 00:20:40,364
Thanks.
272
00:20:40,364 --> 00:20:41,084
Bye.
00:00:10,680 --> 00:00:14,893
Hello, this is Anthony Diana from Reed Smith and welcome to Tech Law Talks.
2
00:00:14,893 --> 00:00:20,157
Today we are continuing a podcast series on AI enabled e-Discovery.
3
00:00:20,157 --> 00:00:31,024
This podcast series will focus on practical and legal issues when considering using AI
enabled e-Disovery with a focus on actual use cases, not just the theoretical.
4
00:00:31,024 --> 00:00:37,750
Joining me today is Dera Nevin from FTI, a good friend of mine and a good friend of Reed
Smith for many, many years.
5
00:00:37,750 --> 00:00:39,071
So welcome, Dera.
6
00:00:39,705 --> 00:00:41,232
Thank so much, Anthony.
7
00:00:41,232 --> 00:00:43,370
Thank you for having me on this podcast.
8
00:00:43,705 --> 00:00:44,105
Excellent.
9
00:00:44,105 --> 00:00:47,398
So, Dera let's talk about AI-enabled discovery.
10
00:00:47,398 --> 00:00:52,162
We know this is the new thing, particularly the use of GenAI and AI agents and all of
that.
11
00:00:52,162 --> 00:00:54,754
And it's obviously a lot of change, right?
12
00:00:54,754 --> 00:00:56,375
A lot of change in the e-discovery field.
13
00:00:56,375 --> 00:00:59,588
It has been pretty static, I would say, for the past 10 years.
14
00:00:59,588 --> 00:01:00,749
And now there's a lot of change.
15
00:01:00,749 --> 00:01:06,644
So I think there's a lot of my clients are really interested in what's happening out
there.
16
00:01:06,644 --> 00:01:10,487
Like what's available, what's good, what's bad, what challenges we have.
17
00:01:10,487 --> 00:01:13,819
So today I really wanted to sort of focus high level.
18
00:01:13,867 --> 00:01:16,190
on what you're seeing, your opinions on things.
19
00:01:16,190 --> 00:01:20,114
So let's start with, know, GenAI and AI agents and the like.
20
00:01:20,114 --> 00:01:28,278
What do you see in terms of the marketplace in terms of where we are today versus where
we're gonna be in the next one or two years?
21
00:01:28,278 --> 00:01:30,549
Okay, well that question's not broad at all, right?
22
00:01:30,549 --> 00:01:37,615
And the great irony of course is that eDiscovery has been pioneering the use of AI for
well over a decade, right?
23
00:01:37,615 --> 00:01:43,257
And so what's happening now is a completely different category of technology coming into
the mix.
24
00:01:43,909 --> 00:01:44,327
and various flavors of it.
25
00:01:44,327 --> 00:01:53,236
So, you know, I'm not going to talk about the AI that we've been using for years, even
though I think some people are suddenly realizing, I have been using AI for years, right.
26
00:01:53,236 --> 00:01:54,837
In a Cal model or retire model.
27
00:01:54,837 --> 00:02:02,504
And we'll start to focus on, you know, the generative AI use cases and how that is
starting to be used in e-discovery.
28
00:02:02,504 --> 00:02:04,786
And it's, I think we're still.
29
00:02:05,220 --> 00:02:07,121
at the beginning of it, right?
30
00:02:07,121 --> 00:02:15,239
Even though in the past year it has become much more pervasive and the quality of the
technology baked into commonly available technology, right?
31
00:02:15,239 --> 00:02:21,946
Such as the relativity platforms or the reveal platforms or some of the cloud-based
platforms, the quality is better.
32
00:02:21,946 --> 00:02:26,070
The interesting thing is this is the worst it's ever going to be.
33
00:02:26,070 --> 00:02:28,412
It's only going to get better from here.
34
00:02:28,412 --> 00:02:31,594
So a lot of the use cases that we're seeing are
35
00:02:31,826 --> 00:02:43,309
early stage experimentation, but we are already starting to see trends and we are starting
to identify places where it can have a real meaningful impact, not only on workflows, but
36
00:02:43,309 --> 00:02:46,462
on outcomes, either cost outcomes or speed outcomes.
37
00:02:46,462 --> 00:02:50,436
So I think we should start to dig into how that stuff is starting to show up.
38
00:02:50,476 --> 00:02:53,787
Yeah, and I like to hear you because again, I think I'm looking at it again.
39
00:02:53,787 --> 00:02:56,008
We talk about e-discovery, right?
40
00:02:56,008 --> 00:03:03,128
EDRM, like it's, and we're getting questions about preservation, collection, review,
production, privilege.
41
00:03:03,128 --> 00:03:09,664
mean, there's so many different areas where it's in use, but obviously there's a lot of
risk because it's new, right?
42
00:03:09,664 --> 00:03:18,407
And so, and I think, look, I think just a level set for the most part, actually, at least
for the major ones, you have to be careful about going.
43
00:03:18,407 --> 00:03:26,780
you know, too far afield of the major players, but things like security, it's not using
public ChatGPT I think that's been pretty much resolved.
44
00:03:26,780 --> 00:03:36,666
I know you still have to do your due diligence, but in terms of other things, like where
of the EDRM model, where are you seeing or beyond that, but where are you seeing it?
45
00:03:36,666 --> 00:03:38,717
Where's the most opportunity today, right?
46
00:03:38,717 --> 00:03:40,694
If somebody used it today, what should they say?
47
00:03:40,694 --> 00:03:43,535
Okay, I'm going to start, let me start experimenting.
48
00:03:43,535 --> 00:03:46,882
What's the, what's the type of, you know, process or tools you think
49
00:03:46,882 --> 00:03:47,762
Yeah.
50
00:03:48,843 --> 00:03:54,456
You know, a place where I am seeing good results, right?
51
00:03:54,456 --> 00:03:57,498
And by good results, I'm not saying perfect results, right?
52
00:03:57,498 --> 00:04:04,391
But a place that I am most interested in using it is at the earlier stages of the whole
process.
53
00:04:04,551 --> 00:04:15,927
So starting to interpose generative AI at the earlier stage of discovery, including kind
of as an ECA type mechanism.
54
00:04:15,927 --> 00:04:21,982
So I know a lot of people just want to go straight to the end, review my documents for me,
right, at the very end of the process.
55
00:04:21,982 --> 00:04:30,189
But we've started to see some really good results when we get a lot of inbound
documentation that we don't have a good sense of what its contents might be.
56
00:04:30,189 --> 00:04:41,227
Using LLMs to summarize or provide us with inventories or lists of what's inside to give
us a sense of, is this something that should actually go into the discovery workflow?
57
00:04:41,227 --> 00:04:43,506
Should it be processed, right?
58
00:04:43,506 --> 00:04:45,387
And at that, you don't need to be perfect.
59
00:04:45,387 --> 00:04:54,892
You just need to be directionally accurate or even able to interrogate it or learn what
questions or what keywords to ask at a very early stage when you're starting to think
60
00:04:54,892 --> 00:04:57,243
about discovery planning.
61
00:04:57,244 --> 00:05:08,733
There are some great opportunities to reduce volumes, improve speed to understanding and
start to help it, you know, can start to help you understand the nuance of the case a bit
62
00:05:08,733 --> 00:05:09,591
earlier.
63
00:05:09,591 --> 00:05:11,318
And that's a tremendous advantage.
64
00:05:11,318 --> 00:05:11,518
Yeah.
65
00:05:11,518 --> 00:05:17,271
And I think, I, I, and I think this is going to be a sea change because I think, you know,
normally the way this works, right?
66
00:05:17,271 --> 00:05:22,494
Whether it's complaint and investigation and like, usually say, okay, we have to start the
review.
67
00:05:22,494 --> 00:05:27,517
think even though we have had, had ECA tools and like, I don't know how much they've
really been used.
68
00:05:27,517 --> 00:05:28,907
I completely agree.
69
00:05:28,907 --> 00:05:31,959
The speed to knowledge, particularly early on.
70
00:05:31,959 --> 00:05:41,250
And I've used it for a case where you get all your documents and then the first thing you
do is put in the AI tool and you just get, you know, timelines and.
71
00:05:41,250 --> 00:05:43,390
people and it surfaces, who do I need to talk to?
72
00:05:43,390 --> 00:05:45,272
And you can ask him and it depends on the tool.
73
00:05:45,272 --> 00:05:48,164
Not all retools slightly different, but I agree.
74
00:05:48,164 --> 00:05:54,448
think this is going to be best practice is the first thing you do is gather documents, put
in the AI tool.
75
00:05:54,448 --> 00:05:56,579
Cause it doesn't, it doesn't cost that much, right?
76
00:05:56,579 --> 00:05:59,651
It's you're not starting a huge review process, whatever.
77
00:05:59,651 --> 00:06:01,051
And I think it's also important.
78
00:06:01,051 --> 00:06:03,052
It's, it's not reviewers.
79
00:06:03,052 --> 00:06:04,113
This is partner level.
80
00:06:04,113 --> 00:06:05,494
Like once it's done.
81
00:06:05,854 --> 00:06:09,456
The partners that the client can really understand.
82
00:06:09,528 --> 00:06:11,828
the issues of the case or start understanding the issues of case.
83
00:06:11,828 --> 00:06:15,560
And I think that's massive opportunity for everybody.
84
00:06:15,560 --> 00:06:20,652
Massive opportunity, uh a particularly great use case, and this is a file that my
colleague did.
85
00:06:20,652 --> 00:06:25,614
So I can talk about what happened, but I can't talk about it with extreme granularity.
86
00:06:25,634 --> 00:06:31,737
But there was a very significant inbound production from opposing party, right?
87
00:06:31,737 --> 00:06:33,077
Very significant.
88
00:06:33,077 --> 00:06:37,939
I don't want to say it was a data dump, but it was voluminous because of the nature of the
case and the issues in the case.
89
00:06:37,939 --> 00:06:44,922
So we had a very senior associate, an extremely knowledgeable experienced practitioner.
90
00:06:45,022 --> 00:06:55,896
work with our data scientist to train an LLM in the issues of that case and in the law and
in the domain area of that client.
91
00:06:55,896 --> 00:06:57,227
So train an LLM.
92
00:06:57,227 --> 00:07:02,478
And then we had that LLM review a subset of the corpus, right?
93
00:07:02,478 --> 00:07:06,529
So 500 documents to identify what was responsive.
94
00:07:06,529 --> 00:07:11,571
And then we fed that into the traditional TAR model, right?
95
00:07:11,571 --> 00:07:12,441
And then we did an
96
00:07:12,441 --> 00:07:19,273
iterative process to be able to surface documents from the inbound collection.
97
00:07:19,633 --> 00:07:24,014
And the lawyer was so impressed with that process, right?
98
00:07:24,014 --> 00:07:35,238
The degree to which the documents were accurately summarized, correctly reflected the risk
and the issues, surfaced additional information, proposed additional inquiries, and then
99
00:07:35,238 --> 00:07:38,619
the way the TAR model performed as a result.
100
00:07:38,619 --> 00:07:39,751
And because
101
00:07:39,751 --> 00:07:42,011
that they went through four or five or six rounds, right?
102
00:07:42,011 --> 00:07:54,351
So, you know, reviewed about two or 3000 documents, but out of a very significant inbound
corpus to be able to surface what was really necessary to be reviewed to prepare for the
103
00:07:54,351 --> 00:07:55,471
next steps, right?
104
00:07:55,471 --> 00:08:00,767
The depositions and the next steps that they then said, well, you know what?
105
00:08:00,933 --> 00:08:05,317
We're going to do this against our own outbound productions, which was also voluminous.
106
00:08:05,317 --> 00:08:08,521
And while we have all done this, it's all various reviewers.
107
00:08:08,521 --> 00:08:11,504
We have various leads on top of that process.
108
00:08:11,504 --> 00:08:13,917
And there's a certain degree of normalization.
109
00:08:13,917 --> 00:08:24,127
But having that LLM process on top of it with a highly trained specialized LLM, I think is
going to be a game changer as well for true understanding and case development.
110
00:08:24,143 --> 00:08:25,554
And I totally agree.
111
00:08:25,554 --> 00:08:30,939
think, and again, whether it's, and I could certainly see it particularly for
investigations, right?
112
00:08:30,939 --> 00:08:39,468
Where you're rushing to get the production out and you're doing the review, but sometimes
you still don't really, like the people really don't know when you're producing it, say,
113
00:08:39,468 --> 00:08:45,402
okay, it's even, you could obviously use it for QC, but I think the idea of when you get
it, you just run the tool.
114
00:08:45,463 --> 00:08:48,315
So, because look, what I'm going to tell people, which I think is true,
115
00:08:48,315 --> 00:08:49,986
The SEC is going to do it.
116
00:08:49,986 --> 00:08:51,496
The plaintiffs, everyone's going to do it.
117
00:08:51,496 --> 00:08:56,208
Your production set, they're going to put it in LLM and nobody's going to review every
document anymore.
118
00:08:56,208 --> 00:08:57,809
They're just going to use the tool.
119
00:08:57,809 --> 00:09:00,090
You might as well see what they're going to find, right?
120
00:09:00,090 --> 00:09:05,672
Like it's like, you're going to see exactly what they're seeing, which is remarkable in
some ways.
121
00:09:05,672 --> 00:09:11,084
But that also is scary because they're going to have speed to knowledge and they're going
to really understand the case, right?
122
00:09:11,084 --> 00:09:12,695
Or the investigation pretty quickly.
123
00:09:12,695 --> 00:09:17,517
So, yeah, I do think it's going to accelerate both those things, both early case
assessment.
124
00:09:17,847 --> 00:09:21,119
which is helpful, you you know, much more.
125
00:09:21,119 --> 00:09:31,064
And then at the end, when, when things are actually produced, both sides are going to have
a much better understanding of what are the key issues, the good, the bad and the ugly,
126
00:09:31,064 --> 00:09:35,397
like you said, the depositions, it's going to be pretty, it's going to level the playing
field.
127
00:09:35,397 --> 00:09:39,289
Like we're all going to be looking at it's different tools, but it's going to be pretty
close, right?
128
00:09:39,289 --> 00:09:45,042
Where everyone's going to have a pretty same view of the strengths and weaknesses of
allegations and the like.
129
00:09:45,042 --> 00:09:47,040
So it's going to be fascinating.
130
00:09:47,040 --> 00:09:56,706
that's going to be great, I think, because then really it will require both sides to
really form a point of view of the validity of their case and the angles of attack and how
131
00:09:56,706 --> 00:10:03,310
it might be prosecuted earlier on, which will lead to potentially different litigation
strategies.
132
00:10:03,310 --> 00:10:08,022
And I think the interesting thing is I've sat back and reflected on this.
133
00:10:08,022 --> 00:10:11,916
When you're leading a very large team in litigation, you're always having that
134
00:10:11,916 --> 00:10:13,137
multiple perspectives.
135
00:10:13,137 --> 00:10:18,179
And that multiple perspectives on the corpus of data, the evidence is very important.
136
00:10:18,179 --> 00:10:26,443
But having a single view about how does this all tie together, that's what a lot of teams
have really struggled to do is get a uniform view.
137
00:10:26,443 --> 00:10:36,888
And having an LLM be a uniform view that you can then test your own biases, hypotheticals,
scenarios against, I think will just improve strategic decision making.
138
00:10:36,926 --> 00:10:38,446
Yeah, absolutely.
139
00:10:38,446 --> 00:10:44,886
And I do think the other thing that you sort of pointed out, which I do think, I think
this is one of the things I think we're all going to, the industry is going to have to
140
00:10:44,886 --> 00:10:48,326
figure out is like TAR has been very effective, right?
141
00:10:48,326 --> 00:10:49,866
It actually works.
142
00:10:50,566 --> 00:10:55,646
One of the challenges, which you sort of talked about is to train the model takes time.
143
00:10:55,646 --> 00:11:04,225
I do think, like you sort of said, that having GenAI help with that initial training set
and figuring it out will
144
00:11:04,225 --> 00:11:06,077
bring down those costs, right?
145
00:11:06,077 --> 00:11:14,243
And hopefully have like this, like the whole point of TAR was supposed to be a senior
associate, like you just said, would actually be the ones reviewing those documents.
146
00:11:14,243 --> 00:11:21,429
I'm not sure how much that actually happens, because oftentimes clients for cost reasons
were like, well, we'll have someone else do it.
147
00:11:21,429 --> 00:11:22,991
But I do think that'll also help.
148
00:11:22,991 --> 00:11:24,272
And it also speeds it up, right?
149
00:11:24,272 --> 00:11:25,853
They want to get started with the production.
150
00:11:25,853 --> 00:11:29,247
You can say, well, let's do GenAI One, we have the early case assessment.
151
00:11:29,247 --> 00:11:30,278
That'll help quite a bit.
152
00:11:30,278 --> 00:11:32,604
And then two, let's do some training.
153
00:11:32,604 --> 00:11:40,057
and then actually still use TAR for producing the documents and review in part because
it's accepted by the other side, right?
154
00:11:40,057 --> 00:11:41,408
It's court accepted.
155
00:11:41,408 --> 00:11:44,989
You don't have to worry about disclosure and all that kind of crap.
156
00:11:44,989 --> 00:11:48,171
You could just really say, look, we're using TAR.
157
00:11:48,171 --> 00:11:48,981
It's accepted.
158
00:11:48,981 --> 00:11:50,111
Everyone's accepted it.
159
00:11:50,111 --> 00:11:56,434
And you don't have to get into, used GenAI and there's going to be all kinds of issues
about disclosure and what prompt did you use?
160
00:11:56,434 --> 00:12:01,606
And like, we know that's going to be litigated, but you're avoiding all that and you're
still getting the benefit.
161
00:12:01,606 --> 00:12:03,781
some real benefit on the production side.
162
00:12:03,781 --> 00:12:04,422
Okay.
163
00:12:04,422 --> 00:12:08,599
And then anything else that you're seeing, are you seeing anything on the privilege side?
164
00:12:08,599 --> 00:12:13,295
I know it's early days and I know people want to, but what are you seeing on the privilege
side?
165
00:12:13,408 --> 00:12:22,456
I think there's still some room for improvement both in LLM output quality and evaluation
and just also in workflow.
166
00:12:22,456 --> 00:12:27,761
Privilege is one of those areas where people are rightly focused on accuracy.
167
00:12:27,761 --> 00:12:32,775
And a lot of people want to use the LLM to review the document and do an individualized
document assessment.
168
00:12:32,775 --> 00:12:35,787
And I've tended to avoid that.
169
00:12:35,787 --> 00:12:37,016
uh
170
00:12:37,016 --> 00:12:48,232
We may get there at some point, but I think where we've seen greater success and this
continues to be a work in progress is really triaging the documents LLM as a
171
00:12:48,232 --> 00:12:48,943
countermeasure.
172
00:12:48,943 --> 00:12:55,286
So, know, like when we're putting it through TAR, we're often using keywords or
supplementation to do an initial screening, and then you're putting eyes on a lot of that
173
00:12:55,286 --> 00:13:00,710
stuff, but using LLM as a secondary check to also evaluate the documents.
174
00:13:00,710 --> 00:13:01,690
if you're...
175
00:13:01,918 --> 00:13:09,172
traditional method in your LLM both agree that the document is likely to be privileged,
you put it into a human-based workflow and do the analysis.
176
00:13:09,172 --> 00:13:17,047
If they both agree it's not privileged, it's probably not, where there is a conflict
between the two that goes in through a screening mechanism.
177
00:13:17,047 --> 00:13:27,043
So we've started to experiment with that and the early indications is that that's gonna be
a very promising mechanism to streamline the workflow and triage documents more
178
00:13:27,043 --> 00:13:28,354
effectively.
179
00:13:28,398 --> 00:13:35,247
into different workflows, which itself can improve the overall speed of the privilege
process.
180
00:13:35,508 --> 00:13:37,151
We're also then using those.
181
00:13:37,151 --> 00:13:47,079
this, the other key point for that, which I think is key for our clients out there, from a
cost perspective, because it's usually when does the law firm have to look at something
182
00:13:47,079 --> 00:13:50,431
versus when can you use your contract attorneys?
183
00:13:50,431 --> 00:13:56,716
They're doing privilege, but if you have the contract attorney says privilege and the tool
says privilege, you do it.
184
00:13:56,716 --> 00:13:58,367
You could obviously say, you know what?
185
00:13:58,367 --> 00:14:00,048
I'm not going to have the law firm look at it.
186
00:14:00,048 --> 00:14:01,893
Again, that's a risk decision.
187
00:14:01,893 --> 00:14:09,098
You may not want to do it individually, maybe do a sample to make sure it's working, but I
see that as a huge cost save and it obviously helps with the risk as well.
188
00:14:09,098 --> 00:14:19,336
And that's what we're talking about, triaging and then where there's conflict, that's when
you have the associate at the law firm at whatever $100 an hour, $800 an hour, look at and
189
00:14:19,336 --> 00:14:20,948
say, okay, yeah, this is the right decision.
190
00:14:20,948 --> 00:14:28,633
I do think that's certainly going to be the next phase for privilege and then we'll see
how it develops over time, but I tend to agree with you there.
191
00:14:28,718 --> 00:14:29,138
that's right.
192
00:14:29,138 --> 00:14:38,195
So we're seeing some really good techniques, again, training that LLM to understand the
case, understand the context, giving it instructions about privilege in context of the
193
00:14:38,195 --> 00:14:42,378
determinations are being made with a specifically trained LLM.
194
00:14:42,378 --> 00:14:49,416
We're also starting to use it with certain types of document summaries to facilitate the
generation of the privilege logs.
195
00:14:49,416 --> 00:14:52,085
A lot of the stuff that is automated.
196
00:14:52,386 --> 00:14:53,757
Not everything can be automated.
197
00:14:53,757 --> 00:15:06,892
So supplementing some of that automation process with LLMs is also, again, improving
preliminary outputs and then improving the speed at which a final log entries can be
198
00:15:06,892 --> 00:15:07,925
drafted.
199
00:15:07,925 --> 00:15:09,519
So help with that process.
200
00:15:09,519 --> 00:15:15,419
you seen any use or good uses of AI agents yet in the discovery process?
201
00:15:15,670 --> 00:15:18,486
You know, I'm seeing a lot of experimentation.
202
00:15:18,486 --> 00:15:23,430
I haven't seen anything where I'm prepared to say that this is ready.
203
00:15:23,430 --> 00:15:25,459
A good news, ready.
204
00:15:25,459 --> 00:15:27,682
impression is it's everyone's looking at it.
205
00:15:27,682 --> 00:15:36,021
And obviously there seems to be, there should be good use cases for it, but I haven't seen
yet someone say, here's an AI agent that can do QC or whatever it is.
206
00:15:36,021 --> 00:15:37,207
I'm not gonna say, yeah.
207
00:15:37,207 --> 00:15:47,655
Yeah, so we are seeing, I wouldn't say like quite agentic behavior, but we're seeing LLMs
train to perform tasks that I think will be taken over by agents.
208
00:15:47,695 --> 00:15:51,538
Again, I think we're at the worst that this technology is.
209
00:15:51,538 --> 00:16:01,434
And I think even more than deploying the technology, what I've really been focused on is
defining the use cases where we are going to get a meaningful.
210
00:16:01,434 --> 00:16:06,887
an accurate result and that there's a way that we can check it and benchmark it against
our expectations.
211
00:16:06,887 --> 00:16:12,655
That's always what is if we start to have discomfort with the process, can we explain
what's going on?
212
00:16:12,655 --> 00:16:14,237
Can we document what's going on?
213
00:16:14,237 --> 00:16:26,005
And to not have it substitute foundational decision making where it really shouldn't be,
where it really is implementing instructions or providing output for further human review
214
00:16:26,005 --> 00:16:28,187
tend to be good use cases.
215
00:16:28,187 --> 00:16:32,408
where like a high degree of benefit can obtain, right?
216
00:16:32,408 --> 00:16:33,428
absolutely.
217
00:16:33,428 --> 00:16:41,708
Yeah, and I think we've learned a lot of lessons learned from when, you know, technology
sits a review came out or whatever, even though we got the courts, there was a lot of
218
00:16:41,708 --> 00:16:45,308
people had bad experiences because of all the challenges.
219
00:16:45,348 --> 00:16:47,508
And then people stopped adopting it.
220
00:16:47,508 --> 00:16:56,108
I do think I think to your point originally, I think we're going to see a large adoption
of TAR now just because people are focused on and realize, OK, you know what?
221
00:16:56,108 --> 00:16:59,372
We made a bad experience 10 years ago, but we now know.
222
00:16:59,372 --> 00:17:02,572
the good, bad, and the ugly about it and how best to use it.
223
00:17:03,452 --> 00:17:04,742
So, see what we can only hope.
224
00:17:04,742 --> 00:17:10,948
people adopt TAR but not realize that it's AI because they now associate AI with
generative AI, right?
225
00:17:10,948 --> 00:17:14,491
So now they see TAR as like automation, which is kind of funny.
226
00:17:14,491 --> 00:17:19,036
So I also wanted to talk about like one of my favorite use cases, right?
227
00:17:19,036 --> 00:17:20,947
And you know, this is actually very important.
228
00:17:20,947 --> 00:17:27,322
There's still a lot of cases that deal with historical records or handwritten records.
229
00:17:27,896 --> 00:17:31,407
And really handling handwritten records can be challenging, right?
230
00:17:31,407 --> 00:17:42,380
You know, are a lot of uh things that will OCR that documentation, but when you get into
cursive writing, the quality of the OCR can be really quite terrible, right?
231
00:17:42,380 --> 00:17:43,791
Really, really quite terrible.
232
00:17:43,791 --> 00:17:55,694
So we've been innovating with some processes to not only use a different mechanism of OCR
using AI enablement to improve the quality of the OCR.
233
00:17:55,780 --> 00:18:07,255
But to then run an LLM on top of that OCR to predict likely words and improve the overall
quality, it is remarkable.
234
00:18:07,255 --> 00:18:08,335
It is remarkable.
235
00:18:08,335 --> 00:18:17,840
I've had a lot of cases with handwritten notes, historical data, a lot of medical notes
where looking at the record, I can't make out the handwriting.
236
00:18:17,840 --> 00:18:23,882
And then when I see the AI, the LLM generated post OCR cleanup.
237
00:18:23,939 --> 00:18:25,921
I'm like, now I know what this is saying.
238
00:18:25,921 --> 00:18:28,084
That's obviously what it's saying.
239
00:18:28,084 --> 00:18:37,397
And then to have all of that available within indexes for searching is transformative in
cases where there's large volumes of documentation.
240
00:18:37,397 --> 00:18:38,506
uh
241
00:18:38,506 --> 00:18:49,441
I've heard also from from others, like particularly where you have a lot of PDFs or
designs like, you know, where that is something that type of data, the like, because
242
00:18:49,441 --> 00:19:00,826
obviously, AI can work on images that does a really good job using image technology, not
necessarily the LMS, but the image technology, which is fascinating to me.
243
00:19:00,826 --> 00:19:04,418
m So it'd be again, I think we're early days and everybody's going to figure this out.
244
00:19:04,418 --> 00:19:05,256
But I think
245
00:19:05,256 --> 00:19:11,796
As we all said, there are so many use cases right now that people can, I think the theme
is you can use it now, right?
246
00:19:11,796 --> 00:19:14,336
There's no reason not to start using it.
247
00:19:14,456 --> 00:19:19,136
And again, I think everything that we've talked about is relatively low risk, right?
248
00:19:19,136 --> 00:19:23,936
We're not getting into, which I keep hearing, the other side is never going to let me.
249
00:19:23,936 --> 00:19:27,896
Like most of the stuff we just talked about, you don't need the other side to agree to
anything, right?
250
00:19:27,896 --> 00:19:29,456
It's early case assessment.
251
00:19:29,456 --> 00:19:34,118
Again, you may have to get approval to run the AI tool on their production, but that's
a...
252
00:19:34,118 --> 00:19:37,270
security, confidentiality issue, you just deal with that.
253
00:19:37,270 --> 00:19:43,354
But it's not getting into what are my prompts, how am I using it, which I think is really
beneficial.
254
00:19:43,354 --> 00:19:48,779
I do think, and again, I think we're starting to see like this year is gonna be the use of
GenAI.
255
00:19:48,779 --> 00:19:51,181
think almost every case is gonna start using it.
256
00:19:51,181 --> 00:19:51,474
No.
257
00:19:51,474 --> 00:19:55,465
some cases where there's a huge amount of video audio finding stuff, right?
258
00:19:55,465 --> 00:20:04,167
Like, I mean, I do this, like when I make my own videos to entertain my family, you know,
I'll you know, say, Hey, here's a whole bunch of videos.
259
00:20:04,167 --> 00:20:08,359
Find me where somebody says X or find me where somebody says Y right?
260
00:20:08,359 --> 00:20:15,007
Like with audio and video technology, which increasingly is prevalent in discovery
corpuses.
261
00:20:15,007 --> 00:20:19,199
I mean, it's gonna make a difference to our ability to get through that stuff quickly.
262
00:20:19,287 --> 00:20:20,018
Okay.
263
00:20:20,018 --> 00:20:21,969
Well, Dera thank you so much.
264
00:20:21,969 --> 00:20:22,729
I think this was great.
265
00:20:22,729 --> 00:20:26,242
I think it'd be a really good summary of where we are and where we're heading.
266
00:20:26,242 --> 00:20:30,745
And like I said, I think the theme is if you're not using GenAI, start using it.
267
00:20:30,745 --> 00:20:33,549
So thank you very much and we'll talk soon.
268
00:20:33,549 --> 00:20:36,080
And everyone else, please, I hope to hear from you.
269
00:20:36,080 --> 00:20:38,062
And obviously this podcast is a part of a series.
270
00:20:38,062 --> 00:20:40,014
So hope you join us next time.
271
00:20:40,014 --> 00:20:40,364
Thanks.
272
00:20:40,364 --> 00:20:41,084
Bye.