このエピソードについて
Diagnostic tests used in research and clinical practice are essential to evidence-based medicine because they shape how clinicians diagnose disease, interpret tests, and make patient care decisions. In this episode of This Is Why with Dr. Busti, Dr. Busti explains how the precision and accuracy of diagnostic studies are designed, validated, and critically appraised using clinically relevant examples and practical reasoning.
This Is Why understanding sensitivity, specificity, predictive values, and likelihood ratios matters far beyond memorization. Dr. Busti walks through how clinicians determine whether a diagnostic test is useful, how bias affects study validity, and why gold-standard reference testing is critical in medical research.
Topics Covered:
- Diagnostic tests in clinical medicine
- Precision and accuracy of diagnostic tests explained
- Bias and error in diagnostic studies
- Considerations about the Gold Standard or Reference Tests
The goal = make medical education easy and clinically relevant.
👉 Get more with a free membership at https://www.thisiswhy.health/
- Access free downloads from our videos
- Access deep dive content from Dr. Busti
- Organize content via playlists & collections
- Join live Q&A
- Receive member newsletters
- Coupons & discounts for exam prep resources
👍 If this helped you, please like, subscribe, and share it with a classmate or colleague. That will help this new channel continue producing free, high-yield medical education content.
🔔 Don’t forget to turn on notifications so you don’t miss upcoming lectures in pharmacology, medical rounds, and more!
#DiagnosticStudies #Diagnostictests #EvidenceBasedMedicine #EBM #DrBusti
Speaker:
Anthony Busti, MD, PharmD, MSc, FAHA, FNLA, is a licensed healthcare professional and medical educator with over 30 years of experience in clinical practice and academic teaching. He has trained and practiced as a nurse, pharmacist, and physician, bringing a uniquely comprehensive perspective to patient care and medical education.
Dr. Busti is dedicated to advancing evidence-based medicine and helping clinicians understand the underlying “why” behind clinical decisions to improve patient outcomes.
About This Channel:
This content is created by Anthony Busti, MD, PharmD, MSc, FAHA, FNLA, a board-certified physician with training at Johns Hopkins School of Medicine and University of Oxford and a medical educator for healthcare professionals and students. All material is based on current medical literature and evidence-based guidelines that align with principles of evidence-based medicine (EBM) and Evidence-Based Healthcare (EBHC).
Disclaimer:
This content is for educational purposes only and is not medical advice. It does not replace individualized evaluation, diagnosis, or treatment. Always seek the advice of a qualified health provider with questions about a medical condition and never delay care because of educational content.
メモを表示 🔗
文字起こし 🔗
00:00:00.239 --> 00:00:04.639
Okay, guys, well, welcome back to our diagnostics test series.
00:00:04.719 --> 00:00:09.839
Now, this is part of our larger EBM and biostatistics series at This Is Why.
00:00:09.919 --> 00:00:17.519
This lecture is going to focus on the concepts of a diagnostic test's precision and its accuracy.
00:00:17.679 --> 00:00:32.719
Okay, now in order to conduct research using new diagnostic processes, or in your review of diagnostic studies, we all need to understand how a particular test we order actually performs.
00:00:32.960 --> 00:00:46.560
Again, this is one of those topics that helps us understand how diagnostic tests were developed and performed so that we can properly use them and interpret the results in their right context.
00:00:46.880 --> 00:00:48.320
And with that, I'm Dr.
00:00:48.479 --> 00:01:08.000
Bustaye, your host on This Is Why, a show that's dedicated to helping you to understand why you do what you do and applying the proper context of the evidence so that we can make evidence-based decisions for the patients we serve every day.
00:01:08.239 --> 00:01:14.719
Now you can find the full playlist of this series on YouTube or at thisiswhy.
00:01:16.640 --> 00:01:18.480
This is part of a larger topic.
00:01:18.560 --> 00:01:23.280
So if you need other parts of the conversation, that's what they're there for.
00:01:23.439 --> 00:01:34.799
It's oriented for both the researcher and the clinician because whether you're a researcher or clinician, we both need to understand what we're reading or publishing and how to interpret it.
00:01:34.879 --> 00:01:36.640
Okay, so this is part of a series.
00:01:36.799 --> 00:01:40.319
Part one is before the study, part two is at the completion of the study.
00:01:40.480 --> 00:01:43.120
Part three, which is where we are predominantly at.
00:01:43.280 --> 00:01:54.319
Part one, remember, is that statistical analysis plan, setting up your groups, the type of data, which statistical analysis, how many patients.
00:01:54.560 --> 00:02:01.439
At the end of the study, you do descriptive analysis and you also do inferential statistics.
00:02:01.760 --> 00:02:06.319
But we are on part three, which is really focused in on those diagnostic tests.
00:02:06.480 --> 00:02:15.199
And if you need the big overview of the topic, then that's where there is an overview lecture that I put together for that.
00:02:15.360 --> 00:02:21.360
But we're on the measurement of accuracy and performance right now.
00:02:21.520 --> 00:02:21.840
Okay.
00:02:22.240 --> 00:02:29.759
Following this, we'll cover some things like sensitivity specificity, again, other measures of performance of that test.
00:02:30.159 --> 00:02:38.800
I'll introduce some of that concepts here, and then we'll hit on later topics on ROC curves, positive and negative predictive values, likelihood ratios.
00:02:38.960 --> 00:02:43.520
Here's the thing: all of those topics can stand alone, including this one.
00:02:44.560 --> 00:02:52.479
But I do pull in and I don't always fully cover all the other components because otherwise the lectures get very long.
00:02:52.639 --> 00:03:02.560
So when we think about precision and/or what is really reliability, we're talking about the consistency, the reproducibility of a test result.
00:03:03.280 --> 00:03:03.599
Okay.
00:03:04.479 --> 00:03:07.039
Does not mean it's always accurate, though.
00:03:07.280 --> 00:03:14.240
It's consistent, it's reproducible, but it may not always be accurate, okay, or true.
00:03:14.319 --> 00:03:17.199
And that's going to talk about accuracy in a minute.
00:03:17.439 --> 00:03:22.400
So this is really about that reliability, precision and reliability.
00:03:22.639 --> 00:03:30.159
If the test experiences random error, this will reduce obviously its precision.
00:03:30.319 --> 00:03:33.520
Okay, so that's reliability and precision.
00:03:33.840 --> 00:03:34.159
Okay.
00:03:34.800 --> 00:03:36.719
What about accuracy?
00:03:36.960 --> 00:03:51.039
Now this reflects the trueness of the test, gives the measurement as close to the actual or true value as compared to the gold standard or sometimes referred to as the reference standard.
00:03:51.280 --> 00:04:14.319
Okay, now we do calculate it, and I'm going to show you how to calculate that going back to our two by two table for diagnostic tests as it relates to diseases, where the accuracy is calculated by taking the true positives plus the true negatives, and you divide that by the total number in the samples.
00:04:14.479 --> 00:04:14.800
Okay.
00:04:15.199 --> 00:04:18.480
And again, I'm going to show you that visually here in just a second.
00:04:18.959 --> 00:04:27.120
We also need to think about systematic error in a study design because that can reduce the accuracy of a test result.
00:04:27.519 --> 00:04:38.000
Okay, so when we talk about precision and accuracy, we have to consider it in the context of bias, an error in the study design that could be introduced.
00:04:40.160 --> 00:04:45.519
Now let's pause and let's talk a little bit about calculations and interpretation.
00:04:45.759 --> 00:04:50.079
I brought this table up before in our overview lecture.
00:04:50.399 --> 00:05:04.399
This really comes up more in our calculation of sensitivity, specificity, our negative, our positive and negative predictive values, and also the likelihood ratio, but it comes up, and I'm introducing it again here because it's part of accuracy as well.
00:05:04.480 --> 00:05:11.439
Remember, I told you those, the true positives and the true negatives divided by the total number of samples.
00:05:11.519 --> 00:05:17.120
And that's what this this is a two by two, where the top and the columns represent the disease.
00:05:17.279 --> 00:05:21.360
It's known, it's diagnosed by the gold or reference standard.
00:05:21.519 --> 00:05:28.879
Whereas on the left are the rows represent the actual tests that we're looking at understanding how it performs.
00:05:29.199 --> 00:05:35.920
In the context of precision, is it reliable, reproducible, and is it also accurate?
00:05:36.079 --> 00:05:42.720
Okay, and so we hope to fall in the true positive or true negative box.
00:05:43.040 --> 00:05:49.120
We don't want to be in the other ones that creates false alarms or missed diseases.
00:05:49.920 --> 00:05:54.000
So when you build one of these, and this is just an example of one, okay.
00:05:54.160 --> 00:06:01.600
This is showing you where for the disease itself, there's a the sensitivity and specificity.
00:06:01.680 --> 00:06:10.480
So sensitivity is calculated by this direction, and the specificity is down this column, okay?
00:06:11.759 --> 00:06:13.279
And you can see those in the formulas here.
00:06:13.360 --> 00:06:17.839
And again, I'm not going to spend time discussing sensitivity specificity right at this moment.
00:06:18.000 --> 00:06:30.800
But I wanted you to see that in the building of this two by two, where you add up the totals in the columns, and then also over here in the rows, they should match up.
00:06:30.959 --> 00:06:34.879
And that this number right here is the total number of samples.
00:06:35.199 --> 00:06:44.000
And that and the true positive and true negative become the formula for accuracy that I mentioned earlier.
00:06:44.240 --> 00:06:47.600
So you can calculate it, or they should calculate it for you.
00:06:48.000 --> 00:06:53.439
Should be provided to you when you're reviewing a paper about a diagnostic test.
00:06:54.240 --> 00:06:54.560
Okay.
00:06:55.439 --> 00:07:00.800
Typically this stuff is just pre-produced, but if you're given the data, you could do it yourself.
00:07:01.120 --> 00:07:04.160
And if the paper didn't do it, you should consider doing that.
00:07:04.399 --> 00:07:04.639
Okay.
00:07:05.120 --> 00:07:09.600
So I wanted you to see the math and how that is done.
00:07:11.519 --> 00:07:12.800
So how is this applies?
00:07:12.959 --> 00:07:14.639
Like, you know, what so what?
00:07:14.800 --> 00:07:16.000
Why do we care?
00:07:16.240 --> 00:07:19.920
Well, you have to think about the equation of error.
00:07:20.000 --> 00:07:25.839
Remember, I talked about that bias, the systematic error that can get introduced.
00:07:26.000 --> 00:07:32.720
We want to make sure that's not happening in the study that we're evaluating on a particular test because I've got to decide whether I'm going to use it or not.
00:07:32.959 --> 00:07:33.279
Okay.
00:07:34.079 --> 00:07:41.519
So the measurement of something has truth, which is influenced by a good study design.
00:07:42.079 --> 00:07:42.399
Right?
00:07:42.560 --> 00:07:53.360
It's also when we're reading about the paper, we're using a critical appraisal tool, uh C A T to assess for the truth.
00:07:53.519 --> 00:07:56.160
We want to know, is it a real value?
00:07:56.480 --> 00:08:04.480
And we're trying to discern bias, which the researcher certainly can introduce, right, by the way, the methodology is done.
00:08:04.639 --> 00:08:07.600
And then we're looking for that systematic error as a reader.
00:08:07.680 --> 00:08:12.639
And that's what the critical appraisal tools are really helping us to also dive into.
00:08:12.959 --> 00:08:18.319
But when we get to this random error over here, there's chances and some variability and size.
00:08:18.399 --> 00:08:21.199
Well, that's influenced by the size of the samples.
00:08:21.439 --> 00:08:26.959
And then our obviously our confidence interval p values start to reflect that.
00:08:27.199 --> 00:08:34.480
But look stepping back, and we can consider these three components as it relates to a measurement of something.
00:08:34.639 --> 00:08:43.919
We have truth, we have the potential risk for bias being introduced by the researcher, and then we have the random error itself that can occur.
00:08:44.159 --> 00:08:53.919
And so when you look at those, again, the target of rely uh the precision, remember it's reproducible.
00:08:54.000 --> 00:08:58.080
It might be inaccurate, but it's reproducible, it's always inaccurate.
00:08:59.120 --> 00:09:02.960
We obviously want precision with accuracy.
00:09:03.519 --> 00:09:19.759
And so when you start thinking about random error and bias, we want low random error over here, and we want low risk for bias so that our precision and our accuracy are where they're supposed to be, the true result.
00:09:21.679 --> 00:09:34.480
But as you have bias being introduced over this direction, you can see that it can change the precision, the reliability of the results.
00:09:34.960 --> 00:09:41.679
If you have high degree of error, well, it's not very accurate, is it?
00:09:42.480 --> 00:09:44.000
It's all over the place.
00:09:44.559 --> 00:09:49.360
And the degree of that uh risk as it relates to bias is also there.
00:09:49.440 --> 00:10:07.120
So that's what this is trying to help you to see in your mind visually when you're studying about an or you're reviewing a paper about a new test, or if you're the researcher doing it, you want to try to get to that one top left corner target.
00:10:07.360 --> 00:10:19.600
So we can judge a new test's accuracy and later on some of its specificity and sensitivity against a gold standard or that reference standard that defines the true disease status.
00:10:20.080 --> 00:10:27.120
Remember, not all, as I mentioned in the overview, not all gold standards or reference standards are perfect.
00:10:27.360 --> 00:10:32.159
In fact, when you develop new diagnostic tests, they may actually perform better.
00:10:32.320 --> 00:10:35.360
And so then you have to ask the question is those are those false.
00:10:37.440 --> 00:10:39.440
Are those are they false positives?
00:10:41.120 --> 00:10:46.480
So remember the gold standard is not always is not always going to be golden.
00:10:46.559 --> 00:10:49.919
So that's where I was just saying it's imperfect, right?
00:10:50.480 --> 00:10:56.799
Um, and so you want to make sure that you have checked these considerations.
00:10:57.200 --> 00:10:59.840
You have imperfect cultures for an infection.
00:11:00.000 --> 00:11:08.080
You maybe the way they were collected, there's contamination, the clinical diagnosis uh without tissue confirmation, right?
00:11:08.559 --> 00:11:12.960
Older or less sensitive imaging modalities, okay.
00:11:13.200 --> 00:11:19.600
Um doing an open mouth x-ray versus a CT scan of a cervical spine, very different, right?
00:11:20.080 --> 00:11:22.559
Subjective reads, okay.
00:11:23.120 --> 00:11:35.840
So when imperfect gold standard bias is being introduced, we think of things like when the reference standard misclassifies the patients and the new test performance becomes then biased by that.
00:11:36.159 --> 00:11:36.480
Okay.
00:11:36.879 --> 00:11:45.840
So for everyday clinicians, you know, reading the literature, we're we're gonna ask ourselves, how good is the gold standard even in this study?
00:11:46.480 --> 00:11:52.480
Was it applied the same way to every patient in the study?
00:11:52.960 --> 00:11:55.840
And is there any better reference available?
00:11:56.000 --> 00:12:08.960
So you always want to know, and you go back to the question of a PERT a pyro structure in your evaluation trying to acquire information, acquire the best study, you're looking for papers that did that.
00:12:09.200 --> 00:12:19.120
They use the proper gold standard or reference test for the patients in the question being looked at.
00:12:19.679 --> 00:12:26.879
Otherwise, this new test may not perform and represent some degree of bias.
00:12:28.159 --> 00:12:35.360
And so that's why critical appraisal tools like the quadus analysis is done to help guide us through that process.
00:12:35.440 --> 00:12:42.960
And again, I introduced that in the overview lecture when we talked about how to review a paper and what are the tools.
00:12:43.279 --> 00:12:51.759
All right, a test that looks excellent versus a weak gold standard may disappoint in also real clinical practice.
00:12:52.080 --> 00:12:59.679
So, core concepts on this particular topic: precision is about reliability, reproducibility.
00:12:59.919 --> 00:13:03.440
How consistent and reproducible are those results?
00:13:03.679 --> 00:13:05.519
They may not always be accurate.
00:13:05.679 --> 00:13:13.200
We want them to be, but that's where accuracy that is the trueness of the results from that test.
00:13:13.440 --> 00:13:24.559
And we want to make sure that we are doing our best to merge both of those and remove bias and random error from that process in the measurement.
00:13:25.120 --> 00:13:35.759
Okay, guys, this topic can be seen in part of the full diagnostic playlist again over at thisiswide.health, okay, and part of our EBM series and collections.
00:13:35.919 --> 00:13:37.600
You can also find them on YouTube.
00:13:38.000 --> 00:13:38.480
All right.
00:13:38.960 --> 00:13:41.200
If you're not a subscriber, please consider joining.
00:13:41.360 --> 00:13:42.559
We really would appreciate that.
00:13:42.720 --> 00:13:43.919
Also, some support.
00:13:44.080 --> 00:13:51.840
We'll give it a thumbs up, maybe post a comment, share some experience, additional perspective that we can all learn from each other.
00:13:52.000 --> 00:13:53.759
So thank you for joining me.
00:13:53.840 --> 00:13:54.240
I'm Dr.
00:13:54.320 --> 00:14:01.759
Bust Eye, and we'll see you next time on the next topic as it relates to evidence based medicine and diagnostic tests.