WEBVTT
1
00:00:00.120 --> 00:00:04.360
Welcome to Data Science Dot Show, the podcast where data
2
00:00:04.440 --> 00:00:09.880
science meets executive leadership. We talk with TOPAI and analytics
3
00:00:09.919 --> 00:00:14.240
thought leaders about turning data into real business impact. If
4
00:00:14.240 --> 00:00:17.800
you're a C level leader or data expert shaping the future,
5
00:00:18.239 --> 00:00:19.359
you're in the right place.
6
00:00:20.359 --> 00:00:23.480
Imagine waking up to a twelve percent drop in online
7
00:00:23.519 --> 00:00:26.760
conversions and no one can tell you why the alerts
8
00:00:26.760 --> 00:00:30.440
were on the dashboard was green, but customers stopped buying.
9
00:00:30.960 --> 00:00:35.000
What if your observability made that difference visible sooner? I'm
10
00:00:35.079 --> 00:00:39.359
Merco Peters. We explore how data science and AI drive
11
00:00:39.479 --> 00:00:44.079
real business impact, straight from the leaders making it happen Today.
12
00:00:44.439 --> 00:00:48.399
I want to talk about observability built for decisions, not
13
00:00:48.600 --> 00:00:53.359
telemetry for its own sake. If you run technology, product
14
00:00:53.560 --> 00:00:57.679
or risk at scale, this episode is about turning alerts
15
00:00:57.840 --> 00:01:02.039
into assurance you can act on. Picture. The typical vantage
16
00:01:02.439 --> 00:01:05.599
a chief data officer at a global bank owning a
17
00:01:05.680 --> 00:01:11.439
dozen models fraud, credit decisions, churn one, silent degradation and
18
00:01:11.560 --> 00:01:16.439
initiative stall. Regulators ask questions, customers get the wrong outcome.
19
00:01:17.000 --> 00:01:21.359
Observabilities stopped being an engineering checklist the day those failures
20
00:01:21.439 --> 00:01:25.000
hit the balance sheet. Each layer gives you different signals,
21
00:01:25.239 --> 00:01:29.400
and each signal needs a business translation. A schema change
22
00:01:29.480 --> 00:01:33.120
matters to engineers to a CFO, it matters only if
23
00:01:33.120 --> 00:01:37.640
it raises false approvals or cuts conversion. So always ask
24
00:01:38.200 --> 00:01:43.920
which KPI will this signal move and by how much. Concretely,
25
00:01:44.319 --> 00:01:49.120
at the data layer, monitors schema changes, missingness and upstream
26
00:01:49.200 --> 00:01:54.640
latency At the feature layer, watch distributional shifts and correlation
27
00:01:54.799 --> 00:02:01.400
breakdowns at the prediction layer, track confidence calibration and prediction volume.
28
00:02:02.120 --> 00:02:07.439
At the outcome layer, measure conversion, revenue per user, error rates,
29
00:02:07.480 --> 00:02:12.360
and complaints. Map each technical signal to a business consequence.
30
00:02:12.919 --> 00:02:16.960
This alert ties to X percent revenue risk or Y
31
00:02:17.159 --> 00:02:24.039
regulatory exposure or Z customer experienced degradation. Translate telemetry, don't
32
00:02:24.159 --> 00:02:28.400
let it stay. Technical noise alerts without an escalation play
33
00:02:28.520 --> 00:02:32.759
are noise. Start with a small set of SLOs three
34
00:02:32.840 --> 00:02:37.879
to five that mirror technical health and business impact examples.
35
00:02:38.400 --> 00:02:43.039
Model availability greater than ninety nine point five percent, calibration
36
00:02:43.319 --> 00:02:48.599
within agreed bounds, and outcome KPI variants under a threshold.
37
00:02:49.080 --> 00:02:53.400
When an SLO breaches, severity matters, soft warnings go to
38
00:02:53.400 --> 00:02:56.759
the model owner. Hard breaches trigger an incident. A run
39
00:02:56.800 --> 00:03:01.159
book and executive notification build compound run books for the
40
00:03:01.240 --> 00:03:05.520
handful of patterns that actually cause outages. Most incidents come
41
00:03:05.560 --> 00:03:10.800
from a few sources pipeline failure, sudden population drift, code regressions,
42
00:03:10.919 --> 00:03:15.159
or broken feedback loops, not some exotic edge case. Each
43
00:03:15.240 --> 00:03:19.919
pattern needs a compact checklist, triage, a quick mitigation like
44
00:03:19.960 --> 00:03:23.439
a fallback model or input validation, and an owner for
45
00:03:23.560 --> 00:03:24.840
post mortem and fixes.
46
00:03:25.400 --> 00:03:28.439
The goal is reduced time to safe state, not to
47
00:03:28.479 --> 00:03:32.719
patch everything in the first hour picture this an ab rollout,
48
00:03:32.840 --> 00:03:37.439
and overnight the recommendation model starts showing strange top end changes.
49
00:03:38.000 --> 00:03:42.840
Early feature distribution alerts flag anomalies. Because the team tied
50
00:03:42.840 --> 00:03:46.439
that alert to CTR on recommended widgets, they paused the
51
00:03:46.479 --> 00:03:50.199
new feed applied to transformation and avoided a twelve percent
52
00:03:50.280 --> 00:03:55.479
drop in conversion small translation, big impact. Another case, a
53
00:03:55.560 --> 00:04:00.000
lending model drifted after a regional policy change. Outcome monitory
54
00:04:00.199 --> 00:04:05.000
spotted rising default rates before internal confidence metrics moved. Early
55
00:04:05.080 --> 00:04:10.199
detection prevented regulatory escalation and allowed a controlled retune. You'll
56
00:04:10.240 --> 00:04:15.120
hear the trade off debate sensitivity versus noise, coverage versus cost.
57
00:04:15.719 --> 00:04:18.439
Someone in finance will ask how much will this cost
58
00:04:18.480 --> 00:04:22.199
to investigate, and you should have the answer. A quick
59
00:04:22.399 --> 00:04:25.920
blunt rule. If an alert cost ten thousand dollars to
60
00:04:26.000 --> 00:04:29.879
investigate and only prevents one thousand dollars in risk, kill
61
00:04:29.920 --> 00:04:34.839
it exactly. Not every alert should exist. Reduce human load
62
00:04:34.959 --> 00:04:40.240
with automated triage, anomaly scoring, aggregated signals, and preliminary root
63
00:04:40.319 --> 00:04:45.000
cause hints. Let machines do the noisy first pass. Humans
64
00:04:45.079 --> 00:04:49.680
focus on decisions. When teams ask build versus buy, ask
65
00:04:49.879 --> 00:04:53.639
three questions. Does the tool expose the signals we need
66
00:04:53.839 --> 00:04:57.399
tied to our KPIs? Can it plug into our incident
67
00:04:57.480 --> 00:05:02.000
and audit workflows? Will it scale cost effectively across our estate.
68
00:05:02.680 --> 00:05:07.160
Many organizations land on hybrid vendor telemetry for baseline detection
69
00:05:07.519 --> 00:05:10.759
and in house layers that map signals to domain specific
70
00:05:10.920 --> 00:05:16.160
SLOs and playbooks. Ownership matters, Product model teams own day
71
00:05:16.199 --> 00:05:20.680
to day monitoring and first line remediation. A central model
72
00:05:20.720 --> 00:05:25.959
reliability or mL OPS platform owns tooling, SLO templates, and
73
00:05:26.000 --> 00:05:30.199
the audit trail, Legal and compliance step in where outcomes
74
00:05:30.279 --> 00:05:36.279
touch regulation make the RACI explicit. Who declares an incident,
75
00:05:36.720 --> 00:05:41.879
who runs the playbook, who updates stakeholders? Who approves rollback
76
00:05:42.639 --> 00:05:46.439
that clarity removes delay at the moment you need speed.
77
00:05:47.120 --> 00:05:52.399
Measuring ROI is simple if you start small, track incident frequency,
78
00:05:52.720 --> 00:05:57.759
meantime to detection and meantime to remediation. Convert those into
79
00:05:57.800 --> 00:06:01.839
dollars by estimating revenue at risk per incident and investigation
80
00:06:02.000 --> 00:06:07.519
cost pilot first scale after you see measurable wins. Auditability
81
00:06:07.720 --> 00:06:11.879
is non negotiable. Every alert should capture raw telemetry, a
82
00:06:11.959 --> 00:06:16.759
triage summary, timestamps, and decision rationale. Store it in an
83
00:06:16.800 --> 00:06:21.360
immutable incident record for compliance and rapid post mortem. Automate
84
00:06:21.439 --> 00:06:25.800
evidence collection where you can logs, model versions and data
85
00:06:25.839 --> 00:06:30.199
snapshots so post mortems are factual and fast. For the
86
00:06:30.240 --> 00:06:33.480
next twelve to twenty four months, focus on three moves.
87
00:06:33.920 --> 00:06:37.639
Mandate a small set of SLOs tied to business KPIs,
88
00:06:38.079 --> 00:06:41.600
require an auditible run book for high severity alerts, and
89
00:06:41.759 --> 00:06:45.560
fund a hybrid monitoring pilot that proves ROI before a
90
00:06:45.560 --> 00:06:49.040
big buy. Two quick actions you can implement next quarter.
91
00:06:49.439 --> 00:06:52.120
Run a one week audit across your top ten models
92
00:06:52.160 --> 00:06:56.439
to map telemetry to business consequences, and require each model
93
00:06:56.439 --> 00:06:59.000
owner to publish a one page run book with an
94
00:06:59.120 --> 00:07:05.639
SLO baseline the central team validates final thought. Observability is
95
00:07:05.680 --> 00:07:09.079
not a dashboard or a checkbox. It's a discipline that
96
00:07:09.120 --> 00:07:14.600
links signals to outcomes, enforces auditable response, and gives executives
97
00:07:14.600 --> 00:07:18.800
the confidence to scale AI, adopt the layered metric map,
98
00:07:19.120 --> 00:07:24.399
prioritize SLOs that matter, and make incident evidence immusable. If
99
00:07:24.399 --> 00:07:28.120
this landed for you, subscribe to Data Science Dot Show,
100
00:07:28.560 --> 00:07:31.360
share it with your leadership team, and connect with me
101
00:07:31.439 --> 00:07:36.120
on LinkedIn. I'm Merco Peters. That's the difference between models
102
00:07:36.160 --> 00:07:36.800
and value.
103
00:07:37.120 --> 00:07:41.000
Thanks for listening to Data Science Dot Show. If you
104
00:07:41.120 --> 00:07:45.279
found value in this episode, subscribe, share it with your network,
105
00:07:45.600 --> 00:07:48.439
and join us next time as we explore how data
106
00:07:48.480 --> 00:07:52.519
science drives smarter decisions and better business outcomes.
1
00:00:00.120 --> 00:00:04.360
Welcome to Data Science Dot Show, the podcast where data
2
00:00:04.440 --> 00:00:09.880
science meets executive leadership. We talk with TOPAI and analytics
3
00:00:09.919 --> 00:00:14.240
thought leaders about turning data into real business impact. If
4
00:00:14.240 --> 00:00:17.800
you're a C level leader or data expert shaping the future,
5
00:00:18.239 --> 00:00:19.359
you're in the right place.
6
00:00:20.359 --> 00:00:23.480
Imagine waking up to a twelve percent drop in online
7
00:00:23.519 --> 00:00:26.760
conversions and no one can tell you why the alerts
8
00:00:26.760 --> 00:00:30.440
were on the dashboard was green, but customers stopped buying.
9
00:00:30.960 --> 00:00:35.000
What if your observability made that difference visible sooner? I'm
10
00:00:35.079 --> 00:00:39.359
Merco Peters. We explore how data science and AI drive
11
00:00:39.479 --> 00:00:44.079
real business impact, straight from the leaders making it happen Today.
12
00:00:44.439 --> 00:00:48.399
I want to talk about observability built for decisions, not
13
00:00:48.600 --> 00:00:53.359
telemetry for its own sake. If you run technology, product
14
00:00:53.560 --> 00:00:57.679
or risk at scale, this episode is about turning alerts
15
00:00:57.840 --> 00:01:02.039
into assurance you can act on. Picture. The typical vantage
16
00:01:02.439 --> 00:01:05.599
a chief data officer at a global bank owning a
17
00:01:05.680 --> 00:01:11.439
dozen models fraud, credit decisions, churn one, silent degradation and
18
00:01:11.560 --> 00:01:16.439
initiative stall. Regulators ask questions, customers get the wrong outcome.
19
00:01:17.000 --> 00:01:21.359
Observabilities stopped being an engineering checklist the day those failures
20
00:01:21.439 --> 00:01:25.000
hit the balance sheet. Each layer gives you different signals,
21
00:01:25.239 --> 00:01:29.400
and each signal needs a business translation. A schema change
22
00:01:29.480 --> 00:01:33.120
matters to engineers to a CFO, it matters only if
23
00:01:33.120 --> 00:01:37.640
it raises false approvals or cuts conversion. So always ask
24
00:01:38.200 --> 00:01:43.920
which KPI will this signal move and by how much. Concretely,
25
00:01:44.319 --> 00:01:49.120
at the data layer, monitors schema changes, missingness and upstream
26
00:01:49.200 --> 00:01:54.640
latency At the feature layer, watch distributional shifts and correlation
27
00:01:54.799 --> 00:02:01.400
breakdowns at the prediction layer, track confidence calibration and prediction volume.
28
00:02:02.120 --> 00:02:07.439
At the outcome layer, measure conversion, revenue per user, error rates,
29
00:02:07.480 --> 00:02:12.360
and complaints. Map each technical signal to a business consequence.
30
00:02:12.919 --> 00:02:16.960
This alert ties to X percent revenue risk or Y
31
00:02:17.159 --> 00:02:24.039
regulatory exposure or Z customer experienced degradation. Translate telemetry, don't
32
00:02:24.159 --> 00:02:28.400
let it stay. Technical noise alerts without an escalation play
33
00:02:28.520 --> 00:02:32.759
are noise. Start with a small set of SLOs three
34
00:02:32.840 --> 00:02:37.879
to five that mirror technical health and business impact examples.
35
00:02:38.400 --> 00:02:43.039
Model availability greater than ninety nine point five percent, calibration
36
00:02:43.319 --> 00:02:48.599
within agreed bounds, and outcome KPI variants under a threshold.
37
00:02:49.080 --> 00:02:53.400
When an SLO breaches, severity matters, soft warnings go to
38
00:02:53.400 --> 00:02:56.759
the model owner. Hard breaches trigger an incident. A run
39
00:02:56.800 --> 00:03:01.159
book and executive notification build compound run books for the
40
00:03:01.240 --> 00:03:05.520
handful of patterns that actually cause outages. Most incidents come
41
00:03:05.560 --> 00:03:10.800
from a few sources pipeline failure, sudden population drift, code regressions,
42
00:03:10.919 --> 00:03:15.159
or broken feedback loops, not some exotic edge case. Each
43
00:03:15.240 --> 00:03:19.919
pattern needs a compact checklist, triage, a quick mitigation like
44
00:03:19.960 --> 00:03:23.439
a fallback model or input validation, and an owner for
45
00:03:23.560 --> 00:03:24.840
post mortem and fixes.
46
00:03:25.400 --> 00:03:28.439
The goal is reduced time to safe state, not to
47
00:03:28.479 --> 00:03:32.719
patch everything in the first hour picture this an ab rollout,
48
00:03:32.840 --> 00:03:37.439
and overnight the recommendation model starts showing strange top end changes.
49
00:03:38.000 --> 00:03:42.840
Early feature distribution alerts flag anomalies. Because the team tied
50
00:03:42.840 --> 00:03:46.439
that alert to CTR on recommended widgets, they paused the
51
00:03:46.479 --> 00:03:50.199
new feed applied to transformation and avoided a twelve percent
52
00:03:50.280 --> 00:03:55.479
drop in conversion small translation, big impact. Another case, a
53
00:03:55.560 --> 00:04:00.000
lending model drifted after a regional policy change. Outcome monitory
54
00:04:00.199 --> 00:04:05.000
spotted rising default rates before internal confidence metrics moved. Early
55
00:04:05.080 --> 00:04:10.199
detection prevented regulatory escalation and allowed a controlled retune. You'll
56
00:04:10.240 --> 00:04:15.120
hear the trade off debate sensitivity versus noise, coverage versus cost.
57
00:04:15.719 --> 00:04:18.439
Someone in finance will ask how much will this cost
58
00:04:18.480 --> 00:04:22.199
to investigate, and you should have the answer. A quick
59
00:04:22.399 --> 00:04:25.920
blunt rule. If an alert cost ten thousand dollars to
60
00:04:26.000 --> 00:04:29.879
investigate and only prevents one thousand dollars in risk, kill
61
00:04:29.920 --> 00:04:34.839
it exactly. Not every alert should exist. Reduce human load
62
00:04:34.959 --> 00:04:40.240
with automated triage, anomaly scoring, aggregated signals, and preliminary root
63
00:04:40.319 --> 00:04:45.000
cause hints. Let machines do the noisy first pass. Humans
64
00:04:45.079 --> 00:04:49.680
focus on decisions. When teams ask build versus buy, ask
65
00:04:49.879 --> 00:04:53.639
three questions. Does the tool expose the signals we need
66
00:04:53.839 --> 00:04:57.399
tied to our KPIs? Can it plug into our incident
67
00:04:57.480 --> 00:05:02.000
and audit workflows? Will it scale cost effectively across our estate.
68
00:05:02.680 --> 00:05:07.160
Many organizations land on hybrid vendor telemetry for baseline detection
69
00:05:07.519 --> 00:05:10.759
and in house layers that map signals to domain specific
70
00:05:10.920 --> 00:05:16.160
SLOs and playbooks. Ownership matters, Product model teams own day
71
00:05:16.199 --> 00:05:20.680
to day monitoring and first line remediation. A central model
72
00:05:20.720 --> 00:05:25.959
reliability or mL OPS platform owns tooling, SLO templates, and
73
00:05:26.000 --> 00:05:30.199
the audit trail, Legal and compliance step in where outcomes
74
00:05:30.279 --> 00:05:36.279
touch regulation make the RACI explicit. Who declares an incident,
75
00:05:36.720 --> 00:05:41.879
who runs the playbook, who updates stakeholders? Who approves rollback
76
00:05:42.639 --> 00:05:46.439
that clarity removes delay at the moment you need speed.
77
00:05:47.120 --> 00:05:52.399
Measuring ROI is simple if you start small, track incident frequency,
78
00:05:52.720 --> 00:05:57.759
meantime to detection and meantime to remediation. Convert those into
79
00:05:57.800 --> 00:06:01.839
dollars by estimating revenue at risk per incident and investigation
80
00:06:02.000 --> 00:06:07.519
cost pilot first scale after you see measurable wins. Auditability
81
00:06:07.720 --> 00:06:11.879
is non negotiable. Every alert should capture raw telemetry, a
82
00:06:11.959 --> 00:06:16.759
triage summary, timestamps, and decision rationale. Store it in an
83
00:06:16.800 --> 00:06:21.360
immutable incident record for compliance and rapid post mortem. Automate
84
00:06:21.439 --> 00:06:25.800
evidence collection where you can logs, model versions and data
85
00:06:25.839 --> 00:06:30.199
snapshots so post mortems are factual and fast. For the
86
00:06:30.240 --> 00:06:33.480
next twelve to twenty four months, focus on three moves.
87
00:06:33.920 --> 00:06:37.639
Mandate a small set of SLOs tied to business KPIs,
88
00:06:38.079 --> 00:06:41.600
require an auditible run book for high severity alerts, and
89
00:06:41.759 --> 00:06:45.560
fund a hybrid monitoring pilot that proves ROI before a
90
00:06:45.560 --> 00:06:49.040
big buy. Two quick actions you can implement next quarter.
91
00:06:49.439 --> 00:06:52.120
Run a one week audit across your top ten models
92
00:06:52.160 --> 00:06:56.439
to map telemetry to business consequences, and require each model
93
00:06:56.439 --> 00:06:59.000
owner to publish a one page run book with an
94
00:06:59.120 --> 00:07:05.639
SLO baseline the central team validates final thought. Observability is
95
00:07:05.680 --> 00:07:09.079
not a dashboard or a checkbox. It's a discipline that
96
00:07:09.120 --> 00:07:14.600
links signals to outcomes, enforces auditable response, and gives executives
97
00:07:14.600 --> 00:07:18.800
the confidence to scale AI, adopt the layered metric map,
98
00:07:19.120 --> 00:07:24.399
prioritize SLOs that matter, and make incident evidence immusable. If
99
00:07:24.399 --> 00:07:28.120
this landed for you, subscribe to Data Science Dot Show,
100
00:07:28.560 --> 00:07:31.360
share it with your leadership team, and connect with me
101
00:07:31.439 --> 00:07:36.120
on LinkedIn. I'm Merco Peters. That's the difference between models
102
00:07:36.160 --> 00:07:36.800
and value.
103
00:07:37.120 --> 00:07:41.000
Thanks for listening to Data Science Dot Show. If you
104
00:07:41.120 --> 00:07:45.279
found value in this episode, subscribe, share it with your network,
105
00:07:45.600 --> 00:07:48.439
and join us next time as we explore how data
106
00:07:48.480 --> 00:07:52.519
science drives smarter decisions and better business outcomes.