OVER DEZE AFLEVERING
Your coding agent can inspect its own chat, but it cannot see you switching between browser research, documents, local files, and devices.
Screen recordings can turn that invisible work into model context effectively with video native models like Gemini.
In this episode, we show how Screenpipe exposed Christine acting as the context bridge around Codex, and Brandon's workflow for recording focused work sessions, compressing them to 720p at one frame per second, splitting them into 15-minute clips, and asking Gemini for a detailed play-by-play. We also extend the same idea to real phone test rigs, cameras, friction logs, and skills generated from recorded tasks.
Superlinear is a podcast about emerging practices for building with AI.
Hosted by Brandon Kase and Christine Yip.
Homepage: https://superlinear.fm
Twitter: https://x.com/superlinear_fm
LinkedIn: https://www.linkedin.com/company/superlinearfm
Brandon Kase: https://x.com/bkase_
Christine Yip: https://x.com/christinetyip
Music licensed through Soundstripe. Code: O2UT6CCEQNH4KT0D
LAAT NOTITIES ZIEN 🔗
TRANSCRIPTIE 🔗
00:00:54.057 --> 00:01:00.936
Hey everyone. Gemini can take video as a native input and reason over what happens across the recording.
00:01:01.676 --> 00:01:06.977
So this means that we do not have to only use transcripts and screenshots as input for agents.
00:01:07.397 --> 00:01:07.417
We
00:01:07.496 --> 00:01:07.516
So
00:01:07.697 --> 00:01:08.016
can
00:01:08.016 --> 00:01:08.037
I'm
00:01:08.076 --> 00:01:12.917
now give it a recording or video file and the recording itself becomes model context.
00:01:13.700 --> 00:01:13.859
So we
00:01:13.859 --> 00:01:14.000
just gonna
00:01:14.040 --> 00:01:14.319
will extend
00:01:14.319 --> 00:01:14.560
say that
00:01:14.599 --> 00:01:16.719
this idea across anything that
00:01:16.719 --> 00:01:16.739
I'm
00:01:16.760 --> 00:01:17.980
would help an agent see.
00:01:18.700 --> 00:01:26.159
You can think of screen of the phone if you're building an app, a video feed of your desk when you're tinkering with something physical
00:01:26.260 --> 00:01:26.280
not
00:01:26.420 --> 00:01:27.099
like a robot.
00:01:27.099 --> 00:01:27.120
sure
00:01:28.379 --> 00:01:34.939
And in this episode, we will help the coding agent break out of the chat by progressively expanding its
00:01:34.980 --> 00:01:35.000
I'm
00:01:35.060 --> 00:01:36.159
field of view.
00:01:36.939 --> 00:01:45.239
So from the current session to its history to your screen and finally to environments even outside of your computer.
00:01:45.939 --> 00:01:49.439
Then give it hands to reflect or act on what it sees.
00:01:50.629 --> 00:01:58.390
So before we get into the examples, Brandon, what is the simplest useful way to reflect on how we work around an agent
00:01:58.847 --> 00:01:59.447
tell it to.
00:02:00.719 --> 00:02:11.646
So, what do by that whenever you're working with an agent, at any point in time you can just ask, hey can you analyze this session and tell me something that I can do better?
00:02:12.294 --> 00:02:14.455
And it'll tell you. And it's really helpful.
00:02:15.326 --> 00:02:28.383
What can you or what can you do better? Or are is there some way that we can set up tools or scripts or our environment or adjust our agents at MD to improve your performance.
00:02:29.062 --> 00:02:32.043
This is this is the kind of questions that you should be asking your agent from time to time.
00:02:32.622 --> 00:02:39.847
Yeah. And a coding agent is the best the the best tool to ask since, everything is in its context.
00:02:40.735 --> 00:02:41.134
Exactly.
00:02:41.794 --> 00:02:42.414
Exactly. So.
00:02:43.479 --> 00:02:50.560
So, whenever you're trying to get feedback and improve within a session, the best way to do it is to just ask the agent in that session.
00:02:51.360 --> 00:03:02.323
But that's not the only situation or the only way in which it's useful to reflect with your agents. Cuz it's also useful to ask about maybe history across lots of sessions.
00:03:03.823 --> 00:03:04.282
Yeah.
00:03:05.643 --> 00:03:31.562
I This reminds me of the Cloud Code reports that Cloud Code generates automatically. Maybe Code X does this too But those kind of reports, they analyze the session summary across multiple sessions you had, and they're really good at surfacing recurring patterns or repeating behaviors.
00:03:32.272 --> 00:03:48.384
Maybe I can show one So here is the Cloud Code Insights report as you can see it's in this report forty forty-five sessions out of ninety-four were included.
00:03:49.465 --> 00:03:54.564
Probably the very short sessions were skipped and probably sub-agent work is not analyzed.
00:03:55.044 --> 00:04:07.402
So this is a view of the main session rather than the entire the entire workflow think it this is one of the interesting things it gave me a prompt to build a skill
00:04:09.323 --> 00:04:09.342
sorry.
00:04:09.802 --> 00:04:17.603
and it's touching on some steps in my workflow that keeps repeating all the time. So I often clear
00:04:17.603 --> 00:04:17.622
I'm
00:04:18.583 --> 00:04:21.862
context during a session and then I start a new session.
00:04:22.142 --> 00:04:22.382
not sure what
00:04:22.403 --> 00:04:23.083
So it
00:04:23.122 --> 00:04:23.182
you're saying.
00:04:23.322 --> 00:04:23.983
noticed that I
00:04:23.983 --> 00:04:24.062
I'm sorry
00:04:24.122 --> 00:04:24.642
keep doing
00:04:24.642 --> 00:04:24.663
I'm
00:04:24.663 --> 00:04:25.362
that across
00:04:25.403 --> 00:04:25.423
sorry
00:04:26.137 --> 00:04:27.257
more than twenty sessions.
00:04:28.036 --> 00:04:30.596
And I often it also noticed that I do it manually.
00:04:30.617 --> 00:04:40.197
So here it recommends me to do to create a skill called session start and session close. So it doesn't happen it doesn't need to happen manually
00:04:40.257 --> 00:04:40.536
I'm sorry
00:04:40.617 --> 00:04:40.716
anymore
00:04:40.716 --> 00:04:52.338
I'm Yeah, I think I think these kinds of insights are really helpful but the the Cloud Code insight report is fixed and it's just for Cloud, right? So
00:04:52.718 --> 00:04:53.077
Mm-hmm.
00:04:53.798 --> 00:05:06.137
we might also want to ask more specific questions and analyze session history across all of our coding agents if you're like me and you use Codex and Cloud and Pi and play with other tools.
00:05:06.898 --> 00:05:15.158
And so you can just ask your agents to look through the histories of all these different things but they're all in different formats and so it'll be kind of clunky.
00:05:16.117 --> 00:05:20.077
So there's actually a nice tool by Doodle Steen.
00:05:21.206 --> 00:05:27.607
So there's this tool called Coding Agent Session Search or CASS that you can install.
00:05:28.307 --> 00:05:31.987
And I've just pulled up the read me and I'll just read the section of why this exists.
00:05:33.307 --> 00:05:54.846
I'll read snippets of it anyway so basically the problem is that all of these tools, all these different harnesses, create conversation trails, but the knowledge is scattered and unsearchable. And so it's fragmented, there's little cross-agent visibility, et cetera.
00:05:55.706 --> 00:06:19.367
And so what CASS does is it indexes all of your session history from all your different tools and creates a unified knowledge base that's normalized, it's indexed, and it's super-fast at doing specific searches. So I'm going to just show as an example of me having used CASS.
00:06:21.983 --> 00:06:50.115
So I asked I asked CASS to analyze my work in the latest podcast video editor that I'm building to edit this podcast in a way that's less annoying And it noticed that it noticed that applying patches fails five percent of the time, which was over fifty the time.
00:06:51.394 --> 00:06:59.192
So so it's saying and then it suggests something that I can add to agents dot MD to to make this better. So that's nice
00:07:01.271 --> 00:07:02.112
Yeah, I like that it's interactive
00:07:02.112 --> 00:07:04.452
and say that again
00:07:05.687 --> 00:07:16.406
I like that it's interactive, that you can ask follow-up questions, like maybe you could like dive deeper or try to understand like what improvement exactly could be.
00:07:17.036 --> 00:07:17.557
Exactly. So
00:07:17.937 --> 00:07:17.976
That
00:07:17.976 --> 00:07:18.317
this is just
00:07:18.317 --> 00:07:18.377
that's
00:07:18.377 --> 00:07:18.497
a conversation
00:07:18.497 --> 00:07:18.956
like a good improvement
00:07:18.956 --> 00:07:19.357
I'm having.
00:07:19.357 --> 00:07:21.797
compared to the Cloud Code report.
00:07:22.408 --> 00:07:30.327
this CASS tool, you can use it manually, I'm not even gonna show that, but it just lets you search quickly through your history.
00:07:27.867 --> 00:07:36.908
Or you can use it through Cloud Code and you can get insights over your session history across all different kinds of coding agents, which is really cool.
00:07:38.648 --> 00:07:43.327
But let's break out of transcripts.
00:07:44.447 --> 00:07:45.708
Let's break out of transcripts.
00:07:46.264 --> 00:08:24.843
Yeah. I actually tried that out, breaking out of the transcripts So because even if I use transcripts it still cannot really see what I the the the transcript captures things that I did in the chat with the coding agent. But my workflow consists of more things than just working in the chat with the coding agent. Like I might open different documents, I might browse, I might do some research in other tools. So I tested another tool called ScreenPipe.
00:08:25.535 --> 00:08:27.714
I think that tool as well, right, Brendan?
00:08:28.502 --> 00:09:04.289
I tried out an early version of ScreenPipe over a year ago, I think it's pretty cool, like ScreenPipe records what you do on your computer and indexes it so that you can search and ask questions about what happened on the screen later. And and that's cool, cuz it can it can capture this kind of information that you're saying that we can later surface to agents. So I haven't actually used it since it briefly, since it came out a year ago. I know the team's been working on it a lot. So what did you find what what did you do with ScreenPipe
00:09:05.830 --> 00:09:22.450
Yeah. So I was building this workflow to prepare podcast episodes for our podcast. And I had ScreenPipe on the whole time.
00:09:23.409 --> 00:09:37.620
And so this is here is like here's a good example of what ScreenPipe noticed and it's something that Codex did not notice because Codex cannot
00:09:37.620 --> 00:09:37.759
Can you
00:09:37.759 --> 00:09:37.860
look
00:09:37.860 --> 00:09:38.230
read it out?
00:09:38.230 --> 00:09:39.899
beyond the chat. Yes.
00:09:40.340 --> 00:09:40.519
Can you read
00:09:40.519 --> 00:09:40.620
So ScreenPipe
00:09:40.620 --> 00:09:41.539
it out for the listeners?
00:09:41.980 --> 00:10:00.899
here observes a sequence where I was in a finder to find a Markdown file. And then it saw that I moved to Chatty PT in my browser to ask questions about it.
00:10:01.840 --> 00:10:12.840
And then I went back to local files in finder to inspect another file or folder that I downloaded after my conversation with Chet TBT.
00:10:13.919 --> 00:10:18.559
And then I returned to the coding agent and gave it the files that I just downloaded.
00:10:19.320 --> 00:10:36.139
So basically what here's ScreenPipe is saying and is observing is while you were comparing and transferring material between a planning conversation a Google Docs outline, local Markdown files, and then the workflow package that you used for the coding agent.
00:10:37.259 --> 00:11:01.652
And it did say like, well probably this cross-checking and moving across different tools was necessary but maybe the repeated navigation was not. And it flagged that the missing capability is that the coding agent, and I was using Codex at the time, Codex lacks direct, convenient access to Google Docs and Jet-CBT planning context.
00:11:02.591 --> 00:11:10.851
And it suggested that a read-only integration or exported local snapshot would have eliminated the need to visually shuttle between the tools, what I did.
00:11:11.873 --> 00:11:17.472
And if I scroll down more to like, OK, what should you move into the hardness of the coding agent?
00:11:17.493 --> 00:11:19.393
What should you automate in a workflow?
00:11:19.913 --> 00:12:08.990
ScreenPipe told me, well, give it read-only imports for tools like Whisperflow, Granola, Google Docs, Jet-CBT context, all those different tools that I used and that I switched between manually So yeah, I felt that was a good use of ScreenPipe so I would not use ScreenPipe for things that coding agents are already good at answering I did also try asking ScreenPipe like, hey what should I improve around my workflow and I was more thinking, well, I was more thinking about I would get back like really good advice about, well, this is here's where you can implement tests in a better way, this is how you can avoid these kind of errors.
00:12:05.509 --> 00:12:24.529
But actually that was, Codex was much stronger at it because Codex actually has all your context in its coding session and I guess coding agents are also trained and built in a way that knows the best way of doing harness engineering.
00:12:25.230 --> 00:12:27.889
But what was but ScreenPipe was good at
00:12:27.970 --> 00:12:27.990
Well
00:12:28.169 --> 00:12:40.590
questions like, hey, what happened every time I left CodeX and then later returned? Or what was I looking for, checking or creating or transferring outside of CodeX and it treated the whole
00:12:40.590 --> 00:12:41.149
I'd put that I'm
00:12:41.149 --> 00:12:41.470
computer
00:12:42.190 --> 00:12:42.370
I'm justiminating
00:12:43.110 --> 00:12:45.370
as the outermost harness around CodeX
00:12:45.629 --> 00:12:45.649
You're
00:12:48.715 --> 00:13:05.434
So yeah yeah. So when I looked at the work around CodeX and at these observations from ScreenPipe, I realized I was acting as a retrieval layer or like the context bridge where I was moving context from one tool to another tool.
00:13:06.115 --> 00:13:19.254
And when these kind of activities recur ScreenPipe can signal what capabilities is helpful to implement that the harness does not have yet.
00:13:20.274 --> 00:13:31.523
Yeah this is this is like a good way to surface like the drudgery of white-collar work.
00:13:28.302 --> 00:13:58.397
That's like a phrase that I'm hearing nowadays. Like anything that's repetitive and boring that we're doing in a in a way that we don't even realize like in this case switching between a couple different programs and copy pasting stuff between each other these kinds of insights that we get from a tool like ScreenPipe helps us notice them and so that we can improve our workflow and get rid of these repetitive tasks and just focus on the stuff that matters, which is really nice
00:13:59.501 --> 00:14:18.081
Yeah. And in this example the first the first thing that I would solve is the retrieval. So give Codex reliable read-only access to other tools where I was manually gathering context but I did notice that in ScreenPipe the recording was still very sparse and event driven.
00:14:18.782 --> 00:14:39.394
So it was useful to that ScreenPipe flagged where I was gluing different contexts from different windows together but it was not measuring like every action. So it's still a little bit felt like there's something missing.
00:14:39.955 --> 00:14:55.615
And I hypothesize that it's probably more useful to reason use ScreenPipe to reason over days of usage and finding patterns there and searching for it because it's because the nature of it's more sparse and event-driven recording.
00:14:56.134 --> 00:15:32.366
what I want is insights in the moment. So when I'm focused on some work, when I sit down and I say, OK, I wanna I wanna work for ninety minutes and I wanna do something. I wanna achieve a goal. I call this like a focused work session I wanna record what I do during that session and get detailed insights When I when I have these focused work sessions, I want to be able to record exactly what I'm doing and understand and I want an agent to be able to understand what's happening with like a play-by-play.
00:15:33.167 --> 00:16:00.625
I want it to know what I'm doing at every moment and from that information give me give me feedback about what I can improve or what I can do better or when I was distracted these kinds of things and and so I think I don't know when this was like late last year around winter time Gemini added support for natively ingesting video tokens.
00:16:01.462 --> 00:16:03.962
OK? Which is the coolest thing ever. So
00:16:05.003 --> 00:16:06.822
what does that mean, natively ingesting
00:16:06.822 --> 00:16:07.023
Yeah.
00:16:07.923 --> 00:16:09.523
video tokens? Is that more efficient?
00:16:09.523 --> 00:16:09.763
Yeah.
00:16:09.822 --> 00:16:11.202
Like what what's the benefit?
00:16:12.552 --> 00:16:16.552
because we should talk about how LLMs understand things.
00:16:17.851 --> 00:16:32.351
Right? So they they take as input and output tokens, not bytes and tokens are, well, they're usually like a few bytes long.
00:16:33.048 --> 00:16:47.067
they cannot natively tokenize video so in order to actually understand what's happening in a video, they have to take that video and break it apart into a series of text and image tokens,
00:16:48.100 --> 00:17:13.880
And that are you saying that step could be lossy because you you rely on the agents to capture all the relevant context? But maybe, in a corner of the screen recording there's something else happening that the agent doesn't capture if it converts the format from video file to a few sentences in text format.
00:17:14.688 --> 00:17:50.832
If you have a specific frame of a video that's important, then it is more precise and it's better for your agent to just grab that frame and put that frame into into its context and process it but if you want the agent to reason about the full video, and yeah it might miss like little details here and there, but but it'll have much more rich information per token than than breaking down the video into frames and looking at specific frames one at a time if it could if it if it can natively consume video, right
00:17:52.291 --> 00:17:52.571
That's
00:17:52.571 --> 00:17:52.612
and
00:17:52.612 --> 00:17:53.372
so cool.
00:17:53.833 --> 00:17:54.212
Mm-hmm.
00:17:54.212 --> 00:17:58.913
in that context what kind of video do we want to send Gemini? It's a screen recording.
00:17:59.940 --> 00:18:17.291
Because I'm using my computer to do work and so I want to record what I'm doing, everything that I'm doing, not just in the agent, not just the chat with the agent but everything, and all that rich information of a focus session we can give to Gemini.
00:18:15.811 --> 00:18:45.561
OK but if you if you just naively like go into Quicktime and do a screen recording and try and send the to Gemini. It's not gonna work because you can't you can't send huge videos like too much data the the LLMs like like large video files will take up more tokens in the context window and we only have we have a million tokens with Gemini. So but that
00:18:45.561 --> 00:18:45.682
Mm-hmm.
00:18:45.682 --> 00:18:47.261
that gets filled up fast with video data.
00:18:48.082 --> 00:19:01.981
So so let me tell you the precise workflow that I use to screen record and get information about my workflows and I'll give an example of the information from a workflow from today.
00:19:02.721 --> 00:19:03.221
How does that sound?
00:19:03.301 --> 00:19:03.642
Yeah.
00:19:04.200 --> 00:19:04.279
OK.
00:19:04.279 --> 00:19:05.180
Yeah, I wanna see it.
00:19:06.500 --> 00:19:26.356
OK. So so the first thing that you do, number one you open Quicktime and you do screen record. That's not really an easy way for me to show this because it's like in the menu of your of your Mac but just open open Quicktime, new screen recording, start, record your whole screen.
00:19:27.156 --> 00:19:27.356
OK?
00:19:28.237 --> 00:19:35.323
So you do that you do that at the beginning and then you do what you wanna do until you're done.
00:19:33.863 --> 00:19:35.663
Maybe you've worked for ninety minutes.
00:19:36.487 --> 00:19:38.346
OK and then you stop
00:19:38.346 --> 00:19:38.487
Is
00:19:38.507 --> 00:19:38.508
the
00:19:38.508 --> 00:19:38.626
that
00:19:38.626 --> 00:19:38.646
screen
00:19:38.646 --> 00:19:38.807
what you
00:19:38.807 --> 00:19:38.866
recording.
00:19:38.866 --> 00:19:39.787
did as well?
00:19:40.416 --> 00:19:42.076
That's what I do.
00:19:40.416 --> 00:19:42.076
That's what I yeah. this is
00:19:42.076 --> 00:19:42.156
OK.
00:19:42.156 --> 00:19:43.656
this is my workflow that I'm describing.
00:19:44.076 --> 00:19:44.176
OK.
00:19:44.176 --> 00:19:44.356
OK?
00:19:44.656 --> 00:19:44.717
All right.
00:19:44.717 --> 00:19:54.497
So I start a screen recording when I'm done with my with my focus session I'll stop the screen recording.
00:19:52.797 --> 00:19:56.916
There's a button on the top in the menu bar that's like a stop icon. You just click on that.
00:19:57.676 --> 00:20:11.757
And that creates a large video file on your computer and you can save it so then what I do is I process that file to make it lower resolution and lower the frame rate.
00:20:12.488 --> 00:20:28.708
So the idea is we wanna get as much information as we can we wanna get as much information as we can into Gemini and it doesn't need to see a hundred twenty frames per second of me moving my mouse around. That's not really high high quality information.
00:20:29.428 --> 00:20:32.627
I find that one frame per second's enough.
00:20:33.607 --> 00:20:40.259
So that's just once a second a frame. Now over ninety minutes that's still a lot of information.
00:20:38.559 --> 00:20:40.259
Right
00:20:40.960 --> 00:20:41.299
Mm-hmm.
00:20:41.299 --> 00:20:58.599
and then I resize down to seven twenty P. I found that's a good that's a good size where enough details of the screen are visible so like the text is still visible it's also smaller so
00:20:58.599 --> 00:20:58.700
Yeah,
00:20:58.700 --> 00:20:58.799
it
00:20:58.819 --> 00:20:58.820
and
00:20:58.820 --> 00:20:58.920
saves
00:20:58.920 --> 00:20:58.940
you're
00:20:58.940 --> 00:20:59.059
saves
00:20:59.059 --> 00:20:59.440
working on
00:20:59.440 --> 00:20:59.539
tokens.
00:20:59.539 --> 00:21:01.779
a on a Macbook, right? A Macbook Pro?
00:21:02.480 --> 00:21:02.740
I'm on a
00:21:02.740 --> 00:21:02.759
Not
00:21:02.759 --> 00:21:02.880
MacBook
00:21:02.880 --> 00:21:02.960
a huge
00:21:02.960 --> 00:21:03.460
Pro, yeah. So
00:21:03.460 --> 00:21:03.519
screen
00:21:03.519 --> 00:21:03.960
I natively,
00:21:03.960 --> 00:21:05.000
monitor screen.
00:21:05.000 --> 00:21:05.240
yeah,
00:21:05.599 --> 00:21:05.799
Mm-hmm.
00:21:06.039 --> 00:21:12.680
I natively record at whatever, like four K basically. So
00:21:12.900 --> 00:21:13.059
Mm-hmm.
00:21:13.059 --> 00:21:26.680
so I downsize that to seven twenty P and you can do that with FFM PEG. You just ask your agent to use FFM PEG to resize to seven twenty P and and and re-encode at one frame per second.
00:21:27.400 --> 00:21:37.579
So that then you have a small video file, a smaller video file it's still too big. If it's ninety minutes long it's still too big. So the next thing that I do is I chop it up into fifteen minute slices.
00:21:38.497 --> 00:21:39.896
A fifteen minute slice
00:21:39.896 --> 00:21:40.257
Mm-hmm.
00:21:40.257 --> 00:21:46.477
at seven twenty P one frame per second is small enough that that you can just send it to Gemini.
00:21:47.436 --> 00:21:47.557
And so
00:21:47.557 --> 00:21:48.057
Mm-hmm.
00:21:48.436 --> 00:21:52.208
I'm gonna show an example of a session I did today.
00:21:52.928 --> 00:22:03.067
This is something that I that I use a lot I actually should use it more, but I've done it maybe ten or twenty times and this is just one from today.
00:22:04.228 --> 00:22:20.768
So I well, you'll see what I did because the LLM will tell you. So here's so what I'm sharing right now is the Gemini web UI. So I attached my video file.
00:22:21.887 --> 00:22:23.968
In this case it was a ten minute ten minute recording.
00:22:24.627 --> 00:22:31.188
And I said, this is a screen recording, help me reflect on my workflow, and then I sent a bunch of other questions to ask.
00:22:31.988 --> 00:22:41.228
was I how was I distracted, what was I doing, what learnings could you pull out? Then I said, give me a play-by-play of the beats of what I did during this session and answer my questions.
00:22:42.167 --> 00:22:44.268
And you get it.
00:22:44.928 --> 00:23:25.948
Check it out so first I reviewed the content outline for this episode I looked at some other coding sessions I did in ChatGPT and in Cloud I fixed a bug in the podcast video editor I asked Cloud Code to research mining session history so that we could so that I could prepare for this episode and I did a bunch of other stuff and from that Gemini was able to determine that my objective was to dogfood this particular content piece on giving agents eyes that we're talking about right now.
00:23:26.887 --> 00:23:35.127
And during this session I was actually setting up CASS and starting to play with it.
00:23:35.847 --> 00:24:03.387
And Gemini noticed that I was pretty focused and that I was not setting up CAS efficiently because I was struggling with the CLI and what I should have done is sent there there's a command built into the tool you can just feed into CLOJ to understand it. So it gave me a constructive suggestion for how to improve my workflow.
00:24:06.261 --> 00:24:13.720
I I try with my focused work sessions to be focused, to have a goal, and I
00:24:13.720 --> 00:24:13.799
Mm-hmm.
00:24:13.799 --> 00:24:43.997
wanna know the the times that I got distracted, what I was distracted by, and how I could improve my workflow going forward if I when I get distracted in that way. And and that that's the kind of thing of like switching off of my terminal, going into the browser. And if I'm researching something related to the work I'm doing, that's not a distraction. But if I'm going on Hacker News that's probably a distraction. And the LLM is able to tell me these things. So
00:24:45.196 --> 00:24:46.977
Did you see that in the past? Like did
00:24:47.416 --> 00:24:47.537
Yes
00:24:47.896 --> 00:24:54.257
did the did Gemini tell you like hey you're you're actually distracted when you're working.
00:24:54.257 --> 00:25:14.196
Yep, exactly. And and it surfaced the kind of the kind of workflow improvements that that you saw with ScreenPipe where it's saying, oh, instead of jumping between all these different products, maybe you could find a way to, get all the information in one place. That's like a common mistake that people make
00:25:15.797 --> 00:25:15.896
Like
00:25:15.896 --> 00:25:16.037
but I
00:25:16.317 --> 00:25:16.376
right.
00:25:16.376 --> 00:25:17.916
think Yeah.
00:25:18.037 --> 00:25:21.656
And did you change anything after, that observation from Gemini?
00:25:23.777 --> 00:25:36.297
Yeah, I've been steadily improving I think like my focused work sessions are more focused and more productive because I'm doing these feedback loops.
00:25:37.241 --> 00:25:50.221
I was more self-aware about, one I was self-aware to not be distracted, cuz also I you can feel the kind of like big brother watching you and you'll get in trouble if you if you do it wrong.
00:25:50.241 --> 00:26:10.676
That's sort of one thing. But yeah just doing more let me think I think there like a lot of basic stuff just it a lot of things boil down to stop doing things manually and ask your agent to do it for you.
00:26:10.676 --> 00:26:10.916
Mm-hmm.
00:26:11.957 --> 00:26:12.037
Right?
00:26:12.037 --> 00:26:12.336
What are the
00:26:12.336 --> 00:26:12.457
So
00:26:12.457 --> 00:26:12.896
top things
00:26:12.896 --> 00:26:13.116
don't
00:26:13.317 --> 00:26:18.876
that you've changed in the way you work by using the insights that you got from Gemini
00:26:19.660 --> 00:26:20.740
I'm distracted less often.
00:26:21.759 --> 00:26:32.910
I instead of copy pasting stuff between lots of different tools, I'm using Markdown files within with my coding agent.
00:26:33.501 --> 00:26:44.442
Instead of googling for stuff, I'm asking my coding agent to search. I'm using skills like last thirty days which is a nice skill for
00:26:44.842 --> 00:26:45.481
Seven six eight
00:26:45.582 --> 00:26:45.902
searching
00:26:45.902 --> 00:26:45.922
nine
00:26:46.241 --> 00:26:47.461
for interesting things
00:26:47.461 --> 00:26:47.481
zero
00:26:47.842 --> 00:26:49.261
across Reddit and X and YouTube
00:26:49.261 --> 00:26:49.561
two one eight
00:26:49.902 --> 00:26:50.402
and things like that
00:26:50.701 --> 00:26:50.721
one
00:26:51.541 --> 00:26:53.772
basically I'm getting more
00:26:53.772 --> 00:26:53.833
seven
00:26:53.833 --> 00:26:53.843
and
00:26:53.843 --> 00:26:53.873
zero
00:26:53.913 --> 00:26:59.313
more of the stuff that's happening outside of my agent into my agent. I think that's like a big category of stuff.
00:27:00.583 --> 00:27:01.403
And this is based on
00:27:01.403 --> 00:27:01.583
And
00:27:01.643 --> 00:27:05.363
the insights and feedback you got from Gemini after it analyzed
00:27:05.462 --> 00:27:05.663
Yeah.
00:27:05.663 --> 00:27:06.242
how you work.
00:27:06.482 --> 00:27:15.542
Yeah, it's it's something that it's obvious in hindsight. All these things are obvious in hindsight, even though any of these things, like once we know about them they're
00:27:15.542 --> 00:27:15.702
Yeah.
00:27:15.702 --> 00:27:30.542
obvious, but I think them being reflected back to you by someone who's by this like you know magical entity that's observing that helps you that helps you be aware of these things and then encourages you to take action to improve them
00:27:32.339 --> 00:27:42.940
I guess the magic is that it can because it observes more things, it now has the ability to like take action or help you take action.
00:27:43.720 --> 00:27:50.599
And yeah, I can totally imagine that even I'm like sometimes when I'm building I'm just an automatic pilot and I don't really stop
00:27:50.640 --> 00:27:50.779
Mm-hmm.
00:27:50.779 --> 00:27:58.619
to think and improve the workflow and it's helpful to give the AI eyes to help observe it
00:27:58.680 --> 00:27:58.839
Mm-hmm.
00:27:58.839 --> 00:27:59.920
for me as well.
00:28:00.768 --> 00:28:00.887
And you
00:28:00.907 --> 00:28:00.909
No,
00:28:00.909 --> 00:28:00.948
know,
00:28:00.948 --> 00:28:01.028
that
00:28:01.107 --> 00:28:01.268
I've
00:28:01.268 --> 00:28:01.548
totally
00:28:01.607 --> 00:28:01.667
I've
00:28:01.667 --> 00:28:01.847
makes
00:28:01.847 --> 00:28:01.848
been
00:28:01.848 --> 00:28:02.607
sense. Mm-hmm.
00:28:03.428 --> 00:28:22.087
I've been I've been capturing the outputs from Gemini and like putting them in a folder and then I can also ask questions across different focus sessions it's a little lossy cuz it's not reasoning over the exact video at that point but I'm like distilling the video down to the information I get from Gemini
00:28:22.087 --> 00:28:22.188
Mm-hmm.
00:28:22.188 --> 00:28:35.208
and then reasoning over that and I've even I started building a tool to do this automatically for me but I got distracted by my wedding or something so maybe I should pick that up again.
00:28:36.688 --> 00:28:51.748
Yeah, that would be interesting I would be curious and if someone wants to try this workflow or this the way how you use Gemini, would you recommend that they do it in the same way as you, like record the whole focused work session?
00:28:52.347 --> 00:28:58.228
Or should they just record their most repeated part of the workflow what do you think
00:28:59.017 --> 00:29:31.325
Yeah, I think a way to think about this is like you get you can get information in fifteen minute chunks. So I think it's still useful to to send like an hour or two hours worth of information in fifteen minute chunks But yeah, if you if you have a particular workflow that you do that you're looking for feedback on or that you wanna try and automate but you don't know how then you can, just record that particular part.
00:29:32.457 --> 00:29:32.916
Yeah. Because
00:29:32.916 --> 00:29:33.297
And.
00:29:33.336 --> 00:29:36.297
I could Yeah. No? Sorry. Go for it.
00:29:37.176 --> 00:29:38.609
Yeah, Go ahead
00:29:38.650 --> 00:29:39.210
Oh, OK.
00:29:39.970 --> 00:30:23.622
No, I was I just wanted to say that I could imagine if you have a repeated representative workflow, part of the workflow that you do all the time it could you could be curious like how you can improve it because it could be quite impactful to the to to how you work and it also sounds like you cannot it's probably too much context if you want to feed Gemini like a week worth of video recordings and tell it to surface the recurring patterns. So basically to shrink that context you give to Codex you can do it intentionally by just identifying, OK, this is what I do all the time, how do I prove this?
00:30:24.362 --> 00:30:24.721
Mm-hmm.
00:30:25.501 --> 00:30:37.701
Yeah. Or what I do, which is like distill videos down using Gemini and then take those distillations of all those moments and put them in a folder and then asking Codex or Claude to analyze it.
00:30:39.664 --> 00:30:40.184
Does that make sense?
00:30:40.384 --> 00:30:46.424
Wait, how do you distill it? Like I you mentioned how you cut it into like fifteen minute parts.
00:30:46.724 --> 00:30:46.904
Mm-hmm
00:30:47.065 --> 00:30:48.125
Do you distill it further?
00:30:50.196 --> 00:30:57.156
By distill I just mean copy and pasting the output from Gemini into a folder.
00:30:58.497 --> 00:31:02.957
Right. And then it's a log of your patterns and recurring behavior.
00:31:03.702 --> 00:31:25.994
Yeah well specifically I will ask Gemini for very detailed play-by-play of every moment of this video and I store that. So that's a cuz you can imagine like a fifteen minute video is you know ten megabytes or so, fifteen megabytes of information that is illegible to cloud and codex
00:31:26.174 --> 00:31:26.575
Mm-hmm.
00:31:27.575 --> 00:31:47.640
and if I give that to Gemini and I ask for that like play-by-play of what I did in every moment, as I showed before, you get a nice a few pages of text that describes what was happening in detail and then I can store that and that's whatever a thousand tokens, two thousand tokens worth of information so
00:31:47.640 --> 00:31:47.700
Gotcha.
00:31:47.700 --> 00:31:54.208
I can store weeks worth of those in one agent's context to you
00:31:54.327 --> 00:31:55.067
Right. And you say
00:31:55.067 --> 00:31:55.167
know
00:31:55.167 --> 00:31:55.347
like
00:31:55.688 --> 00:31:55.788
understand
00:31:55.788 --> 00:31:55.968
illegible
00:31:55.968 --> 00:31:56.928
patterns and things
00:31:57.781 --> 00:32:16.382
And we when you say illegible to Codex and Claude, you're talking about things that happen outside of the chat, right? So maybe working in a browser, going to forums other patterns that turn that that show that you're making a detour.
00:32:17.367 --> 00:32:57.057
Yeah. specifically this screen recording video is unreadable directly by these tools. Yes, they can if you if you tell them to understand it, what it what Cloud Encode X will do is they'll they'll break down the videos into frames and then try and read the specific frames to get a sense of what's going on in the video, but it's more lossy that one way to put it is like the distillation if you if you ask Claude, given this fifteen minute video, tell me what's happening beat by beat and you ask Gemini the same thing, the answer from Gemini is much more detailed because Gemini can natively understand the video.
00:32:57.615 --> 00:32:57.974
Yeah.
00:32:59.220 --> 00:33:24.500
So, it's really interesting because in this all the example that you mentioned, you changed the object of reflection to something different. Like, previously the object was the transcript or the session summary. And now the object of reflection became the whole recording and we are looking at how people behaved during the entire act of building overall and not just in that chat with the coding agents.
00:33:24.980 --> 00:33:25.559
That's pretty cool
00:33:26.559 --> 00:33:39.059
Yeah, it's so cool. And we can use that information to automate the meaningless, repetitive drudgery away and just leave people with the most interesting parts of their work.
00:33:39.815 --> 00:33:55.634
And I think that is so important because why would you wanna be doing the thing that is repetitive and basic? Then you're not growing, you're not learning you're not you're not achieving everything that you can achieve, So
00:33:56.835 --> 00:33:56.894
Mm-hmm.
00:33:56.894 --> 00:33:59.934
so we need to automate more. So let's talk more about
00:33:59.934 --> 00:34:00.055
Yeah.
00:34:00.055 --> 00:34:00.234
that
00:34:01.214 --> 00:34:15.554
I do I do want to offload more of the less meaningful task to the agents. So I do a version of this too So I'm building this app that will run on my phone.
00:34:16.255 --> 00:34:21.375
And in that workflow I connect phones, my two old phones, as test rigs.
00:34:22.054 --> 00:34:39.934
And then my coding agents can control those phones, it can run the features of the app that it just built, and then observe what happens on the real device. And both phones they have different like specs so you can like see how it performs and whether there are any errors.
00:34:36.875 --> 00:34:53.235
And that way you give your agents not only eyes to look outside of their chat but also you can give them hands in that in that environment where that app will run and act on reflect and act on what they see
00:34:53.831 --> 00:34:59.731
Yeah, I think that's I think that's so that's so that's so powerful.
00:34:57.431 --> 00:35:07.972
Like I think a lot of people when they're working on apps they have a good feedback loop set up with an emulator or simulator like an Android emulator or an iOS simulator.
00:35:08.652 --> 00:35:20.818
But but they but maybe they don't realize you can just plug in a phone and like you said then that allows the computer to control it and if the computer can control it then your agent can control it.
00:35:21.204 --> 00:35:26.885
And actually that reminds me of this example from Peter Steinberger, the creator of OpenClaw.
00:35:28.065 --> 00:35:49.465
Here he gave his agent a webcam access so he could end-to-end test some voice wake command on an ESP thirty-two open-claw node. So the agent would repeatedly test the thing by saying, hi ESP, and then it would watch what happens on a physical device and continue debugging it
00:35:50.012 --> 00:35:50.873
So cool.
00:35:51.032 --> 00:36:03.152
so it, yeah. In the same way, it was giving the agent eyes outside of the chat the chat and it gave the agent hands to act on it.
00:36:05.737 --> 00:36:19.777
anytime you're building something in the real world you can connect a camera, have the camera point at that thing and whatever you're building, if you can control it by, connecting to something then
00:36:21.257 --> 00:36:21.858
Exactly.
00:36:21.858 --> 00:36:22.518
Oh, look, you have another example.
00:36:22.518 --> 00:36:25.938
And this is, yeah, I have another example, like here you here
00:36:26.177 --> 00:36:26.208
Well
00:36:26.498 --> 00:36:27.418
you see another
00:36:27.418 --> 00:36:28.157
I'd put
00:36:28.898 --> 00:36:28.938
someone
00:36:28.938 --> 00:36:28.958
that
00:36:29.117 --> 00:36:29.657
else who
00:36:29.657 --> 00:36:29.677
I'm
00:36:29.717 --> 00:36:30.057
have mounted
00:36:30.057 --> 00:36:30.277
I'm just I'm
00:36:30.498 --> 00:36:30.577
a
00:36:30.577 --> 00:36:30.588
visiting
00:36:30.588 --> 00:36:43.057
desk camera to an agent so it could inspect these external devices like Remarkable Tablet and also ESP thirty-two without human taking photos all the time.
00:36:40.757 --> 00:36:43.057
And actually I do this often
00:36:43.057 --> 00:36:43.077
your
00:36:43.438 --> 00:36:45.277
too sometimes when I see things that happen
00:36:45.277 --> 00:36:45.297
site
00:36:45.478 --> 00:36:52.677
outside of the chat window or even outside of my computer. I just keep taking pictures. But even that you could automate away
00:36:52.757 --> 00:36:52.777
to
00:36:54.018 --> 00:36:54.157
like
00:36:54.378 --> 00:36:54.398
put
00:36:54.538 --> 00:37:05.677
you said previously don't spend your human time doing these meaningful not meaningful steps but give your agent the eyes to do it itself
00:37:06.715 --> 00:37:20.103
And I bet, a lot of these workflows you can get something that works well enough by piping the webcam footage into something that like Clutter Codex controls and it can take screenshots and or look at different frames.
00:37:20.563 --> 00:37:35.483
But I think what's really interesting that not enough people are doing is having an intermediate step of sending that video data to Gemini and getting the insights out of Gemini and then feeding that back into your coding agent of choice. I think that's
00:37:35.882 --> 00:37:36.402
What is the difference
00:37:36.402 --> 00:37:36.702
the.
00:37:37.342 --> 00:37:39.643
of skipping that step and including that step?
00:37:40.603 --> 00:37:58.222
If you skip that step then you're you're losing more information from the video data before it goes into cloud or codecs, because cloud and codecs can't natively understand the video data. They have to be, they have to break it apart into frames and they have to sample those frames in some way.
00:37:59.143 --> 00:38:06.262
But Gemini can take a continuous video stream, if it's low enough quality, low enough frame rate,
00:38:06.483 --> 00:38:06.902
Mm-hmm.
00:38:07.063 --> 00:38:10.202
even long, streams of video and natively
00:38:10.202 --> 00:38:10.643
Mm-hmm.
00:38:10.643 --> 00:38:14.722
process them, natively break them into tokens and understand them. And so
00:38:14.882 --> 00:38:15.702
Like in these examples
00:38:15.702 --> 00:38:16.182
and so you get.
00:38:16.182 --> 00:38:17.483
where you have a webcam
00:38:18.483 --> 00:38:18.762
Mm-hmm.
00:38:18.902 --> 00:38:29.702
recording something outside of what's happening on your computer, what do you have an example what Gemini would catch and other tools would miss
00:38:30.302 --> 00:38:36.663
I'd have to try it to like know for sure but but I have I have I have tried this before
00:38:36.663 --> 00:38:36.842
I don't
00:38:36.842 --> 00:38:37.202
on
00:38:37.202 --> 00:38:37.463
know.
00:38:39.003 --> 00:38:40.663
on screen recordings.
00:38:39.003 --> 00:39:08.983
And the the information that Cloud and Codex get out of the video is just less it's less detailed. And so the the same thing would happen for any video file. So any video of the real world as well. You'd get you'd get less precise information or the information that you get would use more of the context of Cloud and Codex and there would be less context available to solve whatever problem came up in that video, if that makes sense.
00:39:09.222 --> 00:39:09.463
Mm-hmm.
00:39:10.242 --> 00:39:20.483
OK. Yeah. That helps giving an idea of the difference the what you would get out of it of adding in, injecting that step of feeding it to Gemini.
00:39:21.565 --> 00:39:22.164
Yeah. So it
00:39:22.164 --> 00:39:22.204
Cool
00:39:22.204 --> 00:39:36.425
looks like, screen recordings, cameras and other device outputs like my like the the phones, the old phones, the test rig, they can all, there're all ways to give your agent eyes outside of the chat.
00:39:37.244 --> 00:40:07.751
And I guess in a in a case of the test rig, if you make it controllable, you can even give your agent hands to act on it. And I guess having more eyes, like you said and even like the way of translating what it sees to what it interprets can reveal more patterns and problems that maybe never occur to you when working on it or never appear in an agent's chat so
00:40:07.751 --> 00:40:08.211
Yeah, and
00:40:09.692 --> 00:40:09.911
Mm-hmm.
00:40:10.032 --> 00:40:40.722
this I guess like throughout this conversation we've mostly been talking about having moments of time where you as the operator, the human, is going and instructing your agents or doing some action to get this feedback I guess the first example of the agent acting on its own with its hands is the one that Christine was just saying where the agent could control the the phone manually or some of those examples that we saw on Twitter.
00:40:41.282 --> 00:40:47.001
But you can also have your agent proactively surface feedback as it's doing its work.
00:40:47.001 --> 00:40:47.402
Mm-hmm.
00:40:49.733 --> 00:40:54.072
Actually, Cloud just introduced a new feature a few days ago
00:40:55.413 --> 00:40:55.773
Mm-hmm.
00:40:56.722 --> 00:41:05.503
which I'll talk through. So Cloud Code can proactively draft feedback when it finds bugs within Cloud on its own.
00:41:06.050 --> 00:41:06.289
OK?
00:41:06.909 --> 00:41:07.230
Mm-hmm.
00:41:07.230 --> 00:41:07.250
So
00:41:07.489 --> 00:41:07.510
Mm.
00:41:08.157 --> 00:41:22.695
this is cool, and and actually this is something that Lovable's been doing for six months, I guess. So Lovable also does this where the Lovable agent
00:41:22.755 --> 00:41:22.804
Mm-hmm
00:41:23.655 --> 00:41:34.914
when a user is making a web app, it has a way to give feedback to the developers to Lovable, to the company that surfaces feedback when there's issues.
00:41:32.534 --> 00:41:34.914
Which is good because then
00:41:34.914 --> 00:41:35.335
Mm-hmm. Mm.
00:41:35.974 --> 00:41:40.135
proactively the team can triage them and fix those errors.
00:41:40.840 --> 00:41:52.940
There's nothing about this workflow that's specific to an application getting the feedback versus you as the user of the agent. And there's been there's been a lot of examples of these kinds of
00:41:53.059 --> 00:41:53.199
Mm-hmm
00:41:53.199 --> 00:41:57.840
of projects, but I'll just share one that came up earlier this month.
00:41:58.902 --> 00:41:59.342
Frog.
00:42:00.342 --> 00:42:11.552
Automated friction logging for agents your agents hit paper cuts all day and then they work around them and they forget. And this frog tool turns them into tract issues.
00:42:12.351 --> 00:42:47.722
But you don't you don't actually need a tool you can just tell your agent you can have this in your agents dot MD for example to record anytime it runs into friction anytime it wants to complain write that down to a file call it vent dot MD or complain dot MD or friction dot MD and and have your agent just write down things that it encounters and then you can have I guess at a later can go and process that or you can build automation to automatically triage that information as it comes in.
00:42:49.163 --> 00:42:59.382
So the lesson is like don't stop at just giving your agent more eyes, more ways to observe, have more observations, but also help it to act on it
00:43:00.603 --> 00:43:00.943
Give it the hands
00:43:00.943 --> 00:43:01.163
I and
00:43:01.702 --> 00:43:03.503
to observe by itself
00:43:04.472 --> 00:43:04.612
Yeah.
00:43:04.612 --> 00:43:05.132
in the moment.
00:43:06.501 --> 00:43:24.181
To observe and act on it, right? Act on what it observed. Like either it's like building these bug reports, logging them turn them into skills like what we saw earlier but there's probably, you could also give your agent computer use
00:43:25.523 --> 00:43:25.922
Mm-hmm.
00:43:26.043 --> 00:43:43.284
access to computer use because then there's also if you observe if you give the agent eyes to observe things outside of the chat history if you give it I could imagine if you give it computer use also do things outside of the chat history it will make it much more efficient and impactful
00:43:44.610 --> 00:44:07.177
Yeah, computer use is expensive and slow, but it's useful if it automates something that you struggle in some other way to automate so it's definitely it's a good tool and it's better than doing it by yourself manually, but it's not as good as being able to automate that task with a CLI tool, for example.
00:44:07.847 --> 00:44:08.146
Mm-hmm.
00:44:08.867 --> 00:44:13.646
Only when there's no CLI tool then computer use is the last resort.
00:44:14.827 --> 00:44:14.947
Yeah
00:44:16.967 --> 00:44:20.987
Yeah. So many ways to give your agent more eyes and also more hands.
00:44:22.266 --> 00:44:37.302
Alright So the other day when I opened Clod, this is what I saw it. this is what I saw it said, skip the repeat work, record your screen while you do a task once, and Clod will turn the recording into a skill it can run again.
00:44:37.885 --> 00:44:38.724
So again here we
00:44:39.164 --> 00:44:39.324
Yeah.
00:44:39.324 --> 00:44:39.485
we see
00:44:39.485 --> 00:44:39.545
Yeah,
00:44:39.625 --> 00:44:42.744
this trend where we are giving the agents more eyes
00:44:42.824 --> 00:44:42.844
Yeah,
00:44:43.425 --> 00:44:46.965
outside of just the transcripts or a
00:44:46.965 --> 00:44:46.985
Yeah,
00:44:46.985 --> 00:44:47.014
summary.
00:44:47.014 --> 00:44:47.045
yeah.
00:44:48.005 --> 00:44:49.605
And the agent
00:44:49.605 --> 00:44:49.625
Yeah.
00:44:49.704 --> 00:45:01.605
is is acting on it after getting these observations. So here it turns whatever it observes in the in the screen recording.
00:45:01.784 --> 00:45:05.224
So maybe it's like a repeating behavior or
00:45:05.324 --> 00:45:05.545
That's the
00:45:05.545 --> 00:45:05.565
recurring
00:45:05.565 --> 00:45:05.585
same
00:45:05.684 --> 00:45:06.164
errors
00:45:06.184 --> 00:45:06.186
thing.
00:45:06.186 --> 00:45:06.264
and
00:45:06.264 --> 00:45:06.284
That's
00:45:06.324 --> 00:45:06.764
then it turns
00:45:06.764 --> 00:45:06.784
the
00:45:06.824 --> 00:45:09.005
that into scale to improve the workflow.
00:45:09.813 --> 00:45:15.893
So they this is basically a progression of field of view and an ability to act.
00:45:16.862 --> 00:45:19.643
Where would you stop adding visibility or control though
00:45:21.018 --> 00:45:46.867
Like everything you can take it too far if you add too much recording, too many controls, too many tools, then the agents get confused there's too much context to wade through and there's too many choices to make about what action to take and then our agents become useless. So it's this is one of those things that you wanna add incrementally and you wanna be precise when you're doing a particular task like we've talked about before.
00:45:47.547 --> 00:45:55.306
Only expose the tools that you need and only record the things that you need to complete the task that you have at hand.
00:45:55.882 --> 00:46:01.083
yeah. I guess the takeaway is the coding agent conversation is only one view of the work.
00:46:01.822 --> 00:46:33.123
Give your agent more eyes beyond the chat through session history, screen recordings, cameras or other device output and this broader view that the agent gets can reveal patterns and problems that you did not know to look for. And then, we're safe, give them hands so they can actually act and reflect on the new findings, like through controllable test rigs or computer use or just places to complain so they can act on what they see, test results, and close more of the feedback loop themselves.
00:46:34.242 --> 00:46:41.891
And if you have video data, give that video data to a model that can actually understand the video data, like Gemini.
00:46:41.891 --> 00:46:42.291
Yes.
00:46:43.246 --> 00:46:45.867
Alright, that's it for this episode of Superlinear.
00:46:46.367 --> 00:47:05.927
If you try this, we are curious what did you discover that you can automate and what would you want your agents to see and act on. Tell us in the comments and if you want more real examples of how we are improving our agents' workflows, subscribe so you can catch the next episode.
00:47:06.726 --> 00:47:07.427
See you next time.
00:47:07.987 --> 00:47:08.527
Bye everyone.