O TYM ODCINKU
AI agents don't just write insecure code — they can escape their sandboxes, delete files, and do whatever it takes to complete a task. The security mental model that served us through the cloud era isn't enough anymore. Guy Podjarny, founder of Snyk and CEO of Tessl, made the case at London's AI Security Summit: it's time to stop securing the code and start securing the coder.
Recorded live at the AI Security Summit in London, this episode features conversations with Brian Vermeer (Snyk), Sam Stepanyan (OWASP London), and a full recording of Guy's keynote on why agentic development demands a fundamentally different approach to security.
What we cover:
- Why shadow AI is the new shadow IT — and why CISOs can't secure what they can't see
- Skills as a new supply chain attack surface (malicious, vulnerable, and negligent skills)
- Why more context is not always better — and what the data says about focused skill design
- The OWASP Top Ten for Agentic AI and what it means for teams building today
- Why security must become agentic to keep up with the attackers who already are
- The Context Development Lifecycle (CDLC) and how leading orgs are using it
Links: 🌐 Tessl: https://tessl.io
Subscribe for weekly episodes on AI-native development
What's the biggest security risk your team isn't talking about when it comes to agentic development? Drop it in the comments.
POKAŻ NOTATKI 🔗
TRANSKRYPCJA 🔗
00:00:00.000 --> 00:00:00.440
or with agents.
00:00:00.440 --> 00:00:05.280
There's a behavior called reward seeking which is you ask them to do something and they're really, really, really keen to do it.
00:00:05.280 --> 00:00:15.000
And so they go up and they try to do everything they can, and they escape their sandbox and they delete files and they do whatever it is to please If you cannot commit to the repository, fail.
00:00:15.000 --> 00:00:17.600
Don't like go off and sort of send it in another way.
00:00:17.600 --> 00:00:22.199
we need to move from securing the code to securing the coder, securing the the agent.
00:00:22.199 --> 00:00:31.000
the attack vector, but also the spectrum of what is there is hard to follow and hard to secure because we want to enable AI as a force multiplier.
00:00:31.000 --> 00:00:33.880
But in the meantime we also have to mitigate the risk.
00:00:33.880 --> 00:00:38.039
so we find ourselves in this kind of carrot and stick mode security.
00:00:38.039 --> 00:00:40.759
If it doesn't become a genetic, it will never keep up.
00:00:42.079 --> 00:00:48.600
The AI native dev is a podcast for developers and engineering leads at the cutting edge of AI and a genetic coding.
00:00:48.679 --> 00:00:51.679
Join your hosts, Guy pigeon and me, Simon Maple.
00:00:51.679 --> 00:00:58.439
Every week as we chat with the most exciting voices in AI and tackle the biggest questions facing developers today.
00:00:59.280 --> 00:01:02.280
This is the AI native dev.
00:01:06.560 --> 00:01:11.680
Back in November, we hosted the first ever in-person AI native dev con in New York.
00:01:11.799 --> 00:01:14.760
This June 1st and second, we're bringing it to London.
00:01:14.760 --> 00:01:31.480
It's two days built for aina to developers and engineering teams, one day full of hands on workshops and one day full of practical talks on agent skills, context engineering, agent orchestration and enablement platforms and how teams are actually shipping AI in production.
00:01:31.719 --> 00:01:39.959
Join us at the Brewery in London near the Barbican for all of that, plus networking parties, giveaways and a room full of people.
00:01:40.000 --> 00:01:42.959
Building the future of AI native development.
00:01:42.959 --> 00:01:46.280
You can also join us from anywhere in the world via the live stream.
00:01:46.439 --> 00:01:47.920
As you're listening to this podcast.
00:01:47.920 --> 00:01:51.879
You get 30% off your ticket with Code Pod 30.
00:01:51.959 --> 00:01:55.719
Just head to AI native Dev Königssee and we'll see you in London.
00:02:02.480 --> 00:02:07.400
Hey there, Simon Maple here at the AI Security summit here in London.
00:02:07.519 --> 00:02:12.719
This is a great conference to talk about all things security in AI development.
00:02:12.719 --> 00:02:18.360
There are going to be practitioners, and there are also going to be loads of CISOs and security leaders at this event.
00:02:18.400 --> 00:02:26.199
Now Tesla's here, we've got a booth and we also have Guy here, previous founder of sneak as well as the founder of tassel.
00:02:26.199 --> 00:02:31.879
And Guy is going to be giving a session talking about Security keeping up with a development.
00:02:31.879 --> 00:02:34.759
Let's secure not the code but the coder.
00:02:34.759 --> 00:02:42.680
Because how can we make developers think more securely when using a genetic methods and a genetic tooling?
00:02:42.759 --> 00:02:43.439
Cool.
00:02:43.439 --> 00:02:46.039
Let's see if we can talk to a couple of people while we're here.
00:02:49.639 --> 00:02:54.319
So now I'm here joined by Brian Vermeer, a staff developer advocate at sneak.
00:02:54.319 --> 00:02:56.199
And we go back a long, long way.
00:02:56.199 --> 00:03:00.240
Brian, you've been on the podcast before, and we're now here at your event.
00:03:00.280 --> 00:03:01.680
What's your event called?
00:03:01.680 --> 00:03:04.680
This event is called the AI Security Summit.
00:03:04.680 --> 00:03:05.719
And what is it about?
00:03:05.719 --> 00:03:07.639
Well, it's about AI and security.
00:03:07.639 --> 00:03:14.159
I think the name says it all, but with AI, there comes a lot of new security attack factors.
00:03:14.199 --> 00:03:15.840
And what do you need to think about?
00:03:15.840 --> 00:03:29.919
How do you need to transition from old fashioned to, say, software dependency analysis and software code analysis to now driving your code agent but also things like skills and MCC and that kind of stuff.
00:03:29.960 --> 00:03:33.520
So like shifting gears into that space of security.
00:03:33.560 --> 00:03:34.080
Amazing.
00:03:34.080 --> 00:03:40.520
And you actually gave the intro speech at the at the leadership summit that we just had across the road.
00:03:40.680 --> 00:03:42.560
also at the main event here.
00:03:42.560 --> 00:03:54.680
What would you say the biggest issues that that are on the CSO minds today when they're thinking about their organizations trying to adopt and roll out AI really as fast as they possibly can?
00:03:54.719 --> 00:03:56.879
What do people coming to you complaining about?
00:03:56.879 --> 00:03:59.280
I think it's a very good question.
00:03:56.879 --> 00:03:59.280
I think it's multiple things.
00:03:59.280 --> 00:04:03.639
First of all, people are not aware, like what kind of AI is already in their systems.
00:04:03.639 --> 00:04:09.120
They think they're not using any models or maybe even any NCP servers.
00:04:09.120 --> 00:04:15.120
But in some cases, even your agent can decide to pull in a certain model to get track of something else.
00:04:15.120 --> 00:04:22.120
So things like shadow AI or employees just using their own free ChatGPT account to do things.
00:04:22.120 --> 00:04:33.199
So the attack vector, but also the spectrum of what is there is hard to follow and hard to secure because we want to enable AI as a, as a, as a force multiplier.
00:04:33.199 --> 00:04:35.680
But in the meantime we also have to mitigate the risk.
00:04:35.680 --> 00:04:39.000
And these are things that CISOs are are challenged to buy.
00:04:39.040 --> 00:04:41.680
So it's almost like a problem of discovery.
00:04:41.680 --> 00:04:59.439
And this is really interesting because this kind of like reminds me, if you think back around, you know, I guess where sneak originated from, when we look at dependencies and trying to identify dependencies, what people are using in their organization, a thing was created called a software bill of materials, where we kind of like these are all the dependencies that we use.
00:04:59.439 --> 00:05:04.040
Do you feel like something is equally needed now from a from an AI point of view?
00:05:04.040 --> 00:05:05.160
Like what are the things I like?
00:05:05.160 --> 00:05:08.720
There's things like AI bombs whereby like where is the training data coming from?
00:05:08.720 --> 00:05:09.399
From my models.
00:05:09.399 --> 00:05:13.439
But what about things like context where you mentioned skills and MCC and things like that?
00:05:13.439 --> 00:05:16.439
Do we do we need like a context, building materials.
00:05:16.519 --> 00:05:20.040
Maybe even that the AI build materials is a thing already and we scan for that.
00:05:20.079 --> 00:05:22.439
We can we can create that for companies.
00:05:22.439 --> 00:05:27.600
But I also think like the context or the skills that you pull in, it's basically a dependency as well, right?
00:05:27.600 --> 00:05:38.240
If you look at how software was developed and you pull packages from and from from repository as your your third party entry points, now you have that with skills.
00:05:38.240 --> 00:05:42.079
And so the problem here now is that skills is just text.
00:05:42.120 --> 00:05:42.319
Yeah.
00:05:42.319 --> 00:05:46.240
And now we need to scan that to see if there are no injection in it.
00:05:46.279 --> 00:05:53.319
If if there are no hidden hidden comments in it or even worse, it has a hit a comment in it that is getting copy to your global memory.
00:05:53.319 --> 00:05:55.879
And then even if you delete the skill, then it's still there.
00:05:55.879 --> 00:05:58.720
So it's back to our basics.
00:05:58.720 --> 00:06:10.920
As in, make sure that whatever you ingest or use is validated and vetted and keep track of that, just like we did with our Docker containers and with our dependencies and the code that we wrote.
00:06:10.920 --> 00:06:21.079
So we need to be aware that that things can go south, even though most of the stuff that are agents nowadays are creating is quite cool.
00:06:21.079 --> 00:06:29.639
And of course, we recently announced the integration between Sneak and Tessa, whereby we expose a lot of that data that those scan results from sneak within the Tesla registry.
00:06:29.639 --> 00:06:40.240
So if you are using skill, you can you can identify if there are threats through the sneak as a judge style approach to identify where those potential issues issues are.
00:06:40.279 --> 00:06:50.879
What one piece of advice would you give people who are using AI in an organization today as a super quick way, a super quick way of really reducing risk within their organizations?
00:06:51.120 --> 00:06:53.959
Have a good overview of what are you using currently?
00:06:53.959 --> 00:06:56.759
Because if you don't know that data, you cannot protect yourself.
00:06:56.759 --> 00:07:03.120
And secondly, make people aware of these choices and of the of the attack vectors that are coming in.
00:07:03.120 --> 00:07:20.040
Because I think most people that are not tech savvy and are now using agents to build their own self aware, if you can call it like that, the their own tooling are not aware that if you connect this piece of data to that piece of MCP server, that we can leak data and these kind of things.
00:07:20.040 --> 00:07:22.600
So awareness is the first point. Amazing.
00:07:22.600 --> 00:07:28.040
For those of you who want to learn more about the integration, we actually did a podcast, a full port cast episode.
00:07:28.040 --> 00:07:31.319
So check out the Native Dev Podcast with Brian Vermeer.
00:07:31.360 --> 00:07:33.399
And we did that in Atlanta. Beautiful sunny Atlanta.
00:07:33.399 --> 00:07:35.680
Now we're in kind of cloudy London, but it's all good.
00:07:35.680 --> 00:07:38.160
And enjoy the rest of the conference.
00:07:35.680 --> 00:07:38.160
Brian, thank you very much.
00:07:43.240 --> 00:07:45.399
Look on the X blow for Sam.
00:07:45.399 --> 00:07:46.079
Stefan.
00:07:46.079 --> 00:07:50.639
Who is the the the head of the OST London chapter.
00:07:50.639 --> 00:07:54.279
And you're also on the global board for OST generally, right?
00:07:54.279 --> 00:07:54.839
That's right. Yes.
00:07:54.839 --> 00:07:55.519
Awesome. So.
00:07:55.519 --> 00:07:58.600
So what are you here to learn today at at this event.
00:07:58.639 --> 00:08:00.040
As always. Right.
00:08:00.040 --> 00:08:07.600
We learn every day and AI is moving so quickly and cybersecurity moving so quickly here to learn from all the periods of college.
00:08:07.600 --> 00:08:17.879
And of course for you guys, what is new in the world of AI security, what new solutions are and what new challenges are you currently facing?
00:08:17.879 --> 00:08:21.879
Because there is so much new stuff happening at the moment.
00:08:22.120 --> 00:08:37.919
See also conversations around mythos right allows for everyone, obviously, because I represent all was well what you see how our OWASp top ten for LMS and brand new always top ten for agenda key is impacting people.
00:08:37.919 --> 00:08:57.279
Because I had several conversations days ago at the AWS conference for our products, and there are lots of companies were just trying to get into the genetic AI world and have like zero clue about security issues, which can meet of that.
00:08:58.080 --> 00:09:05.960
You mentioned the top ten, and of course, for those who don't know, is obviously a security company that governs and does a whole bunch of nonprofit nonprofit.
00:09:05.960 --> 00:09:06.960
Sorry. Yeah.
00:09:06.960 --> 00:09:12.120
And provides a whole bunch of best practices and great advice for a ton of different spaces.
00:09:12.120 --> 00:09:18.519
So the top ten for web applications and security and web applications, mobile applications.
00:09:18.879 --> 00:09:20.159
Cloud security approach.
00:09:20.159 --> 00:09:23.600
And you mentioned the top ten now for for AI and Llms.
00:09:23.639 --> 00:09:24.480
What would you say.
00:09:24.480 --> 00:09:35.360
Kind of like the biggest mistakes that people make as as developers or as CISOs when thinking about building using tools and AI.
00:09:36.840 --> 00:09:39.759
The biggest mistake is they completely ignore security issues.
00:09:39.759 --> 00:09:48.320
They jump straight head in, and they connect AI directly to their production system without understanding the consequences of it.
00:09:48.320 --> 00:10:02.159
Yeah, and I think it's due to the huge pressure that everyone's feeling because everyone's trying to jump on the AI train and they're trying to, of course, use the technology and innovate.
00:10:02.200 --> 00:10:12.320
But the fact that a lot of them are ignorant about cybersecurity issues, which is about AI, makes it, of course, quite, quite worried.
00:10:12.360 --> 00:10:22.000
It kind of reminds me the same situation that we had at the e-commerce 25 years ago when everyone said, oh, this new thing called the web.
00:10:22.000 --> 00:10:26.879
We have to do. We have a website and we must have an ecommerce, so we must sell things online.
00:10:26.919 --> 00:10:28.960
The whole digital transformation, right?
00:10:28.960 --> 00:10:30.759
Let's start taking credit cards online.
00:10:30.759 --> 00:10:33.519
And no one thought about security and everyone started.
00:10:33.519 --> 00:10:34.759
Getting hacked.
00:10:34.759 --> 00:10:38.440
SQL injection was number one vulnerability back then and the same thing here.
00:10:38.440 --> 00:10:50.639
But I was saying, oh, let's all use AI and no one thinks about, you know, basic hygiene, things like prompt injection, but also, of course, lots of other security concerns surrounding AI.
00:10:50.679 --> 00:10:58.600
And we do provide OWASp free and open source guidelines and resources and standards for you.
00:10:58.679 --> 00:11:05.080
There was a one of the very important documents we have which I recommend to everyone is a secure AI adoption guidelines.
00:11:05.600 --> 00:11:09.480
So I highly encourage other organizations who are trying to adopt AI.
00:11:09.960 --> 00:11:16.200
Check out this document which is community created the community to it so they can understand how to adopt AI security.
00:11:16.200 --> 00:11:26.679
And you mentioned like the pressures behind businesses who are like not being forced necessarily, but very, very much encouraged to use AI pressure to use AI as fast as they can.
00:11:26.679 --> 00:11:30.840
Because realistically, businesses that don't in 1 to 2 years.
00:11:30.879 --> 00:11:36.240
They're potentially going to be at a very strong, very big disadvantage to those other companies that are moving so fast.
00:11:36.240 --> 00:11:53.600
So from the point of security securities, very often seen historically as a department that can slow down delivery, how how much is that kind of like seen today in the world of AI is that is is security seen as something that is is hampering the advancements of AI in organizations?
00:11:53.639 --> 00:11:55.360
Well, that depends how you approach it.
00:11:55.360 --> 00:11:58.440
And I would always say that no security is an enabler.
00:11:58.440 --> 00:12:03.960
You know, it's like the brakes in your car actually allow the cars to move faster, right?
00:12:03.960 --> 00:12:10.799
So the same thing with AI security world, you can innovate fast and secure manner.
00:12:10.879 --> 00:12:27.480
But all that you need to do, you need to be aware of all the security implications and make sure that your innovation security experiments, they start from their isolated environment, which is secure, right, and disconnected from your production data.
00:12:27.480 --> 00:12:30.639
So if things go wrong, it goes minimal damage.
00:12:30.679 --> 00:12:43.320
And obviously in the isolated environment you can actually learn how to secure it properly for all of us and also other industry guidelines.
00:12:43.320 --> 00:12:49.639
And one very important point, which I'd like you to make is also remember that AI is non-deterministic.
00:12:49.679 --> 00:12:53.080
So let's say one thing today was a little different than tomorrow.
00:12:53.320 --> 00:13:29.919
Another very important thing that not many people are mentioning today is the problem with a genetic identity is the issue that at the moment, the way how AI integrates with other things is that it acts upon humans behalf, which makes things like traceability and audit and logging of actions is very, very difficult because if you grant an AI agent accessing your email box or your behalf and it's going to go inside sending emails, sending out spam on your behalf, or it's going to go and start deleting and accessing or deleting customer records.
00:13:30.639 --> 00:13:32.519
It's not Simon, it's not Simon.
00:13:32.519 --> 00:13:35.240
AI at IO, it's Simon.
00:13:35.240 --> 00:13:36.639
Oh yeah, it wasn't me.
00:13:36.639 --> 00:13:38.720
It was my agent. Yeah, exactly, exactly.
00:13:38.720 --> 00:13:57.080
Now, I was chatting with Brian Vermeer earlier, and he was talking about one of the, one of the, one of the first things that people do should do is understand where organizations are using AI, how big a problem is, almost like the governance or the understanding of where people are using AI in their organizations today.
00:13:57.080 --> 00:14:00.159
And there being too much, almost like shadow AI.
00:14:00.759 --> 00:14:01.759
There's a lot of shadow.
00:14:01.759 --> 00:14:07.080
I believe it's a big problem in obvious Asians where they don't have proper governance.
00:14:07.080 --> 00:14:07.440
Yeah.
00:14:07.440 --> 00:14:32.440
Of the IT projects and if they allow their developers to run free and wild and innovate on their machines and install whatever they like and particularly specific tool and say small organization startups, because I usually work with organizations and highly regulated industries, such as services, where such governance exists because of compliance.
00:14:32.440 --> 00:14:37.200
Regulations are not all industries might have this kind of compliance regulations.
00:14:37.240 --> 00:14:39.559
Yeah. This is why. Researching this.
00:14:39.559 --> 00:14:48.000
And so the actual problem of finding out who is using AI, where is this use, which models are using, how they actually access it.
00:14:48.039 --> 00:14:52.759
It is a little bit problem because if you don't know what you have, you cannot possibly secure it.
00:14:52.759 --> 00:15:00.639
So do we need do we need AI compliance regulations or do we need an upgrade to existing SoC, TOS and things like that with, you know.
00:15:00.679 --> 00:15:05.559
One of the things that I'm seeing will be coming up with something called AI Bill on material.
00:15:03.480 --> 00:15:05.559
Yeah.
00:15:05.559 --> 00:15:06.679
So at the moment we have software.
00:15:07.879 --> 00:15:11.080
So now there's AI bomb emerging.
00:15:11.360 --> 00:15:15.080
And obviously we do have an inclusive you can actually help you with that.
00:15:15.080 --> 00:15:28.600
We have a train up to which allows you to discover your bill of materials based on the model, and also scan your repositories and see where developers are using specific libraries.
00:15:28.600 --> 00:15:31.639
For example, there was recently a.
00:15:31.879 --> 00:15:32.720
Actually.
00:15:32.720 --> 00:15:36.799
Sort of the supply chain that was on a very popular library called glider level.
00:15:36.840 --> 00:15:37.759
Right? Yeah.
00:15:37.759 --> 00:15:40.919
Do you know where in your organization light is in use?
00:15:41.279 --> 00:15:42.440
Yeah, yeah, I use it.
00:15:42.440 --> 00:15:44.960
The vulnerable version, which was hacked. Right.
00:15:44.960 --> 00:15:52.679
We we can now because we have played open source tools which can actually help you to get that inventory.
00:15:52.679 --> 00:15:54.919
But obviously there's a lot of other companies in the sector.
00:15:54.919 --> 00:16:00.480
But the problem is that you need to understand the actual challenge.
00:16:00.480 --> 00:16:05.519
As you mentioned, the shadow AI or shadow IT, you don't have the inventory.
00:16:05.679 --> 00:16:07.360
You will not be able to secure it.
00:16:07.360 --> 00:16:14.480
Do we need the inventories that you mentioned are really great, like the software building materials, the AI building materials, which mostly focuses on models.
00:16:14.559 --> 00:16:22.200
Context is something that is used more and more these days, skills and contacts that lives in developer environments, in projects.
00:16:22.279 --> 00:16:26.039
These are things that are being used to create the code.
00:16:26.080 --> 00:16:29.639
They're very, very key and they're just not being audited right now.
00:16:29.679 --> 00:16:39.440
Do you feel we need is there a space for a context bill of materials or something that should be added to an AI bomb to include what context was used to generate this code?
00:16:39.440 --> 00:16:41.679
I think that's a very interesting suggestion.
00:16:41.679 --> 00:16:47.279
So we have a whole group at our studio which is currently looking at it.
00:16:47.360 --> 00:16:56.279
And one of the things that you mentioned, skills, you actually have a new working group which is creating a new AI agenda skills top ten.
00:16:56.320 --> 00:16:57.759
Oh nice development.
00:16:57.759 --> 00:16:58.720
I need to join this.
00:16:58.720 --> 00:16:59.279
I need to do.
00:16:59.279 --> 00:17:02.799
This because just like all of US projects, it's an open source project.
00:17:03.000 --> 00:17:07.319
So we highly encourage contributions and collaborations from everyone.
00:17:07.359 --> 00:17:09.319
Amazing.
00:17:07.359 --> 00:17:09.319
I will join and I will contribute.
00:17:09.319 --> 00:17:10.319
Thank you very much Sam.
00:17:10.319 --> 00:17:12.319
Always a pleasure and great to see you here.
00:17:12.319 --> 00:17:14.000
Thank you very much. Thank you. Thank you.
00:17:14.000 --> 00:17:17.079
Hey everyone! Hope you're enjoying the episode so far.
00:17:17.079 --> 00:17:30.720
Our team is working really hard behind the scenes to bring you the best guests, so we can have the most informative conversations about a gentle development, whether that's talking about the latest tools, the most efficient workflows, or defining best practices.
00:17:30.720 --> 00:17:34.599
But for whatever reason, many of you have yet to subscribe to the channel.
00:17:34.640 --> 00:17:38.960
If you're enjoying the podcast and want us to continue to bring you the very best content.
00:17:38.960 --> 00:17:41.440
Please do us a favor and hit that subscribe button.
00:17:41.440 --> 00:17:49.079
It really does make a difference and lets us continue to improve the quality of our guests and build an even better product for you.
00:17:49.240 --> 00:17:50.960
All right, back to the episode.
00:17:56.000 --> 00:17:56.480
Great.
00:17:56.480 --> 00:17:58.160
So we've had a couple of chats already.
00:17:58.160 --> 00:17:59.920
We've heard the opening keynote.
00:17:59.920 --> 00:18:05.599
The time now is 1120, which is just about time for the for guy pigeons session.
00:18:05.640 --> 00:18:06.960
Don't secure the code.
00:18:06.960 --> 00:18:07.920
Secure the coder.
00:18:07.920 --> 00:18:10.440
Can security keep up with the tech dev?
00:18:10.440 --> 00:18:15.680
We've had some great chats with Amy here and the team with a whole bunch of people on the on the expo floor.
00:18:15.720 --> 00:18:18.559
Let's go up and see what guy pose sessions like.
00:18:18.559 --> 00:18:22.000
And I want to introduce somebody who is quite dear to me.
00:18:22.000 --> 00:18:24.160
I know this man for a couple of years.
00:18:24.160 --> 00:18:27.160
Guy, please come to the stage.
00:18:27.960 --> 00:18:30.960
Give it a big hand.
00:18:32.559 --> 00:18:35.559
Guy will talk about obviously.
00:18:35.759 --> 00:18:41.319
Development Guy is the founder and CEO of T-cell and also the founder of sneak.
00:18:41.400 --> 00:18:44.359
He was the person who was crazy enough to hire me.
00:18:44.359 --> 00:18:46.839
And I'm still here.
00:18:48.359 --> 00:18:58.799
Guy will talk about definitely talk about things like skills, but also about how a genetic development has evolved and can security actually keep up with this?
00:18:59.119 --> 00:19:03.640
I think it goes well, along with my introduction and the introduction that Joe gave.
00:19:03.640 --> 00:19:10.319
So with without further ado and not killing all the lights on stage, I will leave it up to you. Thank you. Guy.
00:19:10.359 --> 00:19:12.039
Thanks. Hello everyone. Can you hear me?
00:19:12.039 --> 00:19:13.400
Is this coming through?
00:19:13.400 --> 00:19:15.279
Do we have the slides up?
00:19:15.279 --> 00:19:15.799
Cool, cool.
00:19:15.799 --> 00:19:19.400
So, yeah, I'm Gabrielle, or I'm the founder of snake.
00:19:19.480 --> 00:19:20.920
The chairman of the board.
00:19:20.920 --> 00:19:28.279
But I'm here with a different hat here, which is a couple of years ago, I fell in love with AI and went to found Tesla, which is in the AI development space.
00:19:28.400 --> 00:19:35.839
And in that a lot of our work focuses on sort of securing for the agent era or in general.
00:19:36.000 --> 00:19:38.599
How do you develop software in the sort of a genetic era?
00:19:38.599 --> 00:19:43.000
What's the the new paradigm that we have to adapt to?
00:19:43.000 --> 00:19:44.960
And security is an aspect of that.
00:19:44.960 --> 00:20:02.039
And in that a lot of the kind of the core new world moves us from software development and security revolving around the implementation, revolving around the code to revolving around instructions and intent, because we're now driving and guiding agents.
00:20:02.039 --> 00:20:09.359
And so that's kind of the theme of my kind of points over here, is we need to move from securing the code to securing the coder, securing the the agent.
00:20:10.000 --> 00:20:13.119
So it's no secret that AI is transforming software development.
00:20:13.119 --> 00:20:19.480
And while we had a bunch of very good examples here about non software development flows on it, I think software development is a good harbinger.
00:20:19.480 --> 00:20:22.559
Sort of like Canary to say this will happen to all knowledge work.
00:20:22.559 --> 00:20:24.960
So I'll focus very much about that.
00:20:24.960 --> 00:20:33.720
And in software development we've gone from from AI augmented software development that was pioneered by Copilot and then cursor that it's more about your coding.
00:20:33.720 --> 00:20:38.079
And this helps you write code to AI natives of development.
00:20:38.079 --> 00:20:39.279
That is more about delegation.
00:20:39.279 --> 00:20:46.920
It's more about the agents in which you're asking the agent to do a task for you, and it goes on and and performs it better or not.
00:20:46.920 --> 00:20:50.519
And this agent development is where everything has consolidated towards.
00:20:50.519 --> 00:21:00.240
And so today, I think, again, not a controversial statement to say we should focus on what is how do we secure a genetic development as the mental model for us.
00:21:01.240 --> 00:21:02.759
And the development is amazing.
00:21:02.759 --> 00:21:03.880
It's powerful.
00:21:03.880 --> 00:21:12.079
A single person can do so much to it, but it introduces a bunch of new challenges as compared to how we've been dealing with software security before.
00:21:12.519 --> 00:21:15.559
And again, this is security and software development as a whole.
00:21:15.559 --> 00:21:19.240
And I'll mention three big ones here that I'll focus on in this talk.
00:21:19.240 --> 00:21:21.359
One is it's non-deterministic.
00:21:21.359 --> 00:21:25.079
So we use the software that if it compiled once it will compile again.
00:21:25.119 --> 00:21:25.359
Right.
00:21:25.359 --> 00:21:29.599
We're used to to things that that are the risk that we can say this now works I scanned it.
00:21:29.599 --> 00:21:32.359
It didn't have I'll scan it again.
00:21:29.599 --> 00:21:32.359
It does not have the like.
00:21:32.359 --> 00:21:34.400
It's the same findings that I get.
00:21:34.400 --> 00:21:36.799
And that's no longer the case with agents or with agents.
00:21:36.799 --> 00:21:39.759
We have to make do with the fact that these are non-deterministic creatures.
00:21:39.759 --> 00:21:40.839
We have to get statistical.
00:21:40.839 --> 00:21:44.039
We have to say, well, it works nine times out of ten, 99 times out of 100.
00:21:44.039 --> 00:21:46.119
How do we handle that?
00:21:44.039 --> 00:21:46.119
How do we even find out?
00:21:47.319 --> 00:21:52.039
Second is, as mentioned, it revolves around intent or instructions, not code.
00:21:52.039 --> 00:21:56.240
So what is the sort of new unit of software that we need to secure here?
00:21:56.519 --> 00:22:02.319
How do we evolve our security and improve our security of securing the coder?
00:22:02.319 --> 00:22:04.920
So that's a new challenge for us.
00:22:02.319 --> 00:22:04.920
And I'll talk about that.
00:22:04.920 --> 00:22:09.599
And then lastly, as you might have noticed, it changes at a certain kind of rapid clip.
00:22:09.599 --> 00:22:20.440
And so software development as a whole and the need for security and a bunch of other aspects of the knowledge domain are moving and changing faster than ever before.
00:22:20.440 --> 00:22:21.400
So how do we deal with that?
00:22:21.400 --> 00:22:27.119
So there are clearly many aspects of agent development that I don't cover here, but I'll focus on these three.
00:22:28.240 --> 00:22:32.039
So the first and biggest one is the fact that agent dev is non-deterministic.
00:22:32.039 --> 00:22:37.240
And that really reminds me of this sort of DevOps ethos that we have, right.
00:22:37.319 --> 00:22:39.119
In DevOps.
00:22:37.319 --> 00:22:39.119
We were saying if it moves, measure it.
00:22:39.119 --> 00:22:41.279
If it doesn't move, measure it in case it moves, know.
00:22:41.279 --> 00:22:46.640
It's like the most statistical creatures that we have in our disposal today is these servers that are sometimes up and sometimes down.
00:22:46.640 --> 00:22:48.559
They're not really behaving as they should.
00:22:48.559 --> 00:22:54.960
And so in this world you have to say, well, you can't optimize like the, the learning from there was you can't optimize what you can't measure.
00:22:54.960 --> 00:22:57.640
Right. It's a pretty simple statement, but it's a good one to remember.
00:22:57.640 --> 00:23:00.119
So how do we measure agent behavior.
00:23:00.119 --> 00:23:01.480
How do we think about this.
00:23:01.480 --> 00:23:07.480
Well, generally there are a lot of very evolved ways to do evaluations or evals in AI world.
00:23:07.480 --> 00:23:15.680
But typically what you would do is you would create a task for the agent, hey, agent, here is a task and I'll use 11 labs intentionally starting from non-security.
00:23:15.680 --> 00:23:26.759
So 11 labs is a very successful London based kind of text to speech lab or AI lab, and they have an ability to generate music.
00:23:26.759 --> 00:23:31.759
And so we can give it a task, for instance, in this case a dynamic soundtrack generator for a game studio.
00:23:31.759 --> 00:23:36.240
And then you define up front some criteria to say, what does good look like?
00:23:36.559 --> 00:23:39.680
Like what is what is correct implementation of this.
00:23:39.720 --> 00:23:42.559
And then you run that task across the agent.
00:23:42.559 --> 00:23:45.839
And what this plot shows is it shows five different scenarios.
00:23:45.839 --> 00:23:52.240
We have our sort of dynamic soundtrack over here, but I have five others in areas that I ran each of them ten times and I scored it.
00:23:52.240 --> 00:23:55.680
And you can see a few things like already you're a little bit more informed, for instance.
00:23:55.680 --> 00:23:57.920
You see the lines are not very high.
00:23:57.920 --> 00:24:01.359
The scores that the the agent got for this is not very good.
00:24:01.359 --> 00:24:03.480
And that's because the music API is relatively new.
00:24:03.480 --> 00:24:08.359
So it's not well represented in the, in the sort of the model right, in the weights.
00:24:08.359 --> 00:24:10.079
And so they don't really know how to handle it.
00:24:10.079 --> 00:24:11.440
And so that's one problem.
00:24:11.440 --> 00:24:13.400
The second is that you see the dots are all over the place.
00:24:13.400 --> 00:24:15.359
Like we asked the agent to do the same thing again.
00:24:15.359 --> 00:24:18.200
Again sometimes you see kind of like normal scatter.
00:24:18.200 --> 00:24:20.960
Sometimes they'd better, sometimes they'd worse, maybe within a rage.
00:24:20.960 --> 00:24:23.799
And you have some other cases where it kind of hit the mark.
00:24:23.799 --> 00:24:24.039
Right.
00:24:24.039 --> 00:24:28.000
It sort of guessed correctly one time or two times and then it doesn't.
00:24:28.000 --> 00:24:28.680
These others.
00:24:28.680 --> 00:24:31.640
So we have this mess and it's very hard to work with that.
00:24:31.640 --> 00:24:33.559
So clearly said, well, how can I make it better.
00:24:33.559 --> 00:24:35.000
How can I make it work better?
00:24:35.000 --> 00:24:38.960
How do I kind of train the sort of agents, how do I harness them?
00:24:39.079 --> 00:24:41.799
And harness is actually a different word.
00:24:41.799 --> 00:24:43.400
I didn't say that. Ignore that word.
00:24:43.400 --> 00:24:49.920
And then and it's sort of the most common kind of way to do that is with context.
00:24:49.920 --> 00:24:54.039
And the most common unit of context that is used today are skills.
00:24:54.039 --> 00:24:57.720
So these are basically markdown files a little bit more with some structure.
00:24:57.720 --> 00:25:02.880
But these are bits of information for the agent to use to be able to perform this task.
00:25:02.880 --> 00:25:06.000
And we need to remember that knowledge is not the same as intelligence.
00:25:06.000 --> 00:25:12.359
The the models can be brilliant and very, very capable, but if they don't know something, oftentimes they just can't do it.
00:25:12.359 --> 00:25:16.079
Or it's just mighty inefficient for them to do that so they can figure it out.
00:25:16.079 --> 00:25:18.480
But they will be there will be statistics there.
00:25:18.480 --> 00:25:23.839
And so in this case, we have an 11 labs music skill that helps explain the API.
00:25:23.839 --> 00:25:28.359
And so when we run, when we use that we indeed see that for our dynamic soundtrack generator.
00:25:28.359 --> 00:25:30.519
And I'm using Tesla here to do the evaluations.
00:25:30.519 --> 00:25:32.960
You see that without context it didn't do a bunch of things.
00:25:32.960 --> 00:25:33.599
It did.
00:25:33.599 --> 00:25:37.240
It used the deprecated package, it didn't import it correctly and all that.
00:25:37.240 --> 00:25:39.759
And with the skill with the context, it got it done.
00:25:39.759 --> 00:25:44.680
So it got a 50% average versus a 98 without the skill versus 98%.
00:25:44.920 --> 00:25:46.480
So it's good we managed to correct it.
00:25:46.480 --> 00:25:51.880
And if we run this across the line, we see in some cases we just solved it.
00:25:51.920 --> 00:25:56.480
It just basically now is sufficiently good to fairly consistently do the task.
00:25:56.519 --> 00:26:02.880
And in other cases it's still slightly varied, but it's further up, so it is more capable of doing that task.
00:26:02.920 --> 00:26:07.640
So this is the type of way in which you evolve and you improve your agents ability to do something.
00:26:07.640 --> 00:26:08.559
This is a security event.
00:26:08.559 --> 00:26:10.400
So how do we talk about this for security?
00:26:10.400 --> 00:26:11.839
Well, let's take another example.
00:26:11.839 --> 00:26:14.359
Code guard is something that Cisco created and donated.
00:26:14.359 --> 00:26:20.559
And it's basically a bunch of sort of OWASp security kind of rules packaged up into a skill.
00:26:20.559 --> 00:26:28.039
And and it helps kind of these AI developers, these agents develop more secure code.
00:26:28.240 --> 00:26:30.200
So let's put that to the test.
00:26:30.200 --> 00:26:33.759
I created six different evaluation scenarios.
00:26:33.759 --> 00:26:35.799
And I specifically focused on authorization.
00:26:35.799 --> 00:26:38.519
How well did you handle the security of authorization in this code.
00:26:38.519 --> 00:26:46.440
So for instance hey agent, create an access control test suite for a project management API and then created a bunch of scores for it.
00:26:46.440 --> 00:26:48.759
And I did it with and without code Guard.
00:26:48.759 --> 00:26:50.960
And as you would expect, it got better.
00:26:50.960 --> 00:26:59.720
So without Code Guard, just as is, the agent scored 48% on my kind of authorization criteria, sort of a scorecard.
00:26:59.759 --> 00:27:00.759
So not awesome.
00:27:00.759 --> 00:27:05.680
And if it had these instructions nearly 1.6 x the improvement on.
00:27:05.680 --> 00:27:11.359
So it got a fair bit better because there was a bunch of guides about how to sort of code securely, including authorization.
00:27:11.359 --> 00:27:14.000
So that's that's good.
00:27:11.359 --> 00:27:14.000
That's already useful.
00:27:14.000 --> 00:27:26.039
However, if I took all the information in Code Guard that has a lot of security practices in it, and I shrunk it down to just the authentic authentication authorization related bits, which are, I think about sort of 5% of the total content.
00:27:26.039 --> 00:27:28.440
If I remember correctly, it did a lot better.
00:27:28.440 --> 00:27:29.759
It went to 98%.
00:27:29.759 --> 00:27:34.160
And this is this is sort of a core thing to remember, which is more context is not necessarily better. Right.
00:27:34.160 --> 00:27:42.440
If I sat here and I told you 100 things, no matter how sort of brilliant or dumb they were, you will notice you'll give less attention to each one of them than if I told you.
00:27:42.480 --> 00:27:46.440
Three and attention is a scarce resource with humans and with models.
00:27:46.440 --> 00:27:47.920
And so choosing and managing.
00:27:47.920 --> 00:27:57.759
What is it that you say that actually matters that the model doesn't already know, like not wasteful information that there's no point in saying is important and is a part of that competency.
00:27:58.759 --> 00:28:09.440
Taking that a little bit further, you know, if this is the graph that shows the same numbers from before of code guard, not guard, you also might be surprised to hear the different agents respond to the same information differently.
00:28:09.440 --> 00:28:11.200
So this is an example of the same test.
00:28:11.200 --> 00:28:14.400
In fact a whole it's literally the same execution.
00:28:14.400 --> 00:28:16.720
And what you can see is different models respond differently.
00:28:16.720 --> 00:28:20.119
Opus and sonnet at the top they get about the same results.
00:28:20.119 --> 00:28:23.039
Even though opus is more intelligent, it is more expensive.
00:28:23.039 --> 00:28:28.559
So if you use the opus for this specific task, you kind of wasted money and probably time because you could do the same with sonnet.
00:28:28.559 --> 00:28:30.920
And what you can see is codex and cursor.
00:28:30.920 --> 00:28:34.519
They respond even differently and it doesn't matter which one is better.
00:28:34.559 --> 00:28:35.640
Like it does matter eventually.
00:28:35.640 --> 00:28:42.640
But that's not my point, but rather the fact that different agents, again, almost like humans, listen differently.
00:28:42.680 --> 00:28:48.279
And so you want to know that your instructions are effective, are not wasteful for the agents that you are using.
00:28:49.359 --> 00:28:56.079
So that's kind of core point here is if we want to secure agents, we want to use skills to help them secure code or write secure code.
00:28:56.079 --> 00:28:58.839
We have to learn how to measure that.
00:28:58.839 --> 00:29:00.799
And how do you how do you build good context.
00:29:00.799 --> 00:29:03.160
Like how do you how do you evolve it?
00:29:03.160 --> 00:29:05.960
How do you create kind of quality context, you know, how to build.
00:29:05.960 --> 00:29:15.640
So where to talk a little bit about the fact that you generate you create a skill, you evaluate it, and then once you've evaluated, you can optimize it until you kind of get a better and better guidance for that agent.
00:29:15.640 --> 00:29:18.039
And then you need to distribute that or communicate it.
00:29:18.039 --> 00:29:22.400
If you want to use the human mode to the agents and observe what has happened.
00:29:22.400 --> 00:29:23.480
And the observe is important.
00:29:23.480 --> 00:29:29.240
If the evals are kind of like your tests, once you've got something working and you make a modification, how do you know you're not breaking it?
00:29:29.240 --> 00:29:31.119
How do you know you're evolving it?
00:29:31.119 --> 00:29:33.400
How do you know if you can use a cheaper model or not?
00:29:33.400 --> 00:29:37.039
You have to be able to evaluate, but eventually your test will go out of out of sync.
00:29:37.039 --> 00:29:40.119
They would not represent reality if you don't also observe what is happening.
00:29:40.119 --> 00:29:42.559
So we call this the context development lifecycle.
00:29:42.559 --> 00:29:48.079
And as you build good context, you can use that context across the software development lifecycle.
00:29:48.119 --> 00:29:50.920
So we think the CDC is where us humans should live.
00:29:50.920 --> 00:29:53.599
We should be building good context that guides the agents.
00:29:53.599 --> 00:29:56.799
And then we should apply that to the SDLC where the agents should work.
00:29:56.799 --> 00:30:02.559
And the same context, the same instruction can be applied end to end in the development process.
00:30:02.920 --> 00:30:12.480
Same as like a great developer on the team will use the same knowledge to define a product feature, write the code, troubleshoot something, ship it to production, troubleshoot an incident, etc. etc.
00:30:12.519 --> 00:30:16.599
the same knowledge is useful across the board, so the skills represent that knowledge.
00:30:16.599 --> 00:30:22.359
And of course, from a security lens perspective, we can now use that to secure different steps.
00:30:22.599 --> 00:30:26.559
So this is this is how you should write secure code at the beginning.
00:30:26.559 --> 00:30:29.599
This is what I want to audit in the code review to highlight to you.
00:30:29.599 --> 00:30:31.400
This is what I want to get on.
00:30:31.400 --> 00:30:33.839
This is what I want to inspect when an incident occurred.
00:30:33.839 --> 00:30:37.519
So all of these things can be represented in skills that we use across the SRC.
00:30:40.160 --> 00:30:41.599
So this is non-deterministic.
00:30:41.599 --> 00:30:43.839
It's kind of my biggest point to make.
00:30:43.839 --> 00:30:46.400
And as we as we learn how to evaluate those.
00:30:46.400 --> 00:30:52.559
The second point is we've been talking more and more about these skills and we're sort of optimizing these skills and we're developing these skills.
00:30:52.559 --> 00:30:58.359
And I think it's useful to start thinking about skills in terms of their own security as a unit of software.
00:30:58.400 --> 00:30:59.720
Right. It looked like a markdown file.
00:30:59.720 --> 00:31:02.720
They look like a notion document or like a confluence document.
00:31:02.720 --> 00:31:07.839
But the way we process them is we execute them by the agent or the agent execute them.
00:31:07.839 --> 00:31:15.839
So really, I think we're well served, especially from a security lens, to think about them as a unit of software, not just as a piece of text.
00:31:15.839 --> 00:31:18.680
And we have like a lot of indications of that today.
00:31:18.680 --> 00:31:18.839
Right?
00:31:18.839 --> 00:31:30.599
We have the sneak study, and there were very many others that showed, especially in the open claw world, where a lot of skills were malicious, literally attackers putting in things that are trying to make the, the, the agent to do something it shouldn't.
00:31:30.839 --> 00:31:32.640
Here's an example of a malicious skill.
00:31:32.640 --> 00:31:37.440
This is from the Tesla registry scanned by sneak, where it had a bunch of that it down.
00:31:37.880 --> 00:31:43.759
And while most of your wells were just standard blockchain APIs, one of them was suddenly downloading a password protected zip.
00:31:43.920 --> 00:31:45.359
Fishy doesn't sound right.
00:31:45.359 --> 00:31:47.039
Okay, that's potentially a malicious skill.
00:31:47.039 --> 00:31:51.240
There's a bunch of things we can do, clearly imperfect, but we can try to detect malicious skills.
00:31:51.640 --> 00:31:53.559
There are also vulnerable skills.
00:31:53.559 --> 00:31:54.680
What's a vulnerable skill?
00:31:54.680 --> 00:31:58.240
For instance, a skill that uses insecure credential handling.
00:31:58.240 --> 00:32:05.759
It asks the user to put API keys inside, or it makes MCP calls with sort of plain vanilla tokens for it.
00:32:05.759 --> 00:32:08.039
So that's an example of something that is insecure behavior.
00:32:08.039 --> 00:32:12.079
It's vulnerable to exfiltrate some information outside.
00:32:12.599 --> 00:32:20.839
There's also like new types of flaws that you might have in what I this is not an industry term, but what I like to think of as negligence skills.
00:32:20.839 --> 00:32:26.400
So these are skills that do not have some basic safety instructions inside of them.
00:32:26.799 --> 00:32:29.720
Come along like check this into a repository.
00:32:29.720 --> 00:32:32.279
Do not make it a public repository.
00:32:32.279 --> 00:32:35.200
If you cannot commit to the repository, fail.
00:32:35.200 --> 00:32:37.799
Don't like go off and sort of send it in another way.
00:32:37.799 --> 00:32:40.680
So a bunch of these types of examples are very real examples.
00:32:40.680 --> 00:32:43.680
And we have cases where we've had agents.
00:32:43.880 --> 00:32:48.680
There's a behavior called reward seeking which is you ask them to do something and they're really, really, really keen to do it.
00:32:48.680 --> 00:32:55.359
And so they go up and they try to do everything they can, and they escape their sandbox and they delete files and they do whatever it is to please you.
00:32:55.559 --> 00:32:58.279
And so you have to define a little bit of these safety instructions.
00:32:58.279 --> 00:33:01.079
It's kind of similar to what Brian's example was on.
00:33:01.079 --> 00:33:05.079
Right to say do not disclose information that makes it at least less negligent.
00:33:06.039 --> 00:33:09.720
And then again, similar to software, there's a question about supply chain.
00:33:09.759 --> 00:33:11.839
How do you consume these skills today?
00:33:11.839 --> 00:33:15.319
The reality is that people just consume them out of GitHub repo.
00:33:15.359 --> 00:33:17.880
They download them from wherever.
00:33:15.359 --> 00:33:17.880
You have no idea that it happened.
00:33:17.880 --> 00:33:18.920
You have no idea where it's there.
00:33:18.920 --> 00:33:21.200
Then they check them in to their repositories.
00:33:21.200 --> 00:33:26.400
Different agents read them from different places, so you might check them in seven times the different folders within your repo.
00:33:26.400 --> 00:33:27.279
It's not.
00:33:27.279 --> 00:33:28.960
It's early, it's fine, will improve.
00:33:28.960 --> 00:33:30.160
But for now it's not awesome.
00:33:30.160 --> 00:33:32.680
So you have to think a little bit about supply chain.
00:33:31.000 --> 00:33:35.720
So all of those become obvious once you think about skills as units of software.
00:33:36.839 --> 00:33:38.039
So what do we want to do here.
00:33:38.039 --> 00:33:44.000
Like what is what is enterprise grade kind of skill governance skill usage on it.
00:33:44.000 --> 00:33:45.119
This is a nascent space.
00:33:45.119 --> 00:33:47.759
I'll give you the lens on it because this is our world.
00:33:47.759 --> 00:33:52.519
We think you need to think about three different elements of of evolving this piece of software.
00:33:52.559 --> 00:33:54.160
First is governance and security.
00:33:54.160 --> 00:33:55.839
Know what the hell is going on?
00:33:55.839 --> 00:33:59.559
Try to sort of audit the use of it like you've posted skills.
00:33:59.559 --> 00:34:01.160
Did anybody install them?
00:34:01.160 --> 00:34:03.079
Constrain the use of skills.
00:34:03.079 --> 00:34:06.880
So people download skills always through this sort of centralized path.
00:34:06.920 --> 00:34:11.079
Again, not that dissimilar to what you should be doing with NPM libraries or whatever.
00:34:11.079 --> 00:34:12.599
And so you have to know about the governance.
00:34:12.599 --> 00:34:15.400
If you can't do that, you really oftentimes cannot roll out.
00:34:15.400 --> 00:34:19.480
Once you rolled out, you have a need to standardize and allow reuse.
00:34:19.760 --> 00:34:27.800
If three different people created a skill to review code, and a fourth person comes along and says, I want to use a skill like, first of all, where do they find it?
00:34:27.800 --> 00:34:30.360
Second is, how do they know which of the three to change?
00:34:30.360 --> 00:34:35.360
If I created a skill and then one of you came in and proposed the modification, how do you know if it's good or not good.
00:34:35.360 --> 00:34:42.519
If I now 100 people are using this skill and I'm going to make a change to the skill, how do I know that I'm not breaking it?
00:34:42.519 --> 00:34:48.360
And so there's a bunch of this notion of standardization, of reusability that you have to create a measure.
00:34:48.360 --> 00:34:51.840
And then lastly, and this is the holy grail is continuous optimization.
00:34:51.840 --> 00:34:57.800
You want to know that this is that CLC, that we want the continuous optimization.
00:34:57.800 --> 00:34:59.199
You want to observe what has happened.
00:34:59.199 --> 00:35:00.119
Did the agent fail?
00:35:00.119 --> 00:35:08.360
Did the user need to correct the agent and take that information and route that back to be able to evolve the skill, create new eval scenarios?
00:35:08.360 --> 00:35:09.440
And that is the holy grail.
00:35:09.440 --> 00:35:11.880
And the companies that are at the cutting edge are doing this right.
00:35:11.880 --> 00:35:13.519
They are creating that optimization.
00:35:13.519 --> 00:35:15.760
Most organizations are quite far from it.
00:35:15.760 --> 00:35:18.639
And so that's why this is oftentimes the sequence.
00:35:19.800 --> 00:35:24.119
I'd be remiss if I didn't do a little bit of a Tesla plug over here to doing it.
00:35:24.119 --> 00:35:25.760
So that's oftentimes what we help you do.
00:35:25.760 --> 00:35:40.960
We have a platform in which, on one hand from a from a development perspective, we help you collaboratively develop skills that allow developers to discover and install those quality skills and then observe what has happened, learn from that and create new paths.
00:35:41.119 --> 00:35:45.079
And then within that we have these these controls.
00:35:45.079 --> 00:35:52.639
So we have the ability to now make sure with sneak we scan every skill that gets published into the into the registry.
00:35:52.639 --> 00:35:54.679
So you know that it's not malicious.
00:35:54.679 --> 00:35:57.559
Similarly, when you install we have controls about scan from us.
00:35:57.559 --> 00:35:59.519
We have analytics about who's using what.
00:35:59.519 --> 00:36:10.119
And then lastly for the sort of these nascent agent enablement teams, these platform teams, developer experience team AI enablement teams that are that own successful rollout of agents in the organization.
00:36:10.119 --> 00:36:15.199
We give them a bunch of these abilities to eliminate duplicates, drive skill usage, optimize costs.
00:36:15.519 --> 00:36:18.320
I like to say that a genetic development is cheaper than expensive.
00:36:18.320 --> 00:36:22.440
It's like very cheap at the beginning because a single person can do so much and then you get the bill.
00:36:22.559 --> 00:36:23.639
It's not that awesome.
00:36:23.639 --> 00:36:28.360
So you start thinking, as we use agents more and more, what is the sort of the cost?
00:36:29.119 --> 00:36:40.719
And those, of course, you know, a moment with the sneak hat on, we have some amazing other aspects of securing the agent behavior itself and its runtime in Evo, and I'm sure you'll hear more about this over here as well.
00:36:42.360 --> 00:36:52.400
And then to close off, I want to talk about the third bullet, which is a genetic development moves faster than ever and security must become a genetic to keep up.
00:36:52.480 --> 00:36:55.000
Sounds very familiar to me from the sneak early days.
00:36:55.000 --> 00:36:58.440
And when I when I hearken back to what happened, it sort of sneak roots.
00:36:58.719 --> 00:37:12.840
You think if you're a graybeard like me, you think about the change that happened there when we went from waterfall to cloud and some behaviors, some manual processes and such were tolerated in waterfall and were no longer tolerated in cloud.
00:37:12.840 --> 00:37:19.199
The idea that before any piece of software will shift, someone will manually audit it was tolerated.
00:37:19.239 --> 00:37:25.840
You know, like the best teams automated the security scanning of it, but most people manually audited, so the average team did not.
00:37:25.840 --> 00:37:28.679
In cloud. That cannot be the case once you're in DevOps.
00:37:28.679 --> 00:37:31.800
Once you're in that continuous, you have to automate that scanning.
00:37:31.920 --> 00:37:35.400
I think we're facing now the same thing, which is some things are tolerated in cloud.
00:37:35.400 --> 00:37:36.960
The best teams are automating them.
00:37:36.960 --> 00:37:39.159
They're sort of reviewing them.
00:37:36.960 --> 00:37:39.159
They're auto improving them.
00:37:39.159 --> 00:37:40.480
But most teams are not.
00:37:40.480 --> 00:37:43.039
And they're no longer going to be tolerated in agents.
00:37:43.039 --> 00:37:46.039
So there's a little bit of like the future is here, but it's not evenly distributed.
00:37:46.039 --> 00:37:50.480
We should think about all these things that are like the paper cuts that we have, the places in which like, you know what?
00:37:50.480 --> 00:37:57.880
I'm a secure it's like I will triage these vulnerabilities for my developers or I will, you know, maybe we'll only fix the ones that are truly glaring.
00:37:57.880 --> 00:37:59.159
We're not going to fix the rest.
00:37:59.159 --> 00:38:01.079
Many of these things are just no longer an option.
00:38:01.079 --> 00:38:03.639
So you have to think about how do we improve them.
00:38:03.639 --> 00:38:05.559
And there's a long list of those.
00:38:05.559 --> 00:38:13.360
There are many, many, many things that go from nice to have to, must have from indeed prioritization to automating upgrades to detection of supply chain manipulations.
00:38:13.360 --> 00:38:14.599
Like there's just so many things.
00:38:14.599 --> 00:38:17.079
This is really just a tiny sample set.
00:38:17.079 --> 00:38:24.440
And the good news is that for each one of those agents can really, really help in making these automated like it's agents all the way down.
00:38:24.480 --> 00:38:24.639
Right?
00:38:24.639 --> 00:38:30.679
You can build agents upon agents that will do a bunch of these different steps, and that allows us to scale.
00:38:31.039 --> 00:38:35.559
And so we find ourselves in this kind of carrot and stick mode security.
00:38:35.559 --> 00:38:38.280
If it doesn't become a genetic, it will never keep up.
00:38:38.280 --> 00:38:40.519
We will fail.
00:38:38.280 --> 00:38:40.519
The attackers are becoming a genetic.
00:38:40.519 --> 00:38:42.280
They're moving faster than ever.
00:38:42.280 --> 00:38:47.400
The development, like the business, has to be a genetic to be able to to develop and we have to keep up over there.
00:38:47.400 --> 00:38:51.599
So if we don't become a genetic, there's a bit of a, you know, like you're going to be in trouble.
00:38:51.599 --> 00:38:57.920
But if you do, if we do become a genetic in the world, then we can actually fix application security.
00:38:57.920 --> 00:39:09.159
We can actually fix those things that we've long wanted and tried to get developers to do on a consistent fashion, which to me is exciting. So I'm excited by the new future.
00:39:09.159 --> 00:39:09.960
I'm slightly daunted.
00:39:09.960 --> 00:39:12.880
I think there's a lot of need to change just to plug.
00:39:12.880 --> 00:39:15.280
We have a conference in a couple of weeks here in Dev Con.
00:39:15.280 --> 00:39:18.280
It's running here in London June 1st and second.
00:39:18.280 --> 00:39:25.320
If it's all about a genetic development, adoption like real world scenarios actually in kind of organizations that have it, if you want to check it out.
00:39:25.320 --> 00:39:26.440
I think Brian might be speaking.
00:39:26.440 --> 00:39:33.840
And also I think we stole you to a different one in the previous one and a lot of learning on it.
00:39:33.880 --> 00:39:36.480
Would love to see you there if you'd like.
00:39:36.480 --> 00:39:39.280
That's it for me. Thank you.
00:39:39.280 --> 00:39:41.599
What a day. At the AI Security Summit.
00:39:41.599 --> 00:39:44.480
We had some great discussions on on the Tesla booth.
00:39:44.480 --> 00:39:47.079
We had some wonderful chats on the showroom floor.
00:39:47.079 --> 00:39:55.960
We had some great sessions, guy in particular, super enlightening about securing the coder, not the code from AI Security Summit.
00:39:56.000 --> 00:39:56.760
Wonderful day.
00:39:56.760 --> 00:39:57.440
Thank you very much.
00:39:57.440 --> 00:39:59.840
Sneak. And everyone here.
00:39:59.840 --> 00:40:04.480
The AI native dev is brought to you by the package manager for skills and context.
00:40:04.519 --> 00:40:07.639
Your hosts are Guy pigeon and me, Simon Maple.
00:40:07.679 --> 00:40:09.480
Our producer is Tom Dowler.
00:40:09.480 --> 00:40:12.800
The AI native dev is not just a podcast, it's a community.
00:40:12.840 --> 00:40:16.480
And we host monthly meetups at the Tesla offices in central London.
00:40:16.519 --> 00:40:21.480
Visit Tesla IO forward slash community to learn more and I hope to see you there.