00:00:01.679 --> 00:00:03.520
Welcome to another coffee with developers.
00:00:03.600 --> 00:00:06.320
We hear live at the GitHub offices in San Francisco.
00:00:06.400 --> 00:00:09.919
Well, in the library of it, not everybody works with a place like that, I guess.
00:00:10.080 --> 00:00:10.880
I wish.
00:00:12.080 --> 00:00:13.919
So Judia, you and what are you doing here?
00:00:14.240 --> 00:00:14.880
Thanks for having me.
00:00:14.960 --> 00:00:16.000
My name is Julia.
00:00:16.079 --> 00:00:18.719
I'm a product manager on the VS Code team.
00:00:18.879 --> 00:00:27.600
In particular, I support all of the model launches in VS Code and in GitHub Copilot in VS Code and our chat extension.
00:00:28.239 --> 00:00:30.160
So yeah, it's been a fun journey.
00:00:30.320 --> 00:00:32.960
I joined the team last year in August.
00:00:33.119 --> 00:00:35.759
So it's now been roughly over half a year.
00:00:36.240 --> 00:00:40.240
Lots of learnings, lots of growth, and been super exciting so far.
00:00:40.640 --> 00:00:45.920
So you're part of Microsoft and working with GitHub rather than GitHub?
00:00:46.240 --> 00:00:50.640
It's you know, we all consider ourselves very close friends.
00:00:50.880 --> 00:00:57.759
Um so we are in this very lucky position that we work very closely with GitHub, but also with the Microsoft team.
00:00:57.920 --> 00:01:09.439
Um so if you look at my old chart, yes, I like direct in Microsoft, but we have day-to-day conversations, we work very close with the GitHub folks, so it really goes hand in hand.
00:01:09.680 --> 00:01:18.079
So when people talk about VS Code or old school people like me, we're always like, oh, cool new editing feature, cool new code folding, whatever, and these kind of things.
00:01:18.159 --> 00:01:22.400
But with the integration of uh of Copilot, it becomes much more than that.
00:01:22.560 --> 00:01:27.280
It feels like more like an integrated IDE writer now than a text editor.
00:01:27.680 --> 00:01:36.719
How are the numbers of people actually jumping into the fully like chat integrated thing in Visual Studio Code, or is it more like coders that use it like me, command I when I need it?
00:01:37.040 --> 00:01:40.560
I think we are getting a lot of different use.
00:01:41.120 --> 00:01:52.079
Like we have the ones that grew up with VS Code as the pure IDE code editor, who are not that often yet using our chat extension.
00:01:52.159 --> 00:02:07.359
But then we also get the new developers that have a very AI-first mindset who go in and purely want to use the Copilot chat extension, who immediately go in with the intention of I want to use the AI features, the AI functionalities in here.
00:02:07.599 --> 00:02:23.840
Um the numbers are I can't share too many details, but um they are constantly growing, and we see more and more users, also from the more old school, how you kind of called it, and maybe they are more and more using these AI integration more natively.
00:02:23.919 --> 00:02:28.879
Um and we truly see a shift on how developers are working with AI these days now.
00:02:29.599 --> 00:02:35.199
Now, your dedicated role is to talk about the model integration into the V into VS Code.
00:02:35.680 --> 00:02:42.879
And uh I mean it was interesting because at first it started with a few models and then Microsoft announced like we allow any of them in there.
00:02:42.960 --> 00:02:46.719
So, how many are in uh available out of the box right now?
00:02:47.360 --> 00:02:52.080
And how fast do you how long do you get time to actually get acquainted with them before new ones come out?
00:02:52.400 --> 00:02:53.520
Yeah, good question.
00:02:53.680 --> 00:02:56.159
Um so GitHub Copilep has different SKUs.
00:02:56.319 --> 00:02:59.199
So we have enterprise skew, we have an individual screw.
00:02:59.280 --> 00:03:02.639
So it does depend on which is the subscription, which is the license.
00:03:02.719 --> 00:03:07.199
First of all, you've subscribed to, and then depending on this, what are the models that you're allowed to do?
00:03:07.280 --> 00:03:18.159
As an enterprise, um, certain enterprises have certain policies, so it might be that your admin is restricting what are some of the models that you're allowed to see because of their security or ceiler requirements.
00:03:18.319 --> 00:03:25.520
Um in theory, um, we support all of the three big model providers from Google, Anthropic to OpenAI.
00:03:25.680 --> 00:03:41.439
And we also work very closely with them before these model launches to make sure that it really works end-to-end for all of our developers who immediately from day one want to use these new models to make sure it works in VS Code and with our chat extension.
00:03:41.680 --> 00:03:43.680
Um typically it really depends.
00:03:43.840 --> 00:03:51.280
I thought, so I joined the team and I thought, oh, we must have such good processes and like it works very smoothly.
00:03:51.520 --> 00:04:02.080
The reality is that these model providers sometimes get a checkpoint, a model checkpoint that is very promising, and they're like, oh, we have a very exciting new checkpoint.
00:04:02.240 --> 00:04:11.680
Can can you please go and test this and give us some results back and make sure and let us know how this new model checkpoint is doing within your harness.
00:04:12.000 --> 00:04:14.400
So it's very ad hoc sometimes.
00:04:14.560 --> 00:04:32.639
Um typically I would say we probably get like two-ish weeks in advance where we know oh, there might be a new model, um, uh yeah, a new model release coming in, and that's when our engineering team, our offline eval team started to get to work with these new models.
00:04:33.439 --> 00:04:39.519
Now you talked about e bills, what are they and what do people what why are they important?
00:04:40.000 --> 00:04:40.480
Yeah.
00:04:40.720 --> 00:04:42.800
Um so there are different kinds of evals.
00:04:42.879 --> 00:04:47.120
Usually when I talk about evaluations, um we talk about online and offline.
00:04:47.439 --> 00:04:52.399
Um so one of the very interesting things, AI is non-deterministic.
00:04:52.639 --> 00:04:54.639
So you might have experienced this yourself.
00:04:54.720 --> 00:05:01.680
You go in, you open the chat, and you ask one question, oh, um scaffold me a new web API.
00:05:01.839 --> 00:05:19.519
And every time you do this, you're gonna get a different result, which makes it so hard if we think for us to test before a new model gets released, how do we test and how do we evaluate how these new models perform, especially because of their non-deterministic nature?
00:05:19.920 --> 00:05:29.279
So, how it works is we very similar back then we had um tests, we had test scripts.
00:05:29.439 --> 00:05:34.160
We still do this in the evaluation or offline eval space.
00:05:34.240 --> 00:05:35.600
These are called assertions.
00:05:35.839 --> 00:05:53.040
So we write a test case which is scaffolding a web API, and then we have certain deterministic assertions to say, oh, a web API should have, for this specific task, it should have five endpoints or things like that.
00:05:53.199 --> 00:06:08.240
So we write up these assertions, so then we have a way to test if this non-deterministic model is able to achieve all of our assertions, only maybe a certain amount of these, and how many do we need to make sure that it actually works end-to-end.
00:06:08.959 --> 00:06:11.040
So it feels like a good prompt, actually.
00:06:11.199 --> 00:06:17.680
Like instead of just giving it an open-ended question, you're already limited to the things that you want, and then you get better quality results back.
00:06:18.240 --> 00:06:28.480
Well, yeah, I mean, if everyone is, would be as efficient as prompting, but we all know, including myself, sometimes I go in, I ramble, I don't really know what I want the model to do.
00:06:28.560 --> 00:06:37.519
I have a fair amount of an idea what I want it to do, and but not every prompt is always as um straightforward as it probably should be.
00:06:38.480 --> 00:06:44.560
One uh one thing you said was that companies might limit it down to only using a few models or different ones.
00:06:44.720 --> 00:06:53.279
Now, the big thing right now is that everybody has got worried about is like token pricing and like uh our engineers burning through like lots and lots of money for different things.
00:06:53.519 --> 00:06:56.560
Now, different models are differently efficient at different tasks.
00:06:56.720 --> 00:07:03.199
There's lots of third-party software that I see that actually allows you to see your cloud status, for example, in the taskbar.
00:07:03.360 --> 00:07:12.399
Is there an idea to do that inside VS Code to say, like, okay, here's the limit how many tokens you can use, here's which model is the most efficient one for the tasks that you need to do?
00:07:12.480 --> 00:07:16.800
Is there a way to get a token saving mechanism inside VS Code?
00:07:17.600 --> 00:07:19.279
Yeah, that's a very spicy question.
00:07:19.439 --> 00:07:22.240
I'm not allowed to answer all of them just yet.
00:07:22.480 --> 00:07:32.079
Um one of the things, um, so one of the benefits of GitHub Copilot is that we don't just want to limit you to one model provider.
00:07:32.240 --> 00:07:39.839
It truly is you can pick across different model providers whatever model you think works best for your task.
00:07:40.079 --> 00:07:52.560
Um we are, or we do want to be more transparent in sharing some of our own internal benchmarks across these different model providers to give our end users a little bit of a better, roughly fair estimate.
00:07:52.720 --> 00:07:57.360
Hey, maybe for this task, uh it would be super cool to use this model or another one.
00:07:57.680 --> 00:08:07.199
At the end of the day, something including myself, you get to know models and they all have their own certain personalities in a way.
00:08:07.519 --> 00:08:12.240
So you're getting used to a certain style, a certain way, maybe for a certain task.
00:08:12.560 --> 00:08:21.600
And I see people picking models based on tasks that they are more familiar with, the personality or the style that they like.
00:08:21.839 --> 00:08:32.960
So on a lot of use cases, I see, for example, um claude models being used for more web dev, um backend development being more used with the GPT models.
00:08:33.200 --> 00:08:36.960
Um so yeah, it's uh it is a tough question.
00:08:37.440 --> 00:08:39.360
There's also political decisions sometimes.
00:08:39.440 --> 00:08:47.200
I mean, there's lots of models that are from countries you might not want to send your data to, uh, and uh but that seemed to be very, very powerful.
00:08:47.360 --> 00:08:50.720
But it's it's such an overwhelming market.
00:08:51.039 --> 00:08:58.799
I mean, I look at at Hugging Face and I looked, I used to look up like a specialized model for different tasks, and I'm like, cool that I could run locally and things.
00:08:59.039 --> 00:09:01.840
I kind of gave up because there's a new model every two days.
00:09:01.919 --> 00:09:03.759
How do you keep up with the demand?
00:09:04.080 --> 00:09:08.480
Yeah, so now it's part of my job, so I basically have to keep up with the demand.
00:09:08.720 --> 00:09:30.159
Um it is very interesting for me personally, because now that I that we work so closely with the model providers, we can really see the shift every time a new model has, and what are some of the improvements, also so some of what are some of the personality, even like kind of like personality traits for these new models.
00:09:30.399 --> 00:09:32.639
So we now watch more carefully.
00:09:32.799 --> 00:09:43.279
Um but yeah, that it feels almost like we've also reached a saturation of how good is there even going to be a next big jump?
00:09:43.440 --> 00:09:56.399
Um, because right now it feels like GPD 5.5 and the Opus 4.7 one, they're very close in um quality at this point, so I can see why users are also like the quality is really good.
00:09:56.559 --> 00:09:57.919
Do we need a new model?
00:09:58.080 --> 00:10:00.480
Um we do see a lot of improvements every time.
00:10:01.519 --> 00:10:05.840
Do you see that there's a performance improvement or decrease from model to model?
00:10:06.000 --> 00:10:10.480
It feels like, of course, they can do more, but also are they more efficient than they used to be before?
00:10:11.120 --> 00:10:18.960
So so so far, um looking at our internal benchmarks, we've always seen a jump every time we've released a new model.
00:10:19.120 --> 00:10:30.480
Um, some might be smaller, um, for example, 4.6 and 4.7, there was an increase, it was just a little bit smaller than what we've seen with 4.5, 2.4.6.
00:10:30.879 --> 00:10:34.480
But every time there is um uh somewhat of an improvement.
00:10:34.639 --> 00:10:38.960
Um token efficiency is a different um different story.
00:10:39.039 --> 00:10:44.080
We do see uh all kinds of degrees every time a new model is being um released.
00:10:45.120 --> 00:10:50.240
Is there a way to connect a local model that I have already running on my machine with something like Olama?
00:10:50.879 --> 00:11:02.720
So we have in VS Code bring your own model, um, bring your own key, so you can um bring in, you can connect through your own API key to these models, so you can bring any kind of model that you would want.
00:11:03.120 --> 00:11:12.000
Um the nice thing for built-in models like GBD 5.5, if you don't bring your own model, and we do optimize our coding hardness for these models.
00:11:12.080 --> 00:11:17.919
That's why we put, especially before launch, so much time into making sure the system prompt is updated.
00:11:18.080 --> 00:11:26.399
The system prompt works um great with these new models because we have a little bit of time before the launch to make sure it really works end-to-end in hours.
00:11:26.480 --> 00:11:29.759
So that's one of the advantages of using the built-in ones.
00:11:29.840 --> 00:11:32.559
But we you can always bring whatever model you want as well.
00:11:34.000 --> 00:11:36.000
Can you say something about numbers?
00:11:36.399 --> 00:11:38.879
How many people do that, or am I a freak for doing that?
00:11:39.200 --> 00:11:40.799
I actually I actually don't no no no.
00:11:40.879 --> 00:11:41.600
I think there are.
00:11:41.679 --> 00:11:45.360
I actually don't know the correct number, just to be very transparent.
00:11:45.519 --> 00:11:55.519
But we talk a lot and I see this coming up and a lot of times where people are asking, oh no, I want this specific model because for yeah, whatever reason.
00:11:55.679 --> 00:11:57.200
Um and that's a very fair point.
00:11:57.279 --> 00:11:58.879
I just don't know the number, unfortunately.
00:11:58.960 --> 00:12:00.159
So so you're not a freak.
00:12:00.480 --> 00:12:13.120
We are getting the question more and more, and I see a lot of also um issues in now we as code um repo where people are asking constantly about hey, we would like support for these kind of um models or edge cases.
00:12:13.200 --> 00:12:13.600
So yeah.
00:12:14.559 --> 00:12:19.519
As an engineer yourself, how does that how does that that challenge feel?
00:12:19.759 --> 00:12:33.519
Like you seem to be, it seems like you're you're you're beholden to a third-party provider who does the models, and oftentimes models get like, oh, we might release it, we might release it, we might release it, and then like the newest model is the best thing ever.
00:12:33.840 --> 00:12:37.759
And how much time do you get from like as a as an integrator?
00:12:37.919 --> 00:12:42.240
Do you get preview time or do you also get it when it's when it's finally out there?
00:12:42.480 --> 00:12:48.080
Um so we typically get like the two weeks-ish weeks before a model launch.
00:12:48.159 --> 00:12:54.960
And usually what we do during this time of the two weeks is one very important thing, we are optimizing our coding harness for it.
00:12:55.120 --> 00:13:03.279
So, most importantly, we are looking at our system prompt and make sure that it works with the new nuances, how the new model works.
00:13:03.519 --> 00:13:06.240
Um we do this by running offline evolts.
00:13:06.480 --> 00:13:24.240
We constantly, like in this two weeks, we run so many offline evolts, we are tweaking certain system prompts, we are tweaking certain things in our coding harness to make sure the quality keeps going up, the resolution rate of how many of our test cases are being solved just goes up, so we are providing the best quality overall.
00:13:24.480 --> 00:13:36.639
At the same time, or in parallel, we're still doing a lot of internal dog fooding and testing, and which is interesting because these are the two types of feedback the model providers are really looking to us as the integrators.
00:13:36.720 --> 00:13:43.679
They're always asking about um eval numbers and they are also asking about what we kind of call vibe checks.
00:13:44.000 --> 00:13:48.799
Um so they are they love some quotes that we are getting from our internal dog fooders.
00:13:48.960 --> 00:13:57.440
So before model launch, I go in, I have a fun project that I'm currently working on, and I just test the new model and I'm like, do this.
00:13:57.600 --> 00:14:01.120
Sometimes we get some recommendations from the model providers.
00:14:01.279 --> 00:14:06.720
For example, GBD 5.5, they've done a lot of work into more front-end dev work.
00:14:06.799 --> 00:14:10.879
So they said, hey, our research team focused a lot on more front-end stuff.
00:14:10.960 --> 00:14:20.080
So I test a lot a lot of different tasks about front-end web development to see if whatever they claim to be better at, it truly holds off.
00:14:20.399 --> 00:14:23.120
And then we are sharing this kind of feedback with them as well.
00:14:23.200 --> 00:14:34.080
So they get a we're all hoping the model in providers including that the quality holds and um it truly reaches the promises what they are saying.
00:14:34.480 --> 00:14:44.559
Can you do a direct A B test in VS Code that you say, like, okay, I got this task, run both of the models, and then I will take the better result?
00:14:45.679 --> 00:14:48.799
Um so that's where sessions come in in VS Code.
00:14:49.039 --> 00:14:52.799
So um there we have different types of agents.
00:14:52.879 --> 00:14:55.919
We have local agents or um remote agents.
00:14:56.080 --> 00:15:10.159
Remote agents use um GitHub Copilot CLI, so you can just send it off, it does its work, and it comes back, or your local agent, you're in your editor, you see these local changes also being um applied at the same time.
00:15:10.399 --> 00:15:31.759
So typically what I do is I spin off different local sessions and I ask the same prompt across different models, across different um yeah, across different models, and see what they come up with, and then I let faith or whatever I feel like in that moment kind of decide.
00:15:31.840 --> 00:15:34.000
So I kind of do my own A-B testing as well.
00:15:35.279 --> 00:15:40.879
Is there anything I can do as an outside contributor to help with like integrating the new things?
00:15:40.960 --> 00:15:49.600
Is there a different integration in uh in Canary, for example, in VS Code Canary, or is it just that's more about UX of the system itself?
00:15:50.080 --> 00:15:54.879
Um so all of our system prompts for VS Code, because we are open source, they are public.
00:15:55.360 --> 00:16:19.120
Um, and every model provider gets their own system prompt because how these models behave, um so anybody can look at the system prompt, especially once we release a new model, everyone can immediately see what the new model prompt looks like, and you can kind of also tell and see what are some of the changes that has been applied that the model provider has shared, and we are always open for community suggestions.
00:16:19.279 --> 00:16:31.200
So if you have done your local testing and you feel like, hey, with maybe we want to make some or tweaks to the system prompt based on yours, we are always open for these kind of suggestions as well.
00:16:31.759 --> 00:16:41.519
Um so yeah, that's usually what I recommend people to look into and also to familiarize themselves a little bit more with these different models as well.
00:16:42.000 --> 00:16:47.840
What is a request you always get that you think like people this should be something that should be obvious?
00:16:48.000 --> 00:16:52.639
What is a stumbling block that people have that that comes in as a as a question to you?
00:16:52.960 --> 00:16:56.399
Usually the the very first one I get is what's the best model?
00:16:57.039 --> 00:17:11.680
Um and I wish I had this here is the best model, but you already touched on a point where different models are good for different certain tasks, like dev front-end versus back-end dev.
00:17:12.079 --> 00:17:20.880
Um so I typically, including myself or looking at the VS Code engineering team, I always recommend switching models from time to time.
00:17:21.119 --> 00:17:44.960
And also always please use the latest, just because they are better better, and we do see a lot of enterprises still using some of the older models as well, which um it makes sense maybe sometimes for them because of CELA requirements or enterprise requirements, but I always recommend um go ideally to the latest, just because we do know it holds the best of all quality.
00:17:45.519 --> 00:17:51.680
Um so yeah, usually it's the question well which which is the best model I should be using.
00:17:51.839 --> 00:17:55.680
And um, how does it feel that?
00:17:55.759 --> 00:18:08.480
I mean, that might be a spicy question again, but it always cracks me up that there's so many editors out there from different companies that are just a fork of Visual Studio Code with security turned off, or like uh or like special features.
00:18:08.559 --> 00:18:15.279
Like I love that Monaco and the environment that Visual Studio Code is has become the blueprint for all of them.
00:18:15.599 --> 00:18:22.880
But it does feel weird to say, like, oh, you gotta have start with this editor to be a good developer or with this editor.
00:18:23.039 --> 00:18:30.000
Like, how does it feel to be like in the OG of the of the development environment and other people like, yeah, I'm using that one, not touching yours anymore?
00:18:30.720 --> 00:18:37.200
Of course, I'm not gonna lie, of course, it hurts sometimes where you're like, ah, you know, it is a fork of MPS code.
00:18:37.359 --> 00:18:41.599
Um the re but also the reality is we've always been open source.
00:18:41.680 --> 00:18:43.119
The team has always been open source.
00:18:43.200 --> 00:18:48.240
So open source has always been a core part of the team culture of how we think of the product.
00:18:48.559 --> 00:18:56.480
So we always knew there could have been the opportunity, like the possibility of somebody forking our repo and making their own.
00:18:56.640 --> 00:19:03.599
So as much as it does hurt from time to time, we still value the open source community and we also benefit a lot.
00:19:03.680 --> 00:19:06.559
Like we get PRs from the community themselves.
00:19:06.720 --> 00:19:16.799
So there's always a like pros and cons towards being um open source with yeah, people taking that also for their own advantages.
00:19:17.039 --> 00:19:36.640
But overall, we just value the open source community mindset, and I think nothing has changed towards that mindset, and we're still very proud that we we love that open transparency, um and we don't want to give up on it, so there's never been any talks in the team that to making it um close to anything.
00:19:37.279 --> 00:19:40.799
It feels weird when one of the forks get becomes closed source.
00:19:40.960 --> 00:19:42.799
That's really the weird thing about it.
00:19:42.960 --> 00:19:48.319
And I I want to call that out because I worked on VS Code and Edge and Edge inside VS Code.
00:19:48.400 --> 00:19:53.920
There was an extension that we did, and a lot of people feel like they can't contribute from the outside.
00:19:54.079 --> 00:20:07.039
They're either not they feel they're not good enough because Microsoft has thousands of board developers sitting around that will do it anyways, or they feel like, why should I do it for a company like Microsoft or or why should I support that they have enough money?
00:20:07.279 --> 00:20:09.279
It's a great way to learn, I think.
00:20:09.440 --> 00:20:25.039
Contributing to an open source project, especially a big one like that, is a great way to learn as a developer, and it's also every contribution I found, at least in my team back then, was very much appreciated, and there is no dumb contribution that feels simple enough.
00:20:25.440 --> 00:20:26.720
No, totally agree.
00:20:26.960 --> 00:20:50.000
Um, and I think even especially with AI nowadays, um, I mean I joined the team half a year ago, and I feel like I I feel very empowered with AI because the first thing I did is I um phoned via Scode Repo and I asked, explain this to me, and in this way I was able to narrow down how certain things first of all work.
00:20:50.160 --> 00:21:09.440
Um so it's so much easier to grasp a big GitHub repo to better understand how it works, but then also to make these small contributions and start small, pick one issue that is open that the team has somebody assigned to, but we always love these contributions and we are a small team overall.
00:21:09.839 --> 00:21:13.680
Um so it is really seen and it is um appreciated.
00:21:13.839 --> 00:21:20.960
We call out these PRs even in our stand-up from time to time, where we are like, hey, somebody contributed it.
00:21:21.039 --> 00:21:23.519
This is one of the reasons why we are open source.
00:21:23.759 --> 00:21:33.279
Um so I always recommend use now that we have AI, use the power of AI to understand how it works and then make a small first contribution.
00:21:33.359 --> 00:21:35.279
It is seen definitely across the team.
00:21:35.680 --> 00:21:35.920
Cool.
00:21:36.079 --> 00:21:36.960
Well, thanks very much.
00:21:37.119 --> 00:21:38.880
This was insightful, this was good.
00:21:39.119 --> 00:21:41.599
Um are you coming to our event in September?
00:21:41.839 --> 00:21:43.119
I mean I hope so.
00:21:43.279 --> 00:21:47.119
Um so yeah, we haven't finalized times yet, but I would love to be there for sure.
00:21:47.359 --> 00:21:55.359
Yeah, and if people have any questions about models inside uh like AI models inside VS Code, that's the right person to talk to.
00:21:56.000 --> 00:21:56.799
Thanks very much.
00:21:57.039 --> 00:21:57.839
Thank you.