WEBVTT
00:00:00.080 --> 00:00:07.839
In today's episode, what if we could use our enterprise software only through prompts?
00:00:08.080 --> 00:00:14.000
What if we had a framework to evaluate the quality of the output from LLMs and AI?
00:00:14.400 --> 00:00:20.480
And what if we gave total control to AI of our computer?
00:00:56.000 --> 00:00:58.799
I'm gonna cover three short stories.
00:00:59.920 --> 00:01:07.599
The first one will be about a framework to evaluate the output from AI.
00:01:07.760 --> 00:01:14.799
This article that I found is focused on evaluating voice agents, but I think this has implications for AI overall.
00:01:15.359 --> 00:01:27.920
Then we'll cover what if what could happen and what could be the value of leaving your computer open for AI to use.
00:01:28.319 --> 00:01:34.159
And finally, what if we could use our enterprise software look more like a prompt?
00:01:34.480 --> 00:01:35.760
So let's start.
00:01:36.079 --> 00:01:48.400
So I've I haven't read the whole articles, of course, but I'm just using that as let's say inspirational um elements and to discuss the implications on the user experience level.
00:01:48.560 --> 00:02:06.159
Um that is really helpful to me because sometimes I I suffer from um analysis paralysis, so sometimes it's even more helpful to me to only read some of it and discuss what if what would be the implications on a user experience level.
00:02:06.480 --> 00:02:13.360
Okay, so first and foremost, we have this new framework for evaluating voice agents, so it's called Eva.
00:02:14.400 --> 00:02:18.960
And this is an article that was found that I found on Hugging Face.
00:02:19.759 --> 00:02:35.439
Um, and so this is a framework to evaluate end-to-end um converse conversational conversations with voice agents, um, and so that evaluates multi-turn spoken conversations using realistic bot-to-bot architecture.
00:02:35.680 --> 00:02:40.000
And so, what is interesting is that there are several scores here.
00:02:40.159 --> 00:02:55.360
We have accuracy, so we have let's say the physic the physical aspects of your interaction, in this case, accuracy, and we also have experience, and so how was it perceived by the end user?
00:02:57.199 --> 00:03:08.400
And so I find this really interesting in um in the fact that we are combining finally first and foremost.
00:03:08.479 --> 00:03:13.680
I didn't know that we didn't have any let's say framework for evaluating voice agents.
00:03:13.759 --> 00:03:16.080
I honestly thought we had.
00:03:16.240 --> 00:03:28.080
I worked on that for a year, so I reviewed all the literature that is linked to um how can we make conversations with artificial agents more natural?
00:03:28.240 --> 00:03:39.599
And so, for instance, I discovered a whole bunch of things that we humans do that make our conversations natural and acceptable, such as back channels.
00:03:39.680 --> 00:03:41.680
So, this is a thing that I discovered.
00:03:41.919 --> 00:03:55.439
So when you speak to someone, you are awaiting for cues from this person that manifest that they are understanding you, that they're following you, that they that they they degree or not with what you said, but at least that they provide feedback to you.
00:03:55.599 --> 00:04:04.560
And in absence of that, the conversation looks um really let's say artificial or um or maybe uh eerie.
00:04:05.039 --> 00:04:22.959
And so, yeah, there is a need to let's say take some inspiration from this human-to-human interaction and applying it to human to artificial to some extent, because we can go to the to the uh uncanny valley for those who are not familiar.
00:04:23.199 --> 00:04:40.240
It it describes this sensation, this feeling we would have the more or interaction with um artificial agent, let's say mimic the ones from humans, but at the same time they do not have the same capabilities as humans.
00:04:40.319 --> 00:04:47.360
So once we understand that these are artificial, it creates an eerie sensation, which is cold, which is described by the Incanny Valley.
00:04:48.079 --> 00:05:18.399
Anyways, so I found this to be really interesting because I oftentimes find in our frameworks a really strong emphasis on measuring either the experience or so the experience that a user um has with their AI, their agent or their product, or the other side, like measuring the physical aspect.
00:05:18.480 --> 00:05:24.639
I like to call it physical aspect, it's not necessarily physical, it's like the ingredients, the criteria, the attributes that you put in your product.
00:05:24.720 --> 00:05:26.079
That's what I'm referring to.
00:05:26.319 --> 00:05:33.040
And so what I find interesting is that this framework combines both.
00:05:33.439 --> 00:05:54.720
So for instance, it has accuracy, and by accuracy we understand task completion, faithfulness, measures whether the agents' responses were grounded in its instructions, policies, user inputs, and tool call results, speech, fidelity, measures whether the speech system faithfully reproduced the intent text in spoken audio.
00:05:54.959 --> 00:06:10.160
And then we have experience, which ultimately is really interesting because okay, so it's experience, but I don't know to be honest how do they measure that.
00:06:10.319 --> 00:06:29.759
It looks like okay, they measure conciseness, measures whether the agent's responses were appropriately brief and focused for spoken delivery, conversational progression, measures whether the agent moved the conversation forward effectively, and turntaking measures whether the agent spoke at the right time.
00:06:30.399 --> 00:06:34.240
Neither interrupting the user nor introducing excessive silence.
00:06:34.560 --> 00:06:42.800
Okay, so that's really interesting because I don't know exactly how do they measure those the experience ones.
00:06:43.120 --> 00:06:52.240
Is it really measured with end users or is it measured with the designer?
00:06:52.879 --> 00:07:14.399
Um, so at first glance it looks like and it it almost isn't doesn't matter to make for me to make my point here, and I'm gonna have a full episode dedicated to that, to the importance of having three steps when we evaluate our experiences of whatever we conceive.
00:07:14.720 --> 00:07:34.079
It looks like what they call experience, the measure of experience, is not maybe I'm wrong, and I will correct if needed in a f in a next episode, but this is not um measured with end users, or maybe they use it themselves.
00:07:34.319 --> 00:07:40.800
Um yeah, I'm not finding that information right now, anyways.
00:07:41.040 --> 00:07:46.560
So it's interesting because whatever the result, I would say it sparks conversation.
00:07:46.720 --> 00:07:50.800
So when you create something, you need to evaluate it on several fronts.
00:07:51.199 --> 00:08:07.680
So I'd say if you want to elicit an experience, let's say increase in trust or increase in satisfaction, you should know what you should put in your product that will elicit these aspects.
00:08:07.920 --> 00:08:23.839
So you should be really clear in okay, if I put this, let's say this turn taking or this um voice, it will be perceived as more trustworthy, and so my user will trust it more.
00:08:24.079 --> 00:08:45.120
It's like this three-step evaluation framework that I know is not new, but I'm really advocating for let's say focusing on that even more with AI because we are developing products that we don't know a the intended, like let's say the expected outcome it will have on the user experience.
00:08:45.279 --> 00:08:51.279
So it's it's a pretty much like reverse engineering a really good recipe.
00:08:52.159 --> 00:08:59.840
Let's say you really like a soda and you or your users really like your soda, but you don't know why they like it.
00:08:59.919 --> 00:09:14.639
So if you want to reproduce that, or if you or if you know they they like it to some extent, but not to the extent you were expecting, you would like to know what is the delta that what should what should you be working on?
00:09:14.720 --> 00:09:16.000
It's like anything in life.
00:09:16.159 --> 00:09:37.440
If you are producing some outcome and output at your company and you want to improve, you must know where you stand in relationship to the extremes, and then how is your outcome tied to your output so that you can know what you should change in your output?
00:09:37.679 --> 00:09:55.360
So it's the same, it's like we need to define what to put as ingredients inside of the recipe, then we need a way to evaluate the recipe internally before our consumer eats the recipe, and then we need to evaluate with them again and compare that.
00:09:55.519 --> 00:10:14.399
So that's that's the that's why I found this framework interesting, but still it really looks like maybe I'm wrong, but it really looks like they still evaluate that with only the designers or the the the ones who conceive the model, and it would be great to separate to some extent.
00:10:14.480 --> 00:10:18.720
So there are several implications to the framework that I'm proposing, and it's not new, by the way.
00:10:18.799 --> 00:10:30.320
We are doing heuristic analysis since the beginning of time, we do uh usability inspections uh internally with with teams uh before releasing products since the beginning of time.
00:10:30.399 --> 00:10:33.360
I'm just saying we should emphasize that more.
00:10:34.000 --> 00:10:37.519
But I will have a dedicated episode on that.
00:10:38.000 --> 00:10:41.679
Okay, then it's just a reflection to spark conversation.
00:10:41.840 --> 00:10:48.960
Uh, what I saw yesterday, if I'm not mistaken, we can now let Claude use our computer in co-work.
00:10:49.360 --> 00:11:08.080
Uh so it looks like you can connect, you can authorize Cloud to access some of your of your files and folders so that even through the app, once you're away, um you can chat with it and ask it to message you a presentation because you are not ready for your presentation.
00:11:08.240 --> 00:11:17.039
So I'm really I'm really torn between being amazed and at the same time being kind of doubtful.
00:11:18.399 --> 00:11:26.000
Amazed because of the technology, of course, and at the same time, it's a bit like it's a bit the same as what I said yesterday.
00:11:26.320 --> 00:11:35.759
I think right now there's kind of a conflation between between what the technology can do and what we should authorize it to do.
00:11:36.000 --> 00:11:36.639
I don't know.
00:11:36.799 --> 00:11:56.240
I feel like there is a mix at this stage between Yeah, LLMs are here, so because they are here, it means we are opening the gates to a lot of other things, such as controlling my computer, such as recording people like in yesterday's in yesterday's issue.
00:11:56.320 --> 00:12:20.639
If you if you listen to it this episode, we should open the gates to whatever we um we could just because so I don't know, maybe it's really helpful to some people, but I I I I am struggling to understand why because we have the technology which is LLMs, why does that mean we open the gates to everything else?
00:12:20.960 --> 00:12:32.799
Like to me, these are not things that go hand in hand necessarily, it can, but not necessarily, and I see a push towards the end of privacy.
00:12:34.000 --> 00:12:37.759
I don't know, every time a little bit more, the end of privacy.
00:12:38.000 --> 00:12:55.600
So I give the LLM permission to record my voice, whatever I do, I prompt it with my voice, I record my tra my meetings, and I give and I analyze the transcripts automatically with LLM, and it goes into a server and it analyzes everything and it uses that to improve their model.
00:12:55.759 --> 00:13:01.279
Maybe not at all times, but I prefer to assume that it does always.
00:13:01.519 --> 00:13:02.799
That's my rule of thumb.
00:13:02.960 --> 00:13:08.879
Once you have doubts, always assume that your data is being used to improve whatever product they are releasing.
00:13:08.960 --> 00:13:20.320
And I'm not anti, I work in user experience, so I think that this is necessary to some extent, but I think it's good to know where you're leaving your data.
00:13:20.480 --> 00:13:21.679
I think this is important.
00:13:22.159 --> 00:13:27.120
Like users should know where they're leaving their data, and so it's the same with co-work.
00:13:27.200 --> 00:13:29.919
I really I'm not affiliated or whatever.
00:13:30.000 --> 00:13:48.320
I really like cloud, I love what they're doing, I love the products, and I love their their commitment to to improving, but I just don't understand why there is this direction of um, of course, you can be in control, you can use the folders that you give cloud access to.
00:13:48.559 --> 00:13:53.519
Um yeah, I I I don't know.
00:13:53.600 --> 00:13:55.360
Um I'm just leaving that open.
00:13:55.519 --> 00:14:11.360
It's not so much a commentary, maybe it is, but it's more of a reflection, like philosophically speaking, almost, and maybe even more user experience speaking, like on the level of trust.
00:14:12.480 --> 00:14:18.960
Do we really want to leave open or computer it for instance?
00:14:19.039 --> 00:14:40.559
I'm I'm just wondering why is there not maybe there are, I think there are, why is there not more push for private for private LLMs to which we can speak with a self-hosted server, with a self-hosted LLM on your server, which does exactly the same thing.
00:14:40.879 --> 00:14:45.360
And I see a trend over and over and over again.
00:14:45.600 --> 00:14:58.559
This was the case with Google, um, where we tend to value, we tend to value way more convenience than privacy.
00:14:58.799 --> 00:15:10.799
So it's like even if I know that Google is analyzing my data left and right, I know that I know that the product is superior, and I know that the product is superior just because of that.
00:15:12.639 --> 00:15:14.559
And so I'm willing to make the trade.
00:15:14.799 --> 00:15:20.639
And I I I think this is a pattern that is being repeated with LLMs and way, way more.
00:15:20.799 --> 00:15:40.000
It will have way more way more impact because we talk to these LLMs and we give them access to files, and so ultimately I think this is way more impactful in terms of how quickly they can learn.
00:15:41.600 --> 00:16:00.480
So that's why probably so sorry, I'm thinking out loud and I'm and I'm really sharing my thoughts as they come, but that I'm thinking that's probably why there is so much push for this mix between using or LLMs and giving them access to everything.
00:16:00.799 --> 00:16:14.480
Because ultimately they will be superior to private ones, because private ones don't have as much data to be trained on, and so ultimately these commercial ones will be more convenient because they do the job better.
00:16:14.799 --> 00:16:23.600
So it's like it's it's it's not even comparing apples to apples, I feel like comparing a private LLM to a to a commercial one.
00:16:24.159 --> 00:16:34.000
Maybe there is an in-between, which is you are a pro and you can train that so much that um even if it's private, it does the job perfectly.
00:16:34.399 --> 00:16:42.559
But I really would like to see more emergence of private LLMs that do this kind of task.
00:16:42.799 --> 00:16:50.879
Okay, and finally we have a news sharing a startup who wants to make enterprise software more look like a prompt.
00:16:51.519 --> 00:17:09.519
So that's the story of Josh Siroda, who founded the startup Aragon back in August and has just raised 12 million at a hundred million post-money valuation to build an agentic AI operating system for enterprise customers.
00:17:10.000 --> 00:17:14.559
They say the simple thesis is software is dead.
00:17:14.640 --> 00:17:16.960
So this is an article from TechCrunch.
00:17:17.279 --> 00:17:25.920
Sierota says buttons and dialogue boxes and pull-down menus are a thing of the past, and future business will be done by prompt.
00:17:26.480 --> 00:17:36.079
Aragon is attempting to offer the whole suite of business software, Salesforce, Snowflakes, Tableau, and Jira's through LLM interface.
00:17:36.480 --> 00:17:53.680
Sierra, who worked on go-to-market teams at Oracle and Salesforce, admits to suffering a bit of quarter life crisis in the lead up to moving to San Francisco and launching Aragon with a small team from a live workloft across the street from the Giants baseball park.
00:17:54.000 --> 00:17:56.000
So, okay.
00:17:57.759 --> 00:18:03.599
Well, here my thoughts like I'm not speaking even about the product, like Aragon.
00:18:03.680 --> 00:18:07.920
I I haven't I have no idea about the product itself.
00:18:08.640 --> 00:18:12.799
I'm I'm I'm looking at it as I speak.
00:18:13.039 --> 00:18:18.799
So if we go on our website, it's described as a proprietary AI powering the world of bits.
00:18:18.880 --> 00:18:23.680
It's an OS, they say it's enterprise AI OS.
00:18:24.640 --> 00:18:31.920
Um at the foundation of every company are bits, ones and zeros created every second, stored across every system.
00:18:32.400 --> 00:18:34.720
These bits grow exponentially every second.
00:18:34.799 --> 00:18:37.440
Together they make up our entire business.
00:18:38.559 --> 00:18:43.440
No one has ever been able to see all of it, connect all of it, act on all of it until now.
00:18:43.599 --> 00:18:45.519
So that's the vision, apparently.
00:18:46.000 --> 00:18:46.319
Okay.
00:18:47.039 --> 00:18:56.079
So my unbiased this is this is of course um ironical.
00:18:56.400 --> 00:19:06.960
Um my opinion as a user experience researcher, interested in the human, maybe a little bit more than the technology, but also in the technology.
00:19:07.119 --> 00:19:09.839
That's probably why I'm a user experience researcher.
00:19:10.160 --> 00:19:37.359
I would posit that this is normal, what we are seeing, because it's like when you enter a new territory, you need to map this territory, you need to map the extremes, you need to go to all of the edges of this territory so that you can readjust and settle where where this is maybe less shaky or less dangerous or less less uncertain.
00:19:37.599 --> 00:19:50.960
And so I feel we are entering this era of yeah, we have some new toys, new capabilities, which is AI, and we need to experiment with all of it and see what sticks.
00:19:51.920 --> 00:19:54.480
So that's what I'm seeing with these kind of things.
00:19:54.640 --> 00:20:01.599
Um, I'm not saying that this is wrong or right or whatever, I'm just uh describing what I'm seeing.
00:20:01.920 --> 00:20:13.279
So, in my opinion, when we interact with objects, there are so many senses that are that are solicited.
00:20:13.599 --> 00:20:26.400
We have vision, touch, we have um and and and in senses here, so what is integrated in our perception I will also include other things.
00:20:27.519 --> 00:20:39.519
So like memory, feelings, goals, like in this model that I'm describing, even that is an input, I would say.
00:20:40.480 --> 00:20:44.720
And so then you need to make a decision and you need to act on it.
00:20:45.759 --> 00:21:01.359
And once you act, it's the same, you have a thousand ways to do something, and these thousand ways they are competing with each other, with also including your experience.
00:21:01.519 --> 00:21:07.519
So, for instance, if I'm an enterprise software user, I might have a goal.
00:21:07.920 --> 00:21:19.359
Let's say I might input the the the uh how can I say the payroll data of my employees?
00:21:19.519 --> 00:21:26.079
Okay, so I might have to do that, and I have several ways to do it, and this is what we are seeing right now.
00:21:26.240 --> 00:21:33.359
I might do it with voice, I might do it typing, I might do it by clinking, I I might do it by clicking on buttons.
00:21:35.039 --> 00:21:37.759
Ultimately, it will all depend on my experience, also.
00:21:38.240 --> 00:21:45.440
Am I a new payroll specialist or am I an experienced one with 15 years of experience?
00:21:45.759 --> 00:21:55.039
So that also adds to the complexity how regularly do I need to do it, and so on and so forth.
00:21:55.200 --> 00:21:58.000
How to what extent do I trust technology?
00:21:59.039 --> 00:21:59.759
And I would say.
00:22:00.400 --> 00:22:05.039
That it's not so clear to me that we should use only one modality.
00:22:05.279 --> 00:22:11.440
Like to to make it short, my conclusion is it's not so clear.
00:22:11.759 --> 00:22:16.480
Do we really want to only interact with technology with prompting?
00:22:16.799 --> 00:22:19.200
That's my question that I want to leave here.
00:22:20.480 --> 00:22:25.519
Because if it's the case, how restricting would that be?
00:22:26.720 --> 00:22:34.000
Like let me tell you, I developed I'm developing a website recently and I'm using only AI prompts.
00:22:34.079 --> 00:22:43.279
I was I was really torn between using AI prompts only, well sorry, prompts to LLMs, and using a drag and drop builder.
00:22:43.519 --> 00:22:55.920
And I know to some extent how to code, how to code, sorry, how to yeah, how to come up with a website, but I would say it's really basic HTML and CSS, and I'm not a web designer.
00:22:56.240 --> 00:23:06.559
So to get the job done, I was hesitating between drag and drop builder and and um the use of LLMs.
00:23:06.720 --> 00:23:23.599
Knowing that right now on the market, I looked at it, and it looks like if you want to build your website with an LLM and then and then edit it by drag and drop, which would be the most efficient to me, we don't have such product.
00:23:23.759 --> 00:23:24.960
At least I couldn't find any.
00:23:25.119 --> 00:23:28.880
If you happen to know of any, please let me know.
00:23:29.119 --> 00:23:30.799
Um let me know.
00:23:30.880 --> 00:23:34.000
Uh I think on Spotify you can comment on the show.
00:23:34.240 --> 00:23:38.319
So let me know because I couldn't find any, and it's really really frustrating.
00:23:38.559 --> 00:23:45.920
It looks like companies again and again and again, like I think it's not companies, it's it's maybe the the mindset.
00:23:46.319 --> 00:24:05.359
We tend to always put products first before needs, and so that's what I'm seeing probably with this kind of um with this kind of um let's say take that we could have the enterprise software looking more like a prompt.
00:24:05.440 --> 00:24:12.160
So it's it's really it's really a it's really a mindset saying that you can do everything by just prompting.
00:24:12.960 --> 00:24:25.359
Well, if I take the example of website design, I don't know if I need to change one thing, I need to prompt it and wait for its answer, whereas it would be way easier to do it by hand and drag and dropping.
00:24:25.680 --> 00:24:33.680
So I don't know, I'm not sure that I'm not sure that this would work, to be honest.
00:24:33.759 --> 00:24:35.359
I might be wrong, let's see.
00:24:36.160 --> 00:24:40.640
But at least it would be a learning experiment for this company, even if it doesn't work.
00:24:40.960 --> 00:24:50.160
Um but I'm just saying that having only one modality to interact with technology can kind of feels restrictive, to be honest.
00:24:51.279 --> 00:25:02.640
And sometimes, for having spoken with a lot of users throughout the 10-ish years of experience that I have in user experience, I can tell you people want to do things a certain way.
00:25:02.880 --> 00:25:06.400
And if you take that away from them, they are not happy.
00:25:06.640 --> 00:25:09.599
And sometimes we do want to have control over things.
00:25:10.079 --> 00:25:28.400
So if you don't leave your users an exit door or a retry or and that is part of Nielsen heuristics, well, they are not happy, and understandably so, because you're placing barriers between them and the job they have to accomplish.
00:25:28.559 --> 00:25:32.240
And so restricting things to one modality could be one of those barriers.
00:25:32.319 --> 00:25:33.920
I don't know, I'm just saying.
00:25:34.480 --> 00:25:36.720
So that's it for today's episode.
00:25:36.799 --> 00:25:37.759
These three news.
00:25:37.839 --> 00:25:42.000
I hope you liked it, and I hope you learned at least one thing, or at least that it's part conversations.
00:25:42.079 --> 00:25:45.359
I'm super happy to um have people disagree with me.
00:25:45.519 --> 00:25:47.039
Let me know in the comments.
00:25:47.279 --> 00:25:54.079
Um, I don't have the full, full, full knowledge of um what's behind these articles and these news.
00:25:54.240 --> 00:26:03.039
Uh the there may be way more, let's say, smart people around that could um um let's say compliment what I'm saying.
00:26:03.119 --> 00:26:06.960
So if it's the case, please comment and I would learn from that.
00:26:07.119 --> 00:26:08.880
So until then, take care.
00:26:09.039 --> 00:26:10.240
See you tomorrow.
00:26:10.400 --> 00:26:10.799
Cheers.
00:00:00.080 --> 00:00:07.839
In today's episode, what if we could use our enterprise software only through prompts?
00:00:08.080 --> 00:00:14.000
What if we had a framework to evaluate the quality of the output from LLMs and AI?
00:00:14.400 --> 00:00:20.480
And what if we gave total control to AI of our computer?
00:00:56.000 --> 00:00:58.799
I'm gonna cover three short stories.
00:00:59.920 --> 00:01:07.599
The first one will be about a framework to evaluate the output from AI.
00:01:07.760 --> 00:01:14.799
This article that I found is focused on evaluating voice agents, but I think this has implications for AI overall.
00:01:15.359 --> 00:01:27.920
Then we'll cover what if what could happen and what could be the value of leaving your computer open for AI to use.
00:01:28.319 --> 00:01:34.159
And finally, what if we could use our enterprise software look more like a prompt?
00:01:34.480 --> 00:01:35.760
So let's start.
00:01:36.079 --> 00:01:48.400
So I've I haven't read the whole articles, of course, but I'm just using that as let's say inspirational um elements and to discuss the implications on the user experience level.
00:01:48.560 --> 00:02:06.159
Um that is really helpful to me because sometimes I I suffer from um analysis paralysis, so sometimes it's even more helpful to me to only read some of it and discuss what if what would be the implications on a user experience level.
00:02:06.480 --> 00:02:13.360
Okay, so first and foremost, we have this new framework for evaluating voice agents, so it's called Eva.
00:02:14.400 --> 00:02:18.960
And this is an article that was found that I found on Hugging Face.
00:02:19.759 --> 00:02:35.439
Um, and so this is a framework to evaluate end-to-end um converse conversational conversations with voice agents, um, and so that evaluates multi-turn spoken conversations using realistic bot-to-bot architecture.
00:02:35.680 --> 00:02:40.000
And so, what is interesting is that there are several scores here.
00:02:40.159 --> 00:02:55.360
We have accuracy, so we have let's say the physic the physical aspects of your interaction, in this case, accuracy, and we also have experience, and so how was it perceived by the end user?
00:02:57.199 --> 00:03:08.400
And so I find this really interesting in um in the fact that we are combining finally first and foremost.
00:03:08.479 --> 00:03:13.680
I didn't know that we didn't have any let's say framework for evaluating voice agents.
00:03:13.759 --> 00:03:16.080
I honestly thought we had.
00:03:16.240 --> 00:03:28.080
I worked on that for a year, so I reviewed all the literature that is linked to um how can we make conversations with artificial agents more natural?
00:03:28.240 --> 00:03:39.599
And so, for instance, I discovered a whole bunch of things that we humans do that make our conversations natural and acceptable, such as back channels.
00:03:39.680 --> 00:03:41.680
So, this is a thing that I discovered.
00:03:41.919 --> 00:03:55.439
So when you speak to someone, you are awaiting for cues from this person that manifest that they are understanding you, that they're following you, that they that they they degree or not with what you said, but at least that they provide feedback to you.
00:03:55.599 --> 00:04:04.560
And in absence of that, the conversation looks um really let's say artificial or um or maybe uh eerie.
00:04:05.039 --> 00:04:22.959
And so, yeah, there is a need to let's say take some inspiration from this human-to-human interaction and applying it to human to artificial to some extent, because we can go to the to the uh uncanny valley for those who are not familiar.
00:04:23.199 --> 00:04:40.240
It it describes this sensation, this feeling we would have the more or interaction with um artificial agent, let's say mimic the ones from humans, but at the same time they do not have the same capabilities as humans.
00:04:40.319 --> 00:04:47.360
So once we understand that these are artificial, it creates an eerie sensation, which is cold, which is described by the Incanny Valley.
00:04:48.079 --> 00:05:18.399
Anyways, so I found this to be really interesting because I oftentimes find in our frameworks a really strong emphasis on measuring either the experience or so the experience that a user um has with their AI, their agent or their product, or the other side, like measuring the physical aspect.
00:05:18.480 --> 00:05:24.639
I like to call it physical aspect, it's not necessarily physical, it's like the ingredients, the criteria, the attributes that you put in your product.
00:05:24.720 --> 00:05:26.079
That's what I'm referring to.
00:05:26.319 --> 00:05:33.040
And so what I find interesting is that this framework combines both.
00:05:33.439 --> 00:05:54.720
So for instance, it has accuracy, and by accuracy we understand task completion, faithfulness, measures whether the agents' responses were grounded in its instructions, policies, user inputs, and tool call results, speech, fidelity, measures whether the speech system faithfully reproduced the intent text in spoken audio.
00:05:54.959 --> 00:06:10.160
And then we have experience, which ultimately is really interesting because okay, so it's experience, but I don't know to be honest how do they measure that.
00:06:10.319 --> 00:06:29.759
It looks like okay, they measure conciseness, measures whether the agent's responses were appropriately brief and focused for spoken delivery, conversational progression, measures whether the agent moved the conversation forward effectively, and turntaking measures whether the agent spoke at the right time.
00:06:30.399 --> 00:06:34.240
Neither interrupting the user nor introducing excessive silence.
00:06:34.560 --> 00:06:42.800
Okay, so that's really interesting because I don't know exactly how do they measure those the experience ones.
00:06:43.120 --> 00:06:52.240
Is it really measured with end users or is it measured with the designer?
00:06:52.879 --> 00:07:14.399
Um, so at first glance it looks like and it it almost isn't doesn't matter to make for me to make my point here, and I'm gonna have a full episode dedicated to that, to the importance of having three steps when we evaluate our experiences of whatever we conceive.
00:07:14.720 --> 00:07:34.079
It looks like what they call experience, the measure of experience, is not maybe I'm wrong, and I will correct if needed in a f in a next episode, but this is not um measured with end users, or maybe they use it themselves.
00:07:34.319 --> 00:07:40.800
Um yeah, I'm not finding that information right now, anyways.
00:07:41.040 --> 00:07:46.560
So it's interesting because whatever the result, I would say it sparks conversation.
00:07:46.720 --> 00:07:50.800
So when you create something, you need to evaluate it on several fronts.
00:07:51.199 --> 00:08:07.680
So I'd say if you want to elicit an experience, let's say increase in trust or increase in satisfaction, you should know what you should put in your product that will elicit these aspects.
00:08:07.920 --> 00:08:23.839
So you should be really clear in okay, if I put this, let's say this turn taking or this um voice, it will be perceived as more trustworthy, and so my user will trust it more.
00:08:24.079 --> 00:08:45.120
It's like this three-step evaluation framework that I know is not new, but I'm really advocating for let's say focusing on that even more with AI because we are developing products that we don't know a the intended, like let's say the expected outcome it will have on the user experience.
00:08:45.279 --> 00:08:51.279
So it's it's a pretty much like reverse engineering a really good recipe.
00:08:52.159 --> 00:08:59.840
Let's say you really like a soda and you or your users really like your soda, but you don't know why they like it.
00:08:59.919 --> 00:09:14.639
So if you want to reproduce that, or if you or if you know they they like it to some extent, but not to the extent you were expecting, you would like to know what is the delta that what should what should you be working on?
00:09:14.720 --> 00:09:16.000
It's like anything in life.
00:09:16.159 --> 00:09:37.440
If you are producing some outcome and output at your company and you want to improve, you must know where you stand in relationship to the extremes, and then how is your outcome tied to your output so that you can know what you should change in your output?
00:09:37.679 --> 00:09:55.360
So it's the same, it's like we need to define what to put as ingredients inside of the recipe, then we need a way to evaluate the recipe internally before our consumer eats the recipe, and then we need to evaluate with them again and compare that.
00:09:55.519 --> 00:10:14.399
So that's that's the that's why I found this framework interesting, but still it really looks like maybe I'm wrong, but it really looks like they still evaluate that with only the designers or the the the ones who conceive the model, and it would be great to separate to some extent.
00:10:14.480 --> 00:10:18.720
So there are several implications to the framework that I'm proposing, and it's not new, by the way.
00:10:18.799 --> 00:10:30.320
We are doing heuristic analysis since the beginning of time, we do uh usability inspections uh internally with with teams uh before releasing products since the beginning of time.
00:10:30.399 --> 00:10:33.360
I'm just saying we should emphasize that more.
00:10:34.000 --> 00:10:37.519
But I will have a dedicated episode on that.
00:10:38.000 --> 00:10:41.679
Okay, then it's just a reflection to spark conversation.
00:10:41.840 --> 00:10:48.960
Uh, what I saw yesterday, if I'm not mistaken, we can now let Claude use our computer in co-work.
00:10:49.360 --> 00:11:08.080
Uh so it looks like you can connect, you can authorize Cloud to access some of your of your files and folders so that even through the app, once you're away, um you can chat with it and ask it to message you a presentation because you are not ready for your presentation.
00:11:08.240 --> 00:11:17.039
So I'm really I'm really torn between being amazed and at the same time being kind of doubtful.
00:11:18.399 --> 00:11:26.000
Amazed because of the technology, of course, and at the same time, it's a bit like it's a bit the same as what I said yesterday.
00:11:26.320 --> 00:11:35.759
I think right now there's kind of a conflation between between what the technology can do and what we should authorize it to do.
00:11:36.000 --> 00:11:36.639
I don't know.
00:11:36.799 --> 00:11:56.240
I feel like there is a mix at this stage between Yeah, LLMs are here, so because they are here, it means we are opening the gates to a lot of other things, such as controlling my computer, such as recording people like in yesterday's in yesterday's issue.
00:11:56.320 --> 00:12:20.639
If you if you listen to it this episode, we should open the gates to whatever we um we could just because so I don't know, maybe it's really helpful to some people, but I I I I am struggling to understand why because we have the technology which is LLMs, why does that mean we open the gates to everything else?
00:12:20.960 --> 00:12:32.799
Like to me, these are not things that go hand in hand necessarily, it can, but not necessarily, and I see a push towards the end of privacy.
00:12:34.000 --> 00:12:37.759
I don't know, every time a little bit more, the end of privacy.
00:12:38.000 --> 00:12:55.600
So I give the LLM permission to record my voice, whatever I do, I prompt it with my voice, I record my tra my meetings, and I give and I analyze the transcripts automatically with LLM, and it goes into a server and it analyzes everything and it uses that to improve their model.
00:12:55.759 --> 00:13:01.279
Maybe not at all times, but I prefer to assume that it does always.
00:13:01.519 --> 00:13:02.799
That's my rule of thumb.
00:13:02.960 --> 00:13:08.879
Once you have doubts, always assume that your data is being used to improve whatever product they are releasing.
00:13:08.960 --> 00:13:20.320
And I'm not anti, I work in user experience, so I think that this is necessary to some extent, but I think it's good to know where you're leaving your data.
00:13:20.480 --> 00:13:21.679
I think this is important.
00:13:22.159 --> 00:13:27.120
Like users should know where they're leaving their data, and so it's the same with co-work.
00:13:27.200 --> 00:13:29.919
I really I'm not affiliated or whatever.
00:13:30.000 --> 00:13:48.320
I really like cloud, I love what they're doing, I love the products, and I love their their commitment to to improving, but I just don't understand why there is this direction of um, of course, you can be in control, you can use the folders that you give cloud access to.
00:13:48.559 --> 00:13:53.519
Um yeah, I I I don't know.
00:13:53.600 --> 00:13:55.360
Um I'm just leaving that open.
00:13:55.519 --> 00:14:11.360
It's not so much a commentary, maybe it is, but it's more of a reflection, like philosophically speaking, almost, and maybe even more user experience speaking, like on the level of trust.
00:14:12.480 --> 00:14:18.960
Do we really want to leave open or computer it for instance?
00:14:19.039 --> 00:14:40.559
I'm I'm just wondering why is there not maybe there are, I think there are, why is there not more push for private for private LLMs to which we can speak with a self-hosted server, with a self-hosted LLM on your server, which does exactly the same thing.
00:14:40.879 --> 00:14:45.360
And I see a trend over and over and over again.
00:14:45.600 --> 00:14:58.559
This was the case with Google, um, where we tend to value, we tend to value way more convenience than privacy.
00:14:58.799 --> 00:15:10.799
So it's like even if I know that Google is analyzing my data left and right, I know that I know that the product is superior, and I know that the product is superior just because of that.
00:15:12.639 --> 00:15:14.559
And so I'm willing to make the trade.
00:15:14.799 --> 00:15:20.639
And I I I think this is a pattern that is being repeated with LLMs and way, way more.
00:15:20.799 --> 00:15:40.000
It will have way more way more impact because we talk to these LLMs and we give them access to files, and so ultimately I think this is way more impactful in terms of how quickly they can learn.
00:15:41.600 --> 00:16:00.480
So that's why probably so sorry, I'm thinking out loud and I'm and I'm really sharing my thoughts as they come, but that I'm thinking that's probably why there is so much push for this mix between using or LLMs and giving them access to everything.
00:16:00.799 --> 00:16:14.480
Because ultimately they will be superior to private ones, because private ones don't have as much data to be trained on, and so ultimately these commercial ones will be more convenient because they do the job better.
00:16:14.799 --> 00:16:23.600
So it's like it's it's it's not even comparing apples to apples, I feel like comparing a private LLM to a to a commercial one.
00:16:24.159 --> 00:16:34.000
Maybe there is an in-between, which is you are a pro and you can train that so much that um even if it's private, it does the job perfectly.
00:16:34.399 --> 00:16:42.559
But I really would like to see more emergence of private LLMs that do this kind of task.
00:16:42.799 --> 00:16:50.879
Okay, and finally we have a news sharing a startup who wants to make enterprise software more look like a prompt.
00:16:51.519 --> 00:17:09.519
So that's the story of Josh Siroda, who founded the startup Aragon back in August and has just raised 12 million at a hundred million post-money valuation to build an agentic AI operating system for enterprise customers.
00:17:10.000 --> 00:17:14.559
They say the simple thesis is software is dead.
00:17:14.640 --> 00:17:16.960
So this is an article from TechCrunch.
00:17:17.279 --> 00:17:25.920
Sierota says buttons and dialogue boxes and pull-down menus are a thing of the past, and future business will be done by prompt.
00:17:26.480 --> 00:17:36.079
Aragon is attempting to offer the whole suite of business software, Salesforce, Snowflakes, Tableau, and Jira's through LLM interface.
00:17:36.480 --> 00:17:53.680
Sierra, who worked on go-to-market teams at Oracle and Salesforce, admits to suffering a bit of quarter life crisis in the lead up to moving to San Francisco and launching Aragon with a small team from a live workloft across the street from the Giants baseball park.
00:17:54.000 --> 00:17:56.000
So, okay.
00:17:57.759 --> 00:18:03.599
Well, here my thoughts like I'm not speaking even about the product, like Aragon.
00:18:03.680 --> 00:18:07.920
I I haven't I have no idea about the product itself.
00:18:08.640 --> 00:18:12.799
I'm I'm I'm looking at it as I speak.
00:18:13.039 --> 00:18:18.799
So if we go on our website, it's described as a proprietary AI powering the world of bits.
00:18:18.880 --> 00:18:23.680
It's an OS, they say it's enterprise AI OS.
00:18:24.640 --> 00:18:31.920
Um at the foundation of every company are bits, ones and zeros created every second, stored across every system.
00:18:32.400 --> 00:18:34.720
These bits grow exponentially every second.
00:18:34.799 --> 00:18:37.440
Together they make up our entire business.
00:18:38.559 --> 00:18:43.440
No one has ever been able to see all of it, connect all of it, act on all of it until now.
00:18:43.599 --> 00:18:45.519
So that's the vision, apparently.
00:18:46.000 --> 00:18:46.319
Okay.
00:18:47.039 --> 00:18:56.079
So my unbiased this is this is of course um ironical.
00:18:56.400 --> 00:19:06.960
Um my opinion as a user experience researcher, interested in the human, maybe a little bit more than the technology, but also in the technology.
00:19:07.119 --> 00:19:09.839
That's probably why I'm a user experience researcher.
00:19:10.160 --> 00:19:37.359
I would posit that this is normal, what we are seeing, because it's like when you enter a new territory, you need to map this territory, you need to map the extremes, you need to go to all of the edges of this territory so that you can readjust and settle where where this is maybe less shaky or less dangerous or less less uncertain.
00:19:37.599 --> 00:19:50.960
And so I feel we are entering this era of yeah, we have some new toys, new capabilities, which is AI, and we need to experiment with all of it and see what sticks.
00:19:51.920 --> 00:19:54.480
So that's what I'm seeing with these kind of things.
00:19:54.640 --> 00:20:01.599
Um, I'm not saying that this is wrong or right or whatever, I'm just uh describing what I'm seeing.
00:20:01.920 --> 00:20:13.279
So, in my opinion, when we interact with objects, there are so many senses that are that are solicited.
00:20:13.599 --> 00:20:26.400
We have vision, touch, we have um and and and in senses here, so what is integrated in our perception I will also include other things.
00:20:27.519 --> 00:20:39.519
So like memory, feelings, goals, like in this model that I'm describing, even that is an input, I would say.
00:20:40.480 --> 00:20:44.720
And so then you need to make a decision and you need to act on it.
00:20:45.759 --> 00:21:01.359
And once you act, it's the same, you have a thousand ways to do something, and these thousand ways they are competing with each other, with also including your experience.
00:21:01.519 --> 00:21:07.519
So, for instance, if I'm an enterprise software user, I might have a goal.
00:21:07.920 --> 00:21:19.359
Let's say I might input the the the uh how can I say the payroll data of my employees?
00:21:19.519 --> 00:21:26.079
Okay, so I might have to do that, and I have several ways to do it, and this is what we are seeing right now.
00:21:26.240 --> 00:21:33.359
I might do it with voice, I might do it typing, I might do it by clinking, I I might do it by clicking on buttons.
00:21:35.039 --> 00:21:37.759
Ultimately, it will all depend on my experience, also.
00:21:38.240 --> 00:21:45.440
Am I a new payroll specialist or am I an experienced one with 15 years of experience?
00:21:45.759 --> 00:21:55.039
So that also adds to the complexity how regularly do I need to do it, and so on and so forth.
00:21:55.200 --> 00:21:58.000
How to what extent do I trust technology?
00:21:59.039 --> 00:21:59.759
And I would say.
00:22:00.400 --> 00:22:05.039
That it's not so clear to me that we should use only one modality.
00:22:05.279 --> 00:22:11.440
Like to to make it short, my conclusion is it's not so clear.
00:22:11.759 --> 00:22:16.480
Do we really want to only interact with technology with prompting?
00:22:16.799 --> 00:22:19.200
That's my question that I want to leave here.
00:22:20.480 --> 00:22:25.519
Because if it's the case, how restricting would that be?
00:22:26.720 --> 00:22:34.000
Like let me tell you, I developed I'm developing a website recently and I'm using only AI prompts.
00:22:34.079 --> 00:22:43.279
I was I was really torn between using AI prompts only, well sorry, prompts to LLMs, and using a drag and drop builder.
00:22:43.519 --> 00:22:55.920
And I know to some extent how to code, how to code, sorry, how to yeah, how to come up with a website, but I would say it's really basic HTML and CSS, and I'm not a web designer.
00:22:56.240 --> 00:23:06.559
So to get the job done, I was hesitating between drag and drop builder and and um the use of LLMs.
00:23:06.720 --> 00:23:23.599
Knowing that right now on the market, I looked at it, and it looks like if you want to build your website with an LLM and then and then edit it by drag and drop, which would be the most efficient to me, we don't have such product.
00:23:23.759 --> 00:23:24.960
At least I couldn't find any.
00:23:25.119 --> 00:23:28.880
If you happen to know of any, please let me know.
00:23:29.119 --> 00:23:30.799
Um let me know.
00:23:30.880 --> 00:23:34.000
Uh I think on Spotify you can comment on the show.
00:23:34.240 --> 00:23:38.319
So let me know because I couldn't find any, and it's really really frustrating.
00:23:38.559 --> 00:23:45.920
It looks like companies again and again and again, like I think it's not companies, it's it's maybe the the mindset.
00:23:46.319 --> 00:24:05.359
We tend to always put products first before needs, and so that's what I'm seeing probably with this kind of um with this kind of um let's say take that we could have the enterprise software looking more like a prompt.
00:24:05.440 --> 00:24:12.160
So it's it's really it's really a it's really a mindset saying that you can do everything by just prompting.
00:24:12.960 --> 00:24:25.359
Well, if I take the example of website design, I don't know if I need to change one thing, I need to prompt it and wait for its answer, whereas it would be way easier to do it by hand and drag and dropping.
00:24:25.680 --> 00:24:33.680
So I don't know, I'm not sure that I'm not sure that this would work, to be honest.
00:24:33.759 --> 00:24:35.359
I might be wrong, let's see.
00:24:36.160 --> 00:24:40.640
But at least it would be a learning experiment for this company, even if it doesn't work.
00:24:40.960 --> 00:24:50.160
Um but I'm just saying that having only one modality to interact with technology can kind of feels restrictive, to be honest.
00:24:51.279 --> 00:25:02.640
And sometimes, for having spoken with a lot of users throughout the 10-ish years of experience that I have in user experience, I can tell you people want to do things a certain way.
00:25:02.880 --> 00:25:06.400
And if you take that away from them, they are not happy.
00:25:06.640 --> 00:25:09.599
And sometimes we do want to have control over things.
00:25:10.079 --> 00:25:28.400
So if you don't leave your users an exit door or a retry or and that is part of Nielsen heuristics, well, they are not happy, and understandably so, because you're placing barriers between them and the job they have to accomplish.
00:25:28.559 --> 00:25:32.240
And so restricting things to one modality could be one of those barriers.
00:25:32.319 --> 00:25:33.920
I don't know, I'm just saying.
00:25:34.480 --> 00:25:36.720
So that's it for today's episode.
00:25:36.799 --> 00:25:37.759
These three news.
00:25:37.839 --> 00:25:42.000
I hope you liked it, and I hope you learned at least one thing, or at least that it's part conversations.
00:25:42.079 --> 00:25:45.359
I'm super happy to um have people disagree with me.
00:25:45.519 --> 00:25:47.039
Let me know in the comments.
00:25:47.279 --> 00:25:54.079
Um, I don't have the full, full, full knowledge of um what's behind these articles and these news.
00:25:54.240 --> 00:26:03.039
Uh the there may be way more, let's say, smart people around that could um um let's say compliment what I'm saying.
00:26:03.119 --> 00:26:06.960
So if it's the case, please comment and I would learn from that.
00:26:07.119 --> 00:26:08.880
So until then, take care.
00:26:09.039 --> 00:26:10.240
See you tomorrow.
00:26:10.400 --> 00:26:10.799
Cheers.