Paul Swanson: And what we've discovered is parents do not question us about rigor. Parents question us about clarity. You know, if you're getting a normal distribution of achievement, you're not doing a job properly. Because the normal distribution is what nature gives us. Students differ in their aptitudes, and if we treat all students the same, you will get the bell curve. Our job as teachers is to destroy the bell curve by making sure that those who need more opportunities to be successful get those opportunities. Listening to Data Talks, the podcast from Data in Schools, in association with Atlanta International School and the Mona Letra Effect, and sponsored by Veracross and IntelliSchool. I'm Paul Swanson, Director of Innovation and Institutional Research at Atlanta International School and part of the Data in Schools team. And I'm Matthew Savage, Educator, Leader, Consultant, and Coach. Thomas Skusky, PhD, is Professor Emeritus at the College of Education, University of Kentucky. He is a fellow in the American Educational Research Association and widely known for his research in educational measurement and evaluation, student assessment, grading and reporting, and teacher change. Dylan William is Emeritus Professor of Educational Assessment at the UCL Institute of Education in London, but now lives in Florida. He has taught in urban public schools, served in senior leadership in higher education. And pursued a research program focused on effective, scalable teacher professional development. So welcome to our podcast, the place where Data Talks, we listen. Tom and Dylan, so great to have you here on Data Talks. ⁓ let's just jump into the conversation. Tom, one of the core arguments in your work, many of your books, is that grading reform keeps failing because schools skip one of the central questions of what is a grade actually for? And Dylan, your decision driven model takes the same argument for a different direction that we should start with the decision we need to make and then ask what evidence would serve that. Why is it that schools struggle so much to have that intention set up at the beginning? And what does a school look like when it actually is able to do that? Well, I think Dylan and I are very much of the same mind on this particular question in particular. We really believe that too many school leaders have been seduced by a charming and persuasive and very well intentioned consultants that they should be addressing what questions. And so they ask, what's the best way to grade? Or what's the best way to assess? ⁓ what's the best way to teach reading or basic literary skills? ⁓ what's the science of learning? What's the effect size? And too often we bypass fundamental issues, the most crucial being Why are we doing this in the first place? I mean, why do we have grades? What purpose should they serve? ⁓ what decisions do we want to make with assessment results? And what's the best evidence for making those decisions? I think if we could move toward a more fundamental approach to addressing these why questions first, we'd be far better off when we came to addressing the what questions. I think there's another angle as well, which I think too often we fail in all kinds of educational reform. And that is to understand that education is a system. And too often we try to change one component of the system, not realizing it has a knock on effect on the other parts of the system. So Tom and people like Doug Reeves have been saying for twenty, thirty years that giving zeros for missing work is really stupid. And it's absolutely right. And I you know, I say, well, why stop there? Why not give minus twenty if you really want to show who who's boss? So why does it persist? Doug and Tom have been pointing out the flaws in this idea for decades and yet it still persists. And I think it's crucial to realise that every assessment, every piece of information have has an ev evidential and a consequential component. So the the reason that teachers like giving zeros for missing work is it's a sanction. It's a sanction they hope never to impose. But at least it means that it's worthwhile the student turning in something rather than nothing. And if we're going to stop teachers giving us zeros for missing grades, missing work, then we're going to have to give s teachers some other tools to make sure that students do the work. And so too often we try to tweak one component, not realizing it has knock-on effects into all the others. And I think that's for me the major reason that we haven't had much progress over the last twenty, thirty, forty, fifty years. As well as t talking about what the grade is for, I w I wonder whether sometimes there's a a a challenge that we haven't worked out what the thing is that we want to grade. For instance, w what exactly do we mean by progress? What exactly do we mean by by growth? What exactly do we mean by achievement? And now the measurement of well being comes into it all. What exactly do we mean by that? Do you think that's an issue as well, not being clear enough about the thing we want to measure? Just as what the grade actually what purpose it actually serves? I think teachers are clear, but they don't agree. So I was very surprised when I first worked with teachers in the US. I was working with a middle school science teacher, and they were doing a long project, and they were spending some time in the library doing res research. And thirty percent of the points for the project were for behavior in the library. And that really surprised me because that's just so totally alien to the European tradition of assessment where you only assess intellectual merit. But of course, that teacher decided that that was necessary to get the students to behave in the library. And so teachers often very explicit about this. They'll say there's points for this, there's points for this. They're very clear about what the points are being given for. They just don't agree with the teacher in the classroom next door to them. So you're describing like a But the grades have been a proxy for behavior management in both of those instances. The the zero mark if it's not handed in and the ⁓ behavioral component to the criteria, the data assessment becomes ⁓ partly at least a form of managing students' behavior, which sounds kind of perverse ⁓ f from from our end, right? But that's the point about the consequential versus the evidential basis. So Giving zeros for missing work makes no sense at all if you're trying to describe the achievement of the student, but it makes perfect sense if you want to give some adverse consequences for not doing what the teacher wants. Yes. Though you make a very valid point, Matthew, the idea being that we stress grades should be primarily to serve the purpose of communication. We're trying to communicate with students and with parents and families. ⁓ about students learning progress in school. But for many teachers it's seen also as a device for control. you can screw up, you can mess around, you cannot do what I told you you have to do. But the bottom line, I have the grade and I can fail you. And so it's getting away from that aspect of using it to control ⁓ students or to compel them to act in ways that we think are more defensible that gets in our way when it and it alters that perspective. Services device. And I think the communication argument is very powerful and it's very easily addressed by disaggregating the achievement from the behavior. I think most parents would prefer to get a grade for pure achievement and a separate report on the child's behavior. You know, I don't want my very, very smart child getting an A- just because their behavior was poor. I want to know they got an A for achievement and their behavior was lousy. And so I think disaggregating these components and reporting separately would be very valuable. And we could do that within a standards-based grading framework. The point is the standards could be about behaviour as least as much as they're about achievement. And what you alluded to there, Tom, as as well, was the idea that we assess in order to control. And that what we want is control and power. And the greatest weapon in our arsenal is the grade that we can withhold or give or inflate or reduce. ⁓ but it's it's pointing to quite a a a sinister motivation on the part of some, hopefully a minority, but some in the field. Well, that's true. And we we do even when we looked at these various types of greeting criteria that teachers use outside of the recognition of academic achievement. there is this notion of we call learning enablers. Learning enablers represent ⁓ not necessarily measures of learning per se, but they enable learning process. Formative assessments, for example, are learning enablers. Doing your homework is a learning enabler. Participation in class will be a learning enabler. Then we have another set of entirely different characteristics that relate to the social and emotional learning aspects of it. So here's where things like persistence and collaboration, ⁓ a growth mindset would come into play. But there's a third area, and that we've just generally leveled compliance. And compliance is did you do what I told you you had to do? And that's where this control aspect gets in. So did you turn it in on time? Did you behave in class? ⁓ all these things. And and granted those are important, but they're clearly different from how well a student achieves specified academic goals or learning standards. And to be able to indicate those separately, I think is the most important dimension that we can really talk about here. To indicate clearly those are important, but they're they taint a grade that is supposed to reflect how well students have learned and what they're able to do. Aaron Or alternatively, you could actually argue that it's a feature, not a bug. Because the the reason that high school grade point average predicts college freshman performance better than the SAT or the ACT. Apart from at the most elite universities, is because it measures compliance. It measures did you turn stuff in for four years? And that's basically, according to Brian Kaplan, what employers are paying for. We can't find anything that graduates can do that non-graduates can't do. What you get with a degree, the reason that employers are willing to pay 30% more for a graduate, is because there's a reasonable amount of intelligence, a reasonable amount of conscientiousness. But also a quite a large component of conformity. And so I joke that no assessment system makes sense when viewed from the outside. And the grading practices in American schools support grading practices in American colleges, which in turn support what employers want when they actually employ graduates. So again, it comes down to this systemic nature. We can't change one component without changing all the other components, and that's why it's hard and why things bounce back. Because we didn't actually change the things that were impacted by our policy reform. Dylan, you've said in the past that using assessment for summative purposes, so grading, sorting, ranking, etc., gets in the way of the very learning that in theory we're trying to assess. You've also said that telling students their level after every piece of work is, and I think I quote you correctly, ⁓ bizarre and perverse. And that's you've also said that scores are better than grades. And you've you've questioned the very nature of reliable data. What does it even mean for data to be reliable? And is that a grail worth ⁓ chasing? But we're also in a like a ⁓ an ecosystem where grades are going nowhere, ⁓ certainly nowhere soon, at at least. And I wonder what is the re what is the ⁓ ideal relationship between all of those constituent parts? And essentially what should end up on the report card at the end of the whole process. I think the whole point about the report card is very telling, because I do think it's appropriate to report absolute levels of achievement at the end of sequences of learning. The question is how frequently should you do that? And I think once a year is fine in elementary school, maybe once a semester in middle school, and maybe once a quarter in high school. For me. The interesting thing is what kind of assessments should occur between those terminal assessments? And I think you know, we made a mistake using the word summative rather than terminal, because what we actually want is assessments at the end of sequences of learning. But those are not really going to serve any learning purpose. I joke that school report cards in the US are ways of spoiling a child's summer vacation because it arrives too late to do anything. And everybody's forgotten about it by the time the school reconvenes in September. So it's just basically, as Douglas Reeves points out, it's not so much a medical, it's more like a post-mortem. And so we need to start thinking about doing these things separately. So I'm not against grades. I'm against giving grades and feedback designed to improve performance at the same time. So my advice to teachers is do one or the other. The point about levels. That you picked up, ⁓ Matthew. That comes from the English context where there were these national curriculum levels. And the system was designed so that students made one level progress every two years. And we had this bizarre system of reporting levels on each piece of work. And of course it didn't change. For most students they got the same grade or level every time, and we see that in America. We see students getting D's time after time after time. And the reason I think that's so damaging is because it feeds back to that student this idea that you are a D student. It doesn't no, you actually made progress. You now can do things you couldn't do before, but you still get the same grade. And I think that is demotivating to students. And so I think there's a real problem here with just the scale that we use for reporting achievement. If the number that students are getting doesn't change, it feeds into this idea of a fixed mindset. And I think we have to kind of unpack that somewhat and see if we can move towards More appropriate ways of reporting, just to document the progress that students have made. And there's quite a lot of research from around the world that shows that personal reference norms, comparing you with how you used to be, is far more motivating than social reference norms, how you stack up against the rest of the class. And you referenced specifically national curriculum levels there as the levels that we're referring to. But it is very similar now. Lots of international schools will use ⁓ sort of quasi mastery language to describe the level a child is at. And they will talk about a child either being emerging, developing, working out or exceeding. And you have the scenario not dissimilar from what you just described, where a child could, for the entirety of their elementary school ⁓ journey, be emerging like some perpetual kind of pupa or chrysalis. ⁓ Which again seems to me rather entrapment and rather caging th than anything that enables them to kind of flourish and and and grow. Nobody's fooled by this, by the way. You know, the kids look at emerging, mm, developing, mastery. It's one, two, and three. Students convert it straight away to one, two, or three. So it's fooling nobody. We actually have evidence on that. And Dylan's point is exactly on target. The labels Don't matter. here in the state of Kentucky where I live, when some of the reforms were passed early on, they decided that it would be a really corrupt process to give letter grades to elementary children. And so most of the schools in the state decided to use a different sort of categories for these levels of achievement. They decided to use the same categories as we use in a state assessment. So the state assessment classifies student performance at one of four levels. You can score a novice, apprentice, proficient, or distinguished. Several years into the that program, we did a very large scale study where we interviewed parents about how they were perceiving these particular grade categories. Every single parent we interviewed translated it directly, distinguished as an A, proficient as a B. We accomplished nothing. So changing the labels does not change the interpretation. It's bringing meaning to those interpretations. And w this is also to stress Dylan's point as well with report cards, this is an area where technology is working against us. Because almost all schools in the United States at least now are using computerized grading programs. And these computerized grading programs are not developed by educators and are not developed with a sense of what has been accomplished in the research in the area. They're developed by software engineers. And so almost all the grades included on report cards are determined by mathematical algorithms. The typical mathematical algorithm we use is to average. And what this means is you never get a clear picture of where the student is presently. You get this calmidating, cumulative description of where they were over the entire school year. So if a student struggles in the first quarter but does well in the last quarter, that's never noted. Because we are bound to using a mathematical algorithm to determine their grade. One of the biggest things we have to do is to be able to move away from these mechanical, mathematical formulas for determining grades and do something more reasonable, reflecting the purpose of grading, which is to communicate where students are presently, not where they were six months ago. I think the other thing to add to that is that there's always going to be a trade off here. So Tom has eloquently argued in the past for Reflecting where students are at the end of the semester or the marking period. We shouldn't be penalizing a student for a slow start. And with the current grading systems, it's possible to fail the year before you're halfway through it. So some people have argued for end waiting the assessment. But I think there's a problem there too, because unless we create incentives for students to work steadily throughout the whole semester, then boys in particular are going to Do less work believing that they can kind of phone it in at the last minute. You know, they can just pull it all out at the last minute. And that's maybe good for their end of topic or end of course test. But the the evidence from psychology is the quicker you learn something, the quicker you forget it. So I would think we need to move away from looking for the best system to explicitly acknowledge there are trade-offs in any assessment system. And one of the key ones is the trade off between distributed assessment and synoptic assessment. In other words, do we collect information throughout the whole semester or the whole year, or do we leave it all to the very end? And in Europe, it's very common to leave everything to the end, like in the Abitur or the Baccalauree in France or the A levels in England. There's no credit in the bank until you sit in that hall in June and do an exam. In the American system, most states still have grade-based systems where you do a piece of work, you get a an assessment on it, and you get to keep the A, even if you subsequently forget everything you ever knew to get that A in the first place. And so that's an extremely distributed system. Europe uses extremely synoptic systems. I think we have to explore systems where we do both and understand the trade offs between the different kinds of ways of collecting evidence. And it creates huge technical problems. I mean, how do you how do you trade off achievement in November against achievement in June? Because hopefully, certainly on a cumulative subdiet like math, you've made progress. So hopefully the grade you got in November doesn't actually reflect your mathematical capability six months later. I'm curious, Tom, did you wanna add to that one? No, I I agree completely. And I think that ⁓ What I've been trying to help people understand with the area of grading in particular is that there are two foundational premises upon which we have to establish all our policies and practices. The first being that we do not grade students, we grade performance. And because performance is always temporary, then grades must also be seen as temporary, that they have to be reassigned based on how that level performance changes. And and the second aspect, the second premise, I guess, would be that ⁓ the grades do not reflect who you are as a learner, it reflects where you are in your learning journey. And because where you are is always temporary, then grades must be seen as temporary too. That if if we could use these as our two fundamental premises for grading, that that we grade performance, not students, and that they identify where you are in your learning journey, not how well you can learn, that it would allow us to make great progress in terms of really accurately communicating with parents, families, and with students where they are and how well they've done and how they can improve. And that would give us another advantage, because right now teachers don't use gradebooks for report writing. So if I pick up my my my student gradebook, that tells me nothing about what I need to do with that student. If we can move towards a more standards based system where I know that this student understands the gas laws, but it doesn't understand density properties of matter, then that gives me something to work on. And so for me, the crucial thing is to go much finer than most advocates of standards based grading, towards very fine grained assessments that teachers can use to plan instruction, and then, and this addresses Paul's point, and then aggregate them for different purposes. So, if there's a really important conceptual gap in a student's profile, I don't want to give them the credit for that. I don't want to average that out because I know I need to come back to that to make sure the students understand that thing. But if I'm going to give a summative grade, if I'm going to say overall across the whole semester, this is the level of achievement, then I'm going to forgive that because it's not that relevant for a summative purpose. It is relevant for a formative purpose. And the crucial distinction, the crucial asymmetry here is that you can always take fine-grained assessment evidence and aggregate it to serve a summative purpose, you can't take coarse summative data and disaggregate it to identify learning needs. Imagine a system that understands how independent schools work, where the algebra teacher is also the soccer coach, where parents are alumni, where everyone plays multiple roles. Most school software forces you to juggle separate systems for students, parents, alumni, and donors, fragmenting the people you serve into disconnected data points. It's exhausting and it misses what matters most. Veracross is different. The one-person, one record, comprehensive software. That connects every department with real-time data, all centered on the whole person. No silos, just seamless collaboration across your entire school. Veracross.com, trusted by 3,200 schools worldwide. Imagine what you could achieve with data that works for good. Let's dig into that a little bit more about the idea of aggregation and sort of data reductionism. ⁓ because at some point, Depending on the question that you're asking, you may need to have some sort of data reduction. You know, overall, how did this child do in this subject? How did they do in this? How much knowledge do they have in this? But when you're asking other questions, like you suggest here, what do they still need to learn in science? You have to have a much more fine-grained approach. Based on ⁓ our conversation just before this call, Dylan, it sounded like this may be, or there may be some areas here where you and Tom have some disagreement. Tom, I know you've written a lot about ⁓ or against the idea of using percentiles and the idea of having a hundred different ⁓ shades of grading and and different categories and more towards a finite scope of, you know, five or six grades that there can be more consensus around. Is this a point of disagreement between the two, or are you are there parts of agreement here as well? Well, we've had the wonderful opportunity to present together at several conferences, and this issue has come up. I I always begin this conversation with trying to clarify the difference between types of measurement and particularly direct and indirect measures. But a a direct measure is where we can put a measuring device on an object or a person and measure that trait directly. So for example, if I wish to measure students' height. I would have them stand against the hall, I put a book or a a ⁓ ruler across the top of their head, put a mark on the wall. I'd measure from the floor to the mark to measure their height. If I have a direct measure like that, if my measuring scale is only marked off in say half inches, I can measure their height only to nearest half inch. But if it's marked off in quarter inches or or eighth of an inch, I could do it much more accurately. So when you have direct measures, a finely tuned scale does lead to more accurate measurement. But most of the things we measure in education, especially with regard to students, are not direct measures. They're indirect measures. I I cannot measure a student's achievement directly. So what I have to do is measure something else and then make an inference from that or about what that means in terms of their achievements. So I have them answer a series of questions or perform a series of tasks. And then based on how they do that, I I'm make a judgment or an inference about what level of achievement that represents. Measuring intelligence would be the same. And when you move from direct to indirect measures, then those finely tuned scales really tend to break down. It's very difficult for us to get consistency in that measurement. ⁓ Researchers refer to this as interrator reliability, where we have two comparably skilled, knowledgeable, experienced judges look into the same body of evidence coming up with the same grade. And so I argue for keeping those decisions based on what research we have on reliability down to about four to seven categories. But once you move beyond that, it's much more difficult to establish acceptable levels of interrated reliability. equally competent teachers look into the same body of evidence and student learning it can make the exactly the same grade. Before we unleash the fist fight between ⁓ Dylan and Tom on this topic, it reminds me of ⁓ the the coastline paradox, the idea that if you look at the United Kingdom or any other bees of land anywhere with a coastline, that if you measure it in a hundred kilometer measures, it will be this long. I think the UK, if you measure it in a hundred kilometer measures, I think it the coastline of the UK comes out at like two thousand eight hundred kilometers. But if you measure it with a fifty kilometer measure, I think it comes out at three and a half thousand kilometers. If you measure it with a one kilometer measure, it comes out at eight thousand kilometers. What if you were to measure it by a half a millimeter ⁓ rules? The ⁓ scale would be like beyond comprehension. So I suppose what I'm wondering when you're talking, Tom, is what are the benefits and drawbacks Of either, ⁓ which is more in a vertical as accurate ⁓ than the other? Well, I think it's it comes down to the notion of it being a communication device. I had the great advantage in my life and professional career of having some very wonderful teachers. My biaser, the chair of my doctor's dissertation was Benjamin Bloom. ⁓ and so my perspective on this was. fashioned deeply by what I learned from him and working with him. But Bloom had argued early on that we really, in terms of student assessment, need to just decide one of two grades regarding anything they learn, and that's mastery and not mastery. That we basically, in teaching, want our students to learn those things that we have clearly articulated. Benjamin Bloom's mentor was Ralph Tyler, and in 1949, Ralph Tyler Published this brilliant little book called Basic Principles of Curriculum Instruction, where he argued that before you can teach anyone anything, there are two fundamental decisions you have to make. These are the decisions that Dylan has such a great job in emphasizing as well. First, you have to clearly articulate what you want them to learn and be able to do. And then second, you have to decide what evidence you would accept to verify they learned those things. And so what that means is that. That clarification is in specifying the curriculum. We would never design a curriculum with things in it we expect some students not to be able to learn. That that doesn't make sense. I mean, why would you ever put anything in a curriculum that you expect some students not to be able to learn? So that compels teachers to see their job, their professional responsibility, to have all students learn that excellently. And so the idea of a grading scale would be did you master it or did not master it? And if you didn't master it, then what feedback, what information can we offer to help you reach that level of mastery? Bloom was very clear too, and he was very clever in this way. People always would question him about what does mastery really mean. And he knew that no matter what he specified, people were going to argue with him about it. ⁓ he drew the term mastery from the like middle medieval guilds, where you would begin as an apprentice and Then you would become a a journeyman. But then to reach that that upper level, you had to develop a special piece of work that would be evaluated by a a committee of experts called masters. This work was called a masterpiece. And you would they would be judged, they would use that piece to judge whether you were of the caliber or whether it was sufficient to allow you to enter that that that special class of of master. So when people would question Bloom on what mastery means, he would be turned to them and said, Well, do you give grades? And the teacher would Sure. He But again, do the grades really reflect what students have learned? I mean, don't don't tell me you give A's like the top twenty percent that the based on your standing among classmates. Do the grades really reflect student learning? He said, Yes, they do. He said, Fine. Tell me what you expect for that highest level. I don't have to tell you what mastery is, you've already determined that. So if the highest level is an A, what do students have to do to get an A? And just be clear about that. And what we've discovered is parents do not question us about rigor. Parents question us about clarity. And so Bloom emphasized at that time that we just need to be clear about what mastery is, then our job as educators is pretty clear, having all students reach that mastery level. And if you haven't reached it, you know, you know, one student may still need to learn two things really well. Other student may have to learn four things really well. We individualize the feedback in that way, but the criteria for performance is already established, and our goals advost entry set level. Matthew, to come back to the issue about the f the coastline paradox, I it's not a paradox. It's just people who don't understand fractal geometry. So you can actually take what is called Koch's snowflake curve. So you take an equilateral triangle, and then Mark off thirds on each side and build another triangle on that. And then for each of those sides you just made, build another triangle on that. And you keep on doing that forever. And you end up with something with a curve that has that actually fills space in some way. So it's it it the the curve is actually thick. And so that's why th the paradox, as people call it, emerges, is because the finer you measure, the longer you get. But it's not it's not the same problem as the problem of measurement. So I think there are two issues in the problem of measurement that Tom has identified. The first is that different teachers disagree about what they're giving credit for. So there's a very clear example from Northern Ireland where they still have selective schools, they have secondary modern schools and they have elite grammar schools. And if you ask teachers to assess work, then the secondary modern teachers will give far more weight to ⁓ expressive flair. While the tr traditional academic grammar school teachers will give far more weight to spelling, punctuation, grammar usage and mechanics. And so you think you've got a difference of opinion about what this piece of work is worth, but that's because there's a two-dimensional model and these pieces vary in two different ways, and they're giving different weights to different components. So let's let's simplify that. Let's just say we all agree about what it is that gets better when a student gets better. The question then is, how accurately can we determine that? And I've I've got a lot of support for the for the bench basic Benjamin Bloom model. Mastery or not, and that's where the whole the whole work on formative assessment comes from. You make a decision about whether the class is ready to move on. Have they mastered it or not? And I think Larry Ainsworth's principles about is it foundational and is it enduring are really valuable. Is this something that students will need to know before they move on? Like the fact that matter is made of very small particles? Or is it something that's actually less important, like the phases of the moon, which doesn't have anything building on that in the next two or three years? The other point is, is it enduring? Is it something that the adult will be disadvantaged if they don't learn? And so the big idea of do I move on or do I not move on? Very simple decision, it's twofold, mastery, non-mastery. I think that's a very good starting point. Once we go beyond that, I'm less pessimistic than Tom is about the possibility of consensus. So we had a very interesting system in England whereby teachers would give a grade from a scale from A down to G, or U for ungraded, on a portfolio of work across English language and English literature. And what was interesting was, because teachers discussed these portfolios of work because they brought samples of evidence, there was extraordinary consistency across the country about what this portfolio work was worth. And so I think if we are willing to do the exemplification work, we could actually have teachers meaningfully distinguishing at least ten different levels of achievement for something as messy as English language arts. Here's the interesting thing, they won't be able to agree why they agree. So they can agree that this portfolio is worth a C, that one's worth a D, but they will disagree about the reasons they gave that weighting. And so we can have finer grades if we want to, but I think the fundamental problem I have with the ABCD scale is that people tend to treat those grades as if they're perfectly accurate. And so my point is that of course scores suffer from spurious precision. Seventy three isn't really better than seventy two. But I think grades suffer from spurious accuracy. People think that B's are somehow qualitatively better than C's. And my big concern, and I don't know how to do this, but I do want to get people talking about measures having error in education. And I think before you put any weight on an assessment, you need to know how accurate it is. If I've got an electronic thermometer, my baby's got a fever, I give them I I take that temperature and I get ninety-eight point two. Okay, I feel okay, but just to be sure I do it again and now I get a hundred and four. Okay, my faith in that measuring instrument is now shattered. I I n I it's not giving me any kind of stability at all. And so for me, the faith that we put in any assessment should depend on its stability or its accuracy, its reliability to use the classic assessment term. And I think we don't do enough in terms of reporting the accuracy or the reliability of our assessments. And I think it's easier to do that with finer scales. So you can say your child scored 75 plus or minus 10. I don't know how to say your child scored B plus or minus half a grade. So I think that I I'm not saying which side of the argument we should come down on. What I'm saying is there's going to be a trade-off. And the trade-off we're making currently is. by going for grades, I think, plays down at the inaccuracy in these grades, which I think leads to unfortunate consequences. I think just to be able to bring them into discussion and have teachers talk about these sort of things is the main point. When recent years we've done a number of studies where we have actually interviewed parents and families and students about grading, grading issues. And when we ask the question about ⁓ what they consider the most crucial aspect in fairness to grading, or what they consider most unfair, ⁓ it always comes down to the inconsistency in grading policies and practices among teachers in the same school. So I think what what Dylan and I are both talking about is how can we gain greater consistency in this, a greater agreement about what we're looking for, so that we have that idea that in order to help students improve, we need to be able to give them very specific and well targeted feedback on how to make those improvements. And that means that we need to be consistent in the kind of things we're looking for. And then we can alter our prescriptions, our correctives, if you will, based on individual student needs to help them reach that level of mastery. So I I think it more comes down to the notion of consistency. I think we're both we're both aligned with that sort of thing too. But Paul, picking up on the point you made earlier about decisions versus data. I think most of the information that parents get arrives too late. So I encourage schools to do focus groups with parents. What would you like to know about your child? And when would you like to know it? Because the report card arrives too late. So one recommendation that I've made to schools is every two weeks, a parent gets a check sheet which just minus, equals, or plus for each subject, not as good as expected, about at expectation or exceeding expectation. And if you have concerns, come and talk to us. You know, don't write lengthy reports that take ages to get through the process of spell checking and everything else and then get to the parents. Just give parents quick information so that they can quickly intervene. Is it one subject or is a child going off the boil in two or three subjects? Is it a more general problem? But to come back to what would you like to know about your child and when would you like to know it? Start with the decisions you want people to make and then figure out what data will help them make those decisions more effectively. Your school measures more than any one person can see. Attendance in one system, assessment in another, well-being in a third, built by three different companies. Each holds part of the story. IntelliSchool brings them together into one governed foundation and connects the AI your school chooses. Staff ask questions in their own words and see only what their role allows. Nobody's professional judgment gets replaced, it just gets better informed. IntelESchool.co So I'm I'm really fascinated ⁓ and it prompted by what you said about trade-offs, ⁓ Dylan, I'm really fascinated often by what grades fail to capture. I was speaking with one school recently about wouldn't it be great if we gave two grades for each subject? One the examined grade, i.e., in the confines of what what this examination will measure, you are at a this. And then the other one, an holistic grade. based on everything I've seen and the interactions we've had across the subject, way beyond the syllabus, I reckon you're here. Then being more honest about what the two grades mean, rather than every grade essentially being the examined grade. And I wonder what that misses out. And when I think about what things grades miss out, I I know, Tom, from your what we know about grading synthesis during on all that research and and Dylan, your 2024 book, naming the problem precisely. We need better evidence before we can make better decisions. But I wonder about this question: Is there a version of what a child is? Their growth, their belonging, their becoming, that the whole apparatus of grading and reporting is simply not designed to see. And therefore, if we're presenting a report or a transcript, As the truth about this child. If there are enough things that this report, this truth, doesn't actually include, then it renders it less and less truthful as a representation of what this child is. And you talked about universities earlier. Surely, in the best scenario, universities want to know what sort of child this is holistically, rather than what grade they can get in a particular narrow definition of what a subject is. What are we missing out with grading entirely? Well I think that there's two different answers here. One is what are we currently missing out, and what is really almost impossible to do with any grading system. And like we could certainly do a lot better on reporting on the kinds of things that you just talked about, Matthew. We could do a much better job of reporting on resilience, on collaboration, communication. The difficulty is that those kinds of phenomena, those kinds of constructs, are rather more easily gained than traditional achievement ones. So it in the rather unhelpful language of psychometrics, these are called non-cognitive measures, like resilience, like grit, like growth mindset. And the reason that they're difficult to measure is because they can be gained. So A mathematics test works because if you don't know the answer, you can't fake it. But if you're asking students when you're stuck, is it better to think really hard about the problem you've got or to give up? Students know what kind of answer you want. And so it turns out to be very difficult to come up with good assessments that can't be gained. And so, yes, we could broaden out the basis of the assessment. Absolutely. We could do much better to report on things I mean, you know, I mean, ⁓ the American system is probably the worst in the world in terms of the narrowness of its assessments, because every state has standards for math and English language arts, and almost always social studies and almost always science, but they only assess math and they also assess a very small part of English language arts, which is basically reading. And every state's standard has speaking and listening in their ELA standards, but they just don't assess them. And the reason for that, I think, is because we got used to inexpensive testing. So in England, it's routine to spend, in US terms, probably five hundred dollars assessing a student at the age of sixteen across, say, eight different subjects, because then you have constructed response, students writing essays, an English language arts exam might have a literature paper, which is ninety minutes long, and students have two questions to answer. One on a play and one on a poem. Explain how far you think Shakespeare presents Lady Macbeth as a powerful woman. And they have to write for forty-five minutes. And then it's scored by humans. And it produces really good practice in classrooms. And so we spend probably around about $180,000 educating each US child, and we screw it up by saving up a a few hundred dollars on the assessment. So if we were willing to move towards more authentic forms of assessment, then I think we would change what's happening in our classrooms. We'd have better measures of what our students can do, much broader measures of speaking and listening and writing as well as reading. And we'd actually have a positive backwash into what teachers do in classrooms. Sorry, before I throw that to Tom, the the thing that worries me in my conversations with students. in schools of across the world in totally different sociocultural contexts is that almost universally they see the assessment data and the reported assessment data as who they are and what they're worth. So if the grades miss out on or are blind to or simply ignore a lot of stuff, And they simply capture this amount of stuff about that child, and the child is equating their grade with their worth almost as a human being, then we're actually contributing to an exacerbation of students' self image, their self perception, their self worth at a time when that is ⁓ fragile anyway. People are working on that. So Charles Fidel is working on broadening the dimensions that we value in school, and I think that that's very important. But I think we shouldn't underestimate the difficulty of assessing these things. And so it goes back to this fact that you know it's a system. And right now our s systems rely on producing numbers. I mean, I I actually don't think schools should be driven by what universities want. And yet we are driven by treating certainly the last four years of secondary schooling as being a way of sorting and classifying students, rather than producing evidence about them. That student's capabilities. So in England, we had a very strong movement towards records of achievement. The idea was that a student would develop a portfolio of their achievements as they work through school. And I think that we've got models of those. They just keep on getting sidelined by a push for accountability, which needs quantitative, simplistic, and narrow measures for them to function. So we could actually have much broader systems and If we're willing not to produce information that allows schools to be ranked. If we could get rid of that, we could actually say that the purpose of schools is to educate citizens rather than producing numbers, then we could have a very very enlightened system, but it would actually cut across a lot of the political narratives that we currently are being ruled by. National Association of College Admissions Counseling. And every ⁓ two to three years they do a survey of selective college universities in the United States asking admissions departments what criteria they use for granting admission and then granting financial assistance to students coming in. I've been following this for ⁓ about the last twelve years now. To see what's going up and what's going down on the list. And grades are are clearly important. ⁓ I mean in particular they want to know what grades you're getting. in in honors classes or advanced placement classes, you know, it's it it's the rigor of the curriculum that's most important. But what has risen most rapidly in those twelve years, and now is number four on the list, is what they're labeling positive character attributes, where they want to know, do you have a sense of ⁓ empathy and compassion? Can you identify with people who are different from you? ⁓ do you have a sense of persistence? Can you work with others? That it might be different from you. And because there's they're recognizing that these things are not only essential for your success in college-university, but your success in life afterward. So we now have 17 of the states that have established what they call a portrait of a graduate that look at these other, as as Dylan called them non-cognitive attributes that are contributing to student success. And even ⁓ a recent report. looked at the people that have reached amazing levels of success in their in their businesses, CEOs of major corporations, all multimillionaires, and they asked them what their college grade point average was. Now these are men and women that are amazingly successful in their lives and have risen to the highest levels in in in business and industry. What would you guess to be the the college grade point average of these CEOs? Now many of them didn't want to report it, of course, but Across those who did, it came out to be only two point eight. I mean not even a a B average. And what this is saying is there are characteristics and attributes that are contributed to these people's success that seem to be far more important than their academic prowess. And so with that recognition, we're seeing that these things are not not only important for students' success in school, but in their life afterwards, and we need to be thinking about those more in terms of what we're doing in school and what we're recognizing in school. So I th I think your point is very valid. We're certainly moving in this direction. ⁓ it does get us back to, you know, that famous Lord Calvin quote about if in order to improve something, you must first be able to measure it. And if we can't measure it, then we don't know whether we can improve it or not. And and how we measure those things is clearly a challenge. But that's the kind of challenge that I think we need to take on. If we want to increase students' sense of belongingness, if we want to improve their sense of agency or efficacy, ⁓ we have to w find ways of of measuring those things and being able to report to students. ⁓ here's where you are now, you can get better in these ways. What's the difference between the students are highly creative and the one who's not? Well, if we can't describe that difference, you can't help them become more creative. Describing a difference is an essential prerequisite to improvement. But once we have some definition of that and some criteria we use, then we're better able to assist students in making that kind of progress. I think that's true in many domains, but I think it's not true in others. So I've seen some lovely examples in schools where they have what they call the writing wall, where there are samples of writing up on the wall and just seeing students talk very eloquently about, well, I can write like this, but I can't yet write like this. And so I think we do certainly need to have a sense of what is it that gets better when a student gets better. I'm not sure we should always insist on describing it explicitly. As Guy Claxton says, sometimes the best we can do is develop a nose for quality. And so I think accurate summative assessment involves making sure the teachers share a sense of quality. Effective formative assessment involves bringing our students into that same community of practice. So that students know, yes, this is the kind of writing that looks really good. And then the f last part is to be really effective as teachers, teachers need to have both a sense of quality and an anatomy of quality. In other words, they need to be able to help break down how to get from here to there. So if we do this job well, our students will know where they are, they'll know where they want to get to. What they can't figure out is what's the next step. And good teachers, like good sports coaches, Have the ability to break down that journey. What's the next small first step? And I think if we can actually work on that, then yes, by all means, if we can reduce it to a rubric, then we should. If we can tell students what we're looking for, we should absolutely tell them that. I just want to put on the table the idea that there are some places where we can't do that without basically coming up with a very kind of narrow, simplistic view of what success in that domain looks like. Dylan, well I w lots of the schools I work with at the moment, I'm I'm developing a concept I call anti data, which is very much the the data which isn't measurable data and challenges this this idea that in order if if something is important, in order to know it, we have to measure it. And suggesting actually the what you said about Claxton's ⁓ image of the nose, something can be really important and we can tell with our nose, with our senses, with our gut, with our intuit, through all the broad array of ways of knowing that species and civilizations have employed ⁓ for millennia without necessarily being able to measure it. And so maybe it's about what things can we measure and should we measure, and what things do we need to know through different means, I suppose. I think your definition of measurement is narrower than mine. So if we can get teachers agreeing that this portfolio work represents a seven on a ten point scale. I would call that a measurement, even if they can't articulate why they agree. So I I think that if we can't get teachers agreeing about what's good, then we haven't got a any kind of sense of quality. But I think we need to be broadening our notion of what counts as a measurement. And for me, too often we we don't want it to be subjective. We don't want the the score you get to depend on who does the grading. But if we insist on objectivity, then I think we lose too much. So what I'm interested in is intersubjectivity. In other words, that we agree as a community of interpreters, say English teachers in the state of Kentucky, we agree that this is worth a six or two or a C, even though we can't pin down exactly why we agree. If we could do that, that would be a basis of moving forward. Intersubjectivity rather than objectivity. I love that. Well listen, we're gonna move to the last few questions now. And I wonder if we can do these as quickfire or at least just a couple of sentences. So the f the first is from ⁓ someone called Rushi Ma Mainali from Heritage International Experiential School in Gogawan in India. And she says, with intentionally holistic data systems bringing in more subjective and contextual data. How can schools guard against bias creeping into grading decisions while still honoring the whole child? As I've stressed from the time I began investigating these issues of grading, that we must begin by establishing our purpose. I think that those are really fundamental questions that need to be raised at the very beginning of any discussion about what we're trying to do with grading and trying to make it better. What is the purpose? And once we decide that purpose, then we can move toward policies and practices that better align with it. I think John brought up the notion earlier too, and it's absolutely true, that the greatest difficulty we have is the lack of consistency in why we're doing it in the first place. And so when when we establish a purpose, and this is not easy to accomplish, I describe how it requires a a process of what we call disagree and commit. Disagree and commit is often described as a a management strategy used by some of the most successful corporations where you bring together people who have vastly different perspectives on particular issues, but and everybody puts their their perspective on the table and each is honored, but recognizing that we have to work together in this. We have to come together and reach some consensus because we're working as a team in this. The history of disagreeing commit is actually much longer than that. It was originally developed as a way to plan military campaigns, ⁓ where they would bring together these military advisors and and they would they would each have their perspective on how they should go about with this. particular battle or c or campaign, but then recognizing they had to work together so they would they would develop a plan that everybody would collaborate and and cooperate in order to achieve greater success. We have to do the same with grading. So you bring together people with these very different perspectives, but then come up with a purpose statement. And as soon as you have that purpose statement, it then puts you in a position to evaluate your policies and practices. So for example, many schools And teams of teachers and school leaders will get together and decide that they want the grade to really represent a communication tool, to communicate accurately what students have learned and are able to do. The primary audience is parents, families, and students, and they want to reflect where the students ⁓ current level of achievement is at this time. Well, as soon as they include a statement like that, they throw out averaging. Because as soon as you say, I want the grade to reflect where they are now or their current level of achievement, then you can't average. Now if we had started by challenging those policies and saying that we don't want to average anywhere, they they would have been very resistant to it. But if you begin with your purpose to of making clear what you want that grade to represent, then it makes challenging those long held traditions much easier. So I I think if we could begin there and establish that purpose, ⁓ the the politics that are involved, the external pressures that are involved, all of those can be taken into account by returning to your purpose and saying, This is what we'd want to accomplish. Our primary purpose here is communication. Our primary audience for that communication is parents and families and students. And and how do we want that information to be used. If those are clear, then I think our pathway is going to be clear about how we go about implementing those ideas. So the question then is how do we become clear about those ideas? And I think in a word, it's exemplification. So I think schools need to systematically get as many different examples of student work and get consensus around which ones are the good ones, which ones are the not so good ones. And to avoid the issues of bias, what you need to be doing is systematically exploring all the relevant dimensions. A piece of writing that is really good on characterization, but is terrible on spelling and grammar, for example. And just getting people to thrash out those ideas. So we need to have a really broad frame to make sure that we are considering all the possible ways in which student work might vary. And then we're being clear about which ones are important to us and which ones are less important, and developing, at least at the school level, a culture of this is what we think is good work. And then we bring the students into that same culture. So our our next audience question ⁓ comes from Matt Townsley, an associate professor from the University of Northern Iowa. And I think that we've looked at some of the aspects of this question already, but maybe we can just sort of summarize, or you guys can summarize from your perspectives. On which aspects of data assessment and grading research do you most agree? And where do your perspectives differ? I think we've talked about that ⁓ already. I think that Tom's, you know, s Tom has famously said in the US system, do you know why give scores? Do you really need 70 degrees of failure? And clearly you don't. But of course, that's partly because of the perverse system we have where 70 is the passing score. That doesn't happen in other countries. In England, it's often 30 or 35%. And it actually depends on where you want to draw meaningful distinctions. So any assessment is typically most accurate in its middle range. And so if you're really concerned with sorting out passing versus failing, then you really want your passing score to be around about 50. ⁓ so there are different reasons for this. I think it's just a quirk of history that it came from, I think, Mount Holyoke, ⁓ in the United States, that they decided that seventy was a level of mastery that they wanted to adopt and that just got picked up by everybody else. So d d Tom has been, you very eloquent that you don't need seventy degrees of failure. But I'm concerned that we are losing the ability to report the inaccuracy of assessments if we give up on points too early. And maybe the answer is to actually have points later on in the system. Sorry, yeah, points persisting. Maybe only reporting at the end of a semester a letter grade, but having finer scra finer scale data leading up to that. And then we tell parents where their child is in terms of mastery on the standards, for example. So, you know, they are at this point in terms of understanding the gas laws or covalent bonding, rather than giving a bru a grade which may provide assurance to the parents but won't tell them what their child knows and can't yet do. And I would add too that what I find most amazing and most delightful in discussions like this is that ⁓ I mean D Dill and I have very different ⁓ I mean cultural experiences, we have very different levels of training, we have very different professional experiences, but we agree on so many important issues. And I think that this says a lot for our field, that we do have a knowledge base, we do have a s a foundation upon which we can build better practice. We haven't done it very well, but I think that when when two scholars who have looked at the field as deeply as both of we both of us have, in different ways and with different orientations, but come up with so many comparable ideas and suggestions and recommendations on improvement. It's rather astounding ⁓ that we could do this. But the the very positive part of that is that it says we do have an established knowledge base, we we have a direction, we have ways that our system can become better. And so rather than I guess discussing the the few relatively modest minor points upon which we might have subtle d disagreements, I think that the more impressive thing to me is ⁓ how many things Upon which we do agree. And and how wonderful that is and what that says for our field. That we have a direction here. We have a knowledge base that we can use to make education better. We can make grading better. We can make assessments better. ⁓ we know how to do it. Our challenge is how to find ways to ⁓ make that common knowledge base more common practice and have that be incorporated in our field in better and more effective ways. Thank you. ⁓ Jack George, who's an assistant head at Aeglon College ⁓ in the Swiss Alps, w wonders whose data it is or should be? So how far can we move towards a scenario where ⁓ students own through a process of co-creation and negotiation their own data? Do you think that that's possible? Or do you think that the ownership should lie with the experienced and trained? Professional. Let me start with a parallel. I have never seen a high jump athlete failing to clear the high jump bar at six feet and just saying to their coach, Well, that's just your opinion. There's an objective standard there. And both the coach and the athlete knows that they didn't yet manage to clear the bar, and they're going to work on what they need to do to clear that bar. So I think externalizing the standard is the crucial move here. So I think I would like to move towards a standard, a a situation where students agree about the data and their teacher is going to help them improve that performance. So if we think of the teacher as the student's ally. Against the external standard, I think those conversations are much more productive. And I think, yes, we can move towards a situation, and I'm sure, that in the high jump, the coach's data and the students' data are identical. It's what height of bar you manage to clear on what occasion, maybe at what kind of event. But I think we could absolutely seek to do that with educational data by externalizing the standard, by making sure that our students understand what they need to be. aiming for so that the teacher is no longer seen as the judge and jury, but as the coach and the ally. Yeah, and I would agree, as I said, our so much of our evidence now is indicating that it's it's not a question of rigor, it's a question of clarity. And when that clarity becomes paramount for us and what we emphasize, ⁓ then our discussions become much more meaningful about how to help students achieve those particular well articulated learning goals. And one final point here, I'd like to refer to the work of Richard Stiggins. And he pointed out that we need student involved assessment. We should stop thinking about assessment as something that is done to our students, and rather something that's done with our students, involving the students in the process so they can become more active as owners of their own learning. S self-regulating learners in the psychological jargon But basically, students who own their own learning and could take charge of their learning, because I think that's going to be even more important in the future than it has been so far. Thank you, Vaifuji. Brings us now to our final question. So both of you have been working on making change in the field of grading and assessment and these things for decades now, and probably have seen some changes somewhere, probably not as much as you would have liked in many ways. If all of it were to just disappear. If we were to just say, you know what, we can't do this well, let's just not do this at all. No grades, no marks, no levels. What do you think would be lost and what might we actually be able to see? Again, I would I would take us back to the early work of Benjamin Bloom from the nineteen sixties, who was the the person who actually brought the the f the the word formative. To our lexicon in education. ⁓ He actually borrowed it from Michael Scrivan. Michael Scrivan was a ⁓ social psychologist who did a lot of work in evaluation. In 1967, Michael Scriven wrote this brilliant article about program evaluation, stating that one of the problems with so many program evaluations is we waited till the end to find out if things were working or not. And Scrivan had this idea that. Program developers should be gathering evidence as the implementation is taking place to make changes along the way. And if you did that in a very purposeful, meaningful way, and you could make these changes along the way, you're more likely to have a successful program at the end. To identify those differences, she talked about formative evaluation and summative evaluation. Benjamin Bloom and Michael Scriven were quite good friends and and Bloom said, well, let's bring that to education. And so in nineteen sixty eight, the year following that, he he brought this idea of providing students with that kind of formative information. And he stressed that there needs to be three aspects in that feedback. Number one, you have to clearly identify what you want students to learn to be able to do. Number two, among those things, what have they learned very well? And number three, among those things, what have they not learned well enough yet? Bloom also stressed at that time that that information, that feedback alone was insufficient, that we needed to pair with that feedback what he called guidance and direction on how to improve. The idea being that with the feedback, we had to pair what he called correctives provided by the teacher, individually based on students, particularly the the learning difficulties they were having at that time. And so I think that regardless of the w the labels we attach to it. In any effective teaching and learning encounter, teachers have to be gathering evidence, information about whether the students are learning well or not. Then based on information they need to make changes to guide students in ways that are different because the initial experience wasn't successful, Bloom also stressed the most important thing about the corrective is it must be different than original instruction, that it's not reteaching. It isn't going over the same thing a second time, saying it louder and more slowly. It's providing a different kind of engagement, a different kind of involvement for the students, that you already have evidence that the boy you tried it initially didn't work. So I think those things that we're talking about are essential elements of effective teaching and learning. And so even if we got rid of all the words we attach to it now, we'd still have those core elements that are gonna be part of the process and have been from the beginning of time. I agree. Paul, what you suggest cannot happen simply because any good teacher is going to be collecting evidence from their students and making judgments about what to do next. And then the crucial thing about Benjamin Bloom's work was that he s signaled that we need to move away from teaching as a linear process to teaching as a contingent process. In other words, we teach the best we can, but then we find out what it is our students have learned. And if they haven't learned it, we do something about it. And so th the crucial thing, the the wonderful quote by Bloom, you know, if you're getting a normal distribution of achievement, you're not doing your job properly. Because the normal distribution is what nature gives us. Students differ in their aptitudes, and if we treat all students the same, you will get the bell curve. Our job as teachers is to destroy the bell curve by making sure that those who need more opportunities to be successful get those opportunities. And so I don't think that you can ever stop a good teacher from wanting to know what sense did my students make of my teaching? Am I happy with that? What do I need to do before I move on? And so yes, teachers will always do that. Whether they record it, whether they give it a number is a separate issue. But basically, assessment is the bridge between teaching and learning. It is only by assessing that we can figure out whether what we just did worked and make our teaching responsive to our students' needs. What you've both said there ⁓ ha hap happens ⁓ outside of the grade being given. And you could do all the things that both of you have just described, with there being no grades, no marks, no levels. And given the extent to which students fuse their worth with their grade, their mark, or their level, then could the removal of the grade mark and the level, but the retention of all of the formative assessment that you've both described, could that be an ideal scenario, even if it might never happen because governments want want the political football of education, schools want to compete against other schools, ⁓ and universities, etc. I so I guess all of those things could happen without the grade? I think they could, but we can't do it with our existing students. Basically our students are junkies, they are hooked on grades, the teachers are the pushers, and the parents are the codependents. And so yes, let's try To never introduce this process, and I think we could have what you ha what what you're talking about, but it's going to be extraordinarily difficult to get students who are already used to receiving grades as the most meaningful form of feedback on their work to not getting grades. We've tried it, we've had some success, but you're really pushing uphill. It's it's just it's such a countercultural move that it's really difficult to to prevail in this way. I smile a lot there because I I talk a lot about grades being an addiction, ⁓ and that parents and students and teachers are all addicted to grades. So to change that system we've got to look at how would we ⁓ break an addiction ⁓ cycle. So that's made me smile. listen, it's been fascinating to talk to to both of you. I'm sure I speak for Paul when I say I could carry on for many, many hours, although no one else would want to. And lots of questions you've posed will be ones that I take with me into the future as well. Paul? Yeah, absolutely. Thank you so much. It's been a real pleasure. You're welcome. It's been fun. Yes. Thank you, Paul and Matthew. It's all been a my pleasure completely. Dylan, it's always a great honor to work with you. So thank you for that invitation show. Thank you. Tune in next time as we continue our data conversations with educators and thought leaders from around the world.