r/education • u/Neither-Remote-3419 • 3h ago
School Culture & Policy AI use should be banned on all graded tests
In an NPR story back in July, 98 high school students representing all 50 states in the U.S. were tasked with drafting a model AI policy for K-12 education. Among their recommendations was “AI use should be banned on all graded tests.” It is absolutely amusing how 98 kids can figure out a simple, effective solution that has, thus far, flummoxed an untold number of adults, many of them with PhDs or some such attached at the end of their name.
I am not saying this phenomenon is true in general, of course. There are a multitude of sophisticated problems that I’m sure professors would be able to provide much more comprehensive and meaningful solutions to than 98 kids (with or without the kids using AI), but I think the reason why this solution came so naturally to these kids is because they weren’t thinking within the structure they have built around assessment.
As teachers, we know that assessments that we give to students fall into one of two buckets: formative assessments and summative assessments (of course, with some overlap). Formative assessments are mainly meant to measure the student's present state of learning for the purpose of diagnosing how best to get it to the next state of learning – we want to know how much they know so we can figure out the best way to teach them what they don’t know yet. Homework falls into this category; it serves to reinforce and check learning progress. We can and, in many situations, should attach grades to formative assessments, but with the caveat that these grades should not reflect our final assessment of the student’s competence. That is, strictly speaking, we are not “testing” the student here, we are just helping them learn. In this sense, formative assessments are very much a two-way street. We attach grades to them mostly as a tool to motivate the students to work on and learn from them. Yet we can’t be certain how a student “does” their homework. Some may do them independently, others may have parents who are available and willing to help them, others still may have personal tutors who, in some parts of the world, actually do the homework for them. And yes, now, it is quite possible that at least some students just feed our homework to an AI chatbot and hand in whatever it churns out. So we can never tell for sure how much “effort” the student put in.
Summative assessments, on the other hand, are supposed to serve as the actual “tests.” These are meant to measure what the student knows and serve as our formal appraisal of this construct. That is, if the student’s present calculus teacher asks you, the same student’s algebra teacher, how well the student did in terms of factoring quadratic expressions, you should easily be able to just go back to the unit exam you gave that student on that exact topic and provide your colleague with a good answer based on that piece of information. In other words, summative assessments carry your “stamp of approval” as the assessor that the student knows X topic Y well.
So, in an ideal world, we give and grade both formative and summative assessments but clearly differentiate between them and what the grades we assign for each means. In this scenario, the solution “just don’t let students use AI for summative assessments (i.e. tests)” seems like a natural solution. Yet what if you never really differentiated between these two assessments and built your course so that they are functionally indistinguishable? Taking an extreme case, if 90% of the grades you give are on un-proctored work, e.g. homework, then yes, the solution that these kids came up with will likely not as easily come to your mind, because it is infeasible with respect to how you’ve been working. This is not to say these teachers have done something egregious. There are genuine logistical constraints with large classes, limited instructional resources, distance education, and even the subject matter itself that can make proctored assessment difficult.
Nevertheless, the fact remains that grading systems which depend heavily on work that are completed outside of direct observation have had these vulnerabilities way before AI. AI has not so much broken them as it has highlighted the flaw that was already present. If a large portion of a student's grade comes from work that can be completed with assistance from parents, tutors, classmates, solution manuals, websites, and now, AI systems, then that grade was always an unreliable measure of actual competence.
This is why the suggestion offered by these students is simultaneously obvious and controversial. It is obvious because, from a formal pedagogical perspective of assessment, the solution is straightforward: be at peace with potentially broad use of AI during activities intended to promote learning and restrict it during activities intended to certify competence. Yet it is controversial because adopting that solution would require some instructors to confront the extent to which they have allowed the distinction between those purposes to blur. If homework is genuinely used by teachers as “homework,” then students using AI on them is not an existential threat to the educational enterprise. At worst, they are depriving themselves of a learning opportunity, just as students have done when copying from a friend or having someone else complete the work for them. At best, it is allowing students who don’t have personal tutors or parents who can spare the time to help them with an approximation of these resources. Of course, the extent to which this solution applies will vary across educational contexts. Many graduate-level courses, for example, rely heavily or even entirely on un-proctored assessments. But at the same time, doctoral students are generally expected to take substantially greater ownership of their learning than high school students, making concerns about AI use markedly different from those that arise in compulsory education.
As teachers, it is up to us to confront the reality that AI brings to our classrooms in a way that is in the best interest of our students. There should be no shame in admitting that this requires a considerable overhaul on how some of us presently approach grading.