Skip to main content

Generative AI in Education: Better Homework, Worse Learning

Picture

Member for

1 year 10 months
Real name
The EduTimes Editorial
Bio
The EduTimes Editorial

Modified

AI raises homework performance while weakening unaided test results
Learning losses concentrate where AI replaces cognitive effort
Schools need assessment, training and age-sensitive AI rules

In a study of 26,811 secondary school students, six months of using generative AI in homework was associated with 18 percent higher assignment grades and 20 percent lower grades on closed-book exam scores. The distance between the two measurements didn't close over time; it grew as students became more proficient in using the tool. The finding doesn't question the value of AI in the classroom. It challenges the assumption that its integration can remain a technical detail, an application added to an unchanged curriculum. The evidence that speed and learning can move in opposite directions shifts the question from whether technology will be used to how pedagogy will be designed around it.

Generative AI in Education Needs a Mastery Threshold

The difference between students who benefited and students who were harmed by artificial intelligence is not found in access to the tool, since the difference is not simply access to the tool. It is found in whether they retained the time and effort required to work on a topic. The minority of students who continued to devote the same time to the assignments as before learned the same as those who did not use the tool at all. The majority, who completed the assignments in less than 50 minutes and scored extremely high grades, paid the price in the tests. The pattern confirms a distinction that has been made separately in the pedagogical literature, between the use of a tool that retains the student's executive control and a use that abolishes it completely. In the first case, the student makes the argument, uses the tool to challenge it or identify gaps and revises on his own. When executive control is lost, the tool produces the final text while the student merely accepts it

Figure 1: AI adoption raises homework scores while shifting exam performance lower.

The damage was not equally distributed between subjects. In social subjects, where comprehension depends more on personal processing of arguments than on the application of formulas, test scores fell by 27 percent. In STEM subjects the drop reached 22 percent, in English 17 percent and was smaller in Chinese. This distribution reverses a more well-known picture from the labor market, where artificial intelligence tends to raise the performance of less experienced workers more. In the classroom, the damage disproportionately affected students who were already high performers, an indication that the problem is not a lack of competence but a substitution of a skill that already existed.

The resulting principle is simple but often overlooked in lesson planning. A skill can only be assigned to a tool after it has been sufficiently mastered so that the student can judge whether the output of the tool is correct. Without this threshold, delegation is not a delegation of competence but the abandonment of a skill before it is even developed.

Figure 2: Students who preserve homework time largely avoid the exam penalty.

Human-Centered AI Pedagogy Must Preserve Judgment

A framework developed at the Newcastle Business School at the University of Newcastle in Australia, after twenty months of experimentation in entrepreneurship courses, offers a concrete alternative to the logic of simply adding tools. The framework organizes teaching around five axes: content preparation by the teacher with the help of artificial intelligence, personalized learning, where the student collaborates with the tool, active participation in the classroom where the teacher acts as a coach rather than just a source of knowledge, summative assessment, where AI helps with grading but the teacher reviews it and continuous personalized monitoring of progress. The last two pillars of the framework, the documentation of the student's use of a tool and the mandatory reflective recording in a diary, translate the abstract idea of executive control into concrete assessment practice.

The value of this approach lies not in the technology they use but in what it requires the student to demonstrate. By asking for documentation of where and how the tool was used, the educator can assess the problem-solving ability rather than just the final product. By asking for reflection in each lesson, the framework transforms the use of AI into an object of awareness rather than an invisible habit.

AI Use Should Vary by Student Age

A thirteen-year-old child and an eighteen-year-old student do not possess the same metacognitive awareness ability and programs that implement the same policy of using artificial intelligence at all ages overlook this difference. In early secondary school, priority should be given to the development of fundamental skills with minimal assignment to tools, since no one can supervise a process they do not understand. In middle secondary, controlled assignment can be introduced in specific contexts, with clear instruction on when and how to delegate a task. In the last year of high school, the student can undertake a greater degree of executive control, although even then he needs guidance, since managing his own thinking remains a skill under development.

The larger losses among junior students support age-sensitive sequencing, although the study does not identify metacognition as the causal mechanism. A program that does not distinguish between ages treats AI as a single tool with a single effect, while its effect changes radically depending on how much the underlying skill has already been mastered.

Better AI Tutors Cannot Fix Weak Incentives

A plausible counter-argument argues that slowing the adoption of AI is depriving students of access to a tool with real value and that the solution lies in technology rather than pedagogy. A two-year experiment with the Khanmigo coaching tool, explicitly designed to coach students rather than give them direct answers, shows why this argument is not enough on its own. Despite the availability of a tool designed with the right specifications, students who had access to it did not use it with the consistency it would need to perform. The preference remained towards general-purpose tools, which give a more direct and easier response.

This finding shifts the weight from the design of the tool to the design of the incentives around it. A well-designed guidance tool is not enough if the curriculum, evaluation and expectations around it continue to reward only the result. The objection to speed is right that technology offers a real benefit. It is not right that this benefit materializes automatically, without a change in the context in which it is used.

Schools and Workplaces Must Retain Human Accountability

The shift from use policies to executive control frameworks has direct consequences for those who design curricula and examination systems. First, assessment should measure something beyond the final text, such as the ability to verbally defend a position or identify errors in a text that one did not write oneself. Second, teacher training should include explicit instruction about when a skill has reached the threshold that allows for safe assignment, rather than being left to the judgment of each class individually. Third, examination systems that continue to reward speed of completion without comprehension testing work, unintentionally, against any attempt at human-centered planning.

The objection that this slows down the adoption of useful technology doesn't stand up to the evidence itself. Students who kept their working time constant didn't lose access to the tool; they just didn't trade processing time for it. Human-centered AI pedagogy doesn't call for less technology in the classroom. It calls for a clear answer to the question of who retains thought control when the tool is available and that question remains, for now, without a single answer in most education systems.


This article reflects the analytical judgment of The EduTimes Editorial Board and does not constitute policy advice or the official position of any affiliated institution.


References

Dolman, J. (2026) ‘From cognitive offloading to cognitive enhancement’, The AI English Teacher, 25 January.
Kirschner, P.A. (2026) ‘Understanding AI: Offloading or outsourcing thinking?’, kirschner-ED, 13 January.
Pedley-Smith, S. (2024) ‘Cognitive offloading: You can’t outsource thinking’, Accounting Cafe, 31 May, updated 25 August 2025.
Stavissky, Y. (2025) ‘The Human-Centered AI Pedagogical Engagement (HCAI-PE) Framework: A foundational paradigm shift in AI pedagogy’, SCIREA Journal of Education, 10(6), pp. 344–365.
Strömberg, D., Lei, V. and Wu, Y. (2026) ‘The generative AI learning penalty in secondary school’, VoxEU, Centre for Economic Policy Research, 4 September.
Swiss Institute of Artificial Intelligence Research Editorial (2026) ‘Cognitive outsourcing in education: Why AI’s real classroom crisis is verification, not cheating’, SIAI Research, 24 July.
Verhoeven, B. and Hor, T. (2025) ‘A framework for human-centric AI-first teaching’, AACSB Insights, 12 February.

Picture

Member for

1 year 10 months
Real name
The EduTimes Editorial
Bio
The EduTimes Editorial