Ai News & Updates

AI and Student Exam Scores: Why Are Scores Falling 20%? New Study Reveals the Shocking Impact of AI


Quick Answer

A major study on AI and student exam scores, tracking 26,811 Chinese secondary students over 30 months, found that AI use raised homework scores by 18% while cutting closed-book exam scores by 20%. The effect was strongest among students who used AI to finish homework faster rather than to actually learn — and the damage took up to two years to fully show up in high-stakes entrance exams.

If you’re a parent, teacher, or student trying to understand how AI and student exam scores are connected, this article breaks down what the research actually found, why it happened, and what it doesn’t prove yet.


Here’s a strange pattern showing up in classrooms right now: students who use AI the most are turning in near-perfect homework — and then bombing the exam that covers the exact same material.

That’s not a hunch. It’s the finding of a new working paper from researchers at Stockholm University and the University of Hong Kong, and it’s one of the largest, longest-running studies yet on how AI and student exam scores actually interact in real classrooms.

What the Study on AI and Student Exam Scores Actually Found

Economists David Strömberg, Victor Lei, and Yanhui Wu tracked 26,811 students in grades 7 through 12 across a county in central China, from September 2022 to June 2025 — a full 30 months. About 80% of the students used generative AI tools like DeepSeek and Doubao for schoolwork; the remaining 20%, who didn’t, became the natural control group.

The results, first detailed in a working paper published through the Centre for Economic Policy Research and later reported by The Economist, split sharply in two directions:

Homework went up. Scores rose 18% across all subjects, and the time it took to finish an assignment dropped from 64 minutes to 45 minutes — a 30% cut.

Exams went down. The same students scored 20% lower on monthly closed-book exams than their non-AI-using classmates.

By the time students reached China’s national entrance exams, the gap had widened further. The zhongkao (high school entrance exam) showed a 24% decline among heavy AI users, and the gaokao (university entrance exam) fell 18%. According to the South China Morning Post’s coverage of the research, the full “brain drain” effect took roughly two years to appear at its worst.

Why Homework Scores and Exam Scores Split Apart in This AI Study

For decades, doing well on homework was a reasonably reliable signal that a student understood the material and would do well on the test too. This study found the opposite pattern among heavy AI users — strong homework, weak exams.

The researchers point to a simple explanation: a lot of students weren’t using AI to learn the material — they were using it to produce the answer. When the homework is done in 45 minutes instead of 64, and most of that time savings comes from skipping the actual thinking, nothing gets encoded into memory. The exam then tests the thing that never actually got learned.

Not every student saw the same drop, though. The researchers found that students who used AI but still spent roughly the same amount of time on their homework as non-users lost very little ground. The damage was concentrated among students using AI as a shortcut to finish faster — not among students using it as a study aid.

A Real Classroom Example

The pattern isn’t just showing up in large datasets — it’s showing up in real classrooms too. One Brown University professor described a case that mirrors the study almost exactly: his students turned in a take-home midterm with suspiciously strong results — average scores in the high 90s, with half the class earning a perfect score.

Suspecting AI had done most of the work, he switched the final exam to an in-person, closed-book format. The class average collapsed to 48%. Scores that had never dipped below 65% on the take-home version fell well below it in person, and most of the students who’d scored perfectly on the midterm didn’t come close to a perfect score on the final.

It’s Not Just Exam Scores — Memory Takes a Hit Too

A separate, widely cited MIT Media Lab study adds another layer to this. Researchers there had 54 college students write essays using either ChatGPT, a search engine, or no tools at all, while monitoring brain activity with EEG sensors.

The ChatGPT group showed the weakest neural connectivity of the three groups, and in the very first session, 83% of them couldn’t quote a single line from the essay they’d just written minutes earlier. When the same students were later asked to write without AI assistance, their brain activity stayed lower than the group that had never used AI at all — suggesting the effect didn’t just disappear once the tool was taken away.

Together, the two studies point at the same underlying issue from different angles: one shows the outcome (lower exam scores), the other shows a plausible mechanism (weaker memory encoding when AI does the heavy lifting).

What This Means for Grade Inflation

There’s a wider pattern this research connects to. The share of college grades considered vulnerable to AI-assisted cheating — concentrated in humanities and engineering courses — has reportedly risen sharply since ChatGPT’s release, according to reporting on recent grade inflation trends. If take-home and AI-assisted assignments keep getting easier to inflate, exams may increasingly become the only reliable measure of what a student actually knows.

What Parents, Students, and Teachers Can Take From This

The researchers aren’t arguing AI has no place in education — the students who used it as a study aid rather than a shortcut kept most of their learning gains intact. The distinction that mattered most was how AI was used, not whether it was used at all.

A few practical takeaways line up with what the data actually shows:

  • Speed is the warning sign, not the tool itself. Students finishing homework unusually fast, especially on material they previously struggled with, is worth a closer look.
  • Closed-book, in-person assessments still matter. They remain the clearest way to check whether learning actually happened.
  • AI used for explanation beats AI used for answers. Asking AI to explain a concept and then attempting the problem independently preserves far more of the learning benefit than copying a generated answer.

Frequently Asked Questions

Does using AI for homework actually lower exam scores?

According to this study on AI and student exam scores, yes — but the drop was concentrated among students who used AI to finish assignments faster rather than to understand the material. Students who used AI without cutting their study time saw little to no decline.

How long did it take for the negative effects to show up?

Homework and monthly exam effects appeared within six months. The full impact on high-stakes entrance exams took closer to two years to fully emerge.

Is this study specific to China, or does it apply elsewhere?

The data comes from Chinese secondary schools, and the researchers themselves note it’s one study from one country’s school system. Whether the same pattern holds elsewhere hasn’t been confirmed yet, though the mechanism — less effortful practice leading to weaker recall — isn’t tied to any one country.

Can AI ever help students learn instead of hurting them?

Yes. Both this study and the MIT research point to the same distinction: AI used as a tutor or explainer preserves learning, while AI used to generate finished answers undermines it.

Bottom Line on AI and Student Exam Scores

This is one of the largest real-world datasets so far on AI and student exam scores, and the message is fairly direct: homework grades stopped being a reliable signal of learning the moment AI got good enough to do the homework itself. Exams — especially closed-book, in-person ones — are quickly becoming the only honest measure left.

If you want to keep up with how AI is reshaping education, career prep, and exam trends, check out more coverage in our AI news section.


Sources: Centre for Economic Policy Research working paper, South China Morning Post, Fortune, MIT Media Lab


Leave a Reply

Your email address will not be published. Required fields are marked *