Imagine a student finishing a maths worksheet with every answer correct. An AI assistant helped. The work took less time than usual.

What did the student learn?

The worksheet cannot tell us. We need to see what happens when the student meets another problem without the assistant.

That is the distinction I keep coming back to when reading about AI in education. I have given several seminars on the subject for teachers in local communities. “Does AI help learning?” is a natural question to bring into that room. On its own, it leaves too much unspecified.

Two questions make it useful:

  1. What work does the AI do, and what thinking is left to the student?
  2. Are we measuring performance with the tool, or learning that remains without it?

They change how I read the evidence-and what I would ask of a classroom trial.

When the worksheet and the test disagree

In a Turkish high-school experiment, students using a basic GPT-4 assistant performed 48% better on practice problems than a control group. On a subsequent test without assistance, they performed 17% worse. A more structured GPT-4 tutor avoided that penalty, but showed no positive test effect. Those are relative differences, not percentage points. The tutor also received teacher-prepared solutions and guidance on common mistakes. Bastani et al., PNAS, 2025

The distinction matters for anyone looking at a successful demonstration. A correct answer tells us something about the student and the tool working together. It does not, by itself, tell us what the student can now do alone.

It also matters for design. “Give hints instead of answers” is an appealing rule. But this experiment changed several things together. It does not establish that withholding answers alone produced the difference.

My reading is that choosing a model leaves much of the educational work undecided. Someone still has to choose the problems, anticipate misunderstandings, decide when to intervene, and specify what counts as progress. Those choices deserve as much scrutiny as the model name.

A gain is possible. Its scope matters.

A Harvard physics experiment offers a positive result: 194 undergraduates learned more in less time with a custom AI tutor than in the comparison active-learning classes. The tutor used structured activities and prepared solutions. This was a short intervention with immediate testing, not evidence of lasting gains across an entire degree. Kestin et al., Scientific Reports, 2025

I take that result seriously. I would also keep its boundaries attached whenever I quote it.

A gain on an immediate test, retention a month later, and the ability to apply an idea in an unfamiliar situation are different outcomes. A school may care about all three. Evidence for one cannot simply stand in for the others.

The same applies to the comparison. Better than which lesson, for which students, working on which material? Those details are part of the finding. Remove them and a useful result becomes a slogan.

The number needs its context

This is also why I am wary of a single headline effect for “AI in education.” Before using an average to guide a decision, I want to know what was averaged: which learners, which activities, which comparison groups, and which tests.

There is a concrete reason to check the sources, too. Wang and Fan’s widely discussed meta-analysis was retracted in April 2026 because discrepancies undermined the editor’s confidence in its analysis and conclusions. That disqualifies that analysis as support; it does not settle the rest of the literature. Retraction notice

For a teacher choosing how to use an assistant next week, a narrower finding may be more useful than an impressive average. Does the activity resemble their lesson? Does the assessment capture what they want pupils to learn? Could they provide the preparation and support the intervention required?

What I would ask a school to try

I would start with one learning objective, stated without mentioning AI. For example: students should be able to explain why the same operation must be applied to both sides of an equation.

Then design an activity around it. Ask students to attempt a problem and explain a step before consulting the assistant. Use the assistant to discuss that reasoning. Afterwards, give a fresh problem without it and ask for an explanation again.

That is a proposal for a classroom activity, not a claim that this exact sequence has been validated by the studies above. Its value is that it makes the intended learning visible and gives the teacher something concrete to examine.

If feasible, return to the idea later. Did the explanation survive? Can the student recognise the same principle when the problem looks different? A small classroom check will not establish a general causal effect, but it can reveal where a polished piece of work has hidden a misunderstanding.

There are also activities where competent use of AI is itself the objective. In those, assess how students use it: whether they check an answer, recognise a weak explanation, and justify the choices they make. Be explicit about which capability the lesson is developing.

That is what I want from the conversation about AI in schools: a clearer account of what students are practising and how we will know it helped.

A completed worksheet is useful. The question is what the student can do when the next one is blank.

Mehmed Kadrić Founder, MESH Data Solutions

I write about concrete data and software problems, including what failed, what was built, and what I would change next time.

NEED EVIDENCE FROM YOUR OWN DATA?

Turn the checklist into a focused audit.

Request an Audit