This website uses cookies

Read our Privacy policy and Terms of use for more information.

Nearly 1,000 high-school students spent four maths sessions with one of three setups: textbooks, a GPT-4 chat interface, or a guarded tutor that would hint without handing over the full solution.

With GPT-4, the practice sheets looked much better. The surprise came after the laptops closed.

Students who had used the answer-giving version scored 17% below the no-AI group on the unassisted exams. The guarded group did not. Its exam scores were statistically indistinguishable from the control group.

That result has been stuck in my head because we usually judge AI while it is still in the room. We see the faster draft, cleaner code, solved problem. We rarely test what remains after the tool leaves.

🧠 THE BIG IDEA

The school trial took place in Turkey in 2023. Across four sessions, students reviewed a maths topic, practised it with their assigned resources, then sat a closed-book, closed-laptop exam.

The ordinary GPT-4 interface raised practice grades by 48% relative to the control group. The guarded tutor did even better during practice. But only the ordinary version was followed by a statistically significant drop on the unaided exam.

The chat logs help explain why. Students with the ordinary interface often asked for the answer and copied the solution. The guarded tutor was prompted to withhold the full answer, use teacher-provided problem information, and respond with hints. Its students sent more messages and made more attempts of their own.

I nearly turned this into an anti-AI issue. The second tutor ruined that.

The difference came down to how the interface paced the work. One let students skip the struggle, while the other forced them to query, attempt, and iterate.

There are limits here: this was a single high school, one subject, newer models may act differently, and the exams measured short-term learning. Even so, the study catches a slip that is easy to make outside school: improved output can look identical to improved ability while the tool is on the screen.

A quick word from today’s partner.

TOGETHER WITH 1440

TOGETHER WITH 1440

Every headline satisfies an opinion. Except ours.

Remember when the news was about what happened, not how to feel about it? 1440's Daily Digest is bringing that back. Every morning, they sift through 100+ sources to deliver a concise, unbiased briefing — no pundits, no paywalls, no politics. Just the facts, all in five minutes. For free.

WHAT PRACTICE WAS STILL DOING

The other study in this issue is much smaller and much stranger.

Researchers recruited 31 adults to learn categories of computer-generated, car-like objects. Fourteen completed the full protocol. Eleven were included in the final analyses. Those eleven performed at least 30,000 classification trials over five to ten weeks.

After extensive practice, category-related signals appeared earlier and farther back in the visual system. Task-related functional connectivity between a visual region and part of frontal cortex decreased, while connectivity with motor-related areas increased. By the final round, participants also improved on a combined measure of doing the trained task beside a peripheral visual task. The peripheral task did not significantly improve by itself.

The study did not involve AI, and it cannot explain the school result. I kept it because it makes one casual assumption feel less safe: that a repetition becomes worthless as soon as it feels boring.

For an expert, boring may mean mastered. For a beginner, the same practice may still be building category representations, sharpening error detection, and developing an eye for what looks wrong. From the far side of competence, it is very easy to forget which is which.

🔢 ONE NUMBER TO REMEMBER

That was the median time Harvard physics students spent with a carefully structured AI tutor. The comparison class was assumed to spend 60 minutes learning the same material. The AI group earned higher immediate post-test scores.

This tutor was not a blank chat box. It used expert-crafted solutions, sequencing, scaffolding, self-pacing, and targeted feedback. The study did not include a delayed retention test. A shorter session produced more immediate learning in that setup. The minutes tell us very little unless we know what happened inside them.

🔥 DRAW THE LINE

A tool can supply an answer, hold one back, offer a worked example, ask a question, or wait for your attempt. Calling all of that “using AI” hides the choice that probably matters most.

When should AI enter practice on a new skill?

🛠️ TRY IT YOURSELF

The disappearing-tool test (3 minute exercise)

Choose one task you already hand to AI. Before your next prompt, spend three minutes making an ugly first pass. Write what you think the answer is, where you are unsure, and what a good answer must include.

Then use AI as usual.

When you finish, close the result and reconstruct the key move from memory. Add one mistake you could now catch if the tool made it again.

Do not grade the prose. Notice where your explanation collapses once the answer is gone.

SOURCES

I am not giving up the shortcut. I am just going to close the tab once in a while and see what came with me.

See you tomorrow,

Franky

Reply

Avatar

or to participate

Keep Reading