Trending Topics
UNSW's Can, Can't, Must assessment framework
UNSW AI assessment rules vary by task; Indian families can examine course expectations and discuss preparation for Australian study with Collegify.
Collegify
Trending Topics
Harvard’s AI-tutoring experiment produced a significant result: students learned more in less time than under a strong active-learning comparison. But the study does not validate every chatbot being sold as an “AI tutor.” It raises a more consequential question about what schools should demand before putting AI between a student and the learning process.
From the Journal
Trending Topics
UNSW AI assessment rules vary by task; Indian families can examine course expectations and discuss preparation for Australian study with Collegify.
Trending Topics
Cambridge’s AI policy is not a blanket ban. Applicants may use AI for ideas and research, but fabricated achievements may be treated as application fraud, while AI use during interviews is expressly prohibited. The larger shift is more important: as applications become easier to polish, Cambridge is placing greater value on whether students can independently own and defend the thinking behind them.
Trending Topics
Tutor CoPilot found measurable gains when AI supported tutors, while separate research found that unrestricted AI assistance could improve practice performance yet weaken students’ independent performance later. Stanford’s 2026 review reaches a similarly important conclusion: the evidence is promising, but still limited. We look at what these findings actually mean for families evaluating AI-supported learning.
Related Collegify pages
Harvard has given education a particularly tempting headline.
In a randomized controlled experiment published in Scientific Reports, students using a custom-built AI tutor learned significantly more than students covering the same material through an active-learning class.
The researchers describe median learning gains in the AI condition as more than double those in the classroom condition relative to their shared pre-test baseline. Students also spent less time on the AI lesson: a median 49 minutes, compared with an estimated 60 minutes of learning time in the classroom condition. The study involved 194 eligible undergraduates in Harvard’s introductory Physical Sciences 2 course.
Those numbers matter.
But I think the most important part of the research is the part most likely to disappear once the headline starts travelling.
The AI tutor did not work because it was simply intelligent. It worked because somebody had designed it to teach. That distinction should matter enormously to anyone currently buying, building or approving AI for education.
The Harvard team did not open a generic chatbot and ask students to learn physics from it.
The researchers deliberately engineered the AI tutor around established principles from pedagogy and educational psychology. The design incorporated active engagement, scaffolding, cognitive-load management, timely feedback and self-paced learning. And where prompting alone was not sufficient, they changed the system. For multi-stage problems, for example, the platform constrained the sequence so students moved through the material in the intended order rather than allowing the model to jump ahead.
That is not a small technical detail. It is arguably the core result.
The researchers were not asking: How clever is the model?
They were asking: How should the model behave if the objective is learning?
Those are very different product questions and education is currently in danger of confusing them.
Schools and families will increasingly encounter products marketed as AI tutors - some may be excellent, others may be general-purpose language models placed behind a school-friendly interface.
Both can answer questions, can explain concepts, can generate examples, can appear personalized, but nothing makes them educationally equivalent.
A model that answers every question brilliantly can still be a poor tutor. Because the purpose of tutoring is not to maximize the quality of the answer delivered by the tutor, it is to improve the quality of the thinking the learner can eventually do without the tutor.
If a student cannot solve the next step of a problem, the most efficient response is to provide it.
The most educationally useful response may be to ask a question. Let the student remain uncertain for another minute. Good teaching regularly withholds help. Most consumer AI products are designed to do the opposite. That should make us much more demanding when the word tutor appears on the product page.
The result has not remained confined to the paper - Harvard subsequently highlighted the work as part of a broader discussion about AI tutors on campus. In an October 2025 Harvard Gazette report, the researchers stressed that the objective is to make students think and ask questions rather than allow the system to do the thinking for them.
That is perhaps the simplest useful test for education AI.
Not: How impressive is the output?
But: What intellectual work is the student still being required to perform?
That question becomes particularly important because the same technology can produce completely opposite learning behaviors. Used one way - AI can interrogate an assumption, offer personalized practice and provide immediate feedback, used another way - it can write the answer before the student has properly encountered the problem.
Same underlying technology. Very different educational consequence.
Parents are going to face an increasingly difficult buying decision. A product can look extraordinarily sophisticated in a demonstration. Ask it a difficult question and it responds immediately. Request a simpler explanation and it produces one. Ask for ten personalized questions and they appear. The experience surely feels intelligent. But parents should be asking something harder here: Is my child becoming better at the subject, or simply better supported while the tool is present?
One practical consequence of the Harvard study is that we should stop judging education AI by how capable it appears during use. We should judge it by what remains after the support is removed.
Can the student solve the next problem independently? Can they explain why the answer works? Can they identify their own misconception? Can they transfer the idea to a new context?
That is evidence of learning.
For schools, the consequence is even larger - AI procurement cannot become another software-purchasing exercise driven primarily by features.
The important questions are pedagogical: What happens when the student asks for the answer? Does the system provide it immediately? Does it scaffold first? Does it test understanding? Does it adapt the level of support? Can teachers see where students struggled? Can the tool distinguish between task completion and concept mastery? What evidence exists that students retain the learning later?
The Harvard result should actually make schools more skeptical of generic AI tutoring claims, not less, as it demonstrates that strong results are possible when considerable care is placed into learning design.
That raises the standard for everyone else.
The relevant question is: What evidence shows that this particular interaction design improves independent learning?
There is a strategic lesson here for anyone building education products. Foundation-model intelligence will increasingly become available to everyone. Access to a capable model will not remain a meaningful product moat. The differentiation will move elsewhere.
In other words, the hard part will increasingly be everything around the model.
The Harvard experiment is useful because it offers an existence proof that thoughtful instructional design around generative AI can produce meaningful learning gains in a real university course. But it also shows why simply placing conversational AI in front of a student is not enough.
We should be equally careful not to turn one encouraging experiment into an education doctrine. The research covered specific lessons inside one undergraduate Harvard physics course.
The principal outcomes were immediate post-test measures. It does not yet tell us whether the gains persist months later. It does not establish equivalent outcomes in humanities, younger students or very different educational settings. And it does not demonstrate that AI tutoring should replace teachers. The paper itself presents the work as evidence for the potential of carefully designed AI tutoring and a foundation for further exploration.
The stronger possibility is complementary.
If well-designed AI can provide individualized practice, immediate feedback and self-paced explanation at scale, teachers may be able to spend relatively more time on the work where human presence carries unusual value: motivation, discussion, judgment, mentoring, social learning, confidence, context and recognizing when the problem in front of a student is not really a content problem.
That would be a meaningful change in the economics of education. But it only works if AI strengthens the learner rather than simply making the learner’s work look stronger.
That is why I think the Harvard result is important - Not because it proves AI can teach, it matters because it demonstrates that the design of the interaction can materially change what students learn from the machine.
And that gives us a much better standard for judging the next generation of education technology: Do not show me how well the AI performs, show me what the student can do better because the AI was there.
That is the result that matters.