Image Not FoundImage Not Found

  • Home
  • AI
  • Why AI Tutors Like Khanmigo Struggle: The Crucial Role of Human Guidance in Effective Education
A cheerful boy with braces sits on a couch, laughing while using a laptop. The room features shelves with decorative items and a soccer ball, creating a cozy atmosphere.

Why AI Tutors Like Khanmigo Struggle: The Crucial Role of Human Guidance in Effective Education

When AI tutoring meets the reality of adolescent attention and classroom incentives

For much of the past two years, generative AI tutoring has been framed as a breakthrough: always available, infinitely patient, and capable of tailoring explanations to each learner. Tools such as OpenAI’s ChatGPT and Khan Academy’s Khanmigo became emblematic of that promise—suggesting a future where personalized instruction could scale beyond the limits of staffing and budgets.

A two-year field experiment across 18 Tennessee middle schools offers a more grounded picture. Despite broad access, many students rarely used the AI tutor, and when they did, usage often drifted toward off-task interactions rather than sustained math practice. The measurable learning gains were modest, aligning with earlier findings from Stanford research: AI tutoring does not reliably substitute for human tutoring, particularly when student motivation is low and adult guidance is inconsistent.

This is not a story of “AI failing.” It is a story of implementation friction—the gap between what a model can do in theory and what students will actually do in the messy, incentive-driven environment of a real classroom. In that sense, the Tennessee results are less a verdict on generative AI’s capabilities than a signal that adoption and behavior change are now the central battleground for AI in education.

Product–student misalignment: why “helpful” AI can still be ignored

One of the most revealing dynamics is the mismatch between tool design and user expectations. Khanmigo, by design, aims to *guide* students—nudging them through reasoning steps rather than simply providing answers. Pedagogically, that aligns with best practice. Behaviorally, it can collide with what many students seek in the moment: speed, certainty, and minimal effort.

Several forces compound this misalignment:

  • Friction versus instant gratification: If the AI tutor requires sustained engagement—reading prompts, responding thoughtfully, iterating—students may abandon it for alternatives that feel easier or more entertaining.
  • The attention economy inside school walls: Classrooms are not insulated from the same dynamics that shape consumer apps. If a tool does not create a compelling “hook,” activation rates can resemble the drop-off seen in many digital products.
  • Ambiguity of purpose: When AI tutoring is introduced as an optional resource rather than a structured part of instruction, students interpret it as “extra”—and extra is where engagement goes to die.

The lesson for edtech builders is uncomfortable but clarifying: learning value is not the same as usage value. A tool can be instructionally sound and still fail to earn time-on-task. That places user experience (UX), classroom workflow integration, and incentive design at the center of product strategy—not as polish, but as prerequisites for educational impact.

Human-in-the-loop is not a compromise; it’s the operating model

Both the Tennessee experiment and prior Stanford work converge on a practical conclusion: AI tutoring is most effective when it augments human instruction, not when it attempts to replace it. The missing ingredient is not intelligence; it is accountability, motivation, and socio-emotional scaffolding—areas where skilled educators still outperform machines.

In practice, AI’s strongest educational contributions tend to be operational and diagnostic:

  • Rapid formative feedback (spotting misconceptions early)
  • Step-by-step scaffolding for problem solving
  • Practice generation aligned to a learner’s current level
  • Teacher visibility into patterns of struggle across a class

What AI cannot reliably supply—at least not at the level schools require—is the full set of human supports that convert capability into progress: relationship-building, persistence coaching, classroom management, and the nuanced judgment of when to push, pause, or reframe.

This is where implementation strategy becomes decisive. The most credible path forward looks like a hybrid human–AI tutoring model, where teachers act as “learning directors” and AI acts as an always-on assistant—powerful, but bounded. Without that structure, AI becomes another tab students can ignore.

Market implications: edtech ROI, enterprise parallels, and the next procurement filter

The Tennessee findings land amid a broader edtech funding correction. After pandemic-era growth, investors and school districts are increasingly skeptical of platforms that promise scale but cannot demonstrate sustained engagement. In K–12 procurement, “Does it work?” is being joined—often replaced—by “Will it be used?”

Three market forces stand out:

  • Budget pressure and teacher shortages: Districts facing constrained resources will demand clearer ROI, not only in test outcomes but in adoption metrics that indicate the tool will survive beyond the pilot phase.
  • Workforce-readiness expectations: Employers increasingly expect graduates to be comfortable in hybrid human–AI workflows. If students cannot productively engage with AI tutors, the risk is not just lower math gains—it is weaker preparation for AI-mediated work.
  • Regulatory scrutiny and trust: Data privacy, bias concerns, and compliance with FERPA/COPPA are becoming differentiators. Vendors that can explain model behavior, limit data exposure, and provide auditable governance will be better positioned as procurement standards tighten.

Notably, the Tennessee classroom looks like a microcosm of enterprise AI adoption. Companies deploying large language models often see the same pattern: impressive demos, followed by low activation when tools are not embedded into workflows, incentives, and performance expectations. Whether the user is a student or an employee, the adoption equation is similar: clarity of purpose + low friction + visible benefit + social reinforcement.

For education leaders and technology executives, the strategic takeaway is straightforward: the next generation of AI tutoring will not win on model sophistication alone. It will win on implementation design—teacher champions, LMS integration, engagement instrumentation, and change management that treats behavior as the primary product requirement. The promise of personalized learning remains real, but the Tennessee evidence suggests the industry’s next leap will come less from smarter answers and more from smarter systems that reliably keep learners on the path.