← Latest reporting

A school AI moratorium needs an evaluation plan, not a permanent default

New York City public-school guidance limits student-facing AI while the evidence base remains mixed. A pause can reduce immediate risk, but it should define what evidence, safeguards and learning outcomes would justify continuation, redesign or exit.

Policy, Standards and GovernanceWork and Role Change
A staged conceptual classroom scene shows a pause gate, a protected pilot lane and an evidence checkpoint, rendered as paper theatre rather than documentary photography.
Conceptual AI-generated illustration of a time-bounded school AI pause and evaluated pilot; it does not depict a real classroom or policy meeting.

What happened

New York City Public Schools published guidance on artificial intelligence and screen time that constrains student-facing uses and sets expectations for review; a Stanford evidence review finds limited causal K–12 evidence.

Why it matters

A moratorium can become indefinite if it has no baseline, permitted pilot, evaluation design or decision date. Schools need to test safety and educational value without treating novelty or prohibition as evidence.

New York City Public Schools' AI guidance sets expectations for artificial-intelligence use in the school system alongside wider attention to screen time and student safety. A Stanford review of the K–12 evidence base finds a field with heterogeneous studies and limited causal evidence. That combination can justify caution, but it does not tell schools to freeze policy indefinitely.

A moratorium is a control: it can stop uncontrolled procurement, data collection or classroom experimentation while governance catches up. It is not an evaluation result. Without a defined scope and decision date, a temporary pause can become a permanent default even as products, safeguards and educational needs change. Conversely, lifting a pause because tools are popular would be equally weak evidence.

Specify what is paused

The policy should distinguish student-facing instruction, teacher planning, administrative work, accessibility support and research. Risks differ across those uses. A tool that drafts a lesson outline is not equivalent to a system that profiles a child, gives automated feedback or makes a placement recommendation. The pause should also identify prohibited data, age constraints, vendor access rules and whether local experiments require central approval.

Explicit exceptions are important. Schools may need assistive technology, translation or controlled research before the general policy changes. An exception should state the problem, population, safeguards, duration, accountable owner and evidence to collect. It should not become an informal route around the moratorium.

Turn the pause into an evaluation programme

Start with baseline measures before introducing a pilot: learning outcome, teacher workload, student participation, error patterns, accessibility, incidents and distribution across groups. Use a comparison design proportionate to the decision. Where random assignment is impractical, staged rollout or matched classrooms may still be stronger than post-hoc testimonials. Predefine what would count as benefit, unacceptable harm and inconclusive evidence.

Safety evaluation should include privacy, security, age-appropriate design, hallucinated content, bias, dependence, academic integrity and escalation to a qualified adult. Educational evaluation should ask whether the tool improves a defined learning process, not whether students enjoy it or produce more text. Teacher workload must include correction and monitoring, not only preparation time.

The Stanford review is a reason to narrow claims. Limited causal evidence does not prove that all tools fail; it means the system should avoid broad effectiveness claims and invest in better evaluation. Results from one grade, subject or supported pilot should not be generalised automatically. Schools should publish methods and limitations so families and educators can understand what changed.

A decision rule completes the moratorium. On a stated date, evidence should lead to renewal, redesign, limited approval or exit. The authority making that decision should be named, and unresolved uncertainty should be explicit.

Implementation should include families and educators in the review rather than treating consent and communication as an afterthought. Published summaries can explain which uses were tested, which data were processed, what incidents occurred and why the decision rule was met. Procurement contracts should preserve access to logs and evaluation data, allow suspension, and prevent a vendor from redefining success after the pilot. Independent review is especially valuable where a tool affects vulnerable students, special-education support or consequential recommendations.

A moratorium also has costs that should be measured. It may delay accessibility support, push teachers toward unapproved tools or prevent students from learning how to verify AI output. Those are not arguments for automatic approval; they are counterevidence that belongs in the same decision record. The correct comparison is between controlled alternatives, including non-AI options, not between an idealised innovation and a risk-free status quo.

The Skills Atlas can help identify evaluation, verification, privacy and change-management capabilities. A pause creates time; only a structured evaluation converts that time into a defensible policy.