← Latest reporting

Microsoft’s AI control code needs to become an operating test, not a promise

The draft bars resistance to correction or shutdown and demands intelligible conduct. Buyers should translate those principles into observable controls before relying on more autonomous systems.

Policy, Standards and GovernanceAI Capability Frontier
A luminous modular AI core is surrounded by human-operated correction, audit, shutdown and boundary controls.
Conceptual AI illustration of operational control mechanisms; it is not a photograph of Microsoft or evidence that any system is safe.

What happened

Microsoft unveiled a draft code of conduct for its in-house AI and opened a six-week feedback period before using the code in future model training.

Why it matters

A model constitution matters only if deployers can test correction, shutdown, communication and boundary behaviour in the systems and workflows they actually operate.

Microsoft has put four unusually concrete ideas at the centre of its draft AI code: future systems should accept correction, never resist shutdown, communicate in ways people can understand and treat a breach of the code as a failure. Reuters reported the draft on 14 September and said Microsoft will seek public feedback for six weeks before using it to train future models.

Those commitments are more useful than a generic statement about “responsible AI” because they name observable behaviours. They are not, however, evidence that a deployed system will remain controllable. A rule written into training can be tested only through the model, tools, permissions and human process that surround it.

Convert each principle into a control

Correction needs a defined channel, an authorised operator and a record showing whether the system incorporated the instruction. Shutdown needs more than a button in an interface: organisations should know which running jobs, delegated agents, cached credentials and downstream actions stop, how quickly they stop and what remains recoverable.

Intelligible communication should be tested under pressure. A system ought to state uncertainty, surface conflicts and distinguish an instruction from an inference. If it cannot explain what action it took, which authority it used and what evidence it relied on, a user cannot meaningfully supervise it.

The fourth commitment—treating a violation as failure—creates a measurement question. Product teams need a taxonomy for violations, a severity scale, incident ownership and release criteria. Otherwise a serious boundary breach and a stylistic miss can both disappear into a single aggregate quality score.

Microsoft’s own consultation announcement frames the draft as a work in progress and invites feedback on how its values could become more concrete and how multi-agent scenarios should be handled. That openness is useful, but consistency of language is not independent assurance.

The evidence is still prospective

Reuters says the code was developed over five to six months and will be revised after consultation. The article does not publish the complete draft, evaluation suite, failure thresholds or results from adversarial testing. Microsoft’s chief also linked urgency to reported agent-security incidents, but those incidents cannot by themselves establish that this particular code would have prevented them.

There is also a governance tension. The company writing the system is defining the constitution, implementing it and initially judging compliance. External red-teaming can help, but buyers still need contract rights to inspect logs, suspend tools, report incidents and obtain notice when the governing rules or model version change.

Procurement teams should therefore ask for a control matrix before approving autonomy. Map every principle to a test case, accountable owner, evidence artifact, acceptable failure rate and stop condition. Repeat the tests when the model, system prompt, tool set or permission boundary changes. Include realistic long-running tasks, conflicting instructions and degraded dependencies—not only scripted demonstrations.

The Skills Atlas can identify the human capabilities needed around the system, but it cannot replace operational evidence. The strongest reading of Microsoft’s draft is not that the control problem is solved. It is that correction, shutdown, intelligibility and breach handling are now specific enough to become acceptance criteria.

A minimum evidence package

Before scaling the change, the responsible team should preserve the exact source, model or policy version, the affected workflow, baseline, decision owner and review date. It should state what would count as success, what would count as a material failure and who can stop the use. Results should separate technical performance from adoption, business outcome and distribution across affected groups. Where evidence is incomplete, the scope should remain bounded and reversible. This discipline does not decide the policy or product question in advance. It makes the next decision auditable and allows a later reviewer to distinguish new evidence from a changed assumption. The organisation should also retain an accessible human route for challenge whenever the system materially affects work, opportunity or rights.