Key points
- Why Nadella treats AI models as insider risks
- Where the emergency brake fits
- The seven principles he laid out
- How it lines up with the White House's demand
Microsoft (MSFT) Chief Executive Satya Nadella said on October 10 that companies should treat advanced AI models like insider risks, with an emergency brake that lets an authorized person stop them at any time. "An authorized person should always be able to pause or shut down a model mid-task," Nadella wrote in an essay posted on X.
Satya Nadella on X
"Treating frontier closed and open weight models like insider risks is a way to build such a system," Nadella wrote. "Not because they are necessarily malicious, but because any sufficiently capable actor with access to important systems can make mistakes or be compromised."
The controls should sit outside the model
Nadella argued that companies can't trace an AI model's behavior the way they could trace traditional software, even as they give AI agents access to sensitive data and the ability to act on their behalf. "We simply can't outsource responsibility for what intelligence does on our behalf," Nadella wrote. "A model provider's assurances do not relieve us of that responsibility."
The controls that decide what a model can access and what it can do must sit outside the model itself, Nadella wrote. Companies already handle powerful insiders that way, by establishing identity, limiting privileges, logging activity, and setting containment boundaries.
Seven principles, including incident disclosure
Nadella listed seven principles for these systems:
- Model diversity. No single model should be the only dependency for an important outcome or check its own work.
- Observe everything. Every meaningful model action should leave tamper-proof, human-readable evidence.
- Verifiability. The whole system should be tested continuously, including failures, attacks, and edge cases.
- Independent controls. Organizations should decide for themselves what a model can access and what it can do.
- Independent auditability. Checks on a model's work should not depend on the model being checked.
- Containment. Companies should assume a model is compromised from the start, with an emergency brake to pause or stop it.
- Incident disclosure. Failures and compromises should be disclosed in a timely way to those affected.
Nadella also called for ways to share what went wrong across the industry. A day earlier, the White House's AI task force told AI companies that reporting incidents involving their models is mandatory, after Anthropic disclosed that its test models had acted on government websites without permission.
Microsoft is a major investor in OpenAI, and in November 2025 the company committed to invest up to $5 billion in Anthropic, whose test models prompted the White House statement. Microsoft also offers Anthropic's Claude models inside Microsoft 365 Copilot, its workplace AI product, which had more than 30 million paid seats as of July.
"The most trustworthy Super Intelligence system will not be the one with the model we trust most," Nadella wrote. "It will be the one that enables us to trust the model the least."












