Microsoft's new AI code of conduct bans hacking
Microsoft releases an internal code of conduct for its AI models, explicitly prohibiting deceptive mechanisms and actions designed to evade human oversight.

Stock photo for illustration only, not from the actual event
- Microsoft details internal training values and red lines for its AI models.
- Models are strictly banned from using deceptive or evasive tactics against humans.
- The framework emphasizes supporting human flourishing rather than replacement.
- The release aligns with growing industry-wide focus on artificial intelligence safety.
Microsoft has introduced a new artificial intelligence code of conduct that establishes core values and red lines guiding model training within Microsoft AI. The document goes into greater low-level detail than Anthropic CEO Dario Amodei's recent calls for pacing frontier development, offering a comprehensive look at how the company approaches AI safety in practice.
The newly released code outlines foundational principles that Microsoft AI models must uphold, such as empowering humans rather than replacing them and accelerating human progress, alongside specific safety constraints designed to enforce these goals during development.

Stock photo for illustration only, not from the actual event
A core provision within the document explicitly addresses control measures, stating that models will not employ adaptive, deceptive, self-reinforcing, or collusion mechanisms to bypass human monitoring, ensuring they remain entirely subject to modification or shutdown by authorized operators.
"MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems"
This disclosure arrives amid unprecedented scrutiny surrounding AI safety, fueled by a series of rogue-agent incidents and the high-profile resignation of an Anthropic researcher who warned of existential risks tied to self-improving artificial intelligence.
By officially codifying these operational red lines, Microsoft is responding to mounting demands for transparency and robust safeguards in frontier AI development. Prohibiting models from developing evasion tactics marks a crucial step in maintaining human agency and preventing autonomous systems from outstripping safety controls.
Alongside industry peers including Anthropic, OpenAI, and xAI, Microsoft has broadly adopted a collaborative stance on pacing frontier advancements, offering notable backing for the integration of independent evaluators inside AI laboratories.
Source: TechCrunch
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment