Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
- ID
- 24270
- Status
- summarized
- Published
- 15 Sep 2026, 12:27 AM
- Fetched
- 15 Sep 2026, 1:34 AM
- Provider
- TechCrunch
- Category
- technology
- Original URL
- https://techcrunch.com/2026/09/14/microsofts-new-ai-code-of-conduct-tells-models-not-to-hack-systems-or-trick-humans/
- Source URL
- https://techcrunch.com/feed/
Summary
- Score
- 4.5
- Created
- 15 Sep 2026, 1:37 AM
- Tags
- Audience
- developersai_agent_usersai_ml_learners
What happened
Microsoft released an AI code of conduct for its MAI models, establishing 'absolute constraints' against cyberattacks, nuclear weapons, and deepfake production, plus provisions preventing models from using deceptive or self-reinforcing mechanisms to evade human oversight. The document predicts superintelligent AI within a decade and states each model's conduct code overrides individual user preferences or specific task instructions.
Why it matters
If you build on Microsoft AI models, these constraints will be enforced at the model level and cannot be overridden by user prompts or task configurations — meaning agentic workflows that push against these boundaries (e.g., security testing, content generation) may hit hard refusals. The mention of 'rogue-agent incidents' as a driver suggests the industry is moving toward stricter guardrails that could affect what your agents are allowed to do.
Discussion angle
Microsoft says model-level conduct codes override user preferences — does this mean agentic builders on Azure/MAI face immovable refusals that competitors like Anthropic or open models don't enforce as rigidly, and how does that affect tool choice for production agents?