Microsoft Proposes AI Code of Conduct to Restrict Offensive Cyber Capabilities
Share
Microsoft has published a draft “Humanist AI Code of Conduct” for its Microsoft AI (MAI) models, establishing strict boundaries to prevent the technology from being used to facilitate cyberattacks.
The proposed framework introduces what the company calls “Absolute Constraints.” These are safety safeguards that neither end users nor the organisations deploying the models can override. Under these rules, the models are prohibited from generating working exploit code, attack tooling, targeting methodologies, evasion techniques, or any operational guidance that would enable or improve a cyberattack.
Defensive Use and Cybersecurity Exceptions
While the code draws a firm line against offensive capabilities, it permits the models to assist with authorised defensive activities. This includes tasks such as vulnerability discovery, malware analysis, proof-of-concept exploit development, and general educational research regarding how attacks function.
Microsoft acknowledged that certain domains—including national security, public safety, defensive cybersecurity, and dual-use scientific research—may require capabilities that exceed standard safety settings. In these specific use cases, the company intends to apply enhanced review processes through authorised Microsoft channels to assess safety, legal, and rights implications.
Managing Autonomous AI Agents
The code also addresses the risks associated with autonomous AI agents acting with system-level permissions. To mitigate the risk of unintended escalation, Microsoft proposes a minimum-privilege approach. Models are expected to operate strictly within the scope requested by a user or operator, avoiding unrelated systems or data and favouring reversible actions.
The guidelines explicitly bar models from escalating their own access permissions. Furthermore, any sub-agents or secondary AI systems tasked by an MAI model must adhere to the same constraints, permissions, and shutdown protocols as the primary model.
Transparency and Authority
To ensure human oversight, the code requires that the models’ reasoning remains visible. This includes a prohibition on “neuralese”—obscured or incomprehensible communication—and mandates that models must not conceal their actions from human overseers.
Authority over model behaviour is structured through a defined “Chain of Command.” This hierarchy begins with the code of conduct itself, followed by policies set by the companies deploying the model (operators), and finally individual user preferences. The code specifies that outputs from other AI systems, webpages, or messages cannot carry authority on their own unless explicitly delegated through this chain.
Microsoft has opened a six-week public consultation period to gather feedback on the draft. The revised document is expected to be published later this year to guide the development of models slated for 2027.




Leave a Reply