OpenAI board member Zico Kolter says AI cyber capabilities demand new safety measures
The Carnegie Mellon professor who chairs OpenAI's Safety and Security Committee argues the industry's security playbook needs a rewrite
Zico Kolter, an OpenAI board member and Carnegie Mellon University professor, says the AI industry’s growing awareness of what its models can do in cyberspace is forcing a rethink of how security works.
The comments carry weight because of where Kolter sits. He chairs OpenAI’s Safety and Security Committee, the internal body that can hold back a model launch until its risks are dealt with.
What Kolter is actually saying
Kolter’s core point is that AI systems now present a threat landscape traditional software security was never designed to handle. In his view, that landscape includes AI-enhanced cyberattacks and even help with bioweapon design.
Kolter has singled out two risks in particular: prompt injection and agent autonomy. He describes both as problems with no real equivalent in traditional security.
Prompt injection is when someone hides instructions inside content an AI reads, such as a webpage or an email, and tricks the model into following them.
Agent autonomy is the related worry about AI systems that take actions on their own, like browsing, executing tasks or touching other software. The more freedom an agent has, the more damage a single successful trick can cause.
Kolter’s prescription is a layered one. He emphasizes explicit security training for models, ongoing monitoring, filtering mechanisms and proactive testing of AI systems before and after release.
He also pushes back on a popular assumption in the industry. The idea that smarter models will naturally become safer models does not hold up, according to his approach. Capability and safety, in other words, are separate projects that need separate investment.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
The committee with a pause button
Kolter joined OpenAI’s board in August 2024 and has chaired the Safety and Security Committee since then. The committee, known internally as the SSC, has the authority to request delays on major model releases until identified safety concerns are resolved.
That is a meaningful lever at a company whose business depends on shipping new models. It is also a lever used quietly: the SSC keeps confidential any past actions it has taken.
His role gained more formal footing in November 2025. That month, OpenAI reached agreements with state attorneys general that strengthened the SSC’s governance role as the company restructured.
The restructuring moved OpenAI toward a public benefit corporation controlled by a nonprofit foundation. The attorneys general agreements effectively wrote the safety committee’s influence into that new structure, rather than leaving it as a voluntary internal arrangement.
A researcher who builds the attacks he warns about
Kolter is not a pure policy figure commenting from the sidelines. He co-founded Gray Swan AI around 2023-2024, a company focused on AI security and prompt-injection vulnerabilities.
He has also kept up a public presence on the issue. As of October 2026, he continues to take part in safety discussions at international forums, including the World Summit AI in Amsterdam.
What this means
For the AI industry, Kolter’s framing shifts security from an afterthought to a design requirement. If models need explicit security training and constant monitoring, safety becomes an ongoing operating cost rather than a one-time checklist before launch.
The OpenAI arrangement also offers a template others may be measured against. Tying the SSC’s authority to agreements with state attorneys general links safety practice directly to regulatory compliance.
The SSC’s confidentiality means the public has limited visibility into how its power is exercised in practice.