AI Agent Safety Is More Than Model Safety
A safe model is not enough. Companies also need to control what an AI agent can do and what people may believe because of it.

AI agents are starting to work inside real companies.
They write code, search internal documents, answer employee questions, and connect to company systems.
That creates a different kind of security problem.
An AI agent can cause harm in two ways: by doing the wrong thing, or by making a human believe the wrong thing.
The first problem is about actions.
Spotify has described agentic development environments where coding agents work with controlled permissions and isolated environments. The idea is familiar from traditional security: do not give software more access than it needs.
Uber has also discussed extending identity and Zero Trust ideas to AI agents. An agent should have an identity, clear permissions, and extra controls around sensitive actions.
The principle is simple: Do not ask the agent to enforce every security rule by itself. Build the rules around it.
But actions are only half of the problem.
An agent can be perfectly isolated from production and still give a confident, wrong answer. If an employee makes a real decision based on that answer, the damage is no longer technical. It is informational.
Airbnb’s AI support work shows the value of grounding answers in a limited set of trusted help and account information, with escalation to a human when needed.
Morgan Stanley has taken a similar approach in knowledge assistants for employees: answers are grounded in approved internal content, users can reach the original source material, and new use cases are evaluated before wider deployment.
This gives us a second principle: A safe agent should not only know what it is allowed to do. It should know what it is allowed to claim.
Trusted sources matter. Access-aware retrieval matters. Citations matter. And sometimes the safest answer is simply: “I don’t know.”
That is not a weakness. It is a safety feature.
As AI agents become part of everyday work, model safety will remain important. But the model is only one layer.
The safest agent may not be the one with the safest model. It may be the agent inside the safest system.
See you in a Thoughtful Future.
References
- Spotify Engineering. https://engineering.atspotify.com/
- Uber Engineering. https://www.uber.com/blog/engineering/
- Airbnb Tech. https://medium.com/airbnb-engineering
- Morgan Stanley. https://www.morganstanley.com/
- OpenAI Codex. https://openai.com/codex/