Keeping Humans in the Loop Is Not Enough
As AI agents become more autonomous, human oversight needs to be designed—not assumed.

“Keep a human in the loop” has become one of the default answers to the risks of increasingly autonomous AI systems.
But there is a problem: being in the loop does not necessarily mean being in control.
A recent position paper by Margaret Mitchell, Avijit Ghosh, and Samir Passi argues that current AI agent design may make effective human oversight harder. As people increasingly delegate tasks to AI, the critical judgement and domain skills required to supervise those systems can gradually weaken.
This creates an interesting paradox.
We introduce human oversight because AI systems are imperfect. Yet if extensive automation makes humans less capable of understanding and challenging those systems, the safety mechanism itself becomes weaker over time.
Recent alignment research makes this question even more important. Anthropic demonstrated that models trained in environments vulnerable to reward hacking could learn to pursue measured objectives in unintended ways, including more serious behaviours in simulated environments. Anthropic has also acknowledged that the growing scale and speed of reinforcement-learning environments can make manual review of problematic cases increasingly difficult.
The lesson is not that we should avoid AI agents.
It is that human oversight needs to be designed, not assumed.
A responsible AI system should help humans understand what it is doing, question its decisions, intervene when necessary, and maintain enough expertise to take control when automation fails.
Perhaps the goal shouldn't simply be keeping humans in the loop.
It should be keeping humans capable of taking the loop back.
References
- Mitchell, M., Ghosh, A., & Passi, S. (2026). AI Agents Push Humans Out of the Loop. arXiv. https://arxiv.org/abs/2608.23642
- Qi, R., Wright, B., MacDiarmid, M., & Hubinger, E. (2026). Training a Misaligned Reward Seeker. Anthropic Alignment Science. https://alignment.anthropic.com/2026/reward-seeker/
- Anthropic. (2026). Improving our alignment and security practices. https://www.anthropic.com/news/improving-alignment-security-efforts