AI Agents With Authority But No Boundaries: Why They Fail

When a financial services company grants an AI agent permission to act, the question in the room is almost always ‘what can it do?’ The question that rarely gets asked is: ‘Where must it stop, when should it refer a decision back to a human, and what happens when it fails?’ A mid-sized insurance brokerage in Istanbul, operating with 312 staff, ran directly into this gap at the end of 2024. The firm wanted to automate policy renewal workflows through an agent. The agent was given access to customer records, email-sending rights, and the ability to update policy parameters. The outcome: the agent renewed 47 policies that should not have been renewed before their review window and applied incorrect premium calculations to a subset of them. The system logged no errors. The agent completed its task. The failure was that authority had been defined without limits.The central argument here is one that practitioners pushing back against it will find uncomfortable: AI agent failures are predominantly not technical failures. They are governance failures. Organizations tell agents what they can do. They do not define what agents cannot do, when they must pause, and how the system should behave when something goes wrong. Every agent deployment that ignores this distinction becomes an operational risk source within weeks. Trusting an agent is not the same as bounding an agent. In 2025, this distinction remains one that most Turkish organizations have not yet operationalized in a rigorous way.The authority matrix concept is not new. Human organizations have used it for decades: who can approve expenditure up to what threshold, which decisions require senior sign-off, which transactions are frozen when a limit is breached. For AI agents, this framework is either absent in most deployments or buried inside prompt text rather than enforced at the system configuration level. Prompt-based constraint setting is not a reliable governance mechanism in agentic systems. When context windows overflow, when chains grow long, or when unexpected input arrives, those constraints collapse. The IT director of an Ankara-based factoring firm described the problem plainly: ‘We wrote what the agent should not do. We never checked where in the system that instruction was actually being read.’ An authority matrix is not a document. It is a configuration the system reads, enforces, and logs at runtime.Escalation rules are the second missing layer. A well-designed agent architecture does not attempt to complete every task autonomously. At defined conditions, it returns control to a human decision-maker. Consider a customer dispute management agent at a leasing company in Bursa: the agent classifies incoming disputes, matches them against a standard response library, and sends automated replies. If no escalation rule exists, the agent will respond to a high-stakes dispute message — one containing a legal threat, for example — with the same standard template it uses for a routine billing question. There is no technical fault. The task definition was fulfilled. The business consequence, however, can be severe. Escalation rules can be built around classification confidence score thresholds, sentiment analysis alerts, or specific keyword triggers. The critical point is that these thresholds need to be designed, measured, and updated over time. An agent is not a system you configure once and release. It is an operational process requiring ongoing governance.Fail-safe design is the least discussed yet most operationally critical layer. Agent failure is inevitable. A language model context window filling to capacity, an external API timing out, an unexpected data format arriving mid-chain — any of these can cause an agent to abandon a task mid-execution. The question is not ‘why did the agent fail?’ The question is: ‘What did the system do when it failed?’ An accounting software provider in Izmir learned this the hard way with a reconciliation agent it had deployed for clients. When the agent could not retrieve a bank statement feed, it marked the task as ‘completed’ because no distinct exit code had been defined for failure states. Accountants worked with incomplete reconciliations for weeks without realizing the shortfall. Fail-safe design requires differentiating error codes, routing failed tasks to a dedicated queue, and ensuring that queue receives regular human review. This is not a limitation of the technology. It is a requirement of operational maturity.With the EU AI Act entering application phases in 2025, this is no longer purely an operational preference. For high-risk applications — insurance scoring, credit decisions, recruitment tools — human oversight is mandated. Turkish organizations with exposure to EU markets carry this compliance obligation. But there is a more significant point: institutions that use EU AI Act compliance not as a checkbox exercise but as a design framework for authority matrices and escalation logic accomplish two things simultaneously. They achieve compliance and they build operational discipline. Returning to the Istanbul insurance brokerage: when the firm entered its remediation process, it restructured its authority matrix into three categories — full agent authority, agent recommendation requiring human approval, and agent prohibition. That third category is what most organizations skip entirely. Which decisions can never be delegated to an agent under any circumstances? Asking that question before the policy renewal deployment would have prevented 47 incorrect renewals.Three concrete steps for this week. First, list every active agent deployment in your organization and ask whether agent authority is defined in a system configuration or only in prompt text — if it is only in a prompt, that boundary is not reliable. Second, for each agent, identify the specific measurable condition that triggers human escalation; if you cannot name an explicit threshold, the escalation logic does not exist yet. Third, test what happens when the agent fails mid-task: does the system log the failure state distinctly, does the incomplete task surface for review, or does it disappear? The Istanbul brokerage worked through this remediation over 11 weeks. For the organizations that run this audit before the first major incident, that time is spent on design rather than damage control. Authority without boundary is not a feature. It is a design omission waiting to surface.

This article was originally published in Turkish by Gökhan MERCANOĞLU on January 27, 2025. The English edition has been reviewed and edited by the author.


When machine learning model succeeds, it does not merely put more information on a screen; it gives management clearer decisions. Silos decrease, responsibility becomes visible, and measurable progress starts. Therefore, the issue is not tool selection but rebuilding operating discipline through technology.


Gökhan Mercanoğlu
Yapay Zekâ ve Makine Öğrenmesi