What Breaks in the First 90 Days of an Agent Deployment

A sober look at the failure modes that hit enterprise agent deployments in the first 90 days, what they cost to fix, and how to scope around them.

Eric Lamanna7 min read
An industrial control panel with amber warning lamps in a dim server room

Most agent programs look healthy at launch. The pilot cleared its acceptance tests, the vendor's demo environment behaved, and the first week of production shows green dashboards. The failures that matter arrive later, usually between day 20 and day 70, and almost always from the same short list of causes: tool contracts drift, approvers stop reading, retries hide behind success codes, lineage breaks across hops, and permissions quietly accumulate. These are predictable. They belong in the Statement of Work, in the pre-go-live checklist, and in the first-quarter budget.

The industry data suggests how much headroom to leave. A 2025 S&P Global Market Intelligence survey of more than 1,000 enterprises found that 42% of companies abandoned most of their AI initiatives by mid-year, up from 17% in 2024. Deloitte's 2025 trends work put the production share lower still, with only 14% having solutions ready for deployment and 11% actively running them. The gap between a working pilot and a surviving production system is where this article lives.

Tool Drift Is the First Thing to Break

An agent's reliability is a function of its tool registry, not its model. On day one, every tool call matches a schema the agent was trained and evaluated against. By week six, a payments API has added a required field, a CRM has renamed an enum, and an internal microservice has silently moved a timeout from 30 seconds to 10. The agent keeps calling. The calls keep failing in ways that look like model errors.

Salesforce AI Research's CRMArena-Pro benchmark is a useful calibration here: leading LLM agents hit only 58% success on single-step CRM tasks and 35% on multi-step tasks. Those numbers already assume stable tools. In production, tool drift subtracts from both.

Put three things in the SOW. First, a versioned tool registry with contract tests that run on every upstream schema change, not just on agent releases. Second, a deprecation window — 30 days minimum — between a tool's new version appearing and the old version being withdrawn. Third, a named owner for each tool, because an unowned connector drifts until something expensive breaks. Reserve roughly 8–15% of first-quarter engineering hours for registry maintenance; teams that budget zero end up paying it in incident response at two to three times the rate.

Where Post-Launch Incident Volume Accumulates
Where Post-Launch Incident Volume AccumulatesDay 7: 4; Day 14: 11; Day 21: 22; Day 30: 38; Day 45: 64; Day 60: 92; Day 75: 118; Day 90: 141070.51414Day 711Day 1422Day 2138Day 3064Day 4592Day 60118Day 75141Day 90
Illustrative volume of incidents attributable to each cause over the first 90 days. Illustrative: a visual comparison, not measured data.

Approval Fatigue Turns Humans Into Rubber Stamps

Approval gates are the standard answer to agent risk, and they are the first guardrail to decay. Reviewers who see twenty routine approvals an hour learn to click through. By the time an unusual request arrives, the muscle memory is already wrong. One industry panel on 2026 operations reported that legal and compliance agents average a 61% human-in-the-loop intervention rate, while SDR agents sit near 8% — the higher the rate, the faster fatigue sets in.

Design the gate against the actual decision, not the action. Three practical rules:

  • Tiered thresholds. Below a defined blast radius (dollar value, record count, external recipients), auto-approve with full logging. Above it, require a human. Above a second threshold, require two.
  • Diff-based review. Show the reviewer what changes, not what the agent proposes in full. A wire transfer approval that highlights "new counterparty, first payment, amount 4x historical mean" gets read; a 40-line JSON blob does not.
  • Forced friction on anomalies. When an action falls outside the agent's training distribution, insert a mandatory delay and a free-text justification field. Both measurably reduce rubber-stamping.

For sequencing which actions actually warrant a human, the taxonomy in agent approval tiers is a reasonable starting point, and the mechanics of the gate itself belong in a dedicated approval gates design review before go-live.

A rubber stamp above an endless paper ribbon of approval forms
Action-Gated vs Decision-Gated Approvals
Action-Gated vs Decision-Gated ApprovalsReviewer attention per item: 20; Items requiring human review: 95; Anomalies actually caught: 30; Rubber-stamp rate: 70; Context shown with each item: 25Reviewer attention peritem2080Items requiring humanreview9525Anomalies actuallycaught3075Rubber-stamp rate7015Context shown witheach item2585
Illustrative contrast between gating every action and gating by decision risk — which approach preserves attention where? Illustrative: a visual comparison, not measured data.

Silent Retries Corrupt State Faster Than Loud Failures

A loud failure gets paged. A silent retry gets a 200. Agents that call tools through naive HTTP clients will quietly reattempt on timeouts, duplicate writes to downstream systems, and present a clean trace to the operator. The symptoms show up two weeks later as duplicate invoices, double-booked calendar slots, and reconciliation tickets whose root cause is unrecoverable.

The fix is not more retries; it is idempotency keys on every state-changing call, dead-letter queues for calls the agent cannot resolve, and a reconciliation job that compares agent-asserted outcomes against the system of record on a defined cadence. OWASP's agentic guidance ranks Tool Misuse and Exploitation second on its 2026 list precisely because a legitimate tool used in a legitimate way can still produce illegitimate state when retry semantics are wrong.

Budget for a reconciliation harness in the SOW, not as a nice-to-have. Teams that treat it as optional discover the gap when finance does. The underlying patterns — idempotency, compensating actions, replay — are the same ones that make eventually consistent workflows trustworthy at all.

Lineage Gaps Make Incidents Unanswerable

The question an auditor or a regulator asks after an incident is simple: which input, which model version, which tool version, which policy, which approver, produced this specific action. If any link is missing, the answer is "we don't know," which is the wrong answer in a regulated environment. Vectara's open hallucination leaderboard shows the best frontier models hallucinate at roughly 0.7%–1.5% on document-grounded summarization; that is a narrow task and a floor, not a ceiling. On open-ended agent work, the floor is higher, and the audit trail is the only way to reason about what went wrong.

A defensible lineage record ties together, per action: the prompt and retrieved context (hashed if sensitive), the model identifier and version, the tool registry version, the policy bundle version, the input and output of every tool call, the approver identity and timestamp, and the trace ID that links them. Store it immutably. Make it queryable in minutes, not days. The requirements overlap heavily with what to demand in a vendor's audit trail, and most of the gaps show up at the seams between the agent runtime, the tool proxy, and the ticketing system.

Permission Creep Is the Quietest Failure

Agent identities start narrow and grow. A workflow needs a new SharePoint path, so a scope is added. A new reporting job needs read on a warehouse schema, so a role is attached. Six months in, the agent holds a permission set no one person could justify in a single meeting. OWASP's Excessive Agency entry breaks this into three root causes — excessive functionality, excessive permissions, and excessive autonomy — and the first one usually drives the other two.

Require three mechanisms before go-live. Short-lived, task-scoped credentials rather than static service accounts. A quarterly entitlement review that treats the agent as a non-human identity subject to the same joiner-mover-leaver discipline as a human. And an egress allowlist on every tool, enforced at the network layer, so an agent that acquires a surprising capability cannot immediately exercise it against the internet.

How Excessive Agency Compounds
How Excessive Agency CompoundsTools beyond task scope: 60; Credentials broader than needed: 55; Actions without human approval: 4535300Tools beyond t…60Credentials br…55Actions without human approval45
Illustrative: a visual comparison, not measured data.

What to Budget for the First 90 Days

The build budget is not the production budget. A credible first quarter reserves engineering time against each of the failure modes above, not against a single contingency line. Rough shape, as a percentage of the initial build cost:

  • Tool registry maintenance and contract tests: 10–15%.
  • Approval gate tuning and reviewer training: 5–8%.
  • Reconciliation harness and dead-letter handling: 8–12%.
  • Observability, lineage, and incident response: 10–15%.
  • Identity, entitlement review, and egress controls: 5–8%.

That is 38–58% of the build cost, held in reserve for the quarter after launch. Programs that budget less tend to show up in the rollback statistics; programs that budget more rarely need to spend it all, but they keep the option of fixing a problem in a week rather than a quarter. The decisions about where that money sits — in the vendor contract, in a shared platform team, or in the business unit — deserve the same scrutiny as the original pricing model and the underlying agent architecture.

First-Quarter Budget Posture: Reserve vs Rollback Risk
First-Quarter Budget Posture: Reserve vs Rollback RiskNo post-launch reserve: 10; Contingency line only: 25; Partial reserve (ad hoc): 45; Line-item reserve per failure mode: 75; Full 38–58% reserve + owners: 90 → →123451No post-launch reserve2Contingency line only3Partial reserve (ad hoc)4Line-item reserve per failure mode5Full 38–58% reserve + owners
Illustrative placement of common budgeting postures on two axes — read where each posture lands, not the exact coordinates. Illustrative: a visual comparison, not measured data.

A Short Risk Register for Day 90

The useful artifact to leave the executive sponsor with is not a framework. It is a one-page register, reviewed weekly, with named owners and clear thresholds: tool drift incidents per week, approval override rate, silent-retry count detected by the reconciliation harness, lineage completeness percentage, and net new permissions granted since the last review. Each line has a number, a trigger, and an action. When a trigger fires, the response is already written down. That register, not the launch day screenshot, is what distinguishes an agent that survives its first quarter from one that becomes the next rollback case study.

// written by
Eric Lamanna
Director of Business Development

Eric Lamanna is a Digital Sales Manager with a strong passion for software and website development, AI, automation, and cybersecurity. With a background in multimedia design and years of hands-on experience in tech-driven sales, Eric thrives at the intersection of innovation and strategy—helping businesses grow through smart, scalable solutions. He specializes in streamlining workflows, improving digital security, and guiding clients through the fast-changing landscape of technology. Known for building strong, lasting relationships, Eric is committed to delivering results that make a meaningful difference. He holds a degree in multimedia design from Olympic College and lives in Denver, Colorado, with his wife and children.

Put an agent to work, the right way.

Start on Automatic and put the workflow you want to automate in front of engineers who have shipped agents in regulated environments.

Explore services
// the briefing

Agentic AI, in your inbox.

Occasional, high-signal notes on building and operating AI agents — automation patterns, architecture, and governance. No spam.