What to Put in an Agentic AI Pilot Contract Before You Sign

The acceptance criteria, kill-switches, data clauses, and exit terms to negotiate into an agentic AI pilot contract before the first invoice is issued.

Eric Lamanna8 min read
A brass ship's telegraph with its lever held between full-ahead and full-astern positions.

Most agentic AI pilots die twice. Once when the model quietly stops meeting the spec, and again a quarter later, when the invoice arrives and nobody can point to the paragraph that says the vendor owes anything. The contract is where those two deaths are decided, months before either happens.

Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, blaming escalating costs, unclear value, and inadequate risk controls. Those are contract failures, not model failures. What follows is the specific language a buyer should demand before signing a pilot statement of work, and the vendor-friendly defaults worth striking out. If you already have a signed template in front of you, treat this as a redline checklist.

Define Success as a Number, Not a Feeling

The most expensive clause in a pilot contract is the one that describes success as "demonstrating value" or "validating the use case." Both phrases are legally meaningless. You will pay the invoice, the vendor will produce a deck, and the argument about whether the agent works will happen after the money is gone.

Insist on written acceptance criteria that name a baseline, a target, and a measurement window. The baseline is what the current human process delivers today: time per task, cost per task, error rate, first-pass yield. The target is the improvement the pilot must prove. The window is the number of production traffic days over which the measurement is taken, and it should be long enough to catch a bad Monday. Anything shorter than two weeks of live traffic on real cases is a demo, not a pilot.

Mayer Brown's contracting guidance separates performance commitments into four categories: implementation, operational, agent-performance, and business-outcome. A pilot contract that only names implementation deliverables ("deploy agent, deliver documentation") gives you a working agent and nothing else. Pin at least one metric to each of the other three: an operational SLA (uptime, p95 latency), an agent-performance metric (task success rate on a held-out eval set), and a business outcome (hours reclaimed, tickets resolved, dollars processed).

What a Complete Acceptance Schedule Covers
What a Complete Acceptance Schedule CoversImplementation deliverables: 90%; Operational SLAs: 70%; Agent performance metrics: 55%; Business outcome commitments: 30%Named in the SOWLeft to interpretationImplementationdeliverables90%Operational SLAs70%30%Agent performancemetrics55%45%Business outcomecommitments30%70%
Illustrative. Most pilot templates specify implementation clearly and business outcomes vaguely; the gap is where disputes live. Illustrative: a visual comparison, not measured data.

The approval tier for each action belongs in the same schedule. If the agent can only recommend, say so. If it can execute a refund up to $500 without human review, say so, and put the ceiling in writing. An agent that only suggests is a search box, and you should not be paying agent prices for one.

Own the Eval Set Before the First Prompt

Acceptance criteria without a shared test set are theatre. Both sides need to agree, in writing, on the exact cases the agent will be scored against and the pass rate required on each slice.

The eval set should be your data, drawn from real historical cases, split into a training portion the vendor sees and a held-out portion the vendor does not. Score it on the held-out set. Include the ugly cases on purpose: ambiguous inputs, adversarial phrasing, malformed records, the edge cases that generate the tickets your team actually complains about. If the vendor pushes back on adversarial examples, that is diagnostic.

Three defaults worth striking from vendor templates:

  • Vendor-selected test cases. If the vendor picks the eval set, they will pick the ones the model already handles. Replace with "customer-provided evaluation dataset, refreshed monthly."
  • Aggregate pass rate only. A 90% average hides a 40% failure rate on your highest-value slice. Demand per-slice thresholds, particularly for the workflows with regulatory or financial exposure.
  • Acceptance on "substantial conformance." Substantial is whatever the vendor's lawyer says it is at 5pm on the last day of the pilot. Replace with numeric thresholds and a named tiebreaker.

Tie the final invoice to the held-out score. Milestone payments are fine; a full payment schedule that ignores the eval result is not.

A red emergency stop button on a stainless steel panel, guard flipped open, ready to be pressed.

Kill Switches Belong in the Contract, Not the Runbook

Runbooks describe how you would stop the agent. Contracts describe when you have the right to, and what happens next. Both matter, but only the second one survives a disagreement with the vendor.

A pilot contract should name at least four termination triggers that the buyer, not the vendor, controls:

  1. Failure to meet acceptance criteria at the scheduled measurement window, with no obligation to grant a cure period longer than 30 days.
  2. Material change to the underlying model or its behavior without written notice and a regression window. Digital Thought Disruption's clause catalogue argues the AI contract is part of the architecture precisely because silent model swaps invalidate every eval you ran to accept the system.
  3. Security or compliance incident involving your data, or a change in the vendor's subprocessors that you did not approve.
  4. Change of control at the vendor. Procurement guidance now treats termination-for-convenience on change-of-control with preserved pricing as the substantive right; notification alone leaves you informed but not empowered.

Termination-for-convenience for the buyer, with a pro-rata refund for unearned fees, is standard in enterprise services contracts and should not be controversial here. If the vendor refuses it in a pilot, the reason is worth hearing, because the pattern usually repeats at the production contract.

What the Vendor Gets to Keep When It Leaves

The exit clause most enterprises miss is not about termination rights. It is about what the vendor walks out with.

Exit Artifacts a Buyer Should Own
Exit Artifacts a Buyer Should OwnPrompts and tool definitions: 20%; Evaluation datasets and results: 20%; Agent execution traces and logs: 25%; Fine-tuned weights on your data: 15%; Configuration and version manifests: 20%Prompts and tool de…20%Evaluation datasets…20%Agent execution tra…25%Fine-tuned weights…15%Configuration and v…20%
Illustrative division of the artifacts a pilot contract should list as customer-owned deliverables. Illustrative: a visual comparison, not measured data.

Prompts, tool definitions, evaluation datasets, agent traces, and any fine-tuned weights derived from your data should be enumerated as customer property, deliverable in a machine-readable format within a defined transition period. Model-training rights on your inputs and outputs default to "no" unless you explicitly grant them. If the vendor's boilerplate lets them use your logs to "improve the service," that is training on your data with extra syllables.

The Ward and Smith AI clause library recommends including termination rights on data breach and on reclassification of the system as high-risk under emerging AI law. Both are cheap to add during a pilot and expensive to negotiate later. While you are in the exit section, delete any auto-renewal clause. A pilot that renews itself is not a pilot.

Scope, Change Control, and the Cost Ceiling

Agentic pilots have a specific failure mode: the scope expands one Slack message at a time until the eval set no longer resembles the work the agent is doing, and neither side is sure what it agreed to. A short scope clause combined with a hard change-control process is the antidote.

Name the workflows in scope by their process ID or ticket queue, not by adjective. Name the systems the agent is allowed to call, and require a signed change order to add any new tool or data source. Cap token, inference, and infrastructure spend for the pilot at a specific dollar figure, and require written approval to exceed it. Runaway cost is one of Gartner's three named reasons agentic projects get cancelled, and the ceiling is where you catch it.

This is also the place to align on deployment posture. If your data cannot leave the perimeter, the contract should say so and reference the target environment, whether that is air-gapped, VPC, or a controlled SaaS tenancy. The deployment mode trade-off shapes both the acceptance evidence you can gather and the exit artifacts you can extract.

Evidence, Audit, and Decision Lineage

NIST's agentic AI program frames autonomy as a spectrum of authority that requires proportional oversight. The contract is where that oversight becomes enforceable.

Require the vendor to deliver, at pilot close, a complete audit trail for every agent action taken in production: input, tool calls, intermediate reasoning steps at the level the framework exposes them, output, and the identity of any human approver. Require decision lineage for the acceptance evaluation itself: which model version, which prompt version, which tool registry, which eval set. Version identifiers on both the agent and its dependencies are what let you replay a bad decision six months later without a forensic reconstruction. This is the same discipline described in the site's note on audit logging: cheap to build in, impossible to add after the fact.

Where Pilot Budgets Shift When Contracts Get Real
Where Pilot Budgets Shift When Contracts Get RealModel and inference cost: 45%; Integration and data engineering: 25%; Evaluation and evidence work: 10%; Governance and audit tooling: 5%VENDOR-DRAFTED TEMPLATEBUYER-REDLINED CONTRACTModel and inference c… 45%30%Integration and data… 25%40%Evaluation and eviden… 10%20%Governance and audit… 5%10%
Illustrative. A redlined contract usually pushes spend away from raw inference toward the plumbing and evidence that make acceptance provable. Illustrative: a visual comparison, not measured data.

Grant yourself the right to audit these artifacts during the pilot and for a defined period after, and the right to have a third party do it on your behalf. Vendors will sometimes object to third-party audit rights in a pilot. That objection tells you what the production contract will look like.

Vendor-Friendly Defaults Worth Striking

A short list of clauses that appear in most agentic pilot templates and rarely survive a real redline:

  • Perpetual license to customer prompts and outputs for model improvement. Strike, or narrow to aggregated, de-identified telemetry with an opt-out.
  • Vendor's sole discretion on model upgrades. Replace with notice, regression window, and the right to remain on a supported prior version for the pilot term.
  • Liability cap at fees paid. A pilot fee is often small; the damage an agent can do in a connected system is not. Push the cap up for data incidents and unauthorized actions.
  • Acceptance deemed on 15-day silence. Replace with affirmative sign-off tied to the eval result.
  • "Beta" or "preview" carve-outs disclaiming all warranties. If it is too experimental to warrant, it is too experimental to run in production.

None of these are exotic. They are the terms enterprise procurement teams already apply to outsourced build engagements, translated into the language of autonomous systems.

The Signature Test

Before signing, run one exercise. Circle every acceptance criterion, exit trigger, and evidence deliverable in the draft. If any of them read as narrative rather than as a number, a date, or a named artifact, they will not be enforceable when you need them. Send it back.

A pilot contract is not a purchase order with extra pages. It is the operating manual for a small piece of your business that a machine will be running, briefly and under supervision, to prove it can be trusted to run more. Write it that way, and the pilot either pays for itself or ends cleanly. Write it any other way, and you are buying a demo at production prices.

// written by
Eric Lamanna
Director of Business Development

Eric Lamanna is a Digital Sales Manager with a strong passion for software and website development, AI, automation, and cybersecurity. With a background in multimedia design and years of hands-on experience in tech-driven sales, Eric thrives at the intersection of innovation and strategy—helping businesses grow through smart, scalable solutions. He specializes in streamlining workflows, improving digital security, and guiding clients through the fast-changing landscape of technology. Known for building strong, lasting relationships, Eric is committed to delivering results that make a meaningful difference. He holds a degree in multimedia design from Olympic College and lives in Denver, Colorado, with his wife and children.

Put an agent to work, the right way.

Start on Automatic and put the workflow you want to automate in front of engineers who have shipped agents in regulated environments.

Explore services
// the briefing

Agentic AI, in your inbox.

Occasional, high-signal notes on building and operating AI agents — automation patterns, architecture, and governance. No spam.