How to Scope an Agent Proof of Value in Six Weeks

A concrete scoping checklist for a six-week agent proof of value: workflow choice, success metrics, access requirements, exit criteria, and what to cut.

Eric Lamanna8 min read
A measuring tape, stopwatch, and six wooden blocks arranged on an architect's desk, suggesting careful scoping of a short engagement.

A pilot budget has been approved. The CFO wants a number back by next quarter. Procurement is asking for a statement of work. This is the moment where most agentic AI programs are quietly lost, not during deployment and not during contracting, but in the two weeks before the SOW is signed, when the workflow, integrations, and pass/fail numbers get written down in ways nobody can defend ninety days later.

Six weeks is enough time to prove whether an agent belongs in production. It is not enough time to discover the workflow, negotiate data access, and define success. Those decisions have to be made before the kickoff call. What follows is the scoping work that happens between budget approval and SOW signature, written for the buyer who has to issue the document.

So where does the scoping actually start?

Pick a Workflow That Can Fail Safely

The first mistake is picking the workflow that would be most valuable if it worked. The right pick is the workflow where a wrong agent action is cheap to detect and cheap to reverse. These are rarely the same.

A good proof-of-value candidate has four properties. It runs often enough to generate a statistically meaningful sample inside six weeks, which usually means at least a few hundred executions. Its inputs arrive in a structured or semi-structured form the agent can actually read. Its outputs land in a system that keeps an audit trail by default. And a reversal costs minutes, not days: a journal entry can be voided; a wire transfer cannot.

Finance operations tends to meet all four, which is why finance workflows often go first. Invoice coding, vendor onboarding checks, intercompany reconciliations, and expense policy enforcement produce dense volume, structured inputs, and clean reversal paths. Sales operations and IT ticket triage usually qualify. Anything touching clinical decisions, trading, or external customer money does not belong in a six-week POV, regardless of its ROI on paper.

Write the candidate workflow down as a single sentence of the form: the agent reads X from system A, decides Y under policy Z, and writes result R to system B, with human approval required when condition C holds. If that sentence takes more than two lines, the scope is too wide. Split it or pick a smaller one.

Size the Integration Surface Before You Sign

Integration is where six-week pilots turn into fourteen-week pilots. The SOW should name every system the agent will read from and write to, the authentication method for each, and who owns the credentials. If the vendor cannot list those systems in a kickoff diagram during scoping, the integration estimate is a guess.

A defensible POV touches at most three systems of record on the read side and one on the write side. More than that and most of the six weeks goes to API wiring rather than agent behavior. If a required system has no API, or only a legacy SOAP endpoint behind a VPN, that is not a POV blocker but it needs its own line item and its own owner on your side.

Where Six Pilot Weeks Actually Go
Where Six Pilot Weeks Actually GoIntegration and auth setup: 35; Agent logic and tool registry: 25; Evaluation and labeling: 20; Approval routing and guardrails: 12; Readout and documentation: 835%25%20%12%8%Integration and auth setup35 · 35%Agent logic and tool registry25 · 25%Evaluation and labeling20 · 20%Approval routing and guardrails12 · 12%Readout and documentation8 · 8.0%
Illustrative allocation of effort in a typical scoped POV; integration work dominates when scope is not bounded. Illustrative: a visual comparison, not measured data.

Three questions belong in the scoping conversation and should be answered in writing before signature. Which service accounts will the agent use, and are they already provisioned? What is the rate limit on each downstream API, and does the expected pilot volume fit under it? Who approves a schema change in each source system if one is needed mid-pilot? The absence of named owners for any of these is the single most common reason a POV misses its end date.

If the deployment target is air-gapped or VPC-isolated, add two weeks to any mental estimate and move the networking review forward to week zero. The deployment topology is a scoping decision, not an implementation detail.

Define the Numbers That Decide Pass or Fail

A POV without pre-committed pass/fail numbers is not a proof of value. It is a demo with a budget. The exit criteria belong in the SOW, written as thresholds a non-technical reviewer can check against a dashboard.

Four metrics cover most agent workflows. Task completion rate is the share of inputs the agent processes end-to-end without human intervention. Decision accuracy is the share of those completions that match the correct answer on a labeled holdout set. Intervention rate is the share routed to a human by design, which should be deliberate rather than a dumping ground for everything the agent is unsure about. Cycle time is wall-clock from input arrival to final write.

Set the target for each before kickoff, grounded in what the current human process actually achieves. If the baseline accuracy of your AP team on invoice coding is 94 percent, an agent target of 99 percent is a vanity number. 94 percent at a tenth of the cycle time is a defensible one. For context on where frontier agents sit today, Carnegie Mellon and Salesforce research has measured completion rates of only about 30 to 35 percent on multi-step office tasks, and that is on open benchmarks rather than on a scoped enterprise workflow with a tool registry and approval gates.

A half-open steel gate on a narrow path, suggesting an approval checkpoint between automated steps.

The sample size matters as much as the threshold. A 95 percent accuracy claim built on forty test cases is noise. Agree in the SOW on the minimum number of real production inputs the agent must process before the pass/fail judgment is made, and agree on who labels the ground truth. A reasonable floor for most back-office workflows is 300 to 500 processed items with a labeled sample of at least 100.

Setting an Honest Accuracy Target
Setting an Honest Accuracy TargetInvoice coding accuracy: 94; Vendor onboarding checks: 90; Expense policy enforcement: 96; Intercompany reconciliation: 88ActualTargetInvoice codingaccuracy94 ✓Target 94Vendor onboardingchecks90Target 92Expense policyenforcement96 ✓Target 95Intercompanyreconciliation88Target 90
Illustrative: the agent target (bar) is set against the current human baseline (marker), not against 100%. Illustrative: a visual comparison, not measured data.

Decide What the Agent Is Allowed to Do

Scope is not only about which workflow. It is about which actions inside that workflow the agent may take unattended, which require a human approver, and which are forbidden outright. This belongs in the SOW as a three-column list, not as a general principle.

A workable default for a POV: read-only actions are unattended; write actions below a dollar threshold or a reversibility threshold go through approval gates; write actions above that threshold are out of scope for the pilot entirely. The threshold is a business decision, not a technical one. Pick it with the process owner before kickoff, and expect to revisit it in week three when the first unexpected edge cases appear.

A useful discipline from published agent design guidance is to keep the agent's tool surface narrow and well-named during the pilot. A POV agent with eight well-scoped tools is easier to evaluate than one with thirty. Narrowness also simplifies the discussion of which actions actually need human approval, because the list is short enough to review line by line.

Refuse the Scope That Will Sink the Pilot

The hardest part of scoping is saying no. Stakeholders will arrive with adjacent workflows, additional systems, and nice-to-have features. Most of them belong in a phase two document, not in the SOW.

Four categories of scope creep reliably kill six-week pilots:

  • Multi-tenant generalization. Making the agent work for one business unit in six weeks is possible. Making it work for three with different policies is not. Prove it on one and port later.
  • New data pipelines. If the required data does not exist in a queryable form today, building the pipeline is a separate project with its own timeline.
  • UI deliverables. A dashboard for the operations team is a different piece of software. Scope it as a follow-on if the agent clears its pass criteria.
  • Model selection studies. Comparing three foundation models on your workflow is a research exercise. Pick one for the POV, document the swap path, move on.

The scoping document should name these exclusions explicitly. A line reading out of scope: additional business units, new ETL work, end-user UI, model benchmarking is worth more than another paragraph of ambition. Vendors that push back on explicit exclusions are usually trying to protect optionality on their side at the cost of your timeline.

Lay Out the Six Weeks

The weeks themselves should be allocated in the SOW, not left to the vendor's project plan. A workable shape for most pilots:

Weeks one and two are environment access, data sampling, and a working integration to one read source and one write source with a stub decision step. If that is not running by end of week two, the pilot is already behind and the scope needs to be cut, not the timeline extended.

Weeks three and four are the agent itself: tool registry, decision logic, approval routing, and the first end-to-end runs on real data under human supervision. Weeks five and six are shadow mode on live volume, measurement against the agreed thresholds, and a written go/no-go recommendation. Reserve the last two days for the readout, not the last two hours.

Deloitte reporting cited in recent analyst coverage puts production-ready agentic systems at around 11 percent of surveyed organizations, and Gartner has forecast that over 40 percent of agentic AI projects will be canceled by the end of 2027 on cost, value, and risk-control grounds. Those numbers are not an argument against piloting. They are an argument for scoping the pilot so that a no-go verdict in week six is a cheap, honest answer rather than a reputational problem.

How Scope Pressure Builds Across Six Weeks
How Scope Pressure Builds Across Six WeeksWeek 1: 10; Week 2: 25; Week 3: 70; Week 4: 85; Week 5: 55; Week 6: 30010Week 125Week 270Week 385Week 455Week 530Week 6
Illustrative: requests to expand scope tend to peak mid-pilot — plan to say no hardest in weeks three and four. Illustrative: a visual comparison, not measured data.

Write the SOW to Be Killed

A good POV SOW makes it easy to stop. Named exit criteria. A dated readout. A phase-two decision gate with explicit cost and scope. No automatic rollover into a larger engagement. The vendor should be paid for the six weeks of work regardless of outcome, and both parties should treat a disciplined no-go as a successful pilot.

The POV is a feasibility test, not a commitment. The scoping work above is what makes the test honest: a specific workflow with a reversible failure mode, a bounded integration surface, pre-committed numbers measured on a real sample, and a written list of what the agent will not be asked to do. Teams that put those four things in the SOW tend to issue a defensible verdict in six weeks. Teams that leave any of them to be worked out during the pilot tend to issue the same verdict in sixteen.

// written by
Eric Lamanna
Director of Business Development

Eric Lamanna is a Digital Sales Manager with a strong passion for software and website development, AI, automation, and cybersecurity. With a background in multimedia design and years of hands-on experience in tech-driven sales, Eric thrives at the intersection of innovation and strategy—helping businesses grow through smart, scalable solutions. He specializes in streamlining workflows, improving digital security, and guiding clients through the fast-changing landscape of technology. Known for building strong, lasting relationships, Eric is committed to delivering results that make a meaningful difference. He holds a degree in multimedia design from Olympic College and lives in Denver, Colorado, with his wife and children.

Put an agent to work, the right way.

Start on Automatic and put the workflow you want to automate in front of engineers who have shipped agents in regulated environments.

Explore services
// the briefing

Agentic AI, in your inbox.

Occasional, high-signal notes on building and operating AI agents — automation patterns, architecture, and governance. No spam.