
An AI pilot begins with enthusiasm and a working demo. Six weeks later, people disagree about whether it worked. One group remembers the impressive examples, another remembers the failures, and nobody has numbers to settle it. The pilot then either lingers indefinitely or gets scaled on the strength of a good feeling. Both outcomes are expensive.
The cure is to treat the pilot as an experiment with a hypothesis, a measurement plan and an ending, all written down before the first prompt is tested.
You cannot show improvement without knowing where you started. Before the pilot, measure how the task is done today: how long it takes, how often it needs rework, what it costs in staff time and how satisfied the people downstream are. Even a rough sample of a few dozen real cases is better than none. If the task is not measured today, the first deliverable of the pilot is the measurement.
Be honest about the human baseline. People are not perfect either, and comparing an AI system to an imagined flawless employee sets an impossible bar. Compare it to the actual current process, including its errors and delays.
Pick one primary metric that captures the value you hope for, such as time per case, cases handled per person or first-pass accuracy. Then add guardrail metrics that catch harm you would not accept even if the primary metric improves.
Write the target as a range, not a single number, for example the minimum improvement that would justify the cost of running it in production. That threshold is a business judgment, so involve the budget owner.
Demos use clean, favorable examples. Production does not. Build the evaluation set from real cases, including the awkward ones: incomplete records, ambiguous requests, unusual formats and edge cases your team complains about. Keep a portion of the set hidden from anyone tuning prompts, so you have an honest final check.
Where possible, have reviewers grade outputs without knowing whether they came from the system or from a person. It removes a surprising amount of bias in both directions.
Stop rules are the part teams skip, and the part that saves the most money. Agree in advance on conditions that end the pilot early or send it back for redesign.
Also define what success unlocks: which team, which volume, which controls and which budget come next. A pilot without a next step is a demo.
At the end, hold a short review with the numbers on one page: baseline, result, guardrails, cost and a recommendation to scale, revise or stop. Stopping is a legitimate and valuable outcome. It means the organization learned something cheaply. Record what you learned so the next pilot starts smarter.

Finance teams are cautious for good reason. These three starting points keep a human in control while still removing real manual work.

Connecting an AI assistant to company data is easy. Doing it safely takes five controls worth putting in place before the pilot expands.

Time zones, contracting, quality and communication: how a U.S. front door backed by affiliate companies in Lima and San Jose actually works day to day.
A 45-minute working session, no slides.