THE KEY ANSWER
A good pilot has a single recipient, a limited process, representative data, and documented success criteria. Its outcome is a measurement-backed decision: deploy, narrow the scope, or end the experiment.
Start with an experiment card
On a single page, describe the problem, the user, the current method, and the expected change. Add the scope excluded from the pilot. If the assistant is to prepare response drafts, it should not simultaneously change pricing or send messages. A clear boundary allows for quick decision-making when an idea for another feature arises.
Record the most critical unknown. It could be scan readability, the quality of responses from documents, or the team's readiness to use a new tool. Each of these unknowns requires a different experiment. Do not assess adoption based on a technical test without users, nor quality based on a single impressive conversation.
Context and references: Anthropic: Demystifying evals for AI agents
Pilot data must resemble real work
Collect examples of varying difficulty: complete, unclear, contradictory, and those lacking answers. Remove unnecessary personally identifiable information. Preserve the structure that affects the task, such as document layout or message order. Test cases should have an expected result evaluated by someone familiar with the process.
Separate materials used for refining the solution from acceptance materials. If the team improves the system on every example, the final result mainly reflects fitting to a known dataset. Keep fresh cases for acceptance as well. Establish who resolves cases where two people disagree on the correct answer.
Test the workflow, not just the model
Invited users should go through the full path: start the task, check the result, correct an error, and close the case. Observe where they need to ask about the meaning of the interface. Measure time including verification. If it is easier to write the answer manually than to verify a draft, the problem may lie in the presentation method rather than the model itself.
Introduce the ability to report an issue without leaving the task. Ask for the reason for rejection, not just a thumbs down. “Outdated rate” and “wrong tone” require different corrections. Record the configuration version so that after a change, it is known what the user was evaluating. Do not change all elements simultaneously during the experiment.
Acceptance should lead to a specific action
The final report does not need to be long. It should show the result relative to the starting point, the distribution of errors, support costs, and unresolved risks. Add a recommendation and its conditions. “We deploy only for one type of document” is more useful than a general “AI has potential”.
Before moving to production, establish the owner, monitoring, incident handling, and a rollback plan. Separate experimental work from elements required for real data. The team must know when they are running a demo and when they are using a tool that can support a commitment to a client. If measurements are ambiguous, design a shorter supplementary experiment.
WHERE TO START
Bring this into your project.
- Name the single most important unknown.
- Keep a separate set of acceptance examples.
- Observe the full task performed by the user.
- Prepare the decision along with conditions and an owner.
Choose one thing your process is missing today. It's a useful topic for your first conversation with the team.
QUESTIONS AND ANSWERS
Frequently asked questions.
How long should an AI pilot last?
Long enough to verify the key unknown on representative cases. Access to data and systems often impacts the timeline more than building the prototype. It is worth setting the schedule only after checking these dependencies.
What if the pilot does not meet its goal?
Identify the cause: data quality, incorrect scope, technological limitation, or low demand. Narrowing or ending the project protects a larger budget. The failure of an experiment can be a valuable business outcome.
Sources and context
- Anthropic: Demystifying evals for AI agents ↗
Anthropic describes the evaluation of agent results on tasks with defined criteria. The pilot organization presented here is a proposal by ALGOV.
Prepared by the ALGOV team. Current as of September 8, 2026. Examples describe possible scenarios, not results from client projects. How we create our guides.