Why the task comes before the tool
Most failed agent projects don’t fail because of the technology, but because of the task: it was too broad (“support our sales team”), too rare (twelve cases a year) or too costly when it goes wrong (customer communication without oversight). A good first agent has a narrow, recurring task with a clear result that a person can check in a few seconds.
This playbook takes you in five steps from selection to live operation. It promises no savings and no compliance sign-off – both depend on your process, your data and your legal situation. It provides the questions you need to answer beforehand.
1. Three sensible pilots
- Structure CRM notes as a draft
- Input: meeting note or transcript. Output: structured draft for the required fields (occasion, decision-maker, next step, date, objections). The person reviews and saves. Why it suits: high volume, clear target structure, errors are visible and cheap.
- Pre-sorting incoming inquiries
- Input: inquiry by form or email. Output: proposed category (suitable / unclear / unsuitable), missing details and responsible person – as a draft, not as a reply to the customer. Why it suits: rule-based, measurable against later qualification, no external contact.
- Internal knowledge search with sources
- Input: a question from an employee. Output: draft answer with a reference to the specific internal documents it comes from. No source, no answer. Why it suits: useful with every question, errors can be spotted by checking sources, no customer data needed if the document set is chosen accordingly.
2. Selection grid
Rate each candidate on four dimensions. A pilot needs high scores on the first three and a low one on the fourth.
| Criterion | Question | Good pilot |
|---|---|---|
| Volume | How often does the case occur per week? | Frequent enough that setup and operation pay off |
| Rule-basedness | Can you describe in sentences what a good result is? | Yes, with few exceptions |
| Data quality | Is the input complete, structured and accessible? | Mostly yes; gaps are known |
| Cost of errors | What happens if the result is wrong and nobody notices? | Low: internal, correctable, no external impact |
3. Process from intake to operation
Process intake
Describe today’s process in steps with input, output, duration and owner. Note which steps require judgment. The agent takes over rule-based steps; judgment stays with a person.
Permissions and minimal access
The agent gets exactly the access the task needs – read-only where read-only is enough, on the smallest sensible set of data. Before the start, clarify which data is processed, where it is processed and who is responsible. Whether this complies with data protection or compliance rules in your case is not decided by this playbook, but by your review with the responsible parties.
Test set
Before the first run, assemble a test set from real, anonymized or rebuilt cases: normal cases (everyday), edge cases (incomplete, ambiguous, unusually long) and failure cases (wrong language, empty input, contradictory details). For each case, define what a correct result is – before you see the agent’s result.
Human approval
In the pilot, the agent produces drafts. A person approves, corrects or discards them. Every correction is counted and grouped by type. This list is your most important data source for deciding whether the agent stays.
Monitoring
- Number of cases per week and share without correction.
- Type and frequency of corrections.
- Handling time per case before and after (measured, not estimated).
- Failures, timeouts, empty results.
4. Calculate operating costs before the start
An agent costs in three line items: one-time setup (process intake, configuration, test set), ongoing operation (usage fees, maintenance, people’s approval time) and cost of errors (corrections, rework). Calculate all three before you start – and count the approval time honestly: an agent whose drafts need correcting 40% of the time saves little.
A simple calculation scheme is enough for the decision: cases per month × (minutes before − minutes after) ÷ 60 gives the freed-up capacity in hours. Multiplied by an imputed hourly rate, it yields a calculated benefit; the monthly operating costs are deducted from it. Only when this net value is positive can the one-time setup be paid back in months. The minutes “after” include the person’s approval time – not just the agent’s run time.
- Measure the minutes before on at least ten real cases before the pilot starts.
- Enter usage fees as they arise at the expected volume, not at the test volume.
- Record the hourly rate as an assumption and don’t change it afterward to rescue a result.
5. Stop criteria
Before the start, define when the pilot ends – in both directions. If you only formulate the criteria after four weeks, you judge the result by mood instead of data.
- Abort
- Correction rate above a predefined value after four weeks; an error with external impact; operating costs above the planned range; the team works around the agent instead of using it.
- Handover to operation
- Correction rate stable below the target value for at least four weeks; handling time measurably reduced; an owner named for ongoing operation; permissions and data flows documented.
- Expansion
- Only when the first agent is in operation. A second pilot does not start in parallel with the first.
The automation brief below summarizes all points of this playbook on one page: task, criteria, permissions, test set, approval, monitoring, costs and stop criteria. Fill it out before you choose the first tool.
Downloads
To take away
Markdown opens in any text editor or note-taking tool. CSV files are semicolon-separated and open directly in Excel, Numbers, or LibreOffice.

