The short answer
- A suitable pilot has one trigger, bounded inputs, defined outputs, known exceptions, and a named process owner.
- High failure consequence requires stronger review, restricted actions, escalation, logging, and rollback.
- A pilot needs an existing-process baseline and a representative test set before implementation.
- Data protection, security, permissions, provider terms, and retention belong in opportunity selection.
- A technically accurate output can still fail when the surrounding hand-off is slow, confusing, or unactionable.
Choose a workflow, not a fashionable AI capability
An AI opportunity is a bounded workflow change with an owner, baseline, acceptance test, and controlled consequence. Phrases such as ‘automate customer service’ or ‘add an agent’ are solution labels, not pilot definitions. Rewrite them as one trigger, one set of permitted inputs, one output, one recipient, and one decision boundary.
Separate deterministic automation from probabilistic work. Fixed rules, database queries, validation, calculations, permissions, and routing often belong in ordinary software. AI can be useful where the bounded task requires classification, extraction, summarisation, drafting, or interpreting varied language, but the model should not control steps that a reliable rule can handle more safely.
Reject candidates whose current process has no stable purpose or owner. AI does not remove ambiguity; it can execute ambiguity at greater volume. Clarify the policy, source of truth, exceptions, and accountability before deciding whether a model belongs in the workflow.
Score candidates with an evidence-led opportunity matrix
A strong first candidate combines worthwhile operational friction with unusually good controllability. Score evidence, not enthusiasm. Any high-consequence or hard-to-reverse candidate should move down the queue even when its theoretical value is large.
Use a simple three-level rating—strong, uncertain, or stop—and attach the evidence behind every rating. The matrix ranks discovery work; it does not calculate an investment return or certify legal compliance.
| Criterion | Strong pilot signal | Caution or stop signal | Evidence to collect |
|---|---|---|---|
| Business value | Repeated delay, rework, or avoidable handling affects a named outcome | The benefit is novelty, vague productivity, or headcount reduction without a process case | Volume, waiting time, handling effort, error categories, service impact |
| Input readiness | Permitted inputs are accessible, legible, representative, and connected to a source of truth | Inputs are missing, disputed, unlawfully obtained, highly sensitive without controls, or inaccessible | Data inventory, permissions, quality sample, provenance, retention and provider terms |
| Output testability | Reviewers can apply written acceptance criteria consistently | Experts disagree without a resolution rule or success depends on an unobservable judgment | Labelled examples, reviewer agreement, required evidence, unacceptable-output list |
| Failure consequence | Errors are detectable, reversible, contained, and cheap to correct | An error could materially affect rights, safety, employment, finance, access, or reputation | Impact assessment, escalation path, rollback, affected people and legal review |
| Human control | A trained owner can review exceptions with source evidence and authority to intervene | Review is ceremonial, overloaded, or unable to see why the output was produced | Queue capacity, review interface, decision rights, training and override logs |
| System fit | Authentication, integration, logging, monitoring, support, and exit are understood | The pilot relies on hidden manual work, broad credentials, or an unowned provider dependency | Architecture map, threat model, service limits, incident plan and fallback process |
Map, baseline and test the whole workflow
Map the trigger, actors, inputs, systems, rules, language tasks, decisions, exceptions, hand-offs, outputs, and downstream effects. Record the existing handling time, waiting time, rework, error categories, cost drivers, satisfaction signal, and service constraint that the pilot is supposed to change. A pilot without a baseline can demonstrate activity, but not operational improvement.
Build the test set before tuning the solution. Include ordinary cases, rare but important cases, incomplete records, conflicting sources, adversarial input, sensitive data, unsupported requests, policy exceptions, and examples that must escalate. Keep a sealed evaluation set separate from examples used while configuring prompts, retrieval, rules, or models.
Define acceptance at the workflow level. Measure task correctness, evidence completeness, escalation behaviour, unsafe actions, reviewer agreement, latency, recovery, and downstream usability—not only whether a model response sounds plausible. Re-test when the prompt, model, retrieval source, provider, policy, or connected system changes.
- Give every output a source, status, confidence or exception signal where the reviewer needs it
- Minimise access and restrict which external actions the system can initiate
- Log inputs, relevant configuration, outputs, review decisions, overrides, and failures proportionately
- Preserve a fallback route when the model, provider, integration, or reviewer is unavailable
- Set monitoring thresholds, incident ownership, rollback conditions, and a decommissioning path
Illustrative worked example: triage and draft a support request
This example is illustrative and hypothetical; it is not a Blancc client deployment, benchmark, or promised saving. A software company receives support requests through one mailbox. Staff manually identify the product, request type, urgency, and relevant help article before drafting a reply. The first pilot classifies the request, retrieves approved internal guidance, and prepares a draft inside the existing queue.
The system cannot send messages, change accounts, issue refunds, or close cases. A support agent sees the original request, retrieved source, proposed labels, and draft together. Requests involving security, account ownership, payment disputes, personal-data rights, threats, or unsupported evidence always escalate under written rules.
The evaluation compares the existing and assisted workflows on representative historical cases that the business is permitted to use. Measures include classification correctness, unsupported statements, escalation recall, review time, rework, queue delay, and agent acceptance. The company should proceed only if the complete workflow meets its pre-agreed thresholds without weakening service or control.
- Bounded task: classify, retrieve approved guidance, and draft
- Prohibited actions: send, refund, close, alter access, or make a rights-affecting decision
- Human evidence: original request and retrieved approved source remain visible
- Fallback: the existing manual queue continues when automation is unavailable
- Limitation: results from one product, policy set, language mix, and case sample do not transfer automatically
Common AI automation failure modes
The most common failure is selecting a broad, impressive use case before understanding the process. Teams then optimise a demonstration while authentication, permissions, integration, review capacity, exception handling, and operational ownership remain unresolved.
A second failure is treating human review as a universal safeguard. Reviewers can become overloaded, defer to confident outputs, or lack the source evidence and authority required to intervene. Oversight must be designed around consequence, capacity, visibility, escalation, and recorded correction.
- Using live sensitive data before establishing purpose, access, retention, and provider controls
- Testing only clean examples or the same examples used to configure the system
- Measuring model fluency instead of workflow accuracy, safety, and downstream usefulness
- Granting broad tools or credentials when the pilot needs read-only or draft-only access
- Ignoring provider, model, prompt, retrieval, and policy changes after launch
- Declaring success from time saved while error cost, rework, or affected-user outcomes worsen
Practical checklist, method and limitations
This guide adapts lifecycle ideas from NIST’s voluntary AI Risk Management Framework—govern, map, measure, and manage—into a small-pilot selection method. It also uses UK government AI-assurance guidance, ICO data-protection guidance, and NCSC secure-development guidance as constraints. It is not first-party research, legal advice, a security assessment, or a complete compliance framework.[1][2][3][4]
The ICO currently states that parts of its AI guidance are under review following the Data (Use and Access) Act. Applicable duties also depend on sector, data, people affected, geography, purpose, and decision consequence. Obtain qualified legal, data-protection, security, equality, employment, or sector-specific advice when those risks are material.[2]
- Name the workflow owner, affected people, purpose, trigger, input, output, recipient, and prohibited actions
- Establish the existing baseline and calculate whether the operational problem is worth solving
- Inventory data sources, legal basis where required, permissions, provider use, retention, and deletion
- Create representative development and sealed evaluation sets with written acceptance criteria
- Design least-privilege access, review evidence, exception rules, escalation, logging, fallback, and rollback
- Approve measurable thresholds, monitoring, incident ownership, change control, and a stop decision
Sources and further guidance
Blancc uses primary guidance where factual or regulatory context matters. Recommendations remain general and should be assessed against the specific business, audience, product, and risk.
Read Blancc’s editorial standards and corrections policy for source selection, illustrative labels, update dates, software assistance and corrections.
