Choose an automation opportunity by examining a specific piece of work: who needs the result, what makes it difficult today, whether the inputs are usable, and how someone will recognize an unacceptable output. Then select the least complex intervention worth testing. Technical feasibility is one part of the decision; it does not establish organizational value or permission to act.
For digital and IT leaders, AMF recommends a workflow map that separates potential benefit, readiness, consequences, and validation burden. The map below is a proposed decision aid, not a validated scoring model. Its purpose is to make assumptions visible before a team commits to a pilot.
Start with an episode of work
“Automate operations” is too broad to evaluate. “Prepare a suggested category and evidence summary for a newly received internal service request” has an input, a user, an output, and a point at which a person can disagree.
Observe examples with the people doing the task. Include ordinary cases, exceptions, and cases that required someone to ask for more information. Record where time goes: reading, searching, copying, waiting, resolving ambiguity, checking, or obtaining approval. Do not assume a slow process is slow because it needs better text generation. If most of the delay is waiting for an owner to decide, faster drafting may leave the real constraint intact.
Define the smallest useful unit as a sentence: when this event occurs, this person needs this output to make this decision. Keep the consequential action separate. Preparing a recommendation and executing it can have different owners, permissions, and failure costs.
This emphasis on context is consistent with the Map function in NIST’s AI Risk Management Framework 1.0: intended use, business value, risk tolerance, and knowledge limits inform whether to proceed. NIST’s framework is guidance; the worksheet here is AMF’s own synthesis, not a NIST checklist or certification.
Build an opportunity map before a ranking
Use one row per workflow candidate. In each cell record evidence, an estimate with its basis, or “unknown.” Do not average an unknown permission boundary into a favorable overall score.
| Dimension | Evidence to collect and decision it informs |
|---|---|
| Frequency and volume | Representative arrival counts, peaks, seasonality, and eligible cases. Decision it informs: Is there enough recurring work to justify setup and ongoing evaluation? |
| Process friction | Observed touch time, waiting, handoffs, repeated entry, and exception causes. Decision it informs: Does the intervention address the actual delay or merely a visible activity? |
| Input and data readiness | Source owners, access rights, completeness, freshness, and examples with missing fields. Decision it informs: Can the system obtain the evidence needed for this task? |
| Integration surface | Systems read or changed, available interfaces, identity requirements, and test access. Decision it informs: What must work outside the model for the workflow to succeed? |
| Reversibility | A specific undo or correction procedure, including who performs it. Decision it informs: Can a pilot be contained when its output is wrong? |
| Decision risk | Who could be affected, plausible failure consequences, and how quickly errors would be detected. Decision it informs: Which actions need approval or should be excluded? |
| Human judgment | Disputed cases, policy interpretation, relationships, and decisions with no agreed answer. Decision it informs: Which parts should remain with a person? |
| Validation burden | Reference answers, reviewer availability, review time, and cases that cannot be checked reliably. Decision it informs: Can the team determine whether the pilot is useful and safe enough for its intended scope? |
A high-volume task may still be a poor first pilot if nobody can evaluate its answers. A lower-volume task may be valuable if each instance imposes substantial reconstruction work and the result can be checked against authoritative records. The choice should follow the task’s evidence, not a preference for whichever demonstration looks most autonomous.
Compare interventions, including doing less
For each candidate, compare a process correction, deterministic automation, model-assisted preparation, and an agent that selects steps or tools. Ask what additional capability each option buys and what new checking it requires.
If the same field is repeatedly copied between two systems, an integration may be enough. If the problem is interpreting varied language against a maintained policy, model assistance may warrant a trial. If the path changes from case to case and requires gathering additional evidence, tool selection may become useful. These are AMF selection heuristics, not proven decision thresholds.
Anthropic’s December 2024 engineering article distinguishes predefined workflows from agents that direct their own process, and recommends starting with simpler solutions. That is provider guidance, not independent evidence that a particular architecture will outperform the alternatives in your organization.
An illustrative service-request decision
Consider a hypothetical team that manually categorizes internal service requests. The team is considering both preparing routing suggestions and automatically granting access to systems. This is an illustration, not an AMF client case or a tested implementation.
For categorization, the pilot could read a request and an approved service catalog, suggest a category, and cite the catalog entry. A person would accept or change the suggestion. A missing catalog match would return “needs review.” The expected pilot output is an inspectable recommendation, not a closed ticket.
Access granting crosses another boundary. It requires verified identity, entitlement rules, a specific resource, and an authorized decision. A plausible interpretation of the request cannot establish those facts. Under this proposed pilot, the model would have no account-provisioning permission. An existing human process would handle access requests.
The opportunity is therefore narrower than “automate service requests.” It is to test whether preparing a grounded routing suggestion makes reviewers more effective. If reviewers must reread every source and rewrite most suggestions, the system may add work. Record that outcome rather than declaring success because the model generated a category.
Write the pilot agreement
Before implementation, bring the process owner, a frontline reviewer, and the relevant system/data owners together. Agree on the following in plain language:
- Scope: the eligible input, output, excluded cases, and intended reviewer.
- Authority: permitted reads and writes, and who may approve each action.
- Comparison: the current process and a simpler alternative, using the same representative cases and acceptance criteria where practical.
- Evidence: how to judge correctness, unsupported claims, exceptions, reviewer effort, elapsed time, and completed useful work.
- Stop conditions: unacceptable disclosure, unauthorized action, inability to verify results, or a failure pattern the team cannot contain.
- Exit decision: expand, revise, retain as assistance, or stop—with a named owner and a defined review point.
Start with historical or otherwise approved test inputs where possible. Do not mistake the ability to read a record for permission to send it to a model provider. The data owner must settle the permitted handling before the pilot uses that information.
Keep failures visible in the comparison. Report eligible cases, rejected inputs, successful outputs, corrections, and cases returned to people. A result calculated only from accepted suggestions would omit the work spent on rejected ones. Cost assessment should include integration, review, recovery, and ongoing maintenance as well as model calls; this is a proposed accounting boundary, not a savings estimate.
Know what the map cannot decide
This guide has not established universal weights, minimum volumes, or a return on investment threshold. Organizational priorities and failure consequences determine those choices. A reversible pilot is also not proof that wider deployment is appropriate: new users, inputs, systems, and permissions can change the decision.
If the evidence is insufficient, the next useful step may be process observation, data cleanup, or an explicit owner decision. None requires pretending that an agent deployment has already been justified.
Next action
Take one recurring workflow and fill the eight dimensions with a person who does the work. Select the smallest missing fact that could change the decision and collect it before ranking the candidate. The expected result is a bounded pilot agreement—or a defensible reason to improve the process another way.
Provenance
Prepared for Agent Model Fit with material Codex assistance in source inspection, synthesis, and drafting on September 11, 2026. The framework and service-request example are proposed guidance, not observed customer experience. The cited NIST and Anthropic materials were inspected on that date; the Anthropic article is used for conceptual distinctions, not current tooling specifications. Human editorial review was completed by Cresencio on September 11, 2026.