Insights
How to choose your first agent workflow
Compare three candidate tasks using a five-part scorecard, then turn the strongest choice into a bounded pilot.
By Michael Santiago
The first agent project is easier to judge when the work has a clear beginning and a visible result. That sounds obvious until a team starts collecting ideas. One person wants a research assistant, another wants an autonomous sales process, and a third wants help sorting the requests already waiting in a shared inbox. These proposals differ in more than ambition. They ask the team to accept different kinds of uncertainty.
I would compare the work before comparing tools. The scorecard below is a proposed planning method, with a fictional operations team as a worked example. It is meant to help a team choose a pilot it can inspect and learn from. The scores are discussion aids, not measurements of expected financial return.
Start with three actual tasks
Write each candidate as a short sequence: what arrives, what someone does and what leaves the process. “Help with operations” is too broad. “Read an internal supply request and prepare a draft order for a coordinator” gives the team something it can examine. Include the person who handles the work today in this exercise.
For the example team, the candidates are preparing meeting summaries, sorting facilities requests and placing supplier orders. Each involves text, but the consequences differ. A summary can be edited. A misrouted request can delay work. An order can create an external commitment. Those differences should influence the choice before anyone builds a demonstration.
Collect a few representative inputs for each task. Use fictional or appropriately sanitized material for the initial comparison. Include a normal case and an awkward case. If nobody can provide either, the task may be too poorly understood for a first pilot, even if the idea sounds attractive.
Ask whether the task needs an agent
Anthropic recommends starting with simple approaches and distinguishes fixed workflows from agents that choose their own steps. Apply that distinction to the candidate list. If every request follows the same few rules, a form, a script or an ordinary workflow tool may be sufficient. The project should earn its additional complexity.
For meeting summaries, a single generation step may be enough. For facilities requests, the useful challenge might be interpreting incomplete descriptions before preparing a structured draft. Supplier ordering may require several decisions and permissions, but that does not automatically make it the best first project. More capability also creates more behavior to inspect.
Write down what the model contributes that a simpler method would not. If the answer is only that an agent would be interesting to build, keep exploring. A good candidate has a specific need for interpretation, retrieval or flexible decision-making that the team can describe in ordinary language.
Score five properties
Use a scale from zero to two for each property. Zero means the team has little evidence or control, one means the condition is partly met, and two means it is clear enough to test. Record a sentence explaining each score. The explanation matters more than the arithmetic.
First, assess input stability. Does the work arrive through a consistent channel with the information the process needs? A standard request form might score two. A mixture of telephone calls, photographs and secondhand messages might score zero until the intake is improved. Do not assume the agent will repair a process that people cannot describe.
Second, assess output clarity. Can a reviewer recognize a correct result? A draft record with required fields is easier to assess than a general promise to produce strategic insight. Write an example of an acceptable output and a rejected output. If the team cannot explain the difference, define the task more carefully.
Third, assess reversibility. Can a person review or undo the proposed action without substantial consequences? Preparing a draft is a useful starting point. Sending that draft to an external recipient changes the risk. Score the version of the workflow you actually propose to test, rather than an ambitious future version.
Fourth, assess review capacity. Name the person who will inspect the outputs and estimate how much time the trial will demand from them. An available expert who has agreed to review ten examples gives the pilot a stronger footing than an unnamed department that will supposedly handle exceptions later.
Fifth, assess test access. Can the team assemble representative cases and rerun them after changes? A task with only one clean example is hard to evaluate. Include missing information, contradictory inputs and unavailable tools. These cases help the team find the edges of its proposed scope.
Compare the fictional candidates
In our example, meeting summaries receive two points for stable inputs, one for output clarity, two for reversibility, two for review capacity and two for test access. The total is nine. That looks promising, but the team notes that a single summarization step may solve the problem without a multi-step agent.
Facilities request preparation receives one, two, two, two and two, for nine points as well. Its weaker input score reveals useful preparatory work: make location and contact details explicit in the intake. The team chooses this candidate because it wants to test interpretation of varied descriptions while keeping the output in a draft state.
Supplier ordering receives one, two, zero, one and one, for five points. The low reversibility score does not mean the task can never be addressed. It means the proposed version asks too much of a first trial. A later project might begin by preparing an order for approval instead of placing it.
The tie between the first two candidates is instructive. Scores do not replace judgment. They expose where the team has evidence, where it is making assumptions and where a simpler implementation may be enough. Keep those notes with the project brief so the reasoning survives the planning meeting.
Turn the winner into a bounded pilot
For facilities requests, define the starting event as receipt of a saved example submission. Define the output as a draft record containing location, issue description, missing details and a suggested review queue. The coordinator approves the record before any operational action. Exclude emergency handling from the pilot and route such examples directly to the established human process.
Write a stopping rule in advance. The team might pause the trial if the system attempts an unapproved action or repeatedly omits a required field. Also set a limit on how many cases the reviewer will handle during each session. A pilot that overwhelms its reviewer can fail even when the generated drafts look reasonable.
OpenAI's practical guide discusses workflow selection, evaluation and human intervention. Those are useful planning categories; the team still has to translate them into this particular task. Name the reviewer, the sample set and the pause mechanism in the brief. A general statement that a human remains involved leaves too much unresolved.
Decide what would justify the next step
Before running the pilot, state what evidence would support continuing. The team could require that the agreed sample cases produce complete draft records, that all defined exceptions reach the coordinator and that rejected outputs are easy to diagnose. These are example criteria, not universal thresholds. The people responsible for the work should set the actual standard.
After the trial, inspect the failures before expanding scope. Perhaps the agent handled descriptions well but the intake still lacked location details. That finding may justify a better form rather than a more capable model. Perhaps review took longer than the original task. That is useful evidence too.
Choose your own three candidates and complete the five-property scorecard with the person who performs the work. Keep the first version small enough that you can explain every input, action and reviewer. A first pilot is most useful when it produces a clear decision about what to do next.
