Insights
Give your agent a clear human handoff
Design the pause, review packet and recovery path that let a person take over without reconstructing the whole task.
By Michael Santiago
An agent that asks for help has only completed half of a handoff. Someone still needs to receive the request, understand what happened and decide how the work continues. If the request lands in an unowned queue or arrives without the relevant context, the human becomes a detective before they can become a reviewer.
A useful handoff is a small operating procedure. It needs a trigger, a context packet, a responsible person and a recovery path. The walkthrough below proposes a design for a fictional internal request assistant. Adapt the fields and limits to the actual work, then rehearse them with the person who will receive the exceptions.
Name the moment that requires help
Start with observable conditions. “Ask when uncertain” is difficult to test by itself. More concrete triggers include a missing required field, conflicting source records, a tool failure or a proposed action beyond the agent's authority. Each trigger should identify the point at which the process pauses and what it is allowed to do while waiting.
In the fictional example, an assistant prepares equipment requests for an office coordinator. A submission asks for three monitors, but the inventory record shows two different delivery locations for the requester. The assistant should not pick the location that sounds more plausible. The defined trigger is conflicting delivery information, and the next step is review by the coordinator.
Other triggers can lead to different destinations. A connection error may belong with a technical owner, while an unclear purchasing policy belongs with the person responsible for that policy. Avoid sending every exception to the same general channel merely because it is easy to configure.
Separate missing information from permission
These two situations often look similar in a chat message but need different responses. Missing information means the process lacks a fact it needs. Missing permission means the action is understood but has not been authorized. Asking someone to supply a location is different from asking them to approve an order.
The review screen should make the distinction visible. For the location conflict, the coordinator chooses or supplies the correct destination. For an approval request, the reviewer sees the exact proposed action and its scope. A general “continue” button can hide that difference and encourage a decision without enough context.
OpenAI's practical guide includes human intervention as part of agent design. For this example, translate that principle into a specific pause before the unresolved action. The implementation should preserve the pending request rather than treating any human reply as unlimited authorization to complete the rest of the workflow.
Build a packet the reviewer can use
A compact context packet should answer six questions: what the user wanted, what the agent has done, what stopped it, which evidence matters, what decision is needed and what happens after that decision. Put the requested decision near the top so the reviewer can orient themselves quickly.
For the equipment request, the packet might begin: “Choose the delivery location for request EQ-104.” It then shows the original request, the two conflicting locations and the source associated with each. It states that no order has been placed. Finally, it provides choices to select a location, request clarification or close the request without proceeding.
Include the minimum history needed to make the decision. A full conversation transcript can be useful as an expandable reference, but it should not be the only explanation. Conversely, a one-line error message is rarely enough. The reviewer needs to understand the state of the work without reconstructing it from scattered logs.
Avoid copying unnecessary personal or confidential material into every notification. A notification can identify the case and link to an appropriately restricted review view. Decide which fields are required for the task, which are optional and which should be omitted from the handoff entirely.
Assign ownership and waiting behavior
A queue needs an owner with authority to make the requested decision. Name the role and arrange coverage for absence. If a request is sent to a shared inbox, define who claims it and how other reviewers can see that it has been claimed. Otherwise two people may make different decisions on the same case.
Choose what the agent does while waiting. In this example, the equipment request remains paused and the requester receives a neutral status message. The assistant does not repeatedly create new cases or retry the blocked action. A reminder, if used, should attach to the same case so the history remains intact.
Set an escalation route for a review that cannot be completed. The coordinator may discover that neither location is current. The procedure should let them ask the requester for clarification and retain ownership of the case. Forwarding the problem without a new owner simply moves the ambiguity to another inbox.
Resume from a known state
After the coordinator selects a location, the system should record the decision and resume the specific paused step. It should check whether anything relevant changed while the request waited. A stale inventory record or a withdrawn request could make the original plan inappropriate even after the location issue is resolved.
This is where duplicate actions deserve attention. If a tool timed out, the agent may not know whether an attempted update succeeded. A person should not be asked to approve a blind retry without seeing that uncertainty. The recovery path may need to inspect the target system first and establish the current state.
For the equipment example, the handoff record should retain the selected location, reviewer and time of decision. The final status should state whether the draft request was completed, returned for clarification or closed. That gives the coordinator a way to check the outcome without reading the entire interaction again.
Test the handoff as a user experience
Rehearse the normal exception first. Give the coordinator a fictional case with the conflicting locations and ask them to work through the review screen without verbal coaching. Note where they pause, which information they search for and whether they understand what their decision will authorize.
Then test an absent reviewer, an incomplete packet and a case that has already been resolved elsewhere. Add a tool failure after approval. Each variation checks a different part of the procedure. A handoff that works only when every person and system is available is incomplete.
OpenAI's agent evaluation documentation describes evaluating workflow traces. For handoff testing, the useful question is whether the recorded sequence shows the pause, the human decision and the resulting action. A final success message alone would not establish that the approval boundary was respected.
Keep a small handoff register
For each exception type, record its trigger, receiving role, required fields and allowed next actions. Include a sample case that the team can rerun after changes. This register gives both builders and operators a common reference when a new exception appears.
Review confusing cases with the people receiving them. Perhaps the agent supplies too much history, the queue mixes unrelated decisions or the available buttons do not match the work. Those are design problems worth fixing directly. More elaborate wording in the assistant's message may not solve them.
Start with one real escalation path. Write a packet, nominate its owner and rehearse the return to work. The handoff is ready for a trial when the reviewer can explain the decision, the authority it carries and the state the process will resume from.
