Anaya Iyer 7 min readMost AI agent workflows stop too early. The agent sends the email, updates the record, submits the request, or creates the ticket. The tool returns a success response, and the workflow marks the task complete.
That confirms execution. It does not confirm the intended outcome.
An email can be sent to an invalid address. A customer record can be updated in the wrong system of record. A refund request can be accepted but later rejected by a payment processor. A ticket can be created without the fields required for routing. In each case, the agent acted successfully at the technical layer while failing at the business layer.
The important change is that agent design is beginning to treat post-action verification as part of the task itself. Completion is no longer defined only by whether a tool call returned without error. It is defined by whether the relevant environment entered—and remained in—the intended state.
What changed: execution and completion are separating
Early agent implementations inherited the semantics of APIs. If an API returned a successful status, the action was considered complete. This was convenient because it produced a clean stopping condition: call the tool, parse the response, report success.
Operational systems are not that clean. Many actions are asynchronous, mediated by other services, or subject to later reversal. A request may be accepted into a queue but never processed. A status may change temporarily and then be overwritten by a synchronization job. A document may be generated but never reach the person expected to approve it.
Agent architectures are therefore adding a distinct verification phase after execution. The pattern has three separate claims:
- Command accepted: the target system received the instruction.
- State changed: the expected object or process entered the intended state.
- Outcome held: the state persisted long enough, or triggered the next required event, to count as business completion.
This separation matters because each claim requires different evidence. A tool response can prove acceptance. A fresh read from the system can usually prove state change. Durable completion may require an event, a downstream record, a human acknowledgment, or a timed recheck.
The practical model: define a postcondition before acting
A reliable agent should know how it will verify success before it performs the action. That means attaching a postcondition to the plan.
A postcondition is a testable description of the state that must be true after execution. “Send the renewal notice” is an action. “The renewal notice is delivered to the current contract owner, and the account timeline contains the message identifier” is a postcondition.
Useful postconditions identify four elements:
- Object: the record, transaction, person, or process expected to change.
- Expected state: the exact condition that should become true.
- Evidence source: the system or event that can establish the condition.
- Verification window: when the evidence should exist and when absence becomes meaningful.
Without these elements, an agent can only report activity. With them, it can distinguish “I did the step” from “the operation achieved its purpose.”
Verification should match the failure mode
Not every action requires the same depth of checking. The right verification method depends on how the workflow can fail.
| Failure mode | Example | Appropriate verification |
|---|---|---|
| Rejected command | A CRM update fails validation | Inspect the tool response and structured error |
| False acceptance | A job is queued but never processed | Read the resulting job state after the expected processing interval |
| Wrong target | A notice is attached to the wrong account | Re-read the target using a stable business identifier |
| Downstream failure | A refund is approved internally but rejected by the processor | Wait for the processor event or settlement state |
| State overwrite | An integration restores an old field value | Perform a delayed recheck or monitor change events |
| Human non-response | An approval request is delivered but ignored | Check acknowledgment or approval before the deadline |
The operational mistake is to use the strongest available confirmation from the executing tool as a universal proxy for success. That evidence often describes only the tool’s local responsibility, not the end-to-end process.
A worked example: changing a customer’s payment terms
Consider an agent asked to change a customer from net 30 to net 45 after commercial approval.
A shallow workflow checks that the billing platform accepted the update. A stronger workflow starts by defining the intended result: the approved customer account should show net 45, future eligible invoices should inherit net 45, and no conflicting account policy should revert the change.
The agent first resolves the customer using an account identifier rather than name matching. It confirms that the approval applies to the same legal entity. It then updates the billing setting and records the approval reference.
Verification happens in stages. An immediate read confirms that the account now shows net 45. A policy check confirms that the customer is not governed by a parent-level term that will overwrite the account setting. If invoice generation is near, the agent can inspect the next draft invoice or schedule a check when it becomes available.
Suppose the immediate read succeeds, but a parent-account synchronization later restores net 30. The original action was technically successful. The outcome failed. A delayed verification detects the reversion and opens a narrowly scoped exception: the approved term conflicts with the hierarchy policy. The human is asked to decide whether to change the parent rule or grant an explicit override.
This is materially better than repeating the update. Blind retries would create churn without resolving the governing conflict.
What it means in practice: workflows need verification budgets
Verification has a cost. Reads consume system capacity. Waiting increases latency. Event subscriptions add infrastructure. Human confirmation creates work. The answer is not to verify everything indefinitely, but to allocate verification according to consequence and uncertainty.
A low-impact, easily reversible action may need only an immediate read-after-write check. A payment, entitlement change, external communication, or compliance record may require independent evidence and a delayed confirmation. Where the system is known to be eventually consistent, an immediate mismatch should trigger a bounded wait rather than an instant failure.
A practical verification budget specifies:
- how many checks the agent may perform;
- which evidence sources are sufficiently independent;
- how long it may wait for asynchronous completion;
- whether it may retry, repair, or compensate;
- when unresolved status must be escalated.
This prevents two extremes: agents that declare victory too early and agents that poll forever because perfect certainty is unavailable.
Recovery must be designed alongside verification
Detecting failure without defining the next move simply produces a better alarm. Each postcondition should connect to a recovery policy.
If the action was never accepted, a retry may be appropriate. If the wrong object changed, the agent may need to reverse the change and stop. If a downstream system rejected the request, the agent should surface the rejection reason rather than resubmit unchanged data. If the result is ambiguous, it should avoid repeating non-idempotent actions until it can determine whether the first attempt took effect.
Compensating actions require particular care. Reversing a CRM field is not equivalent to recalling an external message or reversing a payment. The agent must understand whether compensation restores the prior business state or merely issues another transaction.
The safest pattern is to classify recovery paths before deployment: retry, repair, compensate, wait, or escalate. The agent then selects among approved paths using observed evidence, rather than improvising after a failure.
What remains unresolved
The hardest problem is deciding what counts as evidence of a business outcome. Systems expose technical states more readily than commercial reality. “Delivered” does not mean read. “Approved” does not mean implemented. “Closed” does not mean the customer’s issue is resolved.
There is also a boundary problem. An agent can verify only what its environment makes observable. If downstream events are unavailable, it must either accept weaker evidence or ask a human. More integrations do not automatically solve this; conflicting records can create additional ambiguity about which state is authoritative.
Timing remains equally difficult. Some outcomes become knowable immediately, while others emerge over days or depend on future behavior. A verification window that is too short creates false failures. One that is too long delays intervention.
Finally, outcome definitions can encode contested business policy. Sales may treat a signed order as completion. Finance may require cleared payment. Operations may require successful provisioning. The agent cannot resolve that disagreement through better reasoning alone. Leadership must define which state governs the workflow.
The operating standard to adopt now
For every consequential agent action, require a written postcondition, an evidence source, a verification window, and an approved recovery path. Log execution success separately from outcome success. Measure exceptions by failure stage: command, state transition, persistence, or downstream effect.
This changes the core review question. Instead of asking, “Did the agent use the tool correctly?” ask, “What evidence proves the intended state exists?”
An agent that acts without checking is an automation interface. An agent that verifies, detects divergence, and recovers within defined limits is becoming an operational system.
This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.
From our own rounds
Measured on Agent Oracle, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 27
- Questions per round
- 1
Rate this article
Discussion
Comments are moderated. Read our editorial policy.