Agent Oracle

Field Notes: AI Agents Are Learning to Verify Outcomes, Not Just Complete Actions

Last updated: 10/2/2026

Back to blog
Anaya Iyer avatarAnaya Iyer 7 min read
Cover image for Field Notes: AI Agents Are Learning to Verify Outcomes, Not Just Complete Actions
AI-assisted, human-reviewed. Drafted with AI research tools from public sources and edited by our team. How we build these →

Most AI agent workflows stop too early. The agent sends the email, updates the record, submits the request, or creates the ticket. The tool returns a success response, and the workflow marks the task complete.

That confirms execution. It does not confirm the intended outcome.

An email can be sent to an invalid address. A customer record can be updated in the wrong system of record. A refund request can be accepted but later rejected by a payment processor. A ticket can be created without the fields required for routing. In each case, the agent acted successfully at the technical layer while failing at the business layer.

The important change is that agent design is beginning to treat post-action verification as part of the task itself. Completion is no longer defined only by whether a tool call returned without error. It is defined by whether the relevant environment entered—and remained in—the intended state.

What changed: execution and completion are separating

Early agent implementations inherited the semantics of APIs. If an API returned a successful status, the action was considered complete. This was convenient because it produced a clean stopping condition: call the tool, parse the response, report success.

Operational systems are not that clean. Many actions are asynchronous, mediated by other services, or subject to later reversal. A request may be accepted into a queue but never processed. A status may change temporarily and then be overwritten by a synchronization job. A document may be generated but never reach the person expected to approve it.

Agent architectures are therefore adding a distinct verification phase after execution. The pattern has three separate claims:

  • Command accepted: the target system received the instruction.
  • State changed: the expected object or process entered the intended state.
  • Outcome held: the state persisted long enough, or triggered the next required event, to count as business completion.

This separation matters because each claim requires different evidence. A tool response can prove acceptance. A fresh read from the system can usually prove state change. Durable completion may require an event, a downstream record, a human acknowledgment, or a timed recheck.

The practical model: define a postcondition before acting

A reliable agent should know how it will verify success before it performs the action. That means attaching a postcondition to the plan.

A postcondition is a testable description of the state that must be true after execution. “Send the renewal notice” is an action. “The renewal notice is delivered to the current contract owner, and the account timeline contains the message identifier” is a postcondition.

Useful postconditions identify four elements:

  1. Object: the record, transaction, person, or process expected to change.
  2. Expected state: the exact condition that should become true.
  3. Evidence source: the system or event that can establish the condition.
  4. Verification window: when the evidence should exist and when absence becomes meaningful.

Without these elements, an agent can only report activity. With them, it can distinguish “I did the step” from “the operation achieved its purpose.”

Verification should match the failure mode

Not every action requires the same depth of checking. The right verification method depends on how the workflow can fail.

Failure modeExampleAppropriate verification
Rejected commandA CRM update fails validationInspect the tool response and structured error
False acceptanceA job is queued but never processedRead the resulting job state after the expected processing interval
Wrong targetA notice is attached to the wrong accountRe-read the target using a stable business identifier
Downstream failureA refund is approved internally but rejected by the processorWait for the processor event or settlement state
State overwriteAn integration restores an old field valuePerform a delayed recheck or monitor change events
Human non-responseAn approval request is delivered but ignoredCheck acknowledgment or approval before the deadline

The operational mistake is to use the strongest available confirmation from the executing tool as a universal proxy for success. That evidence often describes only the tool’s local responsibility, not the end-to-end process.

A worked example: changing a customer’s payment terms

Consider an agent asked to change a customer from net 30 to net 45 after commercial approval.

A shallow workflow checks that the billing platform accepted the update. A stronger workflow starts by defining the intended result: the approved customer account should show net 45, future eligible invoices should inherit net 45, and no conflicting account policy should revert the change.

The agent first resolves the customer using an account identifier rather than name matching. It confirms that the approval applies to the same legal entity. It then updates the billing setting and records the approval reference.

Verification happens in stages. An immediate read confirms that the account now shows net 45. A policy check confirms that the customer is not governed by a parent-level term that will overwrite the account setting. If invoice generation is near, the agent can inspect the next draft invoice or schedule a check when it becomes available.

Suppose the immediate read succeeds, but a parent-account synchronization later restores net 30. The original action was technically successful. The outcome failed. A delayed verification detects the reversion and opens a narrowly scoped exception: the approved term conflicts with the hierarchy policy. The human is asked to decide whether to change the parent rule or grant an explicit override.

This is materially better than repeating the update. Blind retries would create churn without resolving the governing conflict.

What it means in practice: workflows need verification budgets

Verification has a cost. Reads consume system capacity. Waiting increases latency. Event subscriptions add infrastructure. Human confirmation creates work. The answer is not to verify everything indefinitely, but to allocate verification according to consequence and uncertainty.

A low-impact, easily reversible action may need only an immediate read-after-write check. A payment, entitlement change, external communication, or compliance record may require independent evidence and a delayed confirmation. Where the system is known to be eventually consistent, an immediate mismatch should trigger a bounded wait rather than an instant failure.

A practical verification budget specifies:

  • how many checks the agent may perform;
  • which evidence sources are sufficiently independent;
  • how long it may wait for asynchronous completion;
  • whether it may retry, repair, or compensate;
  • when unresolved status must be escalated.

This prevents two extremes: agents that declare victory too early and agents that poll forever because perfect certainty is unavailable.

Recovery must be designed alongside verification

Detecting failure without defining the next move simply produces a better alarm. Each postcondition should connect to a recovery policy.

If the action was never accepted, a retry may be appropriate. If the wrong object changed, the agent may need to reverse the change and stop. If a downstream system rejected the request, the agent should surface the rejection reason rather than resubmit unchanged data. If the result is ambiguous, it should avoid repeating non-idempotent actions until it can determine whether the first attempt took effect.

Compensating actions require particular care. Reversing a CRM field is not equivalent to recalling an external message or reversing a payment. The agent must understand whether compensation restores the prior business state or merely issues another transaction.

The safest pattern is to classify recovery paths before deployment: retry, repair, compensate, wait, or escalate. The agent then selects among approved paths using observed evidence, rather than improvising after a failure.

What remains unresolved

The hardest problem is deciding what counts as evidence of a business outcome. Systems expose technical states more readily than commercial reality. “Delivered” does not mean read. “Approved” does not mean implemented. “Closed” does not mean the customer’s issue is resolved.

There is also a boundary problem. An agent can verify only what its environment makes observable. If downstream events are unavailable, it must either accept weaker evidence or ask a human. More integrations do not automatically solve this; conflicting records can create additional ambiguity about which state is authoritative.

Timing remains equally difficult. Some outcomes become knowable immediately, while others emerge over days or depend on future behavior. A verification window that is too short creates false failures. One that is too long delays intervention.

Finally, outcome definitions can encode contested business policy. Sales may treat a signed order as completion. Finance may require cleared payment. Operations may require successful provisioning. The agent cannot resolve that disagreement through better reasoning alone. Leadership must define which state governs the workflow.

The operating standard to adopt now

For every consequential agent action, require a written postcondition, an evidence source, a verification window, and an approved recovery path. Log execution success separately from outcome success. Measure exceptions by failure stage: command, state transition, persistence, or downstream effect.

This changes the core review question. Instead of asking, “Did the agent use the tool correctly?” ask, “What evidence proves the intended state exists?”

An agent that acts without checking is an automation interface. An agent that verifies, detects divergence, and recovers within defined limits is becoming an operational system.

This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.

AI agentsoutcome verificationworkflow automationobservabilityoperations

From our own rounds

Measured on Agent Oracle, from real sessions people played on this site — not a third-party dataset.

Rounds played here
27
Questions per round
1
Play a round and add to these numbers
Share this post

Rate this article

No ratings yet

Discussion

Comments are moderated. Read our editorial policy.