Agent Oracle

Three Myths About AI Knowing When the Work Is Good Enough

Last updated: 10/10/2026

Back to blog
Yuna Park avatarYuna Park 7 min read
Cover image for Three Myths About AI Knowing When the Work Is Good Enough
AI-assisted, human-reviewed. Drafted with AI research tools from public sources and edited by our team. How we build these →

“Make it good” sounds like a quality standard. Operationally, it is not. An AI agent must decide which defects matter, what evidence is sufficient, and when another pass costs more than it improves.

This is a stopping problem disguised as a writing, analysis, or automation problem. Stop too early and the agent leaves material errors. Stop too late and it burns time, introduces needless changes, or delays a reversible decision. Three popular claims obscure how that judgment should actually work.

What “good enough” requires

An agent can only stop responsibly when it has a usable definition of acceptance. That definition has four parts:

  • Required properties: what must be present, such as a recommendation, owner, deadline, or cited evidence.
  • Disqualifying defects: what must not be present, such as unsupported claims, exposed customer data, or an unauthorized commitment.
  • Verification method: how each condition will be tested.
  • Residual-risk rule: which uncertainties may remain and which require escalation.

These elements form a stopping policy. Without one, the agent substitutes proxies: length, polish, similarity to prior work, or the absence of obvious errors. Those signals can help, but none proves that the work is fit for use.

Myth 1: More checking always produces better work

The kernel of truth is straightforward: a second pass often catches omissions, contradictions, and mechanical errors. High-consequence outputs deserve deeper verification than disposable drafts.

The myth begins when review becomes an unlimited loop. Repeated passes do not have equal value. An early pass may discover that a contract summary omitted a termination clause. A later pass may merely replace plain language with smoother language. Worse, generative revision can introduce new defects while fixing old ones: a qualifier disappears, a number is transposed, or a precise recommendation becomes vague.

The right control is not “check until confident.” It is check by failure mode. Each pass should target a distinct risk and produce evidence that can change the disposition of the work.

PassQuestionUseful testStop condition
CompletenessAre required elements present?Compare against a field or requirement listEvery mandatory element is found
CorrectnessAre material claims supported?Trace claims to approved inputsNo unsupported material claim remains
ConsistencyDo sections conflict?Compare names, dates, decisions, and constraintsConflicts are resolved or flagged
PolicyIs the proposed action permitted?Test against authority and compliance rulesNo prohibited action is proposed
PresentationCan the recipient use it?Check format, clarity, and requested toneUsability requirements are met

Consider an agent preparing a renewal recommendation. It should verify usage evidence, commercial constraints, and approval authority before polishing the prose. If those checks pass, a fourth stylistic rewrite may add little. If pricing authority is unclear, no amount of editing makes the recommendation ready.

The practical rule is asymmetric: spend verification effort where failure is costly or difficult to reverse. A tentative internal agenda needs less scrutiny than a customer-facing commitment.

Myth 2: A few examples teach AI what “done” means

Examples are powerful. They communicate tacit preferences that are cumbersome to specify, including structure, level of detail, vocabulary, and the shape of an acceptable recommendation.

But an example combines several things: requirements, conventions, incidental choices, and historical circumstances. An agent cannot reliably know which are binding. If three previous board updates used five bullets, it may infer that five bullets are mandatory. If those updates omitted downside scenarios, it may imitate the omission even when the current decision requires them.

Examples describe prior outputs. Acceptance criteria define the current job. The strongest design uses both, with an explicit priority order:

  1. State non-negotiable requirements and prohibitions.
  2. Provide examples as demonstrations of preferred execution.
  3. Identify which features are intentional.
  4. Explain when deviation is allowed.

Suppose an agent drafts incident summaries. A sample may show a chronological narrative, technical terminology, and a short remediation section. The actual definition of done might be: identify customer impact, establish the confirmed timeline, separate facts from hypotheses, name the incident owner, and list open actions. The sample guides expression; the criteria govern acceptance.

This distinction matters under changing conditions. If the latest incident affects a regulated workflow, the agent may need an additional review path even though no example includes one. Blind similarity would look consistent while being operationally wrong.

Myth 3: Human review guarantees quality

Human review is essential when judgment, accountability, or authority cannot be delegated. A reviewer can recognize political sensitivity, interpret ambiguous evidence, and accept risk on behalf of the organization.

Yet “a human will review it” is not a quality system. Reviewers skim, assume upstream checks occurred, and focus on visible wording rather than hidden factual dependencies. A polished output can receive less scrutiny precisely because it appears complete. Human review also fails when the reviewer lacks the necessary source access or decision rights.

The agent should not hand over a finished-looking artifact and an implicit request to find anything wrong. It should prepare a review surface that directs attention to consequential uncertainty:

  • Decision requested: what the reviewer must approve, reject, or modify.
  • Checks completed: which acceptance tests passed and against what inputs.
  • Known gaps: missing evidence, unresolved conflicts, or stale information.
  • Material assumptions: propositions that could change the recommendation.
  • Risk of proceeding: the likely consequence if an assumption is wrong.

For example, an agent preparing a supplier approval should not merely attach a summary for sign-off. It should state that insurance documentation was verified, data-handling terms remain unresolved, and procurement—not the requester—must accept the exception. The human then reviews the actual decision boundary rather than proofreading the entire packet.

A better model: acceptance gates, not confidence alone

Confidence is useful but insufficient. An agent can be confident because the output resembles familiar patterns, while a mandatory condition remains untested. Conversely, it can be uncertain about wording even though every operational requirement has been satisfied.

A stronger model separates three states:

  • Accepted: all mandatory tests pass, no disqualifying defect remains, and residual risk falls within delegated authority.
  • Conditionally usable: the output can support a reversible next step, but identified assumptions still require validation.
  • Blocked: a mandatory input, permission, or test is missing.

This prevents uncertainty from collapsing into either reckless action or needless paralysis. An agent may conditionally draft a launch plan using an explicitly labeled demand assumption. It should block publication if the required legal approval is absent.

Worked example: stopping a market brief

Imagine an executive asks for a brief recommending whether to enter a new segment. A weak agent writes a polished memo, rereads it, and stops when it sounds persuasive.

A governed agent first defines acceptance: the brief must state the decision, distinguish evidence from inference, describe the customer problem, identify operational dependencies, present a credible alternative, and disclose unresolved assumptions. It then runs targeted checks.

The evidence check finds that customer demand is inferred from sales conversations rather than confirmed purchasing behavior. The consistency check finds that the proposed launch date conflicts with the product roadmap. The authority check finds no issue because the brief recommends a decision rather than executing one.

The correct result is not endless research. It is a conditionally usable brief stating: the opportunity appears plausible; demand evidence is preliminary; the proposed timing requires displacing an existing roadmap item; leadership must decide whether to fund validation before committing delivery capacity. The agent stops because the artifact is now fit for the immediate purpose: making the next decision. It does not pretend the market decision itself is settled.

The operating standard

AI knows the work is good enough only when “good enough” has been converted from taste into a decision rule. More review helps when each pass targets a meaningful failure mode. Examples help when they are subordinate to explicit acceptance criteria. Human review helps when the handoff exposes the decisions and uncertainties that require human judgment.

The decisive design question is not whether the agent feels finished. It is whether further work is likely to change the outcome, reduce material risk, or satisfy a missing requirement. If not, stop. If so, specify the next test. If that test requires authority or evidence the agent does not have, escalate the gap rather than polishing around it.

This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.

AI agentsquality controlstopping rulesacceptance criteriaworkflow design

From our own rounds

Measured on Agent Oracle, from real sessions people played on this site — not a third-party dataset.

Rounds played here
27
Questions per round
1
Play a round and add to these numbers
Share this post

Rate this article

No ratings yet

Discussion

Comments are moderated. Read our editorial policy.