AI: The Decisions People Are Getting Wrong
The biggest AI failures rarely begin with a bad model. They begin with a poorly framed decision about workflow, ownership, risk, economics, or control. Here is a practical framework for choosing and governing AI agents that produce measurable business value.
Lucas AragónAI & creator economyFirst published 9/7/2026 · last revised 9/12/2026 with fresh sources, corrections, and new context. Reader corrections are reviewed and folded into future versions.
Summary
AI strategy is becoming an operating-model decision, not a software-shopping exercise. Leaders often focus on model rankings, sweeping automation targets, or highly visible pilots while neglecting workflow diagnosis, data access, exception handling, security, and accountability. The better approach is to identify a bounded business decision, establish its baseline cost and quality, decide what an AI agent may do without approval, and measure the complete economics—including integration, review, monitoring, and failure recovery. Agent Oracle’s central principle is simple: automate a well-defined operating loop, not an impressive demo. The winning system may combine models, rules, retrieval, APIs, human approvals, and audit logs. Its value comes from improving throughput, revenue, service, or risk-adjusted cost—not from appearing autonomous.
Key takeaways
- Do not start with ‘Where can we use AI?’ Start with a workflow whose delay, labor, inconsistency, or error rate has measurable economic consequences.
- A high-performing model is not automatically a dependable agent. Production systems also require permissions, tools, context, approval rules, monitoring, and recovery paths.
- Evaluate total workflow economics: implementation, inference, integration, human review, security, observability, maintenance, and the cost of incorrect actions.
- Autonomy should be earned in stages. Begin with recommendations, advance to approval-gated execution, and permit bounded autonomy only after evidence supports it.
- Treat data access as a governed capability. Give agents the minimum information and permissions required for each task, with identity controls and auditable logs.
- Assign one business owner to every deployment. Technology teams can operate the platform, but a named operator must own outcomes, exceptions, and process redesign.
- Use model flexibility as leverage. Architect around evaluations and business interfaces so models can be replaced without rebuilding the workflow.
- The strongest early use cases are frequent, bounded, digitally observable, and economically material—such as lead research, proposal preparation, support triage, reconciliation, and compliance evidence gathering.
Explain like I'm 5
Imagine hiring a very fast junior analyst who can read, write, search approved files, and use selected software—but can also misunderstand instructions confidently. You would not hand that person the company bank account on day one. You would provide a job description, examples, limited access, checkpoints, and a manager for unusual cases. An AI agent should be deployed the same way. The model is the worker’s reasoning and language ability; tools are the systems it can use; permissions define what it may touch; evaluations are performance reviews; logs are the work record; and human escalation is management. The business mistake is buying the ‘worker’ before defining the job. The sensible sequence is to choose a valuable task, document what good looks like, restrict authority, test real cases, measure results, and expand responsibility only when performance is dependable.
Deep dive
Wrong decision 1: Choosing a model before diagnosing the workflow
Model selection feels concrete, but it is usually downstream of the real decision: which operating constraint should change? A sales team may say it needs an AI prospector when its actual bottleneck is poor account data, slow legal review, or inconsistent follow-up after discovery calls. An operations team may request a chatbot when the costly problem is manual exception routing. Map the workflow from trigger to outcome. Record volume, cycle time, labor minutes, rework, conversion, error severity, systems touched, and escalation frequency. Select a narrow unit of work—such as researching an account and drafting a call brief—rather than an entire role. Only then can buyers compare models on relevant criteria such as extraction accuracy, tool use, latency, data residency, and cost.
Wrong decision 2: Confusing content generation with agency
A generator returns an answer. An agent pursues a goal through multiple steps: interpreting context, choosing tools, taking actions, checking results, and escalating exceptions. That distinction changes the control requirements. A sales agent that drafts an email creates reviewable content; one that updates CRM fields, selects recipients, and sends messages can affect revenue, reputation, privacy, and contractual commitments. Define an autonomy ladder. Level 1 summarizes or drafts. Level 2 recommends an action. Level 3 executes after approval. Level 4 executes inside explicit limits and reports afterward. Level 5 handles broader objectives with continuous oversight. Most businesses should begin at Levels 1–3. Autonomy is not the objective; dependable economic performance is.
Wrong decision 3: Measuring the demo instead of the operating result
A polished prototype can conceal weak economics. Accuracy on ten curated examples says little about production reliability across incomplete records, unusual customers, changing policies, and unavailable systems. Establish a pre-AI baseline and a production scorecard. For sales, measure research time, qualified meetings, acceptance of suggested next actions, pipeline progression, and unsubscribe or complaint rates. For service, track containment, resolution time, reopen rate, customer satisfaction, and harmful-error frequency. Calculate net value as incremental gross profit plus avoided labor and risk losses, minus build, licensing, inference, integration, review, monitoring, maintenance, and remediation costs. A system saving 1,000 hours is not valuable if employees redo its work or if one unbounded mistake creates a major liability.
Wrong decision 4: Treating security and compliance as a final review
Security architecture should shape the design before a pilot receives real data. Determine whether prompts, retrieved documents, outputs, tool calls, and feedback are retained; where they are processed; whether vendors use them for training; and which subprocessors are involved. Apply least privilege, short-lived credentials, role-based access, encryption, environment separation, and tamper-resistant logging. Defend against prompt injection by treating external text as untrusted input and separating instructions from retrieved content. For regulated or high-impact work, preserve source citations, approval history, model and prompt versions, and the evidence behind actions. Frameworks such as the NIST AI Risk Management Framework, ISO/IEC 42001, and the EU AI Act provide useful structures, but controls must be translated into daily operating procedures.
Wrong decision 5: Assuming automation means removing people
The immediate advantage often comes from redesigning human work, not eliminating it. Agents can assemble evidence, classify requests, prepare options, execute routine updates, and monitor queues. Humans remain valuable where intent is ambiguous, stakes are high, empathy matters, or trade-offs require authority. Design the handoff deliberately: what triggers escalation, who receives it, what context accompanies it, how quickly they must respond, and how the resolution improves future evaluations. If reviewers approve nearly everything, the threshold may be too conservative. If they routinely rewrite outputs, the task definition, context, or model is inadequate. Human review is an instrumented control loop—not a ceremonial checkbox.
The Agent Oracle decision framework
Use six gates before scaling. First, value: name the financial or strategic outcome and its baseline. Second, task: specify inputs, outputs, boundaries, and exception classes. Third, evidence: create a representative evaluation set, including adversarial and rare cases. Fourth, authority: define accessible data, permitted tools, spending or communication limits, and approval points. Fifth, operations: assign an owner, dashboards, incident procedures, fallback behavior, and change management. Sixth, economics: compare verified benefits with full lifecycle cost and downside exposure. Run a time-boxed pilot on live but bounded work, preferably with a control group or phased rollout. Scale only when quality, adoption, reliability, security, and unit economics meet written thresholds. This discipline turns AI purchasing from speculative experimentation into portfolio management.
- 1956The Dartmouth Summer Research Project popularized the term ‘artificial intelligence,’ framing machine intelligence as a research field.
- 1997IBM Deep Blue defeated world chess champion Garry Kasparov, demonstrating powerful but narrowly specialized machine decision-making.
- 2012AlexNet’s ImageNet result accelerated commercial investment in deep learning and GPU-based model training.
- 2017Google researchers published ‘Attention Is All You Need,’ introducing the Transformer architecture that underpins modern large language models.
- 2020OpenAI released GPT-3, making few-shot language capabilities accessible through a commercial API and expanding enterprise experimentation.
- 2022ChatGPT launched on November 30, moving generative AI from specialist circles into mainstream business use.
- 2023NIST released AI RMF 1.0 on January 26; businesses also began shifting from standalone copilots toward tool-using agent designs.
- 2024The European Union adopted the EU AI Act, creating a phased, risk-based regulatory regime; ISO/IEC 42001 adoption also gained executive attention.
- 2025Enterprises increasingly emphasized evaluations, agent orchestration, identity, observability, and bounded execution rather than unrestricted autonomy.
Glossary
- AI agent
- A software system that interprets a goal, plans or selects steps, uses tools, and acts within defined permissions to produce an outcome.
- Agentic workflow
- A controlled business process in which one or more AI components can choose actions, call systems, inspect results, and handle or escalate exceptions.
- Evaluation set
- A representative collection of normal, difficult, rare, and adversarial cases used to measure performance before and after deployment.
- Grounding
- Connecting a model’s response to approved, current evidence such as company documents, database records, or verified external sources.
- Hallucination
- A plausible-sounding but unsupported or incorrect model output. The business impact depends on whether it is displayed, approved, or executed.
- Human in the loop
- A control design in which a person reviews, approves, corrects, or handles selected AI decisions or actions.
- Least privilege
- The security principle of granting an agent only the data access and actions necessary for its current task.
- Prompt injection
- An attack or failure mode in which untrusted content attempts to override instructions, disclose information, or trigger unauthorized behavior.
- Retrieval-augmented generation
- A method that retrieves relevant information from designated sources and supplies it to a model when producing an answer.
- Total cost of ownership
- The complete cost of an AI system, including software, models, implementation, integration, review, monitoring, security, maintenance, and incident response.
FAQs
What is the best first AI-agent use case?+
Choose a frequent, bounded workflow with digital inputs, a measurable outcome, and reversible errors. Account research, support classification, meeting preparation, invoice exception triage, and compliance evidence collection are common candidates. Avoid starting with rare, politically sensitive, or irreversible decisions.
Should we build an agent or buy one?+
Buy when the workflow is standardized and the vendor integrates with your systems and controls. Build when proprietary data, unusual process logic, differentiated customer experience, or strict governance is strategically important. Many companies use a hybrid: a purchased platform with custom tools, policies, and evaluations.
How should ROI be calculated?+
Measure incremental gross profit, labor capacity released, cycle-time gains, avoided losses, and quality improvement. Subtract implementation, licenses, model usage, integration, human review, monitoring, maintenance, retraining, and expected failure costs. Use verified adoption and production volumes rather than theoretical capacity.
How accurate must an AI agent be?+
There is no universal threshold. Required performance depends on error severity, detectability, reversibility, and human oversight. A drafting assistant may tolerate correctable defects; an agent sending payments or making employment decisions requires much stronger controls and evidence.
Can sensitive company data be used safely?+
Potentially, but only after reviewing data flows, retention, training terms, subprocessors, geographic processing, access controls, encryption, deletion, and incident obligations. Minimize data exposure and prevent secrets from entering systems that are not approved for them.
Who should own an AI-agent deployment?+
A business executive should own the outcome and process design. Technology, security, legal, compliance, data, and procurement provide essential controls, but they should not become substitutes for accountable operational ownership.
How long should a pilot run?+
Long enough to cover representative volume and exception types, but short enough to force a decision. Many bounded workflows can be tested in four to twelve weeks if baselines, evaluation cases, integrations, owners, and success thresholds are established first.
What should happen when the agent is uncertain?+
It should follow an explicit fallback policy: ask for missing information, retrieve additional evidence, decline the action, route to a qualified person, or revert to the previous process. Uncertainty must alter behavior rather than merely appear as a score on a dashboard.
Predictions
- Agent procurement will shift from comparing chatbot features to testing complete workflows against buyer-owned evaluation suites and service-level objectives.
- Identity will become a core layer of agent architecture: each agent and delegated action will need traceable permissions, credentials, purpose, and revocation controls.
- Model choice will become more dynamic. Routing systems will assign tasks to different models based on risk, capability, latency, privacy, and cost.
- Boards and insurers will demand evidence of inventory, control ownership, monitoring, incident history, and third-party dependencies for material AI systems.
- Sales organizations will move from generic message generation toward agents that maintain account context, prepare evidence-based next actions, and update revenue systems under policy controls.
- The competitive advantage will migrate from access to foundation models toward proprietary workflow knowledge, clean operational data, evaluation assets, trusted integrations, and organizational adoption.
Risks
- Unbounded action: an agent can send, publish, purchase, delete, or modify records beyond the intended scope. Limit tools, transaction values, recipients, environments, and action rates.
- Prompt injection and data exfiltration: malicious instructions in emails, web pages, or documents can manipulate tool-using systems. Isolate untrusted content and require policy checks outside the model.
- Silent quality drift: model updates, changing data, new products, and altered processes can degrade results. Maintain versioning, regression tests, sampled review, and rollback options.
- Automation bias: employees may accept fluent recommendations without checking evidence. Display sources, uncertainty, and approval responsibilities at the decision point.
- Compliance overreach: teams may deploy AI into employment, credit, healthcare, surveillance, or customer communications without recognizing applicable obligations. Classify use cases before launch.
- Vendor concentration: proprietary APIs, pricing changes, outages, or policy shifts can disrupt critical operations. Use abstraction layers, exportable logs, continuity plans, and tested fallbacks.
- False savings: headcount assumptions can obscure integration and review work or merely transfer effort to another team. Measure realized financial outcomes and process capacity after deployment.
Opportunities
- Revenue operations: combine approved account data, call notes, product usage, and CRM history to identify missing information and recommend evidence-based next actions.
- Sales preparation: generate cited account briefs, stakeholder maps, discovery questions, and proposal inputs while preventing unsupported personalization.
- Service operations: classify requests, retrieve policy-grounded answers, draft responses, perform approved updates, and escalate based on value, sentiment, or regulatory risk.
- Finance: collect invoice evidence, reconcile records, explain variances, and route exceptions while preserving separation of duties and approval thresholds.
- Compliance: map controls to evidence, monitor policy changes, prepare audit packages, and flag missing documentation without delegating final legal judgment.
- Executive operations: synthesize board materials, operating metrics, customer signals, and project risks into traceable decision briefs rather than generic summaries.
- Workflow intelligence: analyze process logs and employee handoffs to locate queues, repeated rework, duplicate entry, and exceptions that offer stronger ROI than visible chatbot projects.
| Pressure | Opening | |
|---|---|---|
| #1 | Unbounded action: an agent can send, publish, purchase, delete, or modify records beyond the intended scope. Limit tools, transaction values, recipients, environments, and action rates. | Revenue operations: combine approved account data, call notes, product usage, and CRM history to identify missing information and recommend evidence-based next actions. |
| #2 | Prompt injection and data exfiltration: malicious instructions in emails, web pages, or documents can manipulate tool-using systems. Isolate untrusted content and require policy checks outside the model. | Sales preparation: generate cited account briefs, stakeholder maps, discovery questions, and proposal inputs while preventing unsupported personalization. |
| #3 | Silent quality drift: model updates, changing data, new products, and altered processes can degrade results. Maintain versioning, regression tests, sampled review, and rollback options. | Service operations: classify requests, retrieve policy-grounded answers, draft responses, perform approved updates, and escalate based on value, sentiment, or regulatory risk. |
| #4 | Automation bias: employees may accept fluent recommendations without checking evidence. Display sources, uncertainty, and approval responsibilities at the decision point. | Finance: collect invoice evidence, reconcile records, explain variances, and route exceptions while preserving separation of duties and approval thresholds. |
| #5 | Compliance overreach: teams may deploy AI into employment, credit, healthcare, surveillance, or customer communications without recognizing applicable obligations. Classify use cases before launch. | Compliance: map controls to evidence, monitor policy changes, prepare audit packages, and flag missing documentation without delegating final legal judgment. |
For professionals
For executives, the practical mandate is to govern AI as a portfolio of operating changes. Require every proposal to include a baseline, accountable owner, workflow map, evaluation set, authority model, security review, expected economics, and shutdown criteria. For entrepreneurs, concentrate on a painful vertical workflow and build trust through integrations, evidence, and controlled execution—not by claiming universal autonomy. Consultants should distinguish diagnostic work from technology selection: quantify the constraint before recommending a platform. Sales leaders should protect brand and consent while using agents to improve research, preparation, CRM hygiene, and follow-through. Operations teams should own exception taxonomies, service levels, and fallback procedures. Buyers should demand architecture diagrams, retention terms, subprocessor lists, audit capabilities, penetration-test evidence, uptime commitments, export rights, model-change policies, and proof from representative tasks. A useful 90-day sequence is: days 1–15, map and baseline the workflow; days 16–30, define controls and evaluations; days 31–60, integrate and test in shadow mode; days 61–75, enable approval-gated production; days 76–90, assess quality, adoption, incidents, and net value. The final decision should be scale, redesign, contain, or stop—not ‘continue piloting’ by default.
Sources & references
- NIST AI Risk Management Framework (AI RMF 1.0)
- NIST Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- European Commission: Regulatory Framework for Artificial Intelligence
- ISO/IEC 42001:2023—Artificial Intelligence Management System
- OECD AI Principles
- OWASP Top 10 for Large Language Model Applications
- Attention Is All You Need
Agent Oracle examines The AI Chief of Staff Playbook through AI agents, workflow automation, sales intelligence, executive decisions, compliance, and measurable business ROI, with practical signals, risks, examples, and a reason for readers to return as the story changes.
Agent Oracle examines AI Agent ROI Scorecards for Small Teams through AI agents, workflow automation, sales intelligence, executive decisions, compliance, and measurable business ROI, with practical signals, risks, examples, and a reason for readers to return as the story changes.
Agent Oracle examines Workflow Bottleneck Mapping With Voice Agents through AI agents, workflow automation, sales intelligence, executive decisions, compliance, and measurable business ROI, with practical signals, risks, examples, and a reason for readers to return as the story changes.
The frontier has moved from impressive chatbots to systems that can plan, call tools and alter business records. For buyers, the decisive questions are no longer about model spectacle but workflow fit, economic value and governable autonomy.
The center of gravity in artificial intelligence is moving from models that answer questions to systems that pursue goals, use tools, and complete workflows. The competitive question is no longer who has a chatbot, but who can redesign work around bounded, observable agency.
A field guide to separating AI capability from AI theater—and turning agents, automation, and human judgment into measurable operating leverage.
From our own rounds
Measured on Agent Oracle, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 27
- Questions per round
- 1