The Questions Worth Asking Before You Commit to AI Automation
A boardroom-ready discipline for testing AI agents, workflow automation, voice systems, and vendor claims before money, data, and operating leverage are put at risk.
Saoirse MulliganBooks & ideasFirst published 10/9/2026 · monitored for updates; the next revision publishes a new version and appears here. Reader corrections are reviewed and folded into future versions.
Summary
AI commitments rarely fail because a model cannot produce an answer; they fail because buyers automate an unstable process, trust a weak business case, or discover governance constraints after deployment. Before approving an AI agent, voice platform, or workflow automation program, leaders should interrogate the decision across value, process readiness, accountability, evidence, integration, security, compliance, and reversibility. The central question is not whether the technology appears intelligent, but whether it can improve a defined operating outcome under real-world constraints. A disciplined commitment process converts enthusiasm into a bounded, measurable decision—and preserves the option to stop.
Key takeaways
- Start with the business outcome and baseline, not the model, demo, or vendor category.
- Ask whether the underlying workflow is stable enough to automate; AI can accelerate disorder as efficiently as it accelerates good process.
- Calculate total operating cost, including integration, evaluation, human review, security work, exception handling, and model usage.
- Define what the agent may decide, what requires approval, and who owns failures before production access is granted.
- Demand evidence from representative tasks and edge cases; a polished demonstration is not a production evaluation.
- Treat data access, retention, subprocessors, regional processing, and incident response as purchase criteria—not post-contract details.
- Prefer staged, reversible commitments with explicit success thresholds, kill criteria, export rights, and termination assistance.
- Measure quality and risk alongside speed and savings; automation volume alone is not evidence of value.
Deep dive
What decision are we actually making?
A proposal to ‘adopt AI’ is not decision-ready. Name the operating commitment precisely: for example, allowing an agent to draft renewal outreach, letting a voice system reschedule appointments, or authorizing software to update CRM records after calls. Then identify the economic result expected—lower cost per resolved ticket, shorter sales-cycle time, increased conversion, reduced rework, or greater service capacity. Record the current baseline before any pilot. Without queue volume, handling time, error rate, escalation rate, revenue yield, and labor cost, claimed improvement cannot be separated from normal variation. Also ask what happens if nothing changes for 12 months. That counterfactual reveals whether the initiative solves an urgent constraint or merely follows a procurement trend.
Is the workflow ready—or merely painful?
Pain does not automatically make a process automatable. Map the trigger, inputs, systems, decision points, handoffs, exceptions, approvals, outputs, and process owner. If employees routinely use undocumented judgment or repair incomplete records, the proposed agent will inherit that ambiguity. Ask which steps are deterministic and should use conventional rules or robotic process automation, and which genuinely require language or probabilistic reasoning. Examine task frequency and variation: a high-volume, bounded workflow usually offers a stronger starting point than an infrequent process with costly exceptions. Before buying software, determine whether deleting a step, changing a policy, improving a form, or integrating two systems would solve the problem more reliably.
What evidence survives contact with production?
Replace generic benchmark claims with an evaluation set drawn from representative work, difficult edge cases, adversarial inputs, and known historical failures. Define acceptance thresholds by consequence. A customer-email draft can tolerate human correction; a refund, account change, or regulated disclosure requires tighter controls. Test end-to-end performance, including retrieval, tool calls, permissions, latency, handoffs, and recovery—not only the language model's response. Request references from customers with comparable volumes, jurisdictions, and systems. Ask vendors to distinguish measured production results from modeled estimates. During a pilot, compare against the baseline and, where practical, a control group. Track false actions, missed actions, unsupported claims, human override rates, and outcomes by customer segment rather than relying on one aggregate accuracy score.
Where does accountability sit?
An agent can execute work, but it cannot accept corporate accountability. Assign an executive sponsor, business process owner, technical owner, security reviewer, data owner, and incident commander. Establish an autonomy ladder: observe only; draft for approval; act within narrow limits; or act broadly with retrospective review. Specify prohibited actions and conditions that force escalation, such as low confidence, authentication failure, vulnerable-customer language, payment disputes, legal threats, or requests involving sensitive data. Logging must show the inputs, retrieved material, model or configuration version, tools invoked, approvals, actions, and outcome. Human oversight should be operationally credible: if one reviewer receives 10,000 daily alerts, the control exists only on paper.
What will this cost after the demo?
Total cost includes licensing, usage charges, implementation, connectors, identity management, data preparation, observability, evaluations, red-team testing, human review, training, change management, exception handling, and ongoing tuning. Voice automation adds telephony, transcription, text-to-speech, carrier, recording, and consent-management costs. Model prices may fall while consumption rises, so test low, expected, and peak-volume scenarios. Calculate contribution using conservative adoption and quality assumptions, then identify the break-even volume and payback period. Benefits should not be double-counted: time ‘saved’ has financial value only if capacity is redeployed, demand is absorbed, overtime falls, or headcount plans change. Revenue uplift needs attribution rather than optimistic multiplication.
Can we govern, secure, and exit it?
Map every data class the system receives, creates, stores, or sends to subprocessors. Ask whether prompts and outputs train vendor models, where processing occurs, how long logs persist, how deletion propagates, and whether customer-managed keys, single sign-on, role-based access, audit exports, and regional controls are available. Conduct threat modeling for prompt injection, poisoned knowledge sources, excessive tool permissions, data exfiltration, impersonation, and denial-of-wallet attacks. Align controls with applicable obligations, including GDPR, sector rules, call-recording laws, and the EU AI Act's phased requirements. Finally, design the exit before entry: require exportable data and logs, documented interfaces, termination assistance, deletion attestations, and a manual fallback. The strongest commitment is one the company can reverse without losing customer continuity, evidence, or negotiating leverage.
- 2016The EU General Data Protection Regulation is adopted, establishing principles such as purpose limitation, data minimization, security, and data-subject rights relevant to automated systems.
- 2018GDPR becomes applicable on May 25, raising the stakes for vendors and buyers processing personal data in Europe.
- 2020NIST publishes Privacy Framework 1.0, giving organizations a structured method for managing privacy risk across products and operations.
- 2022OpenAI releases ChatGPT on November 30, accelerating executive demand for generative-AI pilots and procurement frameworks.
- 2023NIST releases AI Risk Management Framework 1.0, organizing AI risk work around Govern, Map, Measure, and Manage.
- 2023The White House issues Executive Order 14110 on October 30, signaling stronger attention to AI safety, security, privacy, and procurement.
- 2024ISO/IEC 42001 gains prominence as organizations begin implementing certifiable AI management systems around governance and continual improvement.
- 2024The EU AI Act enters into force on August 1, beginning phased obligations based largely on risk classification and system role.
- 2025The EU AI Act's prohibitions, definitions, and AI-literacy provisions begin applying on February 2, while broader requirements continue to phase in.
Glossary
- AI agent
- Software that uses an AI model to interpret context, choose steps, invoke tools, and pursue a defined objective with some degree of autonomy.
- Autonomy boundary
- The explicit limit on what an automated system may read, decide, change, spend, communicate, or approve without human authorization.
- Evaluation set
- A controlled collection of representative, difficult, and risky cases used to measure system behavior before and during deployment.
- Human in the loop
- A design in which a person reviews or authorizes selected outputs or actions; its value depends on timing, workload, competence, and authority.
- Prompt injection
- Instructions embedded in user input or external content that attempt to override an agent's intended rules or induce unsafe tool use.
- Retrieval-augmented generation
- A method that supplies a model with selected information from approved sources at response time, usually to improve relevance and grounding.
- Total cost of ownership
- The full lifecycle cost of a system, including software, usage, integration, controls, people, exceptions, maintenance, and eventual exit.
- Model drift
- A decline or change in system performance caused by shifting data, workflows, user behavior, models, or surrounding business conditions.
- Kill criterion
- A pre-agreed condition—such as excessive error, cost, or risk—that pauses or terminates a pilot or production deployment.
- System of record
- The authoritative source for a business entity or transaction, such as a CRM for accounts or an ERP platform for orders.
FAQs
What is the first question to ask before buying an AI agent?+
Ask which measurable operating outcome must improve and what its current baseline is. If the owner cannot specify volume, cost, quality, cycle time, or revenue performance today, the team is not ready to claim ROI tomorrow.
How long should an AI automation pilot run?+
Run it long enough to encounter representative volume, normal operating variation, and meaningful exceptions—not for an arbitrary number of weeks. A bounded workflow may produce evidence quickly, while seasonal sales or support processes may require a longer observation window.
Should we build, buy, or combine both approaches?+
Buy when the workflow is common and vendor capabilities are mature; build when proprietary logic or differentiation justifies ownership. A hybrid approach often works best: purchase infrastructure, then retain control over data, evaluations, orchestration, and critical integrations.
What ROI threshold should leadership require?+
There is no universal hurdle rate, but the calculation should include full lifecycle cost and a risk-adjusted benefit range. Use conservative, expected, and upside scenarios, then require a payback period consistent with the company's capital constraints and the technology's rate of change.
When is human approval necessary?+
Require approval when an action is difficult to reverse, financially material, legally sensitive, or likely to affect rights, access, safety, or reputation. Lower-risk drafting and classification tasks may use sampling and retrospective review if monitoring is strong.
What security evidence should a vendor provide?+
Request architecture and data-flow documentation, independent assurance such as a relevant SOC 2 report, penetration-test summaries, subprocessor details, incident procedures, access controls, and retention settings. Certifications support due diligence but do not replace testing the proposed configuration and workflow.
How do we avoid vendor lock-in?+
Negotiate export rights for inputs, outputs, logs, evaluations, and configuration in usable formats. Favor documented APIs, modular integrations, portable knowledge sources, and contract terms covering termination support and verified deletion.
What metrics belong on an executive dashboard?+
Track business outcome, quality, exception rate, human override, latency, unit cost, adoption, customer impact, and material incidents. Segment results by workflow and risk level so strong performance in easy cases does not conceal failure in consequential ones.
Risks
- Automating an undocumented or unstable workflow can increase error volume, hide root causes, and make ownership less clear.
- Excessive permissions or weak retrieval controls can expose confidential data and turn prompt injection into unauthorized business action.
- A favorable demo may collapse under production latency, edge cases, accents, system outages, hostile inputs, or incomplete customer records.
- Long contracts, proprietary connectors, non-exportable logs, and bundled model services can create commercial and technical lock-in.
- Poorly designed workforce changes can remove essential judgment, reduce adoption, and leave nominal human reviewers unable to intervene effectively.
Opportunities
- Use commitment questions as a standard investment gate, creating a shared language for operations, finance, security, legal, procurement, and business owners.
- Start with high-volume, bounded workflows such as call summarization, lead routing, knowledge retrieval, or after-call CRM updates, then earn greater autonomy through evidence.
- Build reusable evaluation libraries from historical exceptions and incidents; these become durable operating assets across vendors and model upgrades.
- Instrument workflows before automating them. Process data often reveals that policy simplification or conventional integration produces faster returns than an agent.
- Negotiate from a reversible architecture: portable data, modular tools, independent observability, and model choice can improve resilience and commercial leverage.
Sources & references
- NIST AI Risk Management Framework (AI RMF 1.0)
- NIST AI RMF Generative Artificial Intelligence Profile
- European Commission: Regulatory Framework for AI
- Regulation (EU) 2024/1689 — Artificial Intelligence Act
- ISO/IEC 42001:2023 — Artificial Intelligence Management System
- OWASP Top 10 for Large Language Model Applications
- NIST Cybersecurity Framework 2.0
- IBM Cost of a Data Breach Report 2024
| Short pilot | Staged rollout | Enterprise-wide contract | |
|---|---|---|---|
| Best use | Validate feasibility and baseline lift in one bounded workflow | Expand a proven use case by team, region, or autonomy level | Standardize a mature capability with known demand across functions |
| Initial exposure | Low; narrow data, users, and duration | Moderate; exposure increases only after gates are met | High; commercial scope and organizational dependency begin early |
| Evidence produced | Task quality, integration difficulty, unit-cost estimate | Production reliability, adoption, exception load, realized ROI | Portfolio economics, governance consistency, vendor scale performance |
| Required controls | Sandboxing, test data, acceptance thresholds, kill criteria | Role-based access, monitoring, incident playbooks, change control | Central governance, auditability, resilience testing, contractual assurance |
| Reversibility | High if no automatic renewal and data is exportable | Medium to high when each stage has an exit gate | Often low without termination assistance and portable architecture |
| Primary failure mode | Pilot theater: impressive output with no path to operations | Scaling before edge cases, ownership, and support capacity are ready | Lock-in, shelfware, and standardizing the wrong workflow |
A practical blueprint for turning AI agents into a secure, measurable operating layer for executive decisions, sales execution, workflow diagnosis, and company-wide automation.
A boardroom-ready framework for governing AI agents across risk classification, data access, human oversight, vendor controls, testing, monitoring, and audit evidence.
A boardroom-ready framework for funding AI-agent pilots, measuring their economics, containing risk, and deciding which workflows deserve production scale.
A practical operating model for using AI agents to improve sales responsiveness, consistency, and conversion while preserving consent, judgment, security, and the human credibility behind every customer relationship.
A boardroom map of the vendors, operators, advisers, platforms, and control functions shaping AI agents, voice automation, and enterprise workflows—and a practical way to separate capability from accountability.
A beginner-friendly guide to how businesses create value, organize work, measure results, and decide where AI agents and automation genuinely belong.
From our own rounds
Measured on Agent Oracle, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 27
- Questions per round
- 1