Where AI Culture Programs Go Wrong—and What Operators Should Do Instead
Most AI transformations do not fail because employees dislike technology. They fail because leaders substitute messaging for workflow design, incentives, governance, and credible operating choices.
Theo MarchettiInvestigations editorFirst published 9/19/2026 · monitored for updates; the next revision publishes a new version and appears here. Reader corrections are reviewed and folded into future versions.
Summary
Culture becomes a convenient explanation when an AI rollout stalls: employees are called resistant, managers insufficiently innovative, or teams too attached to old habits. Usually, the harder truth is operational. The new agent changes authority, workload, measurement, risk, or customer ownership without making those changes explicit. Leaders should treat culture not as a campaign but as the accumulated evidence employees receive from incentives, system permissions, management behavior, incident handling, and which work the organization actually rewards.
Key takeaways
- Do not diagnose ‘resistance’ until you have mapped the workflow, incentives, permissions, data quality, and exception burden.
- An AI agent changes the operating model whenever it recommends, decides, communicates, or takes action—not merely when it generates text.
- Executive adoption matters, but visible executive restraint matters too: leaders must obey review, privacy, and disclosure rules themselves.
- Start with bounded workflows that have measurable volume, cost, quality, and escalation paths; avoid vague mandates to ‘use AI.’
- Measure adoption alongside rework, override rates, incidents, customer outcomes, cycle time, and realized financial value.
- Give named owners authority to redesign work; a champion network without decision rights becomes an unpaid communications layer.
- Make safe behavior easier than unsafe behavior through approved tools, usable controls, logged actions, and fast exception handling.
Deep dive
The culture diagnosis is often a management alibi
When a sales copilot, support agent, or operations assistant underperforms, leaders often reach for a cultural explanation. That diagnosis can protect the program design from scrutiny. A representative may ignore an agent because its CRM context is stale, not because the representative fears AI. A support manager may demand manual checks because the agent can issue a refund but cannot interpret an unusual entitlement. An analyst may use an unapproved model because the approved environment lacks document upload or takes days to provision. These are rational responses to system conditions. Before commissioning training or internal communications, inspect where the workflow begins, which records are authoritative, what the agent may do, how exceptions move, and who absorbs errors.
Confusing enthusiasm with a usable operating model
Hackathons and prompt libraries can expose possibilities, but they do not define production work. A functioning operating model names the process owner, technical owner, data steward, risk approver, frontline reviewer, and incident responder. It specifies when the agent acts autonomously, when it asks permission, and when it must stop. For a voice agent, that includes identity disclosure, call recording and consent rules, authentication, transfer conditions, prohibited claims, and post-call recordkeeping. For a sales agent, it includes approved sources, outreach limits, CRM write permissions, and human approval for pricing or contractual language. Culture improves when these choices are stable and legible.
The hidden tax of ‘human in the loop’
Human review is frequently presented as a complete safeguard. In practice, it can move risk and labor downstream. If a model produces 10,000 outputs and reviewers must inspect all of them, automation may merely create a faster queue. Review also degrades when staff are rushed, cannot see the evidence, or assume the model is usually correct. Design review by risk tier: permit low-impact, reversible actions within policy; sample routine outputs; require approval for consequential actions; and route uncertain or anomalous cases to specialists. Track acceptance, edit, override, escalation, and reversal rates. Those signals reveal whether the agent is reducing work or manufacturing supervision.
Incentives speak louder than transformation language
Employees infer priorities from targets and consequences. A service team measured only on handle time may accept weak AI answers; a risk team punished for every incident may block useful experimentation; sellers paid for closed revenue may resist meticulous AI documentation that slows outreach. Align measures across speed, quality, risk, and value. Managers need explicit capacity to redesign roles and retire obsolete steps, not an instruction to add AI on top of existing work. Leaders should also distinguish productivity capture from productivity sharing. If every efficiency gain immediately becomes a head-count threat, staff have a rational reason to conceal opportunities and defects.
Governance should be embedded, not bolted on
A policy PDF cannot control copied customer data, autonomous tool calls, or model updates. Effective governance appears inside procurement, identity and access management, logging, evaluation, release gates, vendor monitoring, and incident response. NIST's AI Risk Management Framework, published in January 2023, organizes work around Govern, Map, Measure, and Manage. ISO/IEC 42001, published in December 2023, provides requirements for an AI management system. The European Union's AI Act entered into force on August 1, 2024, with obligations phased in over subsequent years. Organizations should translate such frameworks into controls appropriate to each use case rather than treating certification or compliance as proof of reliable outcomes.
Replace culture campaigns with operating evidence
Choose a workflow with a baseline: volume, labor time, conversion or resolution rate, error cost, compliance exposure, and customer impact. Run a bounded pilot with a control or credible pre-period, then document who does what before and after. Publish failures as well as wins, including where humans corrected the agent. At each stage—assistive, recommendatory, approval-gated, and bounded autonomy—require evidence that quality, risk, and economics remain acceptable. The durable cultural message is not a slogan. It is repeated proof that leadership fixes broken processes, respects controls, protects people who report problems, and stops systems whose benefits do not justify their costs.
- 2016Microsoft releases Tay and withdraws the chatbot within about 16 hours after abusive interactions, illustrating the cost of unmanaged deployment context.
- 2018Amazon reportedly abandons an experimental recruiting model after finding that historical hiring data produced disadvantages for resumes associated with women.
- 2021NIST publishes its definition of explainable AI principles, sharpening discussion of explanation, meaningfulness, accuracy, and knowledge limits.
- 2022OpenAI releases ChatGPT on November 30, accelerating employee-led adoption and unsanctioned use of generative AI at work.
- 2023NIST publishes AI RMF 1.0 on January 26, organizing voluntary risk management around Govern, Map, Measure, and Manage.
- 2023ISO and IEC publish ISO/IEC 42001 in December, establishing requirements for organizational AI management systems.
- 2024The EU AI Act enters into force on August 1, beginning a phased implementation schedule for risk-based AI obligations.
- 2024NIST releases a Generative AI Profile for the AI RMF, addressing risks including confabulation, privacy, security, and information integrity.
Glossary
- Agentic workflow
- A process in which an AI system can plan or select steps, use tools, and take actions toward an objective within defined boundaries.
- Bounded autonomy
- Authority granted to an agent only for specified actions, data, thresholds, and circumstances, with enforced stopping and escalation rules.
- Shadow AI
- AI services used for work without organizational approval, visibility, or controls, often because sanctioned tools are absent or impractical.
- Human-in-the-loop
- A design requiring human participation at a decision point; its effectiveness depends on time, evidence, competence, and genuine authority to intervene.
- Override rate
- The share of agent recommendations or actions that humans reject, materially edit, reverse, or replace.
- Automation bias
- The tendency to favor a system's output even when contradictory information or warning signs are available.
- Process owner
- The person accountable for end-to-end workflow performance and empowered to alter roles, controls, handoffs, and measures.
- Evaluation harness
- A repeatable set of test cases, metrics, thresholds, and procedures used to assess an AI system before and during deployment.
- Sociotechnical system
- A system whose outcomes emerge from the interaction of people, technology, rules, incentives, information, and organizational structure.
FAQs
How can leaders tell whether the problem is culture or product quality?+
Observe work and compare stated objections with telemetry. If usage falls where latency, missing context, false answers, or duplicate entry rises, repair the product and workflow first; interviews and override data can then identify residual trust or capability issues.
Should AI use be included in employee performance goals?+
Do not reward raw usage counts, prompts, or logins. Tie goals to approved outcomes such as reduced cycle time, maintained quality, complete records, and compliant escalation, while allowing staff to reject AI when it is unsuitable.
Is a center of excellence enough to change behavior?+
A center can provide standards, reusable components, evaluations, and coaching, but it cannot own every process. Business process owners need budgets, decision rights, and accountability for local workflow outcomes.
How should an organization respond to shadow AI?+
First determine what unmet need the unauthorized tool solves. Provide a practical approved route, protect relevant records, and apply proportionate enforcement; blanket punishment can hide use without removing demand.
When is human review necessary?+
Require it where actions are consequential, difficult to reverse, legally sensitive, or outside validated boundaries. For routine low-risk work, sampling and post-action monitoring may provide better control than universal review.
What should be communicated before a pilot?+
Explain the business problem, affected work, data used, agent authority, prohibited actions, evaluation criteria, escalation route, and implications for roles. Be candid about uncertainties and state who can pause deployment.
How long should an AI culture program take?+
Avoid treating culture as a time-boxed campaign. A bounded workflow pilot may show operational evidence within weeks or months, while trust and management habits develop across repeated releases, incidents, and role changes.
Who should be accountable when an agent makes a mistake?+
Accountability should remain with named organizational owners, not the model or frontline reviewer alone. Assign responsibility across process, technology, data, risk, and executive sponsorship before production access is granted.
Risks
- Performative adoption: teams generate visible demos and usage statistics while core workflows, customer outcomes, and unit economics remain unchanged.
- Silent work transfer: automation shifts verification, correction, and exception handling to lower-visibility roles, producing burnout and misleading ROI.
- Control theater: policies require approval or human review, but reviewers lack time, evidence, authority, or reliable audit logs.
- Trust collapse after an incident: opaque leadership communication or punishment of messengers encourages concealment, shadow systems, and delayed escalation.
- Premature autonomy: agents receive broad credentials before tool behavior, rollback, identity, data boundaries, and adversarial failure modes are tested.
Opportunities
- Use agent telemetry—edits, overrides, escalations, latency, failed tool calls, and reversals—as a continuous map of workflow friction.
- Create role-specific AI operating agreements that define authority, disclosure, review, evidence, retention, and escalation in language frontline teams can apply.
- Turn governance into a deployment accelerator through preapproved architectures, risk tiers, evaluation templates, vendor clauses, and reusable controls.
- Share productivity gains through better capacity planning, redesigned jobs, training, incentives, or service improvements, increasing employees' willingness to surface viable automation.
- Treat voice and customer-facing agents as service redesign projects, linking containment and cost metrics to resolution quality, customer effort, complaints, and safe transfer performance.
Sources & references
- NIST AI Risk Management Framework (AI RMF 1.0)
- NIST AI RMF Generative Artificial Intelligence Profile
- ISO/IEC 42001:2023 — Artificial intelligence management system
- EU Artificial Intelligence Act — Regulation (EU) 2024/1689
- OECD AI Principles
- ILO: Generative AI and Jobs—A Global Analysis of Potential Effects on Job Quantity and Quality
- Microsoft Work Trend Index 2024
| Culture campaign | Tool-first rollout | Workflow-first operating change | |
|---|---|---|---|
| Starting question | How do we make people embrace AI? | Which platform can we deploy quickly? | Which measurable workflow constraint should change? |
| Primary owner | HR or communications | IT or vendor team | Process owner with technology, risk, data, and frontline partners |
| Success measures | Training completion, sentiment, logins | Licenses, feature usage, demos | Cycle time, quality, rework, risk, customer outcome, realized value |
| Governance pattern | Policy and awareness | Controls added after configuration | Risk tier, permissions, tests, logs, escalation, and rollback designed in |
| Typical failure | Cynicism and performative compliance | Orphaned tool or unsafe workarounds | Slower start if ownership and baseline data are weak |
| Durable behavior | Low: messaging fades | Medium: depends on utility | Higher: reinforced by redesigned work and incentives |
A practical operating model for deciding what AI agents should own, what humans must retain, and how to build delegation habits that improve speed without weakening accountability.
The durable contest is no longer streaming versus theaters or humans versus AI. It is trusted scarcity versus synthetic abundance—and the operators controlling rights, communities, discovery, and live experiences currently hold the stronger hand.
Unpacking common misapprehensions about travel, this guide leverages an AI-centric lens to dissect how intelligent agents are reshaping everything from logistics to perceived value, offering strategic insights for executives and operational leaders.
From our own rounds
Measured on Agent Oracle, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 27
- Questions per round
- 1