What the Numbers Say About AI Today: An Operator’s Field Guide: Operator Field Guide

AI is simultaneously a fast-growing capital market, a rapidly adopted workplace tool, and an uneven operating capability. The useful numbers separate model progress from enterprise value—and reveal where leaders should invest, measure, and govern.

Sven LindqvistSven LindqvistMarkets & macro
16 min read· Published 8/25/2026 v1 · updated 8/25/2026· 8 views
AI-assisted, human-reviewed. Drafted with AI research tools from public sources, fact-checked and edited by our team, and revised over time based on reader corrections. How we build these →
AIWhat the Numbers Say AboutAI Today: An Operator’sField Guide: OperatorField GuideORIGINAL EDITORIAL GRAPHIC · AGENT-ORACLE
Original cover graphic by Agent Oracle editorial.Background texture: Photo · Unsplash
Tweet Share Post
Living article · version 1

First published 8/25/2026 · monitored for updates; the next revision publishes a new version and appears here. Reader corrections are reviewed and folded into future versions.

Summary

AI has crossed from experimentation into routine business use, but adoption statistics are running ahead of operational maturity. Stanford’s 2025 AI Index reported that 78% of surveyed organizations used AI in 2024, while McKinsey found 71% regularly used generative AI in at least one function—yet material enterprise-wide earnings effects remained far less common. Costs are also moving in two directions: training frontier models demands extraordinary capital, while inference for established capability levels has become dramatically cheaper. For operators, the central question is therefore not whether AI is growing; it is whether a specific workflow can produce measurable gains after integration, supervision, security, and failure costs are counted.

Key takeaways

  • Adoption is broad: Stanford’s 2025 AI Index says 78% of surveyed organizations reported using AI in 2024, up from 55% in 2023.
  • Generative AI is concentrating first in marketing and sales, product and service development, service operations, software engineering, and IT.
  • Private investment in generative AI reached $33.9 billion globally in 2024, according to Stanford—an 18.7% increase over 2023.
  • Capability is becoming cheaper to consume: Stanford estimates GPT-3.5-level inference cost fell by more than 280-fold between November 2022 and October 2024.
  • Frontier-model development remains capital-intensive; Stanford estimates selected 2024 training runs cost tens to hundreds of millions of dollars in compute.
  • Usage is not ROI. Reliable value requires workflow redesign, data access, integration, exception handling, adoption, and outcome measurement.
  • Agents increase potential leverage by executing multistep work, but they also expand the permission, audit, and failure surface.
  • The board-grade metric is risk-adjusted workflow economics—not chatbot seats, demonstrations, prompts, or model benchmark scores.

Explain like I'm 5

Think of AI as a very fast junior colleague. It can read, draft, classify, search, summarize, and sometimes use software, but it can also misunderstand instructions or confidently invent an answer. The newest ‘agents’ go beyond answering a question: they can attempt a sequence such as finding a prospect, researching the company, updating a CRM record, and drafting an email. The big numbers tell us many businesses have hired this digital junior colleague, and the price of common AI tasks is falling quickly. They do not prove the colleague is producing profit. A sensible operator starts with one repetitive workflow, gives the system limited access, checks its work, records time and error rates before and after deployment, and expands authority only when the evidence supports it.

Deep dive

Adoption is real, but the denominator matters

Stanford’s 2025 AI Index, drawing on McKinsey survey data, reported organizational AI use rising from 55% in 2023 to 78% in 2024. Regular generative-AI use in at least one business function rose from 33% to 71%. Those figures establish direction, not depth. ‘Use’ can include an employee drafting copy, a team deploying a support copilot, or a governed system embedded in a revenue process. Executives should ask what percentage of eligible workflows is affected, how often the tool is used, whether outputs enter systems of record, and whether economic outcomes have changed. Seat counts and prompt volumes are leading indicators; cycle time, conversion, cost-to-serve, quality, and risk are operating results.

The economics split between building and buying

At the frontier, model development is becoming an infrastructure contest. Stanford estimates that selected training runs reached extraordinary compute costs: GPT-4 was about $79 million and Llama 3.1-405B about $170 million, using estimated cloud-compute prices rather than audited total project costs. Most companies should not interpret this as a reason to train a foundation model. They should notice the opposite trend downstream: the estimated price of querying a model at roughly GPT-3.5 capability dropped more than 280-fold from November 2022 to October 2024. Falling unit prices make document triage, call analysis, research, and drafting economically accessible—but integration, evaluation, data preparation, and human review frequently dominate the production budget.

Where value is appearing first

McKinsey’s 2025 survey places frequent generative-AI use in marketing and sales, product and service development, service operations, software engineering, and IT. These domains contain abundant language work, digital records, observable handoffs, and high task volume. A sales agent, for example, can research accounts, enrich CRM fields, summarize calls, draft follow-ups, and flag stalled opportunities. The attractive headline is labor saved; the harder value may be faster lead response, more complete CRM data, better manager visibility, and fewer missed commitments. Every deployment needs a baseline: minutes per case, cases per employee, rework rate, service-level attainment, conversion, and the loaded cost of exceptions.

From copilots to agents

A copilot typically proposes; an agent can plan, call tools, update systems, and continue until a stopping condition is met. That difference changes both returns and controls. Automation value rises when the system closes workflow loops rather than merely creating text for a person to transfer elsewhere. So does potential damage. An agent with email, CRM, payment, or production access can propagate a bad assumption at machine speed. High-quality implementations constrain tools, identities, spend, data scope, recipients, and transaction size; log every consequential action; require approval for irreversible steps; and test failure paths before scaling.

Measure the workflow, not the model

Benchmark scores help vendors and technical teams compare capability, but operators purchase outcomes. A practical ROI model starts with annual eligible volume multiplied by the baseline cost per case. It then estimates adoption, automation or assistance rate, time saved, quality change, revenue lift, and avoided loss. From gross benefit, subtract software, model usage, integration, data work, evaluation, monitoring, change management, human review, and expected failure cost. Report a range rather than a single-point promise. A strong pilot uses a control or credible pre-deployment baseline and runs long enough to capture exceptions, seasonality, and learning effects.

Governance is part of throughput

NIST’s AI Risk Management Framework and its Generative AI Profile organize controls around governing, mapping, measuring, and managing risk. The European Union’s AI Act adds a legal, risk-based regime with phased application after entering into force on August 1, 2024. Operators should maintain an AI inventory covering owner, purpose, model, data classes, vendors, jurisdictions, permissions, human checkpoints, evaluation results, incidents, and retirement criteria. This is not paperwork detached from performance. Clear ownership and reusable controls reduce procurement friction, prevent shadow deployments, accelerate incident response, and make successful workflows easier to replicate.

Timeline
  1. 2017
    Google researchers publish ‘Attention Is All You Need,’ introducing the Transformer architecture behind modern language models.
  2. 2020
    OpenAI releases GPT-3, demonstrating that scaled language models can perform varied tasks from prompts.
  3. 2022
    OpenAI launches ChatGPT on November 30, bringing conversational generative AI to a mass audience.
  4. 2023
    Microsoft rolls out Microsoft 365 Copilot and Google introduces Duet AI, later renamed Gemini for Workspace, accelerating workplace distribution.
  5. 2023
    The White House issues Executive Order 14110 on safe, secure, and trustworthy AI on October 30.
  6. 2024
    The EU AI Act enters into force on August 1, beginning a phased compliance schedule.
  7. 2024
    OpenAI, Anthropic, Google, Meta, and others continue rapid multimodal and long-context model releases as enterprise buying expands.
  8. 2025
    Stanford’s AI Index reports 78% organizational AI use for 2024 and $33.9 billion in global private generative-AI investment.
  9. 2025
    Enterprise attention shifts from standalone copilots toward governed agents that call tools and complete multistep workflows.
Figure — milestone track built from the dated events in this article.

Glossary

Foundation model
A broadly trained model that can be adapted or prompted for many downstream tasks, such as writing, classification, coding, or vision.
Generative AI
Systems that produce content—including text, software code, images, audio, or video—from instructions and context.
AI agent
A system that pursues a goal through multiple steps, often selecting tools, reading results, updating state, and taking actions.
Inference
The process and cost of running a trained model to generate a prediction or response.
Token
A unit of text processed by a language model; API suppliers commonly meter input and output usage in tokens.
RAG
Retrieval-augmented generation: fetching relevant documents or records at runtime and supplying them to a model as context.
Hallucination
A plausible-looking output that is unsupported, false, or inconsistent with the supplied evidence.
Evaluation
A repeatable test of system quality, safety, latency, cost, or business performance against defined cases and thresholds.
Human in the loop
A design in which a person reviews, approves, corrects, or handles selected AI outputs or actions.
Shadow AI
AI use that occurs outside approved procurement, security, data-governance, or monitoring processes.

FAQs

How many companies use AI today?+

Stanford’s 2025 AI Index reports that 78% of surveyed organizations used AI in 2024, based on McKinsey survey data. Treat this as evidence of broad adoption, not proof that 78% have mature, enterprise-scale deployments.

Is generative AI already improving profits?+

Some organizations report cost reductions or revenue gains in individual functions, but enterprise-wide earnings impact is less widespread than adoption. Attribution is difficult unless the company establishes a baseline, tracks exposed users or cases, and includes implementation and review costs.

Which functions should deploy first?+

Start where work is frequent, digital, measurable, and reversible: service triage, sales research, meeting follow-up, document extraction, internal search, and software assistance are common candidates. Avoid choosing solely by enthusiasm; score volume, friction, data readiness, downside, and economic value.

Should we build our own model?+

Most firms should buy model access and build proprietary workflow logic, integrations, evaluations, controls, and data advantages around it. Training a foundation model is rarely justified unless the organization has unusual scale, specialized data, talent, infrastructure, and strategic need.

How should AI ROI be calculated?+

Measure changes in cycle time, labor per case, throughput, error and rework, conversion, retention, and avoided loss. Subtract licenses, inference, integration, governance, monitoring, training, human review, and expected incident costs from the attributable benefit.

When does a copilot become an agent?+

The practical boundary is action and autonomy. A copilot mainly recommends or drafts, while an agent selects steps, invokes tools, changes systems, or communicates externally under defined constraints.

Are cheaper models automatically better for ROI?+

No. A lower token price can be overwhelmed by retries, long context, poor output quality, manual correction, or workflow abandonment. Compare total successful-task cost, including latency and review, rather than API price alone.

What security control matters most?+

There is no single sufficient control, but least-privilege access is foundational. Pair it with data classification, identity separation, approved tools, action logs, evaluations, human approval for high-impact steps, and an incident process.

Predictions

{"items":["Through 2027, model capability is likely to commoditize faster than integration quality; defensible advantage may shift toward proprietary workflow data, evaluations, distribution, and trusted execution.","Agent deployments will probably expand first in bounded workflows with clear permissions and stopping rules, rather than in open-ended ‘digital employee’ roles.","Enterprise buyers are likely to demand task-level economics—cost per resolved case, qualified opportunity, or completed order—instead of accepting seat adoption as the main success metric.","Smaller and specialized models may gain share where latency, privacy, reliability, or unit economics matter more than maximum general capability.","AI governance may become a reusable operating platform: inventories, approved components, testing gates, and audit trails that speed safe deployment rather than merely satisfy compliance."}]}

    Risks

    {"items":["Automation can scale plausible errors, unauthorized messages, or incorrect system updates faster than human teams can detect them.","Sensitive customer, employee, financial, or intellectual-property data may leak through prompts, logs, connectors, retrieval systems, or vendor retention settings.","Unmeasured pilots can create ‘productivity theater’: high usage and attractive demonstrations without attributable margin, revenue, or service improvements.","Vendor concentration, model deprecation, pricing changes, and evolving regulation can turn an apparently simple deployment into operational dependency.","Over-automation can weaken judgment, create automation bias, and remove the human context needed for exceptional or high-impact decisions."}]}

      Opportunities

      {"items":["Diagnose workflows before buying tools: map handoffs, wait states, duplicate entry, missing information, and exception queues to find the real constraint.","Use agents to close low-risk loops—such as researching an account, updating CRM fields, and preparing a manager-approved follow-up—not merely to draft isolated text.","Create an evaluation library from real cases, including adversarial and edge cases; this becomes reusable infrastructure for model changes and procurement.","Translate falling inference costs into wider coverage of document, call, ticket, and transaction analysis while retaining risk-based human review.","Package governance as a shared service with approved models, connectors, identity patterns, logging, and deployment templates so business teams can move faster."}]}

        For professionals

        For an investment committee, the appropriate unit of analysis is the workflow cohort, not the AI product. Define eligible transaction volume; baseline time, queue delay, quality, conversion, and loss; then segment cases by complexity and consequence. Estimate assisted, automated, and escalated shares separately. A useful formulation is: risk-adjusted annual value equals attributable labor capacity plus incremental contribution margin plus avoided loss, minus recurring technology, implementation amortization, supervision, change-management, and expected incident costs. Capacity released is not cash unless it reduces spend, avoids hiring, increases throughput, or is deliberately reassigned to valuable work. Use confidence intervals and identify which assumption—adoption, quality, volume, or conversion—drives the case. For architecture and governance committees, treat an agent as a nonhuman principal operating inside a constrained control plane. Separate model reasoning from deterministic policy enforcement; issue task-scoped credentials; allowlist tools and recipients; cap spend and transaction values; validate structured outputs; preserve source evidence; and log prompts, retrieved context, tool calls, approvals, and resulting state changes. Evaluate both component performance and end-to-end business outcomes. Model substitution tests are especially revealing: if several models clear the quality threshold, route work by sensitivity, latency, and total successful-task cost rather than brand preference.

        Three enterprise AI operating modes
        CopilotBounded agentHigh-autonomy agent
        Primary behaviorDrafts, recommends, summarizesExecutes a defined multistep workflowChooses and executes broad plans
        System accessMostly read-only; user transfers outputLimited tools and scoped write accessMultiple systems with consequential write access
        Human checkpointReviews nearly every outputApproves exceptions or high-impact actionsReviews sampled outputs and escalations
        Best-fit workResearch, writing, meeting summariesCRM hygiene, ticket routing, document processingDynamic operations with mature controls
        Value metricMinutes saved and adoptionCost and cycle time per completed caseThroughput, margin, and service-level performance
        Primary riskLow adoption or bad adviceError propagation inside a bounded processSystemic action at scale
        Figure — Operator comparison of assistance, bounded agency, and high-autonomy automation; controls should scale with consequence and reversibility.
        The AI market in four operating numbers
        78%
        Organizations reporting AI use in 2024
        Stanford AI Index 2025, citing McKinsey; up from 55% in 2023
        71%
        Organizations regularly using gen AI in ≥1 function
        Stanford AI Index 2025, citing McKinsey; 2024 survey result
        $33.9B
        Global private investment in generative AI, 2024
        Stanford AI Index 2025; 18.7% above 2023
        >280×
        Estimated decline in GPT-3.5-level inference cost
        Stanford AI Index 2025; November 2022 to October 2024
        Figure — Adoption, investment, cost, and exposure indicators. Survey definitions and modeled estimates should not be treated as audited company results.
        The enterprise AI value system
        Workflow diagnosisModel capabilityEnterprise dataAgent orchestrationEvaluationSecurity and compli…ROI measurementAI operating val…
        Figure — Seven connected disciplines that determine whether model capability becomes controlled business performance.
        Rate this article
        Suggest a correction
        Discussion (0)
        Keep exploring
        Related reads · in AI
        All in AI
        The AI Chief of Staff Playbook: Operator Field Guide

        Agent Oracle examines The AI Chief of Staff Playbook through AI agents, workflow automation, sales intelligence, executive decisions, compliance, and measurable business ROI, with practical signals, risks, examples, and a reason for readers to return as the story changes.

        5 min read
        AI Agent ROI Scorecards for Small Teams: Operator Field Guide

        Agent Oracle examines AI Agent ROI Scorecards for Small Teams through AI agents, workflow automation, sales intelligence, executive decisions, compliance, and measurable business ROI, with practical signals, risks, examples, and a reason for readers to return as the story changes.

        5 min read
        Workflow Bottleneck Mapping With Voice Agents: Operator Field Guide

        Agent Oracle examines Workflow Bottleneck Mapping With Voice Agents through AI agents, workflow automation, sales intelligence, executive decisions, compliance, and measurable business ROI, with practical signals, risks, examples, and a reason for readers to return as the story changes.

        5 min read
        A Field Report From the AI Frontier: The Operator’s Guide to Agents That Actually Work: Operator Field Guide

        The frontier has shifted from impressive chat to dependable action. Here is what operators need to know about agent design, workflow economics, governance, and the difficult path from demonstration to production.

        14 min read
        The Open Questions That Will Define AI Next: An Operator’s Field Guide: Operator Field Guide

        The decisive AI questions are shifting from model intelligence to agent reliability, workflow economics, control, liability, and organizational design. Here is what business leaders should watch—and test—before placing the next large bet.

        18 min read
        AI Agents Are the Consequential Shift: An Operator’s Field Guide: Operator Field Guide

        The center of gravity in artificial intelligence is moving from models that answer questions to systems that pursue goals, use tools, and complete workflows. The competitive question is no longer who has a chatbot, but who can redesign work around bounded, observable agency.

        18 min read
        Have a question about AI? Ask our AI — it pulls from this article and others.
        Chat about AI
        ← All Knowledge