How Science Actually Works—and What AI Operators Should Copy
Science is not a conveyor belt that turns data into certainty. It is a disciplined system for exposing claims to reality—a model AI buyers can use to test agents, automation ROI, security controls, and operational change.
Saoirse MulliganBooks & ideasFirst published 9/10/2026 · last revised 9/11/2026 with fresh sources, corrections, and new context. Reader corrections are reviewed and folded into future versions.
Summary
Science works less like a collection of settled facts and more like an error-correction system: people propose explanations, derive testable expectations, collect evidence, and invite others to challenge the result. Real progress is usually nonlinear; observations are noisy, models are incomplete, incentives matter, and even respected findings can fail replication. For leaders deploying AI agents, that machinery is directly useful: a vendor demo is a hypothesis, a pilot is an experiment, and production telemetry is evidence—not proof. The practical lesson is to replace broad claims about intelligence or ROI with measurable predictions, controlled comparisons, documented assumptions, and explicit thresholds for scaling, revising, or stopping.
Key takeaways
- Science is organized skepticism: credible claims must risk being wrong.
- A hypothesis becomes useful when it predicts an observable outcome under stated conditions.
- Experiments isolate causes; observational studies reveal patterns but usually support weaker causal conclusions.
- One positive pilot does not establish general reliability—replication across users, workflows, and time matters.
- Negative results and anomalies are operational assets because they identify boundaries and hidden variables.
- Statistics quantify uncertainty; they do not rescue weak measurements, biased samples, or moving success criteria.
- AI-agent evaluations should pre-register metrics, use baselines, log interventions, and monitor security as well as productivity.
- Scientific confidence is provisional and graded: evidence accumulates rather than flipping a claim from false to permanently true.
Deep dive
Science begins by making disagreement measurable
The familiar classroom sequence—question, hypothesis, experiment, conclusion—is a useful sketch, but actual science loops backward. Researchers notice a pattern, construct a model, test one implication, revise the instruments or assumptions, and try again. Galileo’s telescopic observations in 1610 did not single-handedly prove heliocentrism; they weakened parts of the old celestial model and joined evidence accumulated by Copernicus, Kepler, Newton, and others. The core move was not simply observing. It was comparing what competing models implied. That distinction matters in AI operations. ‘The agent improves support’ cannot be tested until improvement means something: perhaps median resolution time falls 20% without increasing reopen rates, policy violations, or customer escalation. The claim must also specify the workflow, model version, tools, traffic mix, and evaluation period. A falsifiable statement creates a decision rule; a slogan creates room to reinterpret any result as success.
Experiments separate causes from coincidence
In 1854, physician John Snow mapped cholera deaths around London’s Broad Street pump. The evidence was not a modern randomized trial, but geographic clustering and a natural comparison—including low mortality among nearby brewery workers who reportedly drank beer—supported waterborne transmission over the dominant ‘miasma’ account. Removing the pump handle became an enduring symbol, although the outbreak was already subsiding and germ theory matured later. The example shows science converging through imperfect but discriminating evidence. For an AI sales agent, the equivalent is not comparing this month with last month and crediting automation. Seasonality, pricing, lead quality, staffing, and campaigns are confounders. A stronger design randomly assigns eligible leads to the existing process or the agent-assisted process, keeps routing rules stable, and measures conversion, time-to-first-response, opt-outs, and human rework. When randomization is impractical, matched cohorts, phased rollouts, interrupted time-series analysis, or difference-in-differences can improve inference—but assumptions must be recorded.
Measurement is where many claims quietly fail
Ignaz Semmelweis observed in 1847 that handwashing with chlorinated lime sharply reduced mortality in Vienna’s First Obstetrical Clinic. His evidence met resistance, partly because mechanisms were unclear and his comparisons were imperfect. Later germ theory supplied a stronger explanatory framework. The episode demonstrates that measurements can reveal a useful intervention before theory is complete—and that institutions do not automatically accept threatening evidence. AI programs face equally consequential measurement choices. ‘Containment rate’ may look excellent if a chatbot prevents access to a human by exhausting the customer. Token cost omits integration, review, failures, latency, and security operations. Accuracy on a static test set may say little about production tool use. Operators should define a metric portfolio: business outcome, customer outcome, process efficiency, safety, and total cost. Audit raw samples as well as dashboards, because averages can conceal severe failures in a small, regulated, or high-value segment.
Replication turns an event into knowledge
In 2011, researchers in the OPERA experiment reported neutrinos apparently traveling faster than light. The team publicized the anomaly cautiously and invited scrutiny. By 2012, equipment problems—including a faulty fiber-optic connection—accounted for the result. This was not science failing; it was science’s correction machinery working under public attention. A surprising measurement survived long enough to trigger checks, then yielded to better diagnosis. An AI pilot similarly produces local evidence, not a universal law. Performance may depend on expert reviewers, a clean backlog, unusually cooperative users, cached knowledge, or a specific model release. Replication means rerunning the workflow across teams, regions, demand peaks, languages, and adversarial cases. Version every model, prompt, tool schema, retrieval source, policy, and evaluation set. If the system changes during testing, the evidence applies to a moving target and must be interpreted accordingly.
Institutions help—and introduce their own biases
Peer review asks knowledgeable people to inspect methods and reasoning before publication, but it does not certify truth. Publication bias favors striking positive results; flexible analyses can generate chance findings; status and commercial interests can shape attention. The reproducibility debates of the 2010s led journals and funders to promote registered reports, data sharing, stronger statistical practice, and replication. These controls reduce some failure modes without eliminating judgment. Enterprises need analogous institutions: independent evaluation, change approval, incident reporting, red-team exercises, access reviews, and post-deployment monitoring. The team rewarded for launching an agent should not be the sole authority declaring it safe and profitable. Procurement evidence should distinguish vendor benchmarks from customer-controlled tests. Good governance makes correction inexpensive and visible: failed calls are retained, overrides are categorized, regressions trigger rollback, and employees can report unsafe behavior without being treated as obstacles to adoption.
Evidence supports decisions, not certainty
Science rarely proves that a complex intervention will always work. It estimates effects under conditions, quantifies uncertainty, tests alternatives, and updates confidence when new evidence arrives. A statistically significant result can be commercially trivial; a large estimated gain may remain too uncertain to justify scaling. Conversely, waiting for perfect certainty can forfeit value. Executives should therefore connect evaluation to action thresholds. Scale an agent only if the lower bound of expected value clears deployment and risk costs; pause if severe incidents exceed tolerance; redesign if gains concentrate in one segment; retire it if monitoring shows persistent drift. Bayesian updating offers a useful managerial metaphor: begin with a prior informed by comparable deployments, then update with pilot and production evidence. The scientific posture is neither enthusiasm nor cynicism. It is disciplined willingness to change the operating model when reality disagrees.
- 1543Nicolaus Copernicus publishes De revolutionibus, offering a mathematically developed heliocentric model.
- 1620Francis Bacon publishes Novum Organum, arguing for systematic observation and induction.
- 1687Isaac Newton’s Principia unifies terrestrial and celestial motion with testable mathematical laws.
- 1847Ignaz Semmelweis introduces chlorinated-lime handwashing and documents lower puerperal-fever mortality.
- 1854John Snow maps London cholera deaths, strengthening the case for waterborne transmission.
- 1925Ronald Fisher’s Statistical Methods for Research Workers helps formalize experimental design and inference.
- 1948The British Medical Research Council publishes its streptomycin tuberculosis trial, an early landmark randomized controlled trial.
- 1962Thomas Kuhn’s The Structure of Scientific Revolutions describes paradigm change and normal science.
- 2011–2012The OPERA faster-than-light neutrino anomaly is traced to equipment problems after extensive scrutiny.
- 2015The Open Science Collaboration reports replication results for 100 psychology studies, intensifying reform efforts.
Glossary
- Hypothesis
- A proposed explanation or relationship framed so evidence can count against it.
- Falsifiability
- The property of a claim that allows a conceivable observation to show it wrong or inadequate.
- Operationalization
- Turning an abstract concept—such as agent quality—into specified, measurable variables.
- Control group
- A comparison group that does not receive the intervention, helping estimate what would otherwise have happened.
- Confounder
- A factor related to both the intervention and outcome that can produce a misleading association.
- Randomization
- Assigning units by chance to balance known and unknown differences on average.
- Replication
- Repeating a study or test to determine whether a result persists under comparable or varied conditions.
- Statistical significance
- A measure of how incompatible data are with a specified null model; it is not effect size, importance, or proof.
- Pre-registration
- Recording hypotheses, outcomes, and analysis plans before inspecting results to constrain retrospective storytelling.
- Model drift
- Degradation or behavioral change as data, users, tools, policies, or model components evolve.
FAQs
Does science prove things true?+
In mathematics, proofs follow from axioms. Empirical science instead builds degrees of confidence by testing models against observations; even durable theories remain open to refinement within new domains.
What is the difference between a hypothesis and a theory?+
A hypothesis is a specific testable proposal. A scientific theory is a broader explanatory framework supported by multiple lines of evidence, such as germ theory or evolution—not an unsupported guess.
Why is one successful AI pilot insufficient?+
A pilot may reflect favorable users, temporary oversight, clean data, or a particular model version. Replication across realistic loads, segments, time periods, and failure scenarios tests whether the benefit generalizes.
Is an A/B test always required?+
No. Randomized tests often support cleaner causal inference, but ethical, contractual, or operational constraints may prevent them. Phased rollouts, matched comparisons, time-series designs, and qualitative failure analysis can contribute evidence if their assumptions are explicit.
Does a low p-value mean an intervention works?+
No. It indicates that the observed data would be relatively unusual under a specified null model and assumptions. Leaders still need effect size, confidence intervals, measurement quality, business value, and risk analysis.
Can vendor benchmarks be trusted?+
They can be useful screening evidence, especially when methods and datasets are transparent. They should not replace customer-controlled tests using the organization’s tools, permissions, data quality, adversarial cases, and success criteria.
How should an AI evaluation handle model updates?+
Version the model, prompts, retrieval corpus, tool definitions, policies, and test set. Material updates should trigger regression tests and, for high-risk workflows, renewed approval before full production exposure.
What counts as a negative result?+
A pre-specified effect that fails to appear, or an intervention whose costs or harms outweigh gains, is a useful negative result. Preserving it prevents repeated mistakes and helps define where automation should not be used.
Risks
- Metric gaming: teams optimize containment, activity, or benchmark scores while customer outcomes, revenue quality, or safety deteriorate.
- False causality: before-and-after comparisons attribute changes to an agent despite seasonality, staffing, campaigns, or altered lead mix.
- Pilot theater: intensive human support and curated data create results that cannot survive normal production conditions.
- Silent drift: model updates, retrieval changes, tool permissions, and user behavior invalidate earlier evidence without triggering reevaluation.
- Governance capture: launch owners or vendors control evaluation, suppress negative findings, or redefine success after seeing results.
Opportunities
- Build an evaluation registry that records hypotheses, baselines, owners, metrics, model versions, and scale-or-stop thresholds before each pilot.
- Use randomized or phased rollouts to estimate the causal effect of sales, support, and operations agents rather than relying on testimonials.
- Treat overrides, escalations, security events, and failed tool calls as structured experimental data for workflow redesign.
- Create reusable adversarial and regression suites covering prompt injection, authorization boundaries, hallucinated actions, privacy, and compliance.
- Apply evidence tiers to procurement: vendor claims, sandbox tests, controlled pilots, replicated deployments, and continuously monitored production results.
Sources & references
- The Logic of Scientific Discovery — Karl Popper
- The Structure of Scientific Revolutions — Thomas S. Kuhn
- Estimands and Sensitivity Analysis in Clinical Trials — ICH E9(R1)
- Estimating the Reproducibility of Psychological Science — Science
- ASA Statement on Statistical Significance and P-Values
- NIST AI Risk Management Framework 1.0
- The Broad Street Pump: An Episode in the Cholera Epidemic of 1854 — UCLA
- OPERA Collaboration: Measurement of the Neutrino Velocity with the OPERA Detector
| Before/after rollout | Randomized controlled pilot | Phased rollout with matched cohorts | |
|---|---|---|---|
| Causal confidence | Low: time trends remain plausible | High when assignment and compliance are sound | Medium: stronger than simple before/after, weaker than randomization |
| Operational disruption | Low | Medium: parallel processes required | Medium: rollout sequencing required |
| Best use | Early monitoring or low-stakes exploration | High-volume sales, support, and routing decisions | Teams, regions, or workflows that must launch in stages |
| Main failure mode | Seasonality or concurrent change mistaken for impact | Spillover, noncompliance, or underpowered sample | Unmatched groups and differing time trends |
| Minimum controls | Stable definitions, historical baseline, intervention log | Pre-registered outcomes, random assignment, guardrails | Baseline equivalence, explicit matching, common metrics |
| Executive decision value | Directional evidence | Strongest scale-or-stop evidence | Practical evidence under rollout constraints |
A beginner-friendly guide to using scientific thinking when evaluating AI agents, diagnosing workflows, testing automation, and making defensible business decisions.
Science is shifting from AI as an analytical tool to AI as an active participant in hypothesis generation, experiment design, laboratory execution, and institutional learning. The prize is not merely faster discovery—it is a compounding operating system for research.
A boardroom-clear guide to reusable launch systems: how they work, where the economics hold, which operators lead, and how AI agents can improve aerospace decisions without compromising safety.
From our own rounds
Measured on Agent Oracle, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 27
- Questions per round
- 1