← back to methodeutics

Vicarious Trials

Chapter 12 · Vincenti 1990, Downer 2017, Wang 2020, Leveson 2012

Everything so far has assumed you can run the trial again tomorrow. Strip that assumption and the same accounting produces a different discipline. An aircraft may crash once, destroy the mechanism that failed, and make repetition unthinkable. Inquiry survives by moving the perturbation from the world to the causal model: each candidate history predicts a different residue, and the wreckage kills the histories whose traces are missing.

Chapter 11 gave us an evidence machine for cheap trials. Place a bet against a hypothesis, observe another case, update the bankroll, repeat. Software is the clean substrate. A test costs seconds, the input can be replayed, and the output is mechanically visible. The e-value earns its trajectory one observation at a time.

Aviation breaks every convenience at once. The event is rare. The trial destroys its apparatus. The outcome kills people, so no investigator can reproduce it merely to settle a claim. Important components may burn, deform on impact, or disappear. The evidence arrives after the system has stopped producing it.

The loss of repetition does not end verification. It changes what receives the perturbation. A direct trial varies the system and reads the outcome. A vicarious trial holds the surviving evidence fixed, varies the proposed history, and asks what else each history would have left behind. The world ran the trial once. Investigation recovers its design from the differential traces.


Two crashes and a mechanism that left no mark

United Airlines Flight 585 crashed near Colorado Springs in 1991 after an abrupt loss of control. The flight-data record was sparse, the aircraft was destroyed, and no physical examination produced a decisive mechanism. Weather, pilot action, and a rudder malfunction all remained live. The investigation could reconstruct the descent but could not yet discriminate its histories.

USAir Flight 427 crashed near Pittsburgh three years later with a similar upset. Its recorder supplied a richer trajectory. Simulator runs compared that trajectory against candidate failures, and only a rudder hardover reproduced the recorded motion closely enough to survive. That result moved inquiry toward the main rudder power-control unit, but it did not explain how the unit could reverse a pilot's command.

The mechanism appeared in 1996 when investigators injected hot hydraulic fluid into a cold control unit. Under the thermal differential, a secondary slide could jam while the primary slide overtraveled. The resulting hydraulic flow could drive the rudder opposite to the commanded direction. The test did not replay either crash. It reproduced a candidate mechanism and generated the feature the accident histories needed: an uncommanded reversal that could occur without leaving an obvious physical mark.

A surviving crew supplied another trace that wreckage could not. Eastwind Flight 517 experienced a similar upset in 1996 and landed. Its captain reported stiff rudder pedals, control inputs, and the sequence by which the upset ended. The intact aircraft and living witnesses added consequences predicted by rudder reversal but unavailable in the two fatal cases. In 1999 the NTSB attributed Flight 427 to a rudder movement opposite to the pilots' command, most likely produced by the jam-and-overtravel mechanism. In 2001 it revised Flight 585 to the same probable cause. The later evidence changed what the earlier event could support.

This inquiry was not one experiment repeated three times. Each occurrence exposed a different projection of the mechanism:

Occurrence Trace gained Hypotheses constrained
Flight 585Upset trajectory and wreckageBounded the event; did not settle the mechanism
Flight 427Richer recorder trajectorySelected rudder hardover among simulated failures
Thermal testReproducible valve reversalSupplied a mechanism for opposite-command motion
Flight 517Surviving aircraft and crew reportConnected the mechanism to an in-flight sequence

The evidence accumulated across unlike trials because each one bore on a common causal claim. Their values cannot simply be multiplied as though they were independent e-values. The simulator was conditioned on the recorder data, the thermal test was chosen because the simulator implicated the rudder, and the later incident was recognized through the hypotheses already in view. The dependence is part of the provenance. What composes is the constraint: each trace removes histories the remaining traces cannot restore.


Reverse causal inference

Direct experimentation runs from a candidate cause to its predicted effect. Failure investigation begins with the effect and leads backward. That reversal is Peirce's retroduction in its literal direction, but it needs more discipline than choosing the explanation that sounds best.

Yafeng Wang reconstructs this discipline from five NTSB investigations. Three operations recur:

All three turn compatibility into exposure. A hypothesis that explains only the known endpoint can absorb almost anything. A hypothesis that commits to a dependency, an additional outcome, or an intermediate state can lose. The added consequence is the kill condition.

This is deduction doing the work Chapter 1 assigned it. Abduction proposes a history. Deduction unfolds the traces that history requires. Induction checks those requirements against the record, then simulations and component tests reproduce their consequences. The event runs backward only in the investigator's order of discovery. The modes keep their direction.

observed outcome
    ↓ abduce
candidate histories
    ↓ deduce
differential traces
    ↓ check
killed histories + surviving causal account

The phrase differential trace matters. Smoke supports fire, but it may support a dozen fires equally. It moves no edge among them. A trace earns weight from the hypotheses it separates. Generic fit is cheap; discrimination is knowledge.


The economics selects the inquiry regime

Chapter 9 ranked experiments by information gain per unit cost. It treated the experiment as a choice within a regime. The aviation case reveals that cost also selects the regime itself.

Trial cost is not one number. Money and time matter, but neither distinguishes a slow laboratory assay from a fatal crash. The full cost is a vector:

C(T) = ⟨money, latency, harm, irreversibility, rarity, observability⟩

Those dimensions determine how often reality can answer and how sharply its answer separates hypotheses. Call that capacity falsification bandwidth: the rate at which an inquiry can expose competing claims to discriminating verdicts under its trial economics.

Software usually has high falsification bandwidth. The investigator can perturb one line, run thousands of cases, and inspect every intermediate state. A clinical trial has lower bandwidth because recruitment takes time and treatment can harm. An aircraft loss has lower bandwidth still because the event is rare, destructive, and unavailable by design. Yet a recorder can make one event highly observable, so rarity alone does not decide the bandwidth. Cost limits the number of answers; instrumentation determines their resolution.

The bandwidth selects the inquiry regime:

Trial economics Inquiry regime Primary verdict
Cheap, repeatable, observableSequential experimentExecution and e-value trajectory
Expensive but repeatableDesigned experimentPooled, precommitted comparison
Rare and uncontrolledNatural experimentVariation supplied by the world
Singular or destructiveVicarious trialDifferential traces among histories
Catastrophic and prospectiveSimulation, certification, assuranceSurrogate behavior plus monitored barriers
Neither trialable nor trace-bearingNo empirical inquiryNo earned causal entitlement

The table is not a ladder from weak to strong. A well-instrumented singular event can discriminate mechanisms that a thousand poorly framed repetitions cannot. Nor does low bandwidth license weaker standards by itself. It demands more informative traces, stronger deductions, and an honest residual category for what the evidence cannot settle.

The law is narrower than "expensive trials produce retrospective inquiry." Expensive prospective trials still exist. Wind tunnels, fatigue rigs, destructive structural tests, and certification flights spend heavily to avoid spending lives. Walter Vincenti called these substitutions vicarious trial. Analytical, experimental, and simulated methods shortcut direct trial of the finished artifact. Failure investigation extends the term backward. Both forms purchase information without paying the full worldly consequence.


Reliability accumulated across a fleet

John Downer identifies a paradox in aviation knowledge. Tests cannot directly demonstrate a one-in-a-billion-hour failure rate; doing so would require an impossible duration of failure-free testing. Models reach farther only by combining tested components with assumptions about independence, relevance, and operating conditions. Those assumptions are exactly where unfamiliar failures hide.

Yet mature jetliners often achieve the reliability their assessments predict. Downer explains the result through three institutions working together:

  1. Service history. Large fleets of closely related aircraft accumulate operational hours and rare occurrences that laboratory tests cannot reproduce.
  2. Design stability. Manufacturers delay changes to architectures and materials until other settings have produced substantial experience. Similarity keeps old evidence relevant to new aircraft.
  3. Recursive practice. The industry investigates failures, converts their surprises into narrow design knowledge, and carries those corrections into later aircraft.

None suffices alone. Service without investigation repeats its failures. Investigation without design stability produces lessons for artifacts that no longer exist. Stability without accumulated service preserves an untested design. Together they make engineering knowledge cumulative: operations produce anomalies, investigations produce distinctions, and stable lineages preserve the reach of those distinctions.

This changes the unit of the knower. No engineer possesses the operational history of a fleet. The representation lives across accident reports, maintenance records, certification rules, test rigs, design handbooks, and successive artifacts. Aviation reliability is not only a property verified by an institution. It is knowledge implemented as an institution.

The result does not transfer automatically to every dangerous technology. A small population of heterogeneous systems produces little comparable service history. Rapid redesign makes yesterday's failures less informative about tomorrow's machine. The same formal reliability calculation can therefore carry different epistemic weight in two industries because their recursive practices preserve different amounts of relevant experience.

This section adapts John Downer, "The Aviation Paradox", especially his account of service history, design stability, and recursive practice. Licensed CC BY 4.0. The terminology is Downer's; the synthesis and connection to methodeutics are this book's.


The hypothesis graph under low bandwidth

A cheap-trial hypothesis graph stores a claim, a perturbation, a kill condition, and the observed verdict. Under low bandwidth, the perturbation may be unavailable. The node must instead bind a causal history to the traces that distinguish it from its rivals.

history: rudder moved opposite command
requires:
  - recorded motion consistent with rudder hardover
  - a mechanism capable of hydraulic reversal
  - no pilot input sufficient to produce the trajectory
killed_by:
  - recorder trajectory incompatible with hardover
  - intact control unit unable to reverse under bounded conditions
  - an alternative history predicting the traces more completely

The graph also needs typed edges. Evidence supports a claim. Events cause or enable other events. Barriers block a causal path. A later finding may undercut the warrant connecting evidence to a claim without rebutting the claim itself. Writing all four as reason would return the causal structure to the reader's head.

Edge Question it answers
supports(trace, claim)Why believe this feature occurred?
causes(event, outcome)Which intervention would change the outcome?
blocks(barrier, path)Which safeguard should have broken the sequence?
undercuts(trace, edge)Why does this evidence fail to carry the claimed weight?

An accident report usually gives the reader all four relations in prose. It supplies the chronology and findings, then states probable cause, contributing factors, and recommendations. The integrated graph remains implicit. An explicit graph changes the burden. Every causal edge must name its differential trace. Every proposed barrier must name the path it would interrupt. Every unresolved rivalry stays open.

Backport the residual

A vicarious trial runs the causal graph forward. Let Gt generate predicted evidence Êt, then compare that prediction with the preserved evidence E. Wherever measurement permits, compute the residual:

Gt → Êt
Rt = E - Êt
backport(Gt, Rt) → Gt+1

To backport is to propagate the residual of a vicarious trial backward into the causal graph. The residual may revise a node, edge, timing relation, parameter, or omitted variable. Evidence does more than change confidence in a fixed graph. It changes the representation that generated the failed prediction.

Not every residual is numeric. An absent fracture, an impossible event order, or an unrecorded intermediate state can still identify the part of the graph that failed. Numeric residuals add magnitude and shape; qualitative residuals add constraints. Both must remain attached to the observation and measurement process that produced them.

Residuals between rival chains

Two causal chains do not have a residual directly. Each chain must first pass through the same observation model. Let rival graphs G1 and G2 predict the traces the available instruments would record:

Ê1 = observe(run(G1))
Ê2 = observe(run(G2))
R1 = discrepancy(E, Ê1)
R2 = discrepancy(E, Ê2)

The comparison is between residual signatures, not graph shapes. One history may fit the final position but miss the timing; another may fit the trajectory but predict a fracture that is absent. A useful signature keeps the discrepancies separate by trace, time, and uncertainty rather than collapsing them immediately into one score.

Trace Observed Chain 1 predicts Chain 2 predicts Discriminator
TrajectoryRecorded seriesSeries with uncertainty bandSeries with uncertainty bandNormalized error and temporal shape
Event orderPartial chronologyRequired orderRequired orderConstraint violation
Physical residuePresent, absent, or unknownExpected frequencyExpected frequencyLikelihood or impossible trace
Intermediate stateMeasured or unrecordedRequired rangeRequired rangeCompatibility at instrument resolution

Quantitative traces can use error normalized by measurement uncertainty, likelihood, or another preregistered discrepancy function. Qualitative traces use logical compatibility: required and present, required and absent, forbidden and present, or unobserved. Missing evidence is not a zero residual. It is an open coordinate.

A scalar ranking is justified only after the inquiry states how discrepancies compose. Squared error, likelihood, worst-case violation, and lexicographic safety constraints answer different questions. In a safety investigation, one impossible required trace may properly kill a chain despite an excellent aggregate fit elsewhere.

When both chains predict the same observable signature within uncertainty, existing evidence cannot attribute between them. The difference between their predictions specifies the next measurement obligation: find or invent the instrument, recorder field, component test, or surviving trace that makes the chains diverge.

The economy of research asks: which affordable observation would best tell the remaining explanations apart? Under cheap trials, take that measurement next. Under expensive trials, install the recorder or preserve the trace now so the next rare event can answer it.

Backporting stops at an equivalence class when several graphs predict traces within the available resolution. Choosing one graph beyond that point would move the residual back into the reader's intuition. The honest output is the set of surviving histories and the observation that would separate them.

One loop, several disciplines

The same epistemic loop appears under different professional vocabularies. Each field prepares for failures its current model can express. A consequential event then leaves a residual the model did not absorb. Retrospective inquiry explains that surprise and backports it into future practice.

Discipline Prospective preparation Residual evidence Vicarious trial Backport
Aviation Certification, redundancy, simulator training Recorder data, wreckage, witness sequence Reconstruct and simulate candidate histories Design change, airworthiness directive, new recorder parameter
Site reliability engineering Fault model, redundancy, chaos test, runbook Logs, metrics, traces, operator timeline Replay the incident against competing dependency graphs Postmortem, new alert, revised architecture or runbook
Medicine Differential diagnosis, protocol, monitoring Symptoms, labs, imaging, treatment response Compare disease courses and counterfactual treatments Revised diagnosis, protocol, contraindication or case definition
Structural engineering Load cases, safety factors, inspections Fracture surface, deformation, sensor and maintenance history Reconstruct load paths and reproduce component failure Revised load case, detail, inspection rule or code

SRE makes the residual especially legible. Reliability engineering prepares responses for every failure the current fault model contains. An outage outside that closure creates a residual surprise. The postmortem earns its name only when it explains that residual and changes the model; restoring service and naming a broken component are not enough.

fault model
  → prepared response
  → incident
  → residual surprise
  → postmortem
  → backported fault model

The table is a crosswalk, not an equivalence claim. Medicine cannot always rerun a treatment, software traces are often richer than wreckage, and structural codes aggregate failures across jurisdictions. What transfers is the operation: prospective models absorb known failures; retrospective inquiry converts residual surprise into a model future practice can express.

The recommendation begins a second inquiry rather than closing the first. Explaining how a rudder reversal occurred does not verify that a redesign prevents it. The redesign predicts behavior first in rigs and simulators, then in certification, service, and continuing assurance. Its graph faces forward. Accident reconstruction and safety assurance meet at the intervention they share.


What low bandwidth cannot buy

Vicarious trials do not turn residue into certainty. Several histories may entail every surviving trace. Fire may erase the feature that separated them. Investigators may never have recorded the variable that mattered. A simulation may encode the same assumption as the hypothesis it appears to test. Low bandwidth makes underdetermination visible; it does not remove it.

The proper verdict is therefore graded and scoped. Flight 427's report says most likely because the physical mechanism could occur without leaving an obvious mark. "Probable cause" records an earned causal entitlement, not correspondence recovered from wreckage. The alternative histories lost enough differential trials for action to proceed, while the surviving history remained below formal proof.

Institutions determine which traces survive. Aviation buys future falsification bandwidth through flight-data recorders and maintenance logs, confidential reporting and independent investigation, preserved wreckage and living witnesses. It pays before anyone knows which claim will need the evidence. A punitive reporting system can destroy evidence by teaching witnesses not to produce it. An organization that records only terminal outcomes removes the intermediate states process tracing requires. Instrumentation is prospective epistemology.

The economics of inquiry therefore begins before the inquiry. A system spends in advance on observability because catastrophe is too expensive to query afterward. Black boxes are not passive archives. They are prepaid experiments.


Exercises

💻 marks exercises meant for a keyboard. ★ marks open-ended problems with no single right answer.

12.1 A component is found fractured after a crash. Give two causal histories consistent with the fracture, then name one differential trace that would separate fracture-before-impact from fracture-on-impact. Which reasoning mode generates the histories, which derives the trace, and which checks it?

12.2 Classify each trial by its dominant cost dimension and inquiry regime: a unit test, a bridge load test to failure, a randomized drug trial, an eclipse observation, and an aircraft accident investigation. For each, name one change in instrumentation that would raise falsification bandwidth without making the event more frequent.

12.3 A simulator reproduces an accident trajectory under hypothesis H. List three reasons this fit may fail to discriminate H from its rivals. Rewrite the simulator result as a falsifiable claim whose scope does not exceed what the run established.

12.4 💻 Represent the Flight 427 reconstruction as a directed graph with four edge types: supports, causes, blocks, and undercuts. Give one candidate history a predicted flight-data trajectory, compute its residual against an observed trajectory, and backport the mismatch into the graph. Which node, edge, parameter, or missing variable changes? Then delete the thermal-test node and list which causal edges lose their warrant.

12.5 ★ Choose a high-consequence system you know. Write its trial-cost vector and identify the event nobody may deliberately reproduce. Design three prepaid experiments: one instrument that preserves a trace, one safe surrogate trial, and one reporting rule that keeps human evidence available. State the causal claim each would expose to possible loss.


Further reading


Neighbors