This is the fifth and final post in the diagnosis set, after Root-Cause Analysis Under Uncertain, Heterogeneous Evidence, When Is the Evidence Enough?, and Harder Cases in Diagnosis. Those treated a single vehicle. This one scales up: the same trouble code appears across thousands of cars, and the rest of the evidence — telemetry summaries, warranty claims, service notes, build records — arrives as one big batch.
That sounds like more of the same, but it is a different problem. With one car you have a single observation and you ask “what broke?” With a fleet you have a distribution, and a valid explanation has to account for three things at once: how common the failure is, who it happens to, and when it started spreading. Those extra constraints make impostors much easier to kill — but they hide a new trap, the one that catches every fleet investigation: the same code is rarely the same cause. A common DTC across a fleet is usually a mixture of several mechanisms wearing one mask, and if you diagnose the batch as a single failure you will blend two real causes into one fictitious one.
So the fleet method has a different first move — de-mix the population — and a different last move: turn the diagnosis into a remediation scope (who gets an over-the-air fix, who gets a recall, who gets nothing). In between, the cohort contrast that was a side-check on one car becomes the main engine.
What it covers
Thirteen sections, about twenty-six minutes, in plain language.
§ 1 — From a car to a population. Why a fleet hypothesis must explain prevalence, stratification, and onset — and why that makes diagnosis easier, not harder.
Part A · One code, many causes
§ 2 — The mixture trap. Why one DTC across the fleet is usually several root causes, and what averaging them costs you.
§ 3 — De-mix first. Clustering the batch into homogeneous sub-populations and running the method per cluster, with an algorithm and a block diagram.
Part B · The cohort contrast is the engine
§ 4 — Let the boundary name the cause. Stratifying affected vs. spared by build plant, firmware, lot, and climate — and reading the split.
§ 5 — Confounding & Simpson’s paradox. Why a fleet-level correlation can reverse inside every subgroup, and why stratification is now mandatory.
Part C · Aggregating batched evidence
§ 6 — Two levels, not one. Summarizing per vehicle first, then reasoning across vehicles — and not letting volume manufacture false confidence.
§ 7 — Messy by default. Heterogeneous trust, missing data, and the selection and survivorship biases that warranty batches carry.
Part D · Time and the decision
§ 8 — Read the curve. Prevalence and onset over time, and change-points that pin the cause to a production change or an OTA.
§ 9 — From cause to campaign. Forecasting how many more will fail, and choosing the remediation scope — OTA, recall, or targeted service.
§ 10 — The fleet procedure. The whole method as one flowchart, with the master algorithm.
§ 11 — A worked scenario. A 6,000-vehicle fleet, one charging code, two hidden causes — start to finish.
§ 12 — Practical notes. Field rules for population-scale diagnosis.
§ 13 — References & further reading. Epidemiology, reliability, mixtures, and causal inference, with pointers.
Read it
The monograph lives at its own URL in the warm-paper layout of the series, with a mixture diagram, a stratification tree, a Simpson’s-paradox panel, a two-level aggregation block diagram, prevalence and onset curves, a remediation-decision figure, the fleet RCA flowchart, three algorithms, and a worked 6,000-vehicle scenario.
Where it sits
The diagnosis series, in order:
I · The core method
- Root-Cause Analysis Under Uncertain, Heterogeneous Evidence
- When Is the Evidence Enough?
- Harder Cases in Diagnosis
II · At fleet scale
- Fleet-Scale Root-Cause Analysis — (this post)
- Fleet RCA, Made Rigorous
III · Proof, foundations & prevention
- Proving Cause with Experiments
- Diagnosis on the Vehicle
- The Data Plumbing Behind Fleet Diagnosis
- From Root Cause to No Cause
Full map: The Diagnosis Series index.
← Back to Autonomy