Paradox & Proof Entry #0168 Classified Declassified

The reason a backup can lower reliability

A second unit adds a component, a switch, a test and an assumption. Each is a new way to fail, and one of them fails silently.

No visual record attached The written record below is complete.
Plate 430 — The valve that had never been turned

Intuition test — answer before you read on

Why can adding a backup reduce overall system reliability?

The redundant pump is installed, the changeover valve is added, and the system now has two pumps and one valve that has never been operated. Eleven months later the primary fails, the valve does not turn, and the outage is longer than it would have been with one pump.

What everyone sees

Redundancy is treated as arithmetic: two units, each failing independently, give a combined failure probability equal to the product. That calculation is correct under its assumptions. The assumptions are independence, correct switching and known state of the standby, and installing a backup weakens all three.

What is actually happening

Perrow’s analysis of complex systems argues that added components increase interactive complexity and tight coupling, generating failure modes that are not present in any component and are not anticipated in the design. Sagan examined redundancy in high-hazard organisations and identified three specific mechanisms by which it can reduce reliability: added components create new common-mode failure paths, redundancy encourages higher risk-taking because the margin is believed to exist, and responsibility diffuses when more than one element is nominally responsible for the same function. The standby also has a distinctive property that a primary does not: its condition is unobserved during normal operation, so its failure produces no signal until the moment it is required.

Why it stays hidden

The added risk hides because the failure is counterfactual. A system that survives is credited to its redundancy, while the outage caused by an untested changeover is recorded as a valve fault rather than as a consequence of the design that introduced the valve. And the untested component looks like protection for its entire silent life.

Redundancy adds components, switching and an unobserved state. The standby’s failure produces no signal until the moment it is needed.

Redundancy adds components, switching and an unobserved state. The standby’s failure produces no signal until the moment it is needed.

The hidden part — entry #0168

Collect this card

Redundancy adds components, switching and an unobserved state. The standby’s failure produces no signal until the moment it is needed.

0 / 10,000 collected

Sources & further reading 2
  1. Perrow — normal accidents: living with high-risk technologies
  2. Sagan — the limits of safety: organisations, accidents and nuclear weapons

Circulate this file

Annotations are reserved for archive members.

Sign in to annotate