Mean Time Between Failures (MTBF)

Mean Time Between Failures (MTBF) is the average time a system operates between one incident and the next. MTBF measures how reliable something is; MTTR measures how quickly you recover. Together they define uptime: uptime = MTBF / (MTBF + MTTR).

Definition

MTBF is calculated over a fixed window (usually 30-90 days) by dividing total uptime by the number of incidents. A service with 720 hours of uptime and 4 incidents has an MTBF of 180 hours.

MTBF is the reliability lever engineering owns most directly: better testing, staged rollouts, redundant infrastructure, capacity planning. MTTR is the response lever ops owns: paging, runbooks, automated remediation.

Why it matters

Uptime targets are hit by pushing MTBF up or MTTR down. If your MTTR is already tight, pushing MTBF is the only lever left — that means fewer risky deploys, better tests, and change-freeze windows around peak traffic.