What Security Can Borrow From Reliability Practice

A chart of budget spent against time, with a straight dashed line for the sustainable rate and a steeper curve meeting the exhaustion line early

Security operations has adopted the vocabulary of reliability engineering much faster than the mechanics. Teams have objectives, budgets and blameless reviews in their documents, and underneath them the same arrangement as before: targets with no consequence attached, reviews that are polite rather than blameless, and a queue that absorbs whatever arrives until it quietly stops absorbing.

Three of the ideas do transfer. But each has a mechanism doing the work, and importing the word without the mechanism buys nothing except a shared vocabulary with the platform team.

An error budget is a pre-agreed consequence

The mechanism in an error budget is not the arithmetic. Turning a target into an allowance of downtime is elementary — the calculator on this site does it in one multiplication. What makes a budget function is that the organisation agreed in advance what changes when the allowance is exhausted, and that the change is automatic rather than negotiated. Without that clause it is a number on a dashboard, and those do not alter anybody’s behaviour.

The harder question is which quantity is being budgeted, and this is where most attempts fail immediately. It is obviously unacceptable to budget for breaches per quarter; whoever proposes it is laughed out of the room, correctly, and the whole idea goes with it. The mistake is the choice of quantity. Two of them budget honestly.

The reliability of the operations service itself. A security operations function is a service with a promise attached: someone is watching during these hours, triage begins within this period, this pipeline is processing, this enforcement point is enforcing. Each of those is an availability claim, and each is broken sometimes — by a saturated queue, an unstaffed shift, a collector that stopped, a certificate that expired on a control path. Those failures are within your control and measurable, which is exactly what an error budget needs.

Self-inflicted disruption. Security controls cause outages. Blocking rules catch legitimate traffic, endpoint software destabilises a fleet, credentials are rotated in a way that breaks a downstream consumer. The rest of the organisation experiences this friction as the cost of the security function, and it is normally argued about in anecdotes. Budgeting it converts a permanent low-grade dispute about “security versus velocity” into arithmetic: here is the disruption we consider acceptable to impose, here is what we have imposed, here is what remains.

The second is a constraint the security team places on itself, in public — the reason it earns credibility with the teams on the receiving end, and the reason it is rarely proposed.

The quantity you cannot budget

What does not transfer is anything driven by an adversary. A budget is spendable: you decide how much to consume, and the decision buys speed or cost. Adversary-driven events are not spent, they arrive. The rate is set by someone else, so an allowance of them is not a budget, it is a forecast with a moral hazard attached.

Exposure targets — how much of the estate is in a given state, how quickly a class of weakness is closed — are legitimate and entirely different: they are commitments about your own configuration, measured against your own actions. They should not be called error budgets. Two mechanisms sharing one name is how a vocabulary loses its meaning.

Blameless review is a technique, not a temperature

“Blameless” is widely understood to mean a review conducted kindly. That is not the mechanism. The mechanism is an analytical constraint: you may not explain the outcome by what a person should have done differently, because that explanation terminates the investigation at the exact point where the interesting information starts.

“The analyst should have escalated” is not a finding. The findings are underneath it: what the analyst could see at the time, what the interface presented as normal, what the workload was that hour, what the procedure said, what escalating incorrectly had cost the last person who did it. Every one of those is changeable. The instruction to be more careful is not.

Security work adds two complications that reliability practice does not have to handle.

The first is the presence of an adversary, which offers an ever-available external explanation. “It was a sophisticated attack” is the security equivalent of “human error”: it feels like an answer, it is often even true, and it removes every internal lesson from the review. The same goes for tracing a compromise to the person who clicked something. That describes how access was first obtained, not why one click was allowed to matter that much — which is the question the organisation actually needs answered.

The second is that security incidents attract legal, regulatory and contractual processes, and everybody in the room knows the write-up may later be read by people looking for fault. That pressure produces careful, defensive documents in which nothing is learned, and exhortations to candour do not fix it, because the fear is rational. The structural fix is to separate the artefacts: an internal learning review whose audience and purpose are explicitly improvement, kept distinct from the record produced for accountability and disclosure. Pretending one document can serve both is why so many organisations run reviews that teach them nothing.

Load shedding: the queue is already doing it

A saturated operations function sheds load. It has no choice — arrivals exceed capacity, and something is not done. The only question is whether the shedding is chosen or emergent.

Emergent shedding is the default, and it has a characteristic shape. The newest and loudest items get attention; the oldest age quietly at the bottom of the queue until a clean-up closes them; and whatever requires sustained concentration — the improvement work that would reduce next month’s arrivals — is displaced first, because it is the only thing with no external party asking about it. The function degrades in a pattern nobody designed, and from outside it still looks like coping, because everything nominally has an owner.

Explicit shedding means declaring the classes of work, ranking them, stating which are dropped first when capacity is short, and saying so out loud when it happens. It is uncomfortable and far better, because it converts a silent failure into a visible one. An organisation told “these categories are not being worked this quarter” can fund, defer or accept that; one that is not told assumes coverage it does not have, and finds out at the worst moment.

Capacity: the slack is the product

Underneath sits the least intuitive part, and the one that contradicts how operations functions are normally planned. Where work arrives unpredictably, waiting time does not rise gently with utilisation. It rises slowly while there is slack, then accelerates as utilisation approaches the limit, and near full utilisation it grows without bound — the last increment of efficiency costs an unbounded amount of latency. So a function planned to be fully utilised is planned to have no capacity for the exact event it exists to handle, because incidents are by definition arrivals that were not scheduled. The headroom that absorbs them is indistinguishable from waste on a spreadsheet, which is why it goes first in a cost exercise and why so many functions run permanently on the edge of coping. That headroom is not overhead attached to the service. It is the service; almost everything else could in principle be scheduled.

What does not come across

Reliability practice grew up around signals that are continuous and unambiguous: a request either returned an error or it did not. Security operations works with events that are discrete, ambiguous, adversarially influenced and often only classifiable in retrospect. The dashboards do not port, and building them anyway produces precise-looking numbers about quantities nobody can define.

What does port is the discipline underneath: state the target, agree the consequence before you need it, investigate systems rather than people, decide in advance what gets dropped, and defend the slack. None of that needs a continuous signal. All of it needs the commitment written down while things are calm, which is the genuinely hard part and the reason it so rarely happens.