The Alarm Budget
On the morning of July 24, 1994, an electrical storm moved across the Texaco refinery at Milford Haven and knocked several process units into upset. Over the next several hours the disturbance worked through the fluidised catalytic cracking unit, complicated by a control valve that was shut while the control system indicated it was open. Liquid kept being pumped into a vessel whose outlet was closed. Around twenty tonnes of flammable hydrocarbon eventually escaped, found an ignition source about 110 meters away, and exploded. Twenty-six people were injured, the damage came to roughly £48 million, and the fires burned until Tuesday evening.
The investigation found what such investigations usually find: inadequate maintenance, modifications made without assessing consequences, training that had not prepared anyone for a sustained upset. But the finding that outlived the report is a single sentence in HSE’s guidance on the incident: “In the last 11 minutes before the explosion the two operators had to recognise, acknowledge and act on 275 alarms.”
One alarm every two and a half seconds, for eleven minutes, while the plant tried to say the one thing that mattered: a vessel was full and its outlet was closed.
Receiver capacity is an engineering constraint with a number, not a personal productivity failure.
The last post was about the anatomy of one honest signal — the four properties that let a system tell the truth at the moment it runs out of the ability to do its own job. It ended with a question I said I couldn’t answer: when escalations arrive at volume, how does the receiver know which one is load-bearing? Nearly every alarm in that torrent was, individually, true. The population was not. A flood of true statements can be as opaque as silence.
I should be honest about why I went looking for the answer in a refinery. It was not historical curiosity. Most days I have somewhere between four and ten agents working in parallel — analyses, migration dry runs, review passes, the branch-heavy workload the last few posts have been circling. Each one narrates. Each one escalates. A tool I run keeps an audit trail of their decisions and flags the ones that conflict, and as I write this it shows 274 open conflicts. A conflict, here, is not a compiler error. It is a disagreement between agent decisions that needs a human to decide which claim survives before the work can merge, promote, or become precedent. I noticed the number because it is one short of Milford Haven. It shows me the top five by severity, and most mornings that is all I read.
The week I caught myself approving an agent’s diff without reading it — not skimming it, not reading it — I recognized the gesture. It is the gesture of a clinician clicking through a drug-interaction warning. Every industry that has built a channel into human attention has a name for what happens when the channel exceeds it; ours is arriving under headings like attention management in the age of AI, as if the problem were personal productivity. The diagnosis is actually fifty-five years old. Herbert Simon wrote it down in 1971: “What information consumes is rather obvious: it consumes the attention of its recipients. Hence a wealth of information creates a poverty of attention.”
What I did not appreciate, until the 274 sent me digging, is that one industry took the diagnosis and turned it into a budget.
The number
In 1999, five years after the explosion, the Engineering Equipment and Materials Users Association published EEMUA 191, Alarm systems: a guide to design, management and procurement, with HSE’s contribution and endorsement. It is the founding document of the discipline now called alarm management, and the most striking thing in it is a pair of numbers, reproduced in HSE’s guidance: the long-term average alarm rate during normal operation should be no more than one alarm every ten minutes, and no more than ten alarms should be displayed in the first ten minutes following a major plant upset.
Stop and look at what kind of number that is. It is not a property of the plant. Nothing about the size of the refinery, the count of its sensors, or the complexity of its chemistry enters into it. It is a property of the receiver. A human who must recognise a signal, diagnose it, and act on it has a sustainable throughput, and the guide’s entire posture is that everything upstream — every sensor, every threshold, every priority assignment — must be engineered until the population of signals fits inside that throughput. The plant proposes; the budget disposes.
The ANSI/ISA-18.2 standard, which followed in 2009 and was later adopted internationally as IEC 62682, made the receiver-centricity explicit. An alarm is an indication of a condition requiring a timely response — not a condition worth knowing about, a condition requiring a response. An alarm flood is defined relative to the operator, not the plant: ten or more alarms in ten minutes, per operator, the rate beyond which a human stops managing what arrives. Milford Haven’s last eleven minutes ran at about twenty-five times the flood line overall, or still more than twelve times even if split evenly between its two operators. Even priority is budgeted — roughly 5% high, 15% medium, 80% low — because priority is a claim on the receiver’s preemption, and preemption only works if it is rare. A plant where a third of the alarms are critical is a plant where nothing is.
Underwriting the claim
How do you make a plant with thousands of instruments fit inside one alarm every ten minutes? Not at runtime. The discipline’s answer is a design-time pass called rationalization: every candidate alarm must justify its existence before it is allowed to exist, with a documented operator action, consequence of no response, and time available to respond, the priority falling out of consequence and time mechanically. The HSE guidance is blunt about the entry criterion: “process status indicators should not be designated as alarms,” and the first quick win it lists for an overloaded system is to eliminate or review alarms with no defined operator response. If there is nothing for the receiver to do, it is not an alarm. It is information wearing an alarm’s costume, and every time it sounds, it spends credibility that belongs to the signals that can pay for it.
This is expensive to honor. The HSE sheet estimates a thorough review takes more than one shift per alarm, and not every plant paid — a target that expensive is easier to publish than to hit. But that is what the budget changed. Once the number existed, over budget became an auditable state instead of an operator’s complaint.
There is a control group for this experiment, and it is medicine. Clinical order-entry systems grew their alert populations the way our industry is growing agent escalations now — each alert individually defensible, none underwritten against a receiver budget. A 2006 systematic review in JAMIA found clinicians overriding drug-safety alerts in 49% to 96% of cases. The literature named the mechanism alert fatigue, but the name undersells it. A channel over its budget does not degrade linearly; it trains its receivers — the dynamic the cost of doubt traced through code review, running at clinical scale for two decades. The channel’s credibility is shared infrastructure. The honest signals pay for the noise. I think about the top of that range every time I notice my thumb already moving toward an approve button.
Software has written one of these before
I want to resist the impression that the receiver budget is exotic process-industry lore, because software wrote one, and most of us have worked under it without noticing what it was. Google’s SRE book derives its on-call policy from a measurement: dealing with one incident properly — root-cause analysis, remediation, postmortem, follow-up — takes about six hours. “It follows that the maximum number of incidents per day is 2 per 12-hour on-call shift.” Around that number sits a larger allocation: at least 50% of an SRE’s time goes to engineering, and no more than 25% to being on-call.
Look at the shape of the derivation. The budget is not set by how many pages a human can acknowledge — acknowledging is nearly free, which is exactly the trap. It is set by the cost of metabolizing one signal properly. And when a service exceeds the budget, the prescribed response is back-pressure: the pager load is treated as evidence that engineering work is owed, not that the responder should read faster. That is rationalization in a hoodie. We know how to do this. We did it for the escalation channel called paging, where the producers were services. We have not done it for the channel where the producers are agents, and the production rate is whatever we feel like spawning.
When the budget fails anyway
A budget is a steady-state instrument; break the system badly enough and the flood comes regardless. What that looks like from the receiver’s chair is Qantas flight 32. In November 2010, an A380 climbing out of Singapore suffered an uncontained engine failure, and ECAM — about as rationalized and consequence-ranked as alerting gets — did exactly what it was designed to do, all at once. One of the pilots later described it: “43 ECAM messages in the first 60 seconds after the explosion and probably another ten after that. So it was nearly a two-hour process to go through those items.”
What carried QF32 was not the alerting system. The crew worked the queue item by item, and when ECAM presented a fuel-transfer procedure that would have moved fuel through galleries running across a wing full of shrapnel holes, they evaluated it against their model of the damage and refused: “We elected not to.” The same pilot’s summary of the two hours: “We didn’t blindly follow the ECAMs. We looked at each one individually, analysed it, and either rejected it or actioned it as we thought we should.” Two things made that possible. A model: the crew could simulate the aircraft well enough to overrule its advice. And capacity: the flight deck that day held five pilots instead of the usual three, because a routine check ride had stacked two check captains behind the operating crew. When the budget fails, triage falls to a receiver who has a model of the system and the standing to reject the system’s advice — and nobody acquires that model by watching a queue. That clause is the one that worries me about my own fleet, because watching the queue is increasingly the whole job.
Agents, rediscovering the number
The agent-oversight literature is currently rediscovering all of this from first principles. The AI-control line of work, starting in late 2023, builds its protocols around an explicitly budgeted scarce resource: human auditing. A weaker trusted model scores the suspiciousness of everything an untrusted model produces, and humans audit the top 2% most suspicious — 2% being the budget the protocol is calibrated against. The instructive part is the defer-to-trusted variant: for the suspicious band the audit budget cannot cover, the protocol does not page a human. It throws away the untrusted model’s work and substitutes the trusted model’s instead. The budget is honored by the system changing what it does, not by the receiver reading faster.
The adjacent learning-to-defer literature has been converging on the same constraint from a different direction. DeCCaF frames deferral as cost-sensitive allocation under workload constraints: the question is not simply is the human better than the model, but which expert should receive which case when each expert has finite capacity and different errors cost different amounts. FALCON makes the fatigue assumption explicit, modeling human performance as workload-varying rather than constant. The interesting part is not that either paper solves agent oversight. It is that both refuse the fantasy that a human reviewer is an infinite oracle. Once escalation is treated as spending a bounded resource, the design problem changes: the system has to decide not only whether a case is risky, but whether spending the next unit of reviewer capacity on this case improves safety more than spending it elsewhere. That is the 1999 guide in ML clothing.
What everyday agent tooling has not reproduced is the discipline that followed the numbers. Today an agent system’s escalation behavior is usually a threshold in a config file, owned by whoever last edited it, and nothing in the common stacks I use requires that an escalation have a defined receiver action before it is allowed to exist. By the standard’s entry criterion, most of what agents surface to humans is not escalation at all — it is status narration, the same chat-window sentence the last post read as four small lies. I encountered an error. Let me try a different approach. That is a process status indicator designated as an alarm. It spends the channel’s credibility on telemetry, at machine rate, against receivers whose override statistics — mine included — already look like medicine’s.
The budget and the brief
The four properties make one escalation honest: apply what you can, name the residue, land where context lives, and write the resolution back as fact. What Milford Haven adds is that the channel has a property none of the signals carry individually: a rate, measured against a specific receiver, that the population must fit inside. An interface with honest signals and no budget tells the truth at a volume where truth cannot be metabolized.
The part of rationalization that transfers is the underwriting requirement, inverted. A plant can enumerate its alarm space at design time; an agent’s escalation space is open — the thing it cannot decide this week may have no name yet. But the receiver’s action space is small and enumerable: approve or reject this diff; decide whether Alice’s stale count still supports the budget she wrote against it; choose between these two values; authorize this spend; take the system out of service. Rationalizing escalations for agents probably means cataloguing receiver actions rather than signal types, and admitting an escalation onto the queue only when it can name which action it is requesting. Anything that cannot is telemetry, and telemetry goes to the log. The budget then has to do what the pager budget and the 2% audit budget do: push back into the producer. A fleet that approaches the receiver’s capacity and keeps manufacturing escalations has chosen the Milford Haven failure mode with nicer sorting. The alternatives are all back-pressure — pause the speculative branches, narrow the autonomy gates, defer to trusted paths. Like each primitive that becomes infrastructure, the budget externalizes a piece of the receiver’s hidden labor: the 2 a.m. judgment about which page to answer first.
One sketch is enough to make the difference concrete. If my morning budget is five reads, an admitted escalation has to name one of five receiver actions: approve a diff, reject a diff, resolve a conflict, authorize an external effect, or stop the run. The sixth admitted escalation does not ask me to become more diligent. It pauses new speculative branches until the queue clears, or routes low-consequence work through a trusted path. The number is not sacred. The behavior at the number is.
What I am still figuring out
Whether rationalization survives an open signal space. More than one shift per alarm was payable for a refinery with a bounded instrument list; my fleet invents new escalation shapes faster than I could rationalize them. Cataloguing receiver actions instead of signals is my best answer, but I am not sure the action class carries enough semantics to set priority — in ISA-18.2, priority comes from consequence and time-to-respond, and those live in the signal. Choose between two values can be a formatting question or a campaign budget. If priority cannot be derived from the action class alone, something has to play the role the hazard study plays for a plant, and I do not know what the agent equivalent is.
Whether the inverted U recurses. The obvious response to a saturated human is an agent receiver, and the last post asked what the four properties owe a receiver that is also an agent. The budget version: agent attention is elastic but not free, and an agent receiver under flood degrades too — context pressure instead of fatigue, but a capacity curve all the same. My suspicion is that every tier of receivers has its own number, and a hierarchy that looks like it absorbs the flood is moving the saturated queue up one level, pre-aggregated and harder to audit. Rate headroom bought with provenance. That trade should be priced, not assumed.
Who signs the number. After Milford Haven, the budget had owners — the company in HSE’s case study stood up a steering committee with a senior management champion, and ISA-18.2 later put the alarm inventory under change control. The pager budget has an owner too; that is what an SRE team is. My escalation budget is a constant in a config file. I own it in the sense that nobody else does, which is not the same thing. A budget without a name on it is a number, not a contract, and I notice that every industry that solved this converged on the same answer: a standing body whose job is to say no to new claims on the receiver, with the authority to make the no stick.
The two operators at Milford Haven were not incompetent. They were handed 275 claims on their attention in eleven minutes, in a control room whose displays could not show them the plant, and they did what flooded receivers always do: they fell back on the picture of the system they already had, which said the unit could be kept running. The guidance that came out of their worst morning is unusual among engineering documents because of where it points. It does not say make the plant safer. It says the receiver has a capacity, writes the capacity down as a number, and makes everything upstream answerable to it.
Some version of EEMUA 191 will eventually exist for agent fleets — a published receiver budget, an admission rule for what may claim attention, a flood definition with back-pressure semantics, an owner whose no sticks. The refinery got its numbers the expensive way. I would rather take them secondhand. The tool on my desk says 274. Most mornings I read five.