An alarm flood is formally defined as an operator receiving 10 or more alarms within a 10-minute window, a threshold that ties directly to measurable loss of situation awareness [S5]. On the Milford Haven Refinery incident timeline, operators received 275 alarms in the 11 minutes before the explosion, including the Flare High Level alarm that was missed [S5].
Rationalization, not additional dashboards, is the corrective path: every alarm reaching an operator should be tied to a documented cause, consequence, time-to-respond, and corrective action, and priority must be assigned from consequence severity rather than instrument habit [S1][S4]. Without that discipline, the master alarm database accumulates new tags from every commissioning event while existing nuisance alarms are never removed [S4].
What an alarm flood is, and the threshold that triggers it
An alarm flood is any period in which the alarm arrival rate exceeds the rate at which an operator can effectively acknowledge, interpret, and act [S5]. The widely cited operational definition is 10 or more alarms in 10 minutes, and downstream practice benchmarks the consequence: an operator who is processing more than roughly 300 alarms per day is operating well outside the manageable band defined by EEMUA 191 and ISA 18.2 [S4][S5]. Floods are not abstract metrics; they have a documented accident signature, and chattering, duplicate logic, and low-value notifications are consistently named as the noise that buries the single actionable alarm [S4].
Floods are also recognized in the academic literature as a state in which operators cannot preserve situation awareness because incoming events are paced faster than the human cognitive loop, not because individual alarms are technically wrong [S2]. For related operator-side consequences, see the engineering fixes for alarm fatigue in condition monitoring, which shares the same cognitive-overload mechanism.
Root causes: chattering, process upsets, communication loss, and database drift
Floods cluster around four recurring root causes, all documented in current industry literature. First, chattering alarms from noisy instruments or aggressive deadbands fire repeatedly on a single disturbance and consume a disproportionate share of the alarm budget [S2][S4]. Second, sudden process upsets such as a feed interruption, a controller mode change, or a compressor surge trigger dozens of dependent tags in seconds, and those cascades reveal which alarms are consequences rather than root causes [S2][S3]. Third, communication failures between the DCS, SCADA, and field networks generate retransmitted or stale alarms that present as a flood even when the process is stable [S3]. Fourth, database drift from un-rationalized additions, vendor default configurations, and standing alarms that have not been aged out compounds the first three, so the baseline rate is already elevated before any upset occurs [S1][S4].
The dynamic risk analysis literature treats floods as a control-room safety problem with a specific failure mode: the operator's attention is captured by the highest-rate tags while the lower-frequency but higher-consequence alarm scrolls off the top of the summary display [S2][S7]. A related trade-off, separating alarm load from control bandwidth, is discussed in the full port vs reduced port ball valve selection logic, where the dominant signal also gets buried if the supporting set is over-specified.
The ISA 18.2 rationalization workflow, step by step

ISA 18.2 (Management of Alarm Systems for the Process Industries) prescribes a documented lifecycle from philosophy through identification, rationalization, detailed design, implementation, operation, and ongoing audit, and it requires a master alarm database as the single source of truth [S4]. The practical workflow that maps to that lifecycle, as published in May 2026, has four steps: start with an alarm audit to baseline the current daily alarm rate and identify the worst offenders; define alarm conditions and required operator responses in a workshop with operators, process engineers, and control engineers; use that data to drive prioritization, suppression, and deadband changes; and treat the result as a living program rather than a one-time project [S1][S4].
Within the rationalization step itself, every alarm must be reviewed against four fields: cause, consequence, time-to-respond, and corrective action, and the priority is set by the documented consequence, not by the instrument type or department habit [S4]. Chattering and nuisance alarms are detected against the EEMUA 191 target of roughly 300 alarms per operator per day, and any tag that consistently exceeds its allowable rate is suppressed or re-engineered before it can seed the next flood [S4]. For a related example of how layered data is reduced to a single decision-ready signal, the R-value per inch comparison uses a similar filtering principle on a different physical variable.
Comparison of the main rationalization options against decision criteria
Plants can choose between four rationalization options, each with different cost, speed, and coverage characteristics. The comparison below is drawn from current industry guidance and from the published ISA 18.2 and EEMUA 191 framework [S1][S4][S5].
1) Manual paper-based rationalization, the legacy approach: low software cost, slow (typically 12-18 months for a mid-size plant), high labor cost, and limited ability to keep the database synchronized with live DCS or SCADA changes [S4].
2) Spreadsheet-driven rationalization: low-to-moderate cost, faster than paper, but prone to version-control errors and poor audit trail, which violates the documentation requirement in ISA 18.2 [S4].
3) Dedicated alarm-management software integrated with the control system: moderate-to-high license cost, the fastest path to a synchronized master alarm database, automatic chattering detection, real-time flood detection, and reporting against EEMUA 191 and ISA 18.2 benchmarks [S4].
4) Dynamic suppression and shelving layered on top of any of the above: the lowest-cost operational add-on, but a supplement, not a substitute, for rationalization, and exida notes that suppression alone "may not eliminate [floods] entirely" [S5].
For most process plants the choice reduces to option 3 as the long-term backbone, with option 4 applied in parallel for residual floods during the transition [S4][S5]. The trade-off logic, where one parameter is optimized at the expense of others, also appears in the low vs medium vs high carbon steel selection rules.
HMI design during a flood: what helps the operator

Rationalization prevents floods at the source, but it does not eliminate them, and the remaining floods must be handled in the HMI layer. Three practices have been validated by the Abnormal Situation Management (ASM) Consortium and the Center for Operator Performance (COP) for maintaining situation awareness during a flood [S5]. First, a system overview display that aggregates alarm status by sub-system lets the operator identify which equipment area needs attention and pull a filtered summary for that area only, rather than scrolling a global list [S5]. Second, configuring the alarm summary so the highest-priority alarms stay pinned at the top of the screen; COP research shows this is the single highest-impact change for flood performance because it sacrifices response to low-priority alarms, which is the correct trade-off when prioritization has been done properly [S5]. Third, an alarm timeline display that places events on a time axis so cause-and-effect relationships are visible, and that overlays operator actions, gives the operator a forensic view of the flood [S5].
None of these HMI techniques is a substitute for rationalization, and exida's published position is that the HMI layer is the last line of defense, not the first [S5]. Plants that skip the lifecycle step and rely on dashboard engineering are typically fighting the same flood every quarter.
Who rationalization is for, and where it is not the right tool
Rationalization is for any plant that already has more than the EEMUA 191 manageable alarm rate, which is widely used at roughly 300 alarms per operator per day as a target, or that has had any documented flood event in the past 12 months [S4]. It is also for plants that have added a significant number of tags during a recent capacity expansion, because every new instrument ships with vendor default alarm configuration that usually over-alarms [S1][S4]. It is not the right tool for a brand-new plant still in commissioning, where the right deliverable is the alarm philosophy document and the master alarm database built during the design phase, not a retrofit project [S3][S4].
It is also not the right tool for solving underlying instrumentation problems, and the literature is explicit that a chattering level transmitter or a sticky control valve will keep creating floods regardless of how much rationalization work is done, until the physical asset is fixed [S2][S4]. The pattern, where the right fix is at the right layer, is the same one used in medium die casting machine footprint and crane hook height selection: tooling must match the physical constraint.
Standards, metrics, and what a rationalized system looks like

Two documents define the metric framework: ISA 18.2 for the lifecycle and the documentation requirements, and EEMUA Publication 191 for the alarm-rate targets, the 10-in-10 flood threshold, and the priority distribution [S4][S5]. A rationalized system has a daily alarm rate below the EEMUA 191 target, a recorded cause-consequence-response for every alarm, a defined ownership for every standing alarm, and a logged root cause for every flood event that feeds back into the next rationalization cycle [S4].
For comparison, the supporting hardware layer that produces the alarms in the first place has its own standards: ATEX 2014/34/EU and IEC 60079-x for hazardous-area certification of perimeter alarm and fire alarm control panel equipment, and ISA 18.2 itself for the alarm management lifecycle that sits on top of those instruments. The signal infrastructure that carries the alarms, the marshalled cabinets, the gas alarm controller racks, the lamps and light fittings that populate the annunciator panels, and the lighting equipment and electric lamps that back-light the operator console, is governed by separate equipment standards but feeds the same operator cognitive budget.
Two trackable signals to watch in the next reporting cycle: published plant case studies that quote a sustained alarm rate below the EEMUA 191 target after an ISA 18.2 rationalization, and the release of any updated EEMUA 191 edition that revises the 10-in-10 flood threshold or the daily alarm rate target. Both are concrete, sourceable indicators of where the industry is converging on flood management.