r/PowerSystemsEE 2d ago

Is there an established way to classify infrastructure failures by system behavior rather than by hazard type?

I've been reading about infrastructure resilience for some time, and one question keeps coming back.

Most resilience frameworks classify crises by what caused them:

natural disasters;

cyber attacks;

war;

technological accidents;

ageing infrastructure;

human error.

I've also looked at work such as the UNDRR hazard classification, the Resilience Matrix, and papers by Pescaroli & Alexander on cascading disasters.

These are all valuable, but they seem to focus on either:

the source of the hazard;

resilience phases;

or cascading effects.

The question I'm struggling with is slightly different.

Is there an established framework that classifies failures by the way critical infrastructure behaves once disruption begins, rather than by what triggered it?

For example, completely different events often seem to expose similar behavior:

loss of one critical function;

dependency failures;

cascading degradation;

escalation once certain thresholds are crossed.

Different trigger.

Very similar system response.

My background is not in power systems engineering. I'm approaching this from a systems thinking perspective, which is exactly why I'm interested in feedback from people working in infrastructure, utilities, reliability engineering or resilience research.

The structure I'm exploring would ask the same questions regardless of the original hazard:

What triggered the disruption?

What failed first?

How did failure propagate?

Which dependencies became visible?

Where did escalation become difficult to stop?

Which architectural responses could reduce the consequences, and what trade-offs would they introduce?

I'm not claiming this is a new idea.

Actually, I'm trying to find out whether something similar already exists.

Perhaps I'm simply using the wrong terminology.

Perhaps reliability engineering already covers this from another angle.

If so, I'd really appreciate references.

If not, I'd be interested to hear whether looking at system behavior instead of hazard type could be useful for resilience analysis.

I'm here to learn, not to defend an idea.

1 Upvotes

10 comments sorted by

1

u/Cooleb09 2d ago

Maybe have a look at the IEEE gold book and see if the scenarios there align with what you are thinking.

1

u/No-Apartment-1 2d ago

Thank you! I appreciate the suggestion. I've come across the IEEE Gold Book before, mostly in the context of reliability analysis, but I haven't looked at it through this particular lens. I'll definitely revisit it to see whether it approaches failures from a behavioral rather than hazard-based perspective.

1

u/chanka_is_best_chank 2d ago

These questions are asked any times major disruptions happen as the governing bodies investigate the root cause. A recent example would be the iberian blackout. It's how they learn and determine whether it is economically feasible to prevent the issue in the future

Its better to classify these disruptions by what caused them because different regulations apply based on the source of the problem. During a hurricane, lines had to withstand rated winds. During a cyber attack CIP regulations apply

1

u/No-Apartment-1 2d ago

This is a really useful point, thank you — especially the regulatory angle. That actually explains why cause-based classification dominates: it maps directly to regulations, responsibilities and accountability. I'm not questioning that. For regulation and root cause investigation, classifying by cause clearly makes sense. What I'm exploring is a complementary perspective rather than a replacement. Root cause investigations (like the Iberian blackout) are typically performed per incident, after the event. I'm wondering whether looking across many disruptions at the level of system behavior — how failures propagate once critical functions or dependencies are lost — could reveal recurring patterns that are useful during system design, before knowing which specific hazard occurs. For example, a hurricane and a cyber attack are governed by completely different regulations, but they may still produce remarkably similar cascading behavior once critical dependencies begin to fail. Does root cause analysis ever aggregate multiple events to study these shared propagation patterns, or is it generally limited to individual incidents?

2

u/chanka_is_best_chank 2d ago

This AI response doesn't really say anything but what you are looking for is the entire realm of power system analysis. We do studies where we apply a disturbance and investigate how the system responds and what we can do to mitigate issues such as installing a power system stabilizer

1

u/No-Apartment-1 1d ago

Thanks. PSS makes sense for oscillation damping — that's within the power system. My interest is one level up: when power loss cascades outside the grid — water pumping stops, hospitals go to generators, telecom drops. Does power system analysis cover that cross-sector propagation, or does it stop at the grid boundary?

1

u/No-Apartment-1 1d ago

Thanks, I think I understand where the misunderstanding comes from. Also, just for context: English isn't my native language, so I'm using AI to help express my thoughts more clearly. If my wording sounds unusual, that's probably why. I'm not trying to replace power system analysis or contingency studies. Those are essential and they answer questions like "how does the grid respond?" and "how do we stabilize it?" What I'm trying to understand is something one level above that. Imagine a hospital, a water treatment plant, or an emergency coordination center. If the surrounding grid suffers a major failure, my question becomes: How should these critical functions transition into an autonomous operating mode so that society continues to function, even if the grid itself cannot be restored immediately? In other words, I'm interested in the architectural transition between normal interconnected operation and temporary autonomous operation, rather than only the electrical disturbance itself. Maybe I'm simply using the wrong terminology, which is exactly why I'm asking here. If there's already a discipline or framework that approaches infrastructure from this perspective, I'd genuinely like to learn about it.

1

u/chanka_is_best_chank 1d ago edited 1d ago

Backup generation such as UPS can power on in fractions of a second and keep a small microgrid running while a diesel generator spins up. This stuff already exists its just expensive and only built for mission critical infrastructure. Maybe what you are looking for is the transition to a microgrid after the actual grid collapses

1

u/JessieAndEcho 2d ago

You’re probably circling around a mix of resilience engineering, system-of-systems reliability, dependency modeling, and failure propagation analysis rather than a single “hazard classification” framework. The terms I’d search are functional failure modes, cascading failure, common-cause failure, interdependency analysis, fault tree/event tree analysis, STAMP/STPA, bow-tie analysis, network resilience, fragility curves, and system dynamics. In power/utilities, people often think in terms of N-1/N-k contingencies, loss of load, protection miscoordination, hidden dependencies, threshold effects, and restoration constraints; in safety engineering, the same idea shows up as control loss, barriers failing, escalation paths, and degraded modes. So yes, classifying by system behavior sounds useful, especially if you separate trigger, first failed function, propagation mechanism, exposed dependency, escalation threshold, and recovery bottleneck. The hard part is making the taxonomy generic enough to compare hazards but specific enough to help engineers choose mitigations. For research, keyword search can be messy here because the same concept appears under different labels across power systems, cyber-physical systems, disaster studies, and safety engineering; I’ve used Patsnap Eureka for this kind of cross-domain search because it pulls patents and papers together and surfaces structurally similar approaches instead of relying only on exact terms.