From Cybersecurity to Cyber-Resilience

How should tens of millions of distributed energy resources be connected in a world that cannot be fully defended? — A quantitative comparison of centralized control and distributed autonomy with the Cyber Blast Radius model

Suomi Masuda / I-S3 Co., Ltd. | 7 October 2026

日本語版

Connect 10 million × 5 kW = 50 GW of DER through one control plane and a single cloud compromise moves 35 GW at once; certification that cuts the compromise probability tenfold does not change the outcome. Connect them through independent domains with a device-side output floor and authority caps and the same compromise moves 117 MW. But once the number of attack surfaces and the shared vendor firmware-line channel are counted, software-only distributed autonomy is no safer than centralization. What lowers total risk is an output floor held by a monitor circuit independent of the main firmware, and a cap on the capacity one firmware line can reach remotely (about 1,350 MW in the reference grid); a server-side aggregate cap brings the centralized design to the same order. The question is not how many units are certified, but how many GW one attacker can move at once.

Version 1.0 (2026-10-07) — I-S3 Special Research Project — Author: Suomi Masuda

Summary

1. The question: what do we design when we cannot defend?

Solar inverters, home batteries, EV chargers, heat pumps. Each is a few kilowatts, but in Japan alone residential solar exceeds 3.4 million systems, cumulative PV capacity passed 100 GW, and home batteries reached about 1 million units (data: der_fleet_japan.csv). Add EVs and heat pumps in the 2030s and tens of millions of devices will be remotely adjustable.

Almost all of them connect to a cloud — for monitoring, curtailment, balancing-market participation, firmware updates. Every reason is legitimate. And in 2024–2025, defects on the cloud side were published one after another.

None caused harm. That is not evidence of safety; it means the defects were fixed before publication and that no grid attack through a DER cloud has yet been observed (Figure 12). Attacks on grids themselves already exist: Ukraine's distribution and transmission control systems were attacked in 2015 and 2016 [I05], substation-targeting malware reappeared in 2022 [I08], and pre-positioning in US critical infrastructure including power has been disclosed [I09]. DER clouds are untouched as a matter of sequence, not capability.

Figure 12 Incidents: potential versus realised Figure 12. Published incidents: capacity that could potentially have been moved versus capacity actually moved. Cyber incidents (●) show GW-scale potential with near-zero realised harm — evidence that vulnerabilities were fixed before publication and that no grid attack through a DER cloud has yet been observed, not evidence of safety; attacks on grids themselves already exist. Non-cyber reference events (■) show the consequence when a disturbance of that size actually occurs.
Figure 12 Incidents: potential versus realised Figure 12. Published incidents: capacity that could potentially have been moved versus capacity actually moved. Cyber incidents (●) show GW-scale potential with near-zero realised harm — evidence that vulnerabilities were fixed before publication and that no grid attack through a DER cloud has yet been observed, not evidence of safety; attacks on grids themselves already exist. Non-cyber reference events (■) show the consequence when a disturbance of that size actually occurs.

The standard answer is "defend harder": stronger authentication, encryption, device certification, cloud monitoring. All necessary. But this article asks a different question.

When the defence is breached, how many MW move at once — and who decided that number?

To answer it we define Cyber Blast Radius and compare two ways of connecting DER against the same attack on the same grid model. The conclusion is not fixed in advance; the conditions under which distributed autonomy loses are written with equal weight. Before publication we put two independent review teams to work and asked, from nine perspectives, "where would the best researcher who disagrees with this attack it?" Many of their objections were right, and the numbers changed substantially. Their objections and our responses are published in full (review/red_team_report.md, review/final_response_matrix.md).

2. Security and resilience are different quantities

Security lowers the probability of compromise. Authentication, cryptography, vulnerability management, penetration testing, certification all act here.

Resilience shrinks the consequence of a compromise: how far it spreads, how long recovery takes, what keeps working meanwhile.

They multiply into systemic risk:

\[ \text{Systemic Risk} = P(\text{compromise}) \times \text{Consequence}(\text{Blast Radius}, \text{Grid Sensitivity}, \text{Synchronization}, \text{Recovery Time}) \]

Multiplication matters. Cutting P to one-tenth while Consequence grows a hundredfold raises risk tenfold. And that is what is happening: alongside efforts to lower P, the capacity behind one credential is growing by gigawatts.

3. Why "certified = secure" does not hold

Device certification (Japan's JC-STAR, the EU RED cyber requirements, IEC 62443 product certification) confirms design-time compliance. It lowers P. Three things remain unchanged by the label.

  1. Change in operation. Leaked keys, missed revocations, unpatched firmware, departed employees' credentials.
  2. Outside the device. In Solarman/Deye and SUN:DOWN most defects were in cloud APIs and account management [I01][I02]. IEEE 1547-2018 notes the importance of cybersecurity in §10.9 without requirements; 1547.3-2023 is a guide [S01][S02].
  3. Size of authority. If 10 million certified devices obey the same cloud, an attacker holding that cloud's legitimate authority commands 10 million devices. Certification tells a legitimate command from a forged one; it does not ask whether a legitimate command is large enough to break the grid.

We do not reject certification. The model includes its reduction of compromise probability (p_reduction_cert, default 3×, sensitivity 1–10×) and applies the same factor to both the centralized and the distributed design. What we reject is the inference "certified, therefore the grid problem is solved".

4. From one device to ten million: scale changes the problem

\[ 5\ \text{kW} \times 1{,}000{,}000\ \text{units} = 5\ \text{GW} \]

5 GW is four to five large nuclear units, more than all of Hokkaido's demand on 6 September 2018 (3,090 MW). Behind one credential, that credential weighs as much as the emergency-stop buttons of five nuclear plants.

DER also act simultaneously. Power stations fail one at a time; a cloud command reaches 10 million devices in seconds. What a grid can absorb depends on the speed of a change as well as its size. In a low-inertia grid a few GW of step change drives frequency into the load-shedding band within seconds (Section 13).

So the problem moves from "how do we protect each device" to "what structure do we permit in which one compromise moves N MW at once".

5. Defining Cyber Blast Radius

Definition. Cyber Blast Radius (CBR) is the generation or consumption capacity (MW) an attacker can move simultaneously within the critical time window (from inertial response to full primary reserve, about 30 seconds) after one successful compromise.

UnitDefinitionReference case (10 million × 5 kW = 50 GW)
CBR_deviceOne device0.005 MW
CBR_domainOne aggregator / control-domain credential500 MW (100 equal domains)
CBR_vendorOne firmware line (same signing key, OTA path, cloud)15,000 MW (largest share 30%)
CBR_cloudNational control-plane credential / API50,000 MW
CBR_regionReachable capacity in one grid areaInstalled capacity of the area

This is the "raw", installed-capacity CBR. The capacity that actually moves is reduced by the share of devices that receive remote commands (reach), the per-domain authority cap (A_max), the lowest operating point a remote command can reach (the floor, e), and the probability that the floor actually works (ρ).

\[ \text{CBR}_{\text{eff}} = \text{reach} \times \min(\text{CBR}_{\text{raw}},\ A_{\max}) \times \bigl(1 - \rho\,(1 - e)\bigr) \]

With no device-side constraint (ρ = 0, e = 1), CBR_eff = reach × CBR_raw: 35,000 MW for a national-plane compromise at reach 0.7 (Figure 2).

e is defined as a floor — the operating point that no number of repeated remote commands can push below — not as a per-command limit, because a per-command limit disappears under repetition.

This quantity is not new. NERC CIP-002 has since 2013 rated protection by "MW per cyber asset": generation of 1,500 MW or more controlled by one cyber system is Medium Impact; control centres over 3,000 MW are High Impact [I10]. The EU Network Code on Cybersecurity (NCCS) introduced ECII, an impact index in MW, with the "high impact" threshold at each country's minimum secondary reserve and "critical" at the primary reserve (3,000 MW) [S05]. CBR brings both down from large transmission assets to distribution-level DER and devices.

6. The Systemic Cyber Risk model

6.1 Severity of one compromise

\[ x = \max\!\left(\frac{s \cdot \text{CBR}_{\text{eff}}}{\text{FCR}},\ \frac{\text{CBR}_{\text{eff}}}{\text{FCR} + \text{FRR}}\right) \]

S(x) is the probability of under-frequency load shedding or worse, not of collapse. Load shedding is the grid working as designed to avoid collapse, but to customers it is an outage of millions of households; we take it as the lower bound of "wide-area event". x = 1 is the design absorption limit; in the reference grid load shedding starts around x ≈ 2.2–3.0 (Section 7).

6.2 Annual risk for the whole fleet, all channels

The draft compared the risk of one compromise, p × S(x), between A and B. Reviewers called this "asymmetric compromise units": A's unit was the national plane, B's one of 100 domains. B has 100 attack surfaces, and both A and B share a separate channel — the vendor firmware line. We accepted this and redefined risk as the probability that at least one wide-area event occurs per year across the fleet.

\[ R_{\text{plane}} = (1-\beta)\,\bigl[1 - (1 - p\,S(x_{\text{domain}}))^{n}\bigr] + \beta\, p\, S(x_{\text{all}}) \]
\[ R_{\text{vendor}} = 1 - \prod_{v}\bigl(1 - p_v\, S(x_v)\bigr), \qquad R_{\text{total}} = 1 - (1 - R_{\text{plane}})(1 - R_{\text{vendor}}) \]

p is the annual compromise probability per control unit: base 0.01, range 0.001–0.1. It is an order-of-magnitude placement between the rate at which control-reaching vulnerabilities were published in 2024–25 (several per year) and the rate at which they were exploited with grid intent (zero). It is not a decomposition into P(vulnerability) × P(exploitation) × P(grid intent); that is a stated limitation (Section 23).

Linear metrics that do not saturate are reported alongside: expected affected capacity E[MW/yr] = Σ p × CBR_eff, and recovery exposure [MWh/yr] = Σ p × CBR_eff × T_rec.

The linear metric has an important property: splitting into n domains leaves Σ p × (C/n) × n = p × C unchanged. Splitting does not lower expected affected capacity; only the floor (e, ρ) and reach do. Splitting helps only through the convexity of S(x) — n small disturbances are less likely to cause a wide-area event than one large one.

The exponent k drives the result: k = 1 is linear, k = 4 is near zero below x = 1. We use k = 3 as base and report k = 1–4 throughout (Figure 6 right, Figure 9).

7. Calibrating grid sensitivity

We fit the x scale to literature cases. There are too few points, from grids of different size, to call this calibration in a statistical sense; it is a fit.

CaseSimultaneous changeReference reservexOutcome
France 2023, two nuclear units tripped [R06]2,660 MW3,000 MW (CE FCR)0.8949.88 Hz, no shedding
ENTSO-E reference incident [R08]3,000 MW3,000 MW1.0Design absorption limit
Dabrowski et al. (model) [R09]4,500 MW3,000 MW1.5Below 49 Hz, shedding band (paper's model result; our swing model gives 49.80 Hz and does not reproduce it)
East Japan, 16 March 2022 Fukushima-oki earthquake [R13]>6,000 MW≈1,200 MW (3% of ≈40 GW night demand)≈5UFR operated, >2 GW shed, no collapse
Hokkaido 2018, initial loss [R01]1,160 MW155 MW7.5UFLS operated; further losses led to blackout
SolarPower Europe/DNV critical threshold [R06]10,000 MW3,000 MW3.3Cascading-collapse indicator (report threshold)

The fitted Weibull is the black line in Figure 3, with the classification of our swing-equation model (Section 13: 45 GW demand, H = 3.5 s, FCR 3%, FRR 5%, UFLS at 48.5/48.2/47.9 Hz) as blue steps. A 1,350 MW step (x = 1) reaches 49.51 Hz; 3,000 MW (x ≈ 2.2) reaches 48.55 Hz, just above the first shedding stage; 4,000 MW (x ≈ 3.0) triggers it. S(2.2) = 0.89 and S(3.0) = 0.996, so as "probability of shedding or worse" the two agree. S(1.0) = 0.18 and S(1.5) = 0.50 assign probability to a region the swing model classes as "degraded"; we read this as S absorbing what a single-area model omits (transmission constraints, protection coordination, voltage, initial-condition spread).

Figure 3 Severity calibration Figure 3. Severity function S(x) = 1 − exp(−(x/1.7)³) fitted to literature points (black) with the swing-equation classification of the reference grid as blue steps. S is the probability of under-frequency load shedding or worse. Below x ≈ 0.5 the function is unidentified; the grey band shows k = 1–4.
Figure 3 Severity calibration Figure 3. Severity function S(x) = 1 − exp(−(x/1.7)³) fitted to literature points (black) with the swing-equation classification of the reference grid as blue steps. S is the probability of under-frequency load shedding or worse. Below x ≈ 0.5 the function is unidentified; the grey band shows k = 1–4.

16 March 2022 and 6 September 2018 both have x above 5 and S ≈ 1. One ended in load shedding, the other in collapse. S(x) does not express the difference; additional losses (in Hokkaido, hydro tripped by a line fault) and grid size did. We group both as "shedding or worse".

Limitations first. The model is single-area, frequency only. Transmission congestion, voltage collapse and protection cascades are absent; the Iberian blackout of April 2025 was voltage-driven [R04] and the 2006 European split flow-driven [R03]. And x = ΔP/FCR is not transferable between grids with different inertia or FCR-to-demand ratios: Continental Europe's FCR is about 0.75% of demand, the reference grid's 3%. Carrying an x₀ fitted on European points to a Japanese grid is a stretch; we carry the error in the range of k and leave per-grid calibration as an open problem (research/open_problems.md).

8. Two models, both defined realistically

To be fair we give both sides a "realistic best". The draft's A was "one credential moves everything, pre-2020 design"; reviewers called it a straw man. We define A in two stages.

Model A: centralized control + strong authentication + certification. All DER connect to one DERMS/cloud control plane. Devices are mutually authenticated, encrypted and certified; compromise probability is one-third of uncertified (sensitivity to one-tenth). Commands are executed without device-side validation. The plane is redundant, but credentials, API and keys are logically one.

Model A+: A + server-side aggregate cap. The tools of a real DERMS — tenant isolation, scoped API keys, rate limits, multi-party approval for fleet-wide commands, staged rollout (1% → 10% → 100%), command quotas on HSM-held keys, an independent safety monitor that blocks deviation from the TSO schedule — expressed as "the total capacity that can be moved in one window". Base 1,500 MW per window (the CIP-002 Medium threshold). The cap binds under an API-credential compromise; it does not under a full-backend compromise (including the safety monitor), to which we assign 0.3 p (0.1–1).

Model B: distributed autonomy, upper link advisory. DER belong to 100 independent control domains (independent credentials, keys, tenants; firmware distribution still shared per vendor). Each device observes voltage, frequency, RoCoF and SOC locally and has a constraint engine that accepts, modifies or rejects upper-level commands. The lowest operating point a remote command can reach is 30% of rating (floor e = 0.3), rate of change 20% of rating per minute, remote authority per domain capped at 500 MW. The floor actually works with probability ρ = 0.95 (firmware defects, misconfiguration, maintenance bypass leak 5%). On loss of the upper link devices run on the last valid schedule — a function Japan's curtailment-capable inverters already have in design A [S12], so not a B-specific difference. B's only specific difference is enforcing the floor and the authority cap on the device.

e = 0.3, 20%/min and 500 MW are design variables without numerical basis. Germany's §14a EnWG guarantees controllable loads at least 4.2 kW (40% above 11 kW) when the grid operator dims them, forbids full disconnection, and limits it to two hours [S06] — a legislated floor, equivalent to e ≈ 0.4 for an 11 kW wallbox, enforced in the smart-meter gateway software, not in an independent circuit.

Figure 1 Three architectures Figure 1. The three architectures compared. A: all DER on one national control plane; commands execute without device-side validation. H: ten regional domains, no device constraints. B: 100 independent credential domains, each device holding an output floor, ramp limit and authority cap, with the upper link advisory. Firmware distribution (OTA) is shared per vendor in all three.
Figure 1 Three architectures Figure 1. The three architectures compared. A: all DER on one national control plane; commands execute without device-side validation. H: ten regional domains, no device constraints. B: 100 independent credential domains, each device holding an output floor, ramp limit and authority cap, with the upper link advisory. Firmware distribution (OTA) is shared per vendor in all three.

Intermediate forms are evaluated too: H hierarchical (10 regional domains, no device constraints), F federated (100 domains, none), C cellular (10,000 cells, none — no operational reality, a theoretical floor), B4 single plane + device constraints only, B1 100 domains + floor, no cap, and B variants B5 floor held by independent monitor, B6/B7 B5 + firmware-line reach cap 12%/5%.

9. The Local Autonomy Principle and Bounded Authority

9.1 The compromised command problem

A command arrives by the legitimate path, with a legitimate signature and legitimate authority — and it is the attacker's. Zero Trust (NIST SP 800-207 [S11]) verifies everything, but a command that passes verification is executed. In Solarman/Deye the attacker's commands were legitimate API calls [I01].

Authentication is not useless here, but it is not sufficient. What is needed is a limit on what a device will execute even for a legitimate command.

9.2 The Local Autonomy Principle

Principle. A DER must be able to decide on its own, from locally observable quantities (voltage, frequency, RoCoF, SOC, temperature), an operating point that does not harm the grid, regardless of whether the upper link is present or correct. Upper-level commands are accepted within that local judgement.

This is not distrust of the upper level, whose commands are essential for balancing, markets and congestion. The principle separates survival control from economic control: the former closes locally, the latter is the upper level's job. Grid-forming inverters hold frequency without communication [W04]; microgrid islanding is a precedent for surviving a lost link [S03].

9.3 Bounded Authority and the Command Envelope

ConstraintSymbolMeaningPrecedent
Output flooreLowest operating point reachable by remote command (fraction of rating); cannot be undercut by repetition§14a's 4.2 kW floor
Rate of changeΔP/ΔtUpper limit on speed of changeCurtailment ramp requirements
Simultaneous unitsN_maxUnits one credential may command at once—
Regional totalP_region,maxTotal one domain may move within an areaNCCS ECII, CIP-002 1,500 MW
Local conditionsf, V, RoCoFe.g. reject a curtailment command while frequency is lowIEEE 1547 ride-through

A command is accepted (within the envelope), modified (clipped to it) or rejected (contradicts local observation, or repeatedly outside). Rejections are reported upward.

9.4 What "hardware enforcement" means

The draft placed the last stage "outside software authority, as relays or hardware limits". Inverter designers objected, correctly: the DSP firmware sets the current reference, and slew-rate limits and floor clamps are themselves firmware. Interconnection protection relays can only trip, and a trip is a 100% step — the opposite of a floor. No component "keeps 70% flowing in hardware".

We decompose the term into three:

  1. Physical disabling of the remote-OFF and remote-OTA command classes. A communications module with no write path, a jumper, a DIP switch. Feasible in hardware. The attacker can no longer send "stop" but can still send "reduce".
  2. An output floor held by an independent monitor circuit. A monitor MCU separate from the main MCU watches output with independent current measurement and overrides the main MCU below the floor. This is a second firmware, meaningful only if its update path and signing key are separated from the main firmware. Make it non-updatable and it cannot be fixed; make it updatable and it returns to a common cause. We cannot resolve this dilemma here (open problem). For PV the floor is 30% of available output, not of rating — on a 40%-irradiance day the attacker gets 10%. For batteries and EVs it is a floor on the charge/discharge point.
  3. Rate-of-change limit. Clouds change PV output by 50% in seconds; a monitor circuit cannot know a command's origin; fault-ride-through requires fast recovery. The ramp limit therefore stays in the main firmware and is lost under firmware compromise.

The model represents (1)+(2) as vendor_hw_enforce (share whose floor survives) and treats (3) as s = 1 under firmware compromise. This decomposition is what moved scenario B5b in Section 13 from "contained" in the draft to "load shedding".

10. Failure domains and the Blast Radius Cap

Splitting 10 million devices into 10,000 failure domains gives 1,000 devices, 5 MW per domain — the intuition behind "distribution is safety". The model's result is more involved. All values below are whole-fleet, per year, all channels (architecture_comparison.csv, Figure 4).

ConfigurationCBR_eff per compromise (plane)xCBR_eff of largest firmware lineP(some compromise)/yrRisk: plane channelRisk: vendor channelTotal
A0 single plane, uncertified35,000 MW25.910,500 MW1.0%1.0×10⁻²1.8×10⁻²2.8×10⁻²
A single plane + certification (p/3)35,000 MW25.910,500 MW0.3%3.3×10⁻³6.1×10⁻³9.4×10⁻³
A+ A + server-side cap 1,500 MW1,050 MW0.7810,500 MW0.1% (backend)1.3×10⁻³1.8×10⁻³3.1×10⁻³
H 10 domains, no device constraints3,500 MW2.5910,500 MW9.6%8.9×10⁻²1.8×10⁻²1.1×10⁻¹
F 100 domains, none350 MW0.2610,500 MW63%3.9×10⁻³1.8×10⁻²2.2×10⁻²
C 10,000 domains, none3.5 MW0.00310,500 MW100%5.0×10⁻⁴1.8×10⁻²1.9×10⁻²
B4 single plane + floor (software)11,725 MW3.2610,500 MW1.0%1.0×10⁻²1.8×10⁻²2.8×10⁻²
B 100 domains + floor + cap (software)117 MW0.0310,500 MW63%5.1×10⁻⁴1.8×10⁻²1.9×10⁻²
B3 B + certification (p/3)117 MW0.0310,500 MW28%1.7×10⁻⁴6.1×10⁻³6.3×10⁻³
B5 B + floor held by independent monitor117 MW0.033,518 MW63%5.1×10⁻⁴6.8×10⁻³7.3×10⁻³
B6 B5 + firmware-line reach cap 12%117 MW0.031,407 MW63%5.1×10⁻⁴3.5×10⁻³4.0×10⁻³
B7 B5 + firmware-line reach cap 5%117 MW0.03586 MW63%5.1×10⁻⁴7.5×10⁻⁴1.3×10⁻³
Figure 4 Architecture comparison Figure 4. Fourteen configurations, fleet-wide per year. Left: largest capacity moved by one compromise (solid = control plane / one domain; hatched = largest vendor firmware line). Middle: annual probability of a wide-area event (UFLS or worse) split into the control-plane channel (counting n domains and common cause β) and the vendor channel (cloud/OTA/firmware). Right: total. Domain splitting and the floor cut the plane channel by one to two orders, but the vendor channel bypasses authority caps and a software-only floor dies with the firmware, so software-only B totals about the same as certified A. What lowers the total is a floor held by an independent monitor (B5) and a firmware-line reach cap (B6, B7); a server-side aggregate cap (A+) brings the centralized design to the same order.
Figure 4 Architecture comparison Figure 4. Fourteen configurations, fleet-wide per year. Left: largest capacity moved by one compromise (solid = control plane / one domain; hatched = largest vendor firmware line). Middle: annual probability of a wide-area event (UFLS or worse) split into the control-plane channel (counting n domains and common cause β) and the vendor channel (cloud/OTA/firmware). Right: total. Domain splitting and the floor cut the plane channel by one to two orders, but the vendor channel bypasses authority caps and a software-only floor dies with the firmware, so software-only B totals about the same as certified A. What lowers the total is a floor held by an independent monitor (B5) and a firmware-line reach cap (B6, B7); a server-side aggregate cap (A+) brings the centralized design to the same order.

Five things can be read.

  1. On the plane channel alone, splitting and the floor work. F and C keep one compromise within reserves; B stops at 117 MW. Ten domains (H) are not enough. And splitting multiplies the attack surface — in B some domain is compromised in 63% of years. Even so B's plane-channel risk is about one-seventh of certified A: many small disturbances are less likely to cause a wide-area event than one large one.
  2. The floor alone is not enough. A single plane with device floors (B4) still moves 11.7 GW at ρ = 0.95, e = 0.3; Figure 10 (left) shows x stays above 1.5 even at ρ = 0.99.
  3. The authority cap is meaningless without independent credentials. If one credential yields all 100 domains, the cap "exists" per domain but the total is B4's 11.7 GW (Section 13, B3). The inter-domain β represents the middle ground.
  4. The total is dominated by the vendor channel. A compromise of the largest firmware line (15 GW installed, 10.5 GW reachable) crosses domain caps and erases the floor and ramp limit implemented in firmware. Shared by A and B, it leaves software-only B's total (1.9×10⁻²) equal to A0 and worse than certified A; B3 with certification is about equal to A. Changing the architecture does not change the vendor channel.
  5. What lowers the total is the independent-monitor floor and the firmware-line reach cap. B5 7.3×10⁻³, B6 4.0×10⁻³, B7 1.3×10⁻³. The centralized side reaches the same order with a server-side cap (A+ 3.1×10⁻³).

The design requirement is therefore not "number of domains" but a cap on the capacity reachable from one credential, one key, one distribution path — the Blast Radius Cap, proposed as the policy KPIs "Maximum Remote Authority" and "remote reach per firmware line" (Section 19).

Figure 2 CBR decomposition Figure 2. Decomposition of Cyber Blast Radius by compromise unit for the reference fleet (10 million × 5 kW). Installed (raw) capacity versus effective capacity after reach, authority cap and device floor. The national plane reaches 35 GW; the largest firmware line 10.5 GW regardless of domain splitting.
Figure 2 CBR decomposition Figure 2. Decomposition of Cyber Blast Radius by compromise unit for the reference fleet (10 million × 5 kW). Installed (raw) capacity versus effective capacity after reach, authority cap and device floor. The national plane reaches 35 GW; the largest firmware line 10.5 GW regardless of domain splitting.

11. The centralization paradox and the break-even

The benefits of central control are real: monitoring in one place, fast anomaly detection, one fix delivered everywhere, maximum balancing value. So centralization proceeds.

The paradox is that the efficiency centralization brings and the Blast Radius it brings are the same thing. Moving everything from one place is efficient, and moving everything from one place is dangerous. Only the ceiling — how much can be moved — can be separated. A+ is exactly that ceiling on the centralized side, giving up part of the efficiency (moving more than 1,500 MW at once) for a smaller Blast Radius.

The break-even, comparing single compromises: centralization lowers compromise probability to 1/p_f while multiplying Blast Radius by Y. With the distributed Blast Radius at 10 MW (x = 0.007):

Probability reduction 1/p_fBlast Radius multiplier YCentralized CBRRisk ratio A/B (k = 3)Range (k = 1–4)Linear ratio A/B
1/10×10100 MW10²10⁰–10³1
1/10×1001,000 MW10⁵10¹–10⁷10
1/3×100010,000 MW10⁶10²–10⁸333

Ratios depend on k by orders of magnitude and are given as ranges. Even on the linear metric, a tenfold probability cut with a hundredfold Blast Radius leaves the centralized side ten times worse.

Figure 6 Break-even Figure 6. Break-even between probability reduction and Blast Radius multiplication for a single compromise. Left: risk ratio A/B over the grid of probability reduction and Blast Radius multiplier. Right: the probability reduction a centralized design needs to equal a 10 MW distributed Blast Radius, for k = 1–4. Even at k = 1, 10,000 MW needs a 227-fold reduction.
Figure 6 Break-even Figure 6. Break-even between probability reduction and Blast Radius multiplication for a single compromise. Left: risk ratio A/B over the grid of probability reduction and Blast Radius multiplier. Right: the probability reduction a centralized design needs to equal a 10 MW distributed Blast Radius, for k = 1–4. Even at k = 1, 10,000 MW needs a 227-fold reduction.

Figure 6 (right) shows, for k = 1–4, the probability reduction needed to make a centralized Blast Radius equal in risk to a distributed one. Even at k = 1, making 10,000 MW equivalent to 10 MW needs a 227-fold reduction. Certification is expected to deliver a few-fold.

Above the reserve, severity saturates: with 35 GW movable, cutting probability to a third or a tenth leaves the outcome "collapse". Lowering probability there lowers risk only proportionally.

As an extreme example, compare A: p = 0.1%, 5 GW with B: p = 1%, 10 MW (illustrative; changeable in the simulator). A has x = 3.7, severity ≈ 1, systemic risk 1.0×10⁻³/yr. B has x = 0.007 and severity 8×10⁻⁸ at k = 3, 4×10⁻³ at k = 1. The risk ratio lies between 10¹ and 10⁸ for k = 1–4, and A is 50× larger on the linear metric. Even with a tenfold higher compromise probability, a 500-fold smaller Blast Radius gives lower risk for any k. This is a single-compromise comparison; counting attack surfaces fleet-wide, as in Section 10, narrows it.

12. Correlated failure, common cause, monoculture

"The probability that 10 million devices are compromised together is p¹⁰⁰⁰⁰⁰⁰⁰, so negligible" holds only for independent failures. Devices sharing a cloud, a firmware line, a certificate or an aggregator are not independent.

The standard treatment is the β-factor method [S04]: a share β of each device's failures occur as common cause. IEC 61508 puts β at 0.5–10%. ccf_analytic.csv shows that with β = 0 the probability of 10% failing together is effectively zero; at β = 1% it is 10⁻⁶, at 10% 10⁻⁵. The mean failure rate is unchanged; the tail is.

Figure 7 Common-cause failure Figure 7. Monte Carlo of common-cause failure (vendor cloud, firmware line, aggregator credentials shared by groups of devices). Left: annual exceedance of simultaneous affected capacity. A single vendor (red) is rarely hit but its tail reaches 35 GW. Right: more vendors raise the probability that some vendor is compromised while lowering the 99th-percentile capacity. Diversity limits the size of an event, not its frequency; only with per-line reach caps and a surviving floor does it stay within reserves.
Figure 7 Common-cause failure Figure 7. Monte Carlo of common-cause failure (vendor cloud, firmware line, aggregator credentials shared by groups of devices). Left: annual exceedance of simultaneous affected capacity. A single vendor (red) is rarely hit but its tail reaches 35 GW. Right: more vendors raise the probability that some vendor is compromised while lowering the 99th-percentile capacity. Diversity limits the size of an event, not its frequency; only with per-line reach caps and a surviving floor does it stay within reserves.

Figure 7 (left) is a 200,000-year Monte Carlo with vendor cloud, firmware line and aggregator credentials as common causes. A single vendor (red) is rarely compromised but its tail reaches 35 GW. Eight vendors and 100 domains (orange) raise the frequency and stop the tail at 10.5 GW (largest vendor). Holding the floor in an independent monitor (blue dashed) brings it to 3.5 GW.

Figure 7 (right) is the "diversity paradox": the annual probability of exceeding FCR at least once is 0.023 for one vendor and 0.044 for eight vendors with independent monitors — higher when split, because there are more independent places to breach. But the 99th-percentile affected capacity falls from 35 GW to 3.5 GW. Diversity reduces the size of an event, not its frequency. It is insufficient alone; combined with per-firmware-line reach caps and floors it stays within reserves (20 vendors, 15% top share, independent monitors: p99 1.9 GW, 1.5×FCR exceeded 0.0065/yr).

The true common cause is OTA. B's "100 independent domains" concerns credentials and keys; firmware distribution remains with eight vendors. In Japan the same inverter platform is sold under several brands (OEM supply), so firmware-line share exceeds brand share. CrowdStrike 2024 made a trusted update channel the common cause for 8.5 million machines [I07]; South Australia 2016 saw the fault-ride-through settings of nine wind farms become the common cause of a state-wide blackout [R05]. "Configuration monoculture" has already caused wide-area failures in power.

12.1 Galápagos syndrome and monoculture

There is a Japan-specific path. If Japan-specific certification or communication requirements narrow the field of vendors, the largest firmware line's share rises; from 30% to 60% doubles CBR_vendor from 15 to 30 GW. A certification scheme that unintentionally strengthens monoculture and widens Blast Radius exists in the equations. Whether JC-STAR does so is an empirical question for future market data; we only note the possibility (research/jc_star_and_japan_context.md). Geer et al. called software monoculture a systemic risk in 2003 [W07].

13. Attack simulation: one attack, two endings

We apply the same attack to a single-area swing-equation model (model/grid_sim.py). Grid: 45 GW demand, H = 3.5 s, FCR 1,350 MW (3%, τ = 10 s), FRR 2,250 MW (5%, τ = 90 s), load damping 1.5, UFLS at 48.5/48.2/47.9 Hz at 10% each, generator tripping at 47.5 Hz. Attack: a compromise of the control cloud or a firmware line cuts reachable DER output simultaneously. Devices without a working floor (1 − ρ) and devices whose ramp limit died with the firmware move as a step; the rest at 20%/min. x follows Section 6 (step part against FCR, total against FCR + FRR, larger of the two).

ScenarioOutput lostof which stepxMax RoCoFNadirOutcome
A1 national plane, 50% cut17,500 MW17,50013.02.78 Hz/s47.59 HzLoad shedding
A2 national plane, full OFF35,000 MW35,00025.95.56 Hz/s47.44 HzCollapse
A3 largest vendor cloud, full OFF10,500 MW10,5007.81.67 Hz/s48.20 HzLoad shedding
A4 A+ server-side cap 1,500 MW, API credential compromise1,500 MW1,5001.10.24 Hz/s49.43 HzDegraded (reserves exhausted)
H1 one of 10 domains, full OFF3,500 MW3,5002.60.58 Hz/s48.50 HzLoad shedding
F1 one of 100 domains, full OFF350 MW3500.260.06 Hz/s49.92 HzContained
B1 one of 100 domains, device constraints on117 MW180.05≈050.00 HzContained
B2 national plane, floor on, no cap11,725 MW1,7505.00.69 Hz/s48.20 HzLoad shedding
B3 cap per domain, one credential opens all11,725 MW1,7505.00.69 Hz/s48.20 HzLoad shedding
B4 10 of 100 domains at once (shared IdP)1,173 MW1750.500.03 Hz/s49.96 HzContained
B5a largest firmware line, software-only floor10,500 MW10,5007.81.67 Hz/s48.20 HzLoad shedding
B5b same, floor survives in independent monitor, ramp lost3,518 MW3,5182.60.58 Hz/s48.50 HzLoad shedding
B5b′ hypothetical: floor and ramp both survive3,518 MW5251.50.08 Hz/s49.86 HzContained (reference)
B5c firmware-line reach capped at 12%, floor survives1,407 MW1,4071.040.22 Hz/s49.48 HzDegraded
B5d firmware-line reach capped at 4%, software-only1,400 MW1,4001.040.22 Hz/s49.48 HzDegraded
Figure 5 Attack simulation Figure 5. The same attack — a compromise of the control cloud or a firmware line cuts reachable DER output — applied to a single-area swing-equation model. Left: first 60 s of representative scenarios. Right: affected capacity versus frequency nadir for all scenarios. A1, A3, H1, B2, B5a and B5b reach load shedding; F1, B1 and B4 stay within reserves. In B5b the 30% floor survives but the ramp limit dies with the firmware, so 3.5 GW drops as a step. B5c (firmware-line reach capped at 12%) and A4 (server-side cap 1,500 MW) both land in the 'degraded' band just above the reserve.
Figure 5 Attack simulation Figure 5. The same attack — a compromise of the control cloud or a firmware line cuts reachable DER output — applied to a single-area swing-equation model. Left: first 60 s of representative scenarios. Right: affected capacity versus frequency nadir for all scenarios. A1, A3, H1, B2, B5a and B5b reach load shedding; F1, B1 and B4 stay within reserves. In B5b the 30% floor survives but the ramp limit dies with the firmware, so 3.5 GW drops as a step. B5c (firmware-line reach capped at 12%) and A4 (server-side cap 1,500 MW) both land in the 'degraded' band just above the reserve.

A1 hits 47.6 Hz within two seconds even at a 50% cut, with load shedding just before generator tripping. A3 — one vendor's cloud — also ends in shedding, the structure SUN:DOWN described as "more than a dozen manufacturers above 3 GW each in Europe" [R06].

In B1 the same plane attack stops at 117 MW and frequency barely moves. The difference is not that the compromise did not happen; it does, with the same probability. The difference is the reach of one credential (1/100) and how far a remote command can reduce output (to 70%).

B2, B3 and B5a are where distributed autonomy loses: an unsplit plane, non-independent credentials, a software-only floor. Missing any of the three, B ends like A3.

B5b was "contained" in the draft. As reviewers noted, the ramp limit dies with the firmware; even with the floor, 3,518 MW drops as a step and shedding follows. B5b′ is the reference value if a ramp limit survived in hardware, which current inverters cannot provide. Beyond this point only the reach of one firmware line can lower the outcome: to avoid shedding (x < 1) in the reference grid, about 1,350 MW (11.5% of the fleet) if the floor survives, 3.9% if software-only (Figure 10 right, vendor_share_limit.csv). A hypothetical design keeping both floor and ramp would allow 30.7%, but it does not exist today.

Figure 10 Limits of the floor alone and the firmware-line reach cap Figure 10. Left: trying to lower x with only a device floor e and local-enforcement survival ρ while keeping a single control plane; even at ρ = 0.95, e = 0.3, x ≈ 3.3 remains in the UFLS band. Right: the reach cap for one firmware line (same signing key, same OTA path) so that its compromise stays within reserves, using the Section 6 definition of x (step part against FCR, total against FCR+FRR): about 4% software-only, about 11.5% if an independent monitor keeps the floor, about 31% in a hypothetical design where floor and ramp both survive.
Figure 10 Limits of the floor alone and the firmware-line reach cap Figure 10. Left: trying to lower x with only a device floor e and local-enforcement survival ρ while keeping a single control plane; even at ρ = 0.95, e = 0.3, x ≈ 3.3 remains in the UFLS band. Right: the reach cap for one firmware line (same signing key, same OTA path) so that its compromise stays within reserves, using the Section 6 definition of x (step part against FCR, total against FCR+FRR): about 4% software-only, about 11.5% if an independent monitor keeps the floor, about 31% in a hypothetical design where floor and ramp both survive.

A4 is the centralized answer. An independent server-side safety monitor with an aggregate cap stops an API-credential compromise at 1,500 MW — the same "degraded" band as B5c. What remains is a full-backend compromise.

14. Monte Carlo: what survives when assumptions move

We sampled the uncertain variables of the assumptions table — unit count, capacity, reach, vendor count and shares, compromise probability, certification factor, relative probabilities of vendor-channel and backend compromise, inter-domain β, server-side cap, demand, reserves, severity shape — from LOW–HIGH triangular distributions (20,000 draws). B's design variables (100 domains, 500 MW, e = 0.3, 20%/min) are fixed as a design; the certification factor applies to all configurations. Four configurations: A, A+, B (software-only), Bfull (B + independent-monitor floor + 12% firmware-line reach cap).

ComparisonDistributed side lowerLower by 10×Median ratio
Plane channel only: B vs A100%65%14
Total: B (software) vs A100%5%1.8
Total: Bfull vs A99%46%9
Total: A+ vs A99%0%2.7
Total: Bfull vs A+83%24%3.3
Total: B (software) vs A+33%2%0.7
Figure 8 Monte Carlo Figure 8. Monte Carlo over the LOW–HIGH ranges of the uncertain assumptions (20,000 draws), with B's design variables fixed (100 domains, 500 MW cap, e = 0.3, 20%/min) and the certification factor applied to all configurations. Left: on the control-plane channel alone, distributed B is always lower (median 1/14). Second: adding the vendor channel, software-only B is lower than A but only by a median 1.8×. Third: B with an independent-monitor floor and a 12% firmware-line reach cap is 1/9 of A. Right: compared with a realistic centralized A+ (server-side aggregate cap), the same B is lower in about 83% of draws, median ratio 3.3.
Figure 8 Monte Carlo Figure 8. Monte Carlo over the LOW–HIGH ranges of the uncertain assumptions (20,000 draws), with B's design variables fixed (100 domains, 500 MW cap, e = 0.3, 20%/min) and the certification factor applied to all configurations. Left: on the control-plane channel alone, distributed B is always lower (median 1/14). Second: adding the vendor channel, software-only B is lower than A but only by a median 1.8×. Third: B with an independent-monitor floor and a 12% firmware-line reach cap is 1/9 of A. Right: compared with a realistic centralized A+ (server-side aggregate cap), the same B is lower in about 83% of draws, median ratio 3.3.

The tornado plots (Figure 9) show that p's range is the largest uncertainty everywhere. A is set by p and the certification factor, saturated against capacity variables. A+ adds the server cap and backend ratio. Software-only B responds to p, the vendor-channel ratio and unit count; domain count and authority cap barely move the total. Bfull responds to the floor, grid size and severity shape.

Figure 9 Tornado Figure 9. One-variable sensitivity (tornado) of total fleet risk with design variables fixed. A is set almost entirely by compromise probability and the certification factor; capacity variables do not move it because severity is saturated. A+ adds the server-side cap and the backend-compromise ratio. Software-only B responds to p, the vendor-channel ratio and unit count; domain splitting barely affects the total. With the independent monitor and reach cap, the floor, grid size and severity shape (k, x₀) become influential. The range of p is the largest uncertainty everywhere.
Figure 9 Tornado Figure 9. One-variable sensitivity (tornado) of total fleet risk with design variables fixed. A is set almost entirely by compromise probability and the certification factor; capacity variables do not move it because severity is saturated. A+ adds the server-side cap and the backend-compromise ratio. Software-only B responds to p, the vendor-channel ratio and unit count; domain splitting barely affects the total. With the independent monitor and reach cap, the floor, grid size and severity shape (k, x₀) become influential. The range of p is the largest uncertainty everywhere.

15. When distributed autonomy loses

The hypothesis must be falsifiable. Conditions under which the model says centralization is safer, or distribution does not function:

  1. The fleet is small relative to the grid. If x_A < 1, splitting gains little and operational simplicity wins — e.g. 1 GW of DER on a large grid.
  2. The gain from splitting loses to the multiplied attack surface. n domains cut one Blast Radius to 1/n but multiply compromise opportunities by n. Net gain depends on the convexity of S(x); where x_B is already small, more n brings nothing. Most Monte Carlo cases where A beats B are cases where one B compromise is within reserves but the number of surfaces wins.
  3. The vendor channel is relatively likely. If firmware-line compromise is as likely as plane compromise, splitting does not lower the total; only the independent-monitor floor and reach cap do.
  4. The centralized side can protect its server-side cap. With an independent safety monitor and rare backend compromise, A+ lands in Bfull's order and wins on simplicity.
  5. Firmware is compromised and the floor is software-only. The constraint engine is taken with it (B5a); central detection and mass key revocation may then be faster.
  6. Domains are not independent. Reused credentials, shared CA, shared operator, shared IdP: splitting is nominal and CBR equals the centralized case (B3, β → 1).
  7. Recovery. Firmware fixes travel via vendors at the same speed in A and B; key reissue runs in parallel across B's 100 domains; site visits do not depend on structure. We find no basis for order-of-magnitude recovery-time differences, but simultaneous reconnection is a disturbance regardless (Section 17, L5).
  8. Flexibility value. A 30% floor limits remote reduction to 70%. AEMO's emergency backstop faces the opposite problem — too little remotely reducible capacity [S07]. In Japan the TSO/DSO has the authority and duty to curtail PV to zero (Section 20); the floor's width becomes a quantity agreed with the operator.
  9. The attacker is on the physical side. Substations and protection relays (Ukraine 2015 [I05], Industroyer2 [I08]) make DER design irrelevant.
  10. Severity is near-linear. At k = 1 the difference is a ratio, not orders — yet still a 227-fold probability reduction (Figure 6 right).

We do not hide these. Conversely, where they are not met — GW-scale fleets, concentrated credentials, no cap on server or device, firmware lines concentrated in a few vendors — the conclusion applies as stated.

16. Relation to prior work

The structure is not new; we have re-posed prior work with cyber Blast Radius as the objective.

What this article adds: (1) the CIP-002/NCCS threshold logic decomposed to distribution-level DER and devices; (2) plane and vendor channels summed at fleet level per year, showing numerically when splitting does not lower the total; (3) "hardware enforcement" decomposed into implementable elements, with the reach cap computed on the premise that the floor survives and the ramp limit does not; (4) an open model readers can check with their own numbers, with the reviewers' objections and our responses in full.

17. Physics overrides software authority

In one sentence: however large software's authority, physics (frequency, voltage) and an independent monitor must be able to refuse it.

The implementation is a graceful-degradation ladder:

Figure 11 Degradation ladder Figure 11. The graceful-degradation ladder: how far a device trusts upper-level commands at each level and what remains. L0 normal; L1 doubt (modify and report); L2 ignore (last valid schedule and local control); L3 link lost (autonomous, islanding where possible); L4 local protection (relay trips; independent monitor overrides below the floor); L5 staged recovery with ramp limits.
Figure 11 Degradation ladder Figure 11. The graceful-degradation ladder: how far a device trusts upper-level commands at each level and what remains. L0 normal; L1 doubt (modify and report); L2 ignore (last valid schedule and local control); L3 link lost (autonomous, islanding where possible); L4 local protection (relay trips; independent monitor overrides below the floor); L5 staged recovery with ramp limits.

L4's two elements do not depend on main-firmware authority: relays can only trip, and the floor is kept by the independent monitor. While both remain, "firmware compromise = total shutdown" does not follow. While the floor lives in the main firmware, B5a's ending is unavoidable.

L5 is easily overlooked. Ten million devices reconnecting at once is the same synchronized change as an attack. IEEE 1547's randomized reconnection delay is the existing answer.

18. Architecture comparison and economics

AspectCentralized (A)Centralized + server cap (A+)Hierarchical (H)Federated (F)Distributed, software-only (B)Distributed + independent monitor + reach cap (B6)
Flexibility valueMaximumHigh (≤1,500 MW at once)HighMedium–highMedium (loses the 30% floor)Medium
Max CBR per compromise: plane35,000 MW1,050 MW3,500 MW350 MW117 MW117 MW
Max CBR: largest firmware line10,500 MW10,500 MW10,500 MW10,500 MW10,500 MW1,407 MW
Total systemic risk (/yr)9.4×10⁻³ (certified)3.1×10⁻³1.1×10⁻¹2.2×10⁻²1.9×10⁻² (6.3×10⁻³ certified)4.0×10⁻³
Expected affected capacity (MW/yr)15249455455222152
Behaviour on link lossLast schedule (common)SameSameSameSame + local controlSame + local control
Common causesPlane, firmware linesBackend, firmware linesRegional planes, firmware linesFirmware linesFirmware linesFirmware lines (capped), monitor-MCU supplier
RecoveryFirmware via vendors (common); key reissue at 1 siteSame10 sites100 sites100 sites (parallel)Same
Implementation complexityLow (centre) / high (scale)Medium (monitor independence)MediumMedium–highHigh (device constraint engine)High + device BOM + new test standards

Economics, per year, all order-of-magnitude (research/open_problems.md gives bases and ranges):

ItemWho paysAnnual (50 GW reference fleet)
Lost flexibility value from a 30% floorGrid operator, market (ultimately consumers)35 GW reach × 70% = 24.5 GW, market participation 10–30% × ¥5,000–10,000/kW·yr ≈ ¥12–74 bn
Independent monitor circuit (monitor MCU, independent sensing, override drive)Device owners (passed through)¥1,000–3,000/unit × 200–300 k units/yr ≈ ¥0.2–0.9 bn
New test standard and per-model certificationVendorsNo floor/rejection test exists in JIS C 8961/8962, IEC 62109 or IEC 62116; must be created. Millions of yen per model, years of work
Operating 100 credential domains (HSM, audit, SOC)Aggregators¥ tens of millions × 100 ≈ ¥ several bn
Server-side safety monitor (A+)DERMS operatorIndependent build and run ¥ hundreds of millions to ~1 bn
A's expected blackout costSociety (externalized)p 0.003–0.01/yr × wide-area blackout cost of several to over ten trillion yen ≈ ¥10–50 bn

Two readings. First, the floor's lost value and A's expected blackout cost are the same order; economics alone do not settle it, and the draft's "expected loss exceeds added cost" is withdrawn. Second, the cheapest items — the independent monitor circuit and the server-side safety monitor — each cut total risk to a third to a fifth. The floor's width e is the tuning variable between lost value and risk; 30% is only a starting point.

Incidence aligns everyone toward loosening the cap: owners pay for parts, vendors for certification, aggregators for operations, operators in flexibility, while benefits spread thinly over society. No cap emerges voluntarily; it must be written into connection requirements (Section 19).

19. Grid Code 2.0 and new policy KPIs

Today's grid codes specify one device's effect on the grid: ride-through, power factor, ramp rates. Cyber is handled by separate regimes (device certification, security guidelines). On the transmission side CIP-002 and NCCS look at "MW per compromise"; on the distribution side nobody does.

Grid Code 2.0 proposes writing remote-authority caps and device-side floors into connection requirements, with KPIs that are measurable and auditable.

KPIDefinitionHow measuredAB (software)B6
Cyber Blast Radius (MW)Largest capacity one compromise can move (plane or firmware line)Reach register × floor attributes35,00010,5001,407
Maximum Remote Authority (MW per credential)Cap on remote authority per credential/domainServer-side cap declaration and audit50,000500500
Remote reach per firmware line (MW)Installed capacity reachable through one signing key / OTA pathPer-line capacity register15,00015,000≤6,000 (12%)
Settings-authority reach (MW)Capacity behind credentials that can rewrite schedules, protection settings or clocksSame50,000500500
Device-side command rejection (auditable attributes)Remote-OFF class disabled / independent monitor circuit / OTA path separatedDevice label, type testno/no/nono/no/noyes/yes/yes
Vendor concentration (HHI, top share)Concentration of installed-capacity sharesRegister0.18 / 30%0.18 / 30%≤0.10 / 12%

The draft's "Autonomous Survival Time" is removed: running on the last schedule when the link drops is a function Japan's curtailment inverters have in A as well. "Communication independence ρ" cannot be measured by test and is decomposed into three binary attributes.

The firmware-line reach cap is not a vendor market-share cap. A vendor may sell any number of GW; it must split the capacity reachable through one signing key and one OTA path. Nationality-neutral, not a quota. Enforcement needs a per-firmware-line capacity register, which does not exist today.

Connections to existing regimes, item by item:

20. The question for Japan

Japan's grid is split into 50 Hz and 60 Hz areas linked by about 2.1 GW of frequency converters. East Japan effectively shares one frequency. Hokkaido is DC-linked, frequency-independent, and went to blackout in 2018 from a 1,160 MW initial loss.

The question is one:

In Japan, how many GW can one attacker move at once?

Nobody can answer today. Device counts are known; certified counts can be counted. But "how many MW sit behind one credential" and "how many MW does the largest vendor's cloud reach" are not compiled.

The first thing to measure is Japan's largest existing control plane. Online curtailment by TSO/DSOs sends schedules from a per-area command server to capable inverters, which execute them [S12]. It is not a precedent for B. In our classification it is type H — a per-area single control plane, no device floor, authority down to zero — a configuration Section 10 classes as load shedding. In Kyushu, more than half of PV is reduced within tens of minutes on sunny spring days. How many times the area's primary reserve does that server reach? That number is the first CBR.

TSO/DSOs have the authority and duty to curtail; a device design that "rejects" commands conflicts with it. A reconciliation:

  1. Separate the schedule path from the immediate-command path. Curtailment schedules (next-day hourly output caps) are outside the floor, with per-area aggregate caps and a device-side frequency/voltage check instead (do not execute a scheduled reduction while frequency is low).
  2. Apply the floor and authority cap to the immediate-command path (balancing market, aggregator DR) — the subject of B.
  3. Count settings authority. Capacity behind credentials that can rewrite schedules weighs as much as immediate-command CBR: the attacker need not act within the window; writing zero into tomorrow's schedule and cutting the link suffices.

Answering requires:

  1. Declaration by TSO/DSOs, aggregators and vendors of "remote reach" and "settings-authority reach"
  2. Registration of device attributes (remote-OFF class disabled, independent monitor, OTA separation)
  3. Publication of x against each area's primary and secondary reserve

In a Hokkaido-sized grid the same 1,000 devices give an x more than ten times East Japan's (grid_context_cases.csv). Thresholds must be ratios to each area's reserve, not national constants.

21. Relation to RightOS / RightFlow

I-S3 develops a platform (RightOS/RightFlow) that verifies rights and authority at the device. This article's "constraint engine" shares its structure — a device accepting, modifying or rejecting upper-level commands against its own conditions.

But the article does not presuppose the product. The CBR model, the floor and the degradation ladder are design principles valid for any vendor's implementation. Whether I-S3's implementation meets them should be measured by this article's KPIs — above all "independent monitor circuit" and "OTA path separation". A software-only constraint engine, by this article's results, does not lower total risk.

22. Using the simulator: your own numbers

The numbers here are a reference case. Readers' grids, fleets and threat estimates differ. The Cyber Blast Radius Simulator below implements the Python model's equations in JavaScript; change an input and the three columns A, A+ and B update.

Cyber Blast Radius Simulator

Inputs: unit count, average capacity, remote reach, vendor count and top share, independent domains, authority per domain, local autonomy on/off, floor, ramp limit, floor survival, compromise probability, certification factor, relative vendor-channel and backend probabilities, share whose floor survives in an independent monitor, inter-domain β, server-side cap, recovery time, demand, primary and secondary reserve, inertia, severity shape.

Outputs: total capacity, max CBR per compromise (plane, firmware line), x, severity, frequency response (plane and firmware-line compromise), probability of some compromise per year, systemic risk (plane channel, vendor channel, total), expected affected capacity, recovery exposure, HHI.

Python–JavaScript agreement is checked by simulator/parity_test.js on seven cases. Equations, code and the assumptions table are public; disagree with an assumption, change the row and recompute.

23. Falsification conditions and limitations

The claim is wrong if any of the following is shown:

  1. Severity is nearly flat in x (a grid where ten times the reserve does not lead to shedding); then CBR differences are not risk differences. Consistency with Section 7 would need to be shown.
  2. Compromise probability differs by orders of magnitude by structure (a central plane 1/1000 as likely as a distributed domain); the break-even could then reverse (Figure 6).
  3. The independent-monitor floor is unimplementable or bypassable; B5a becomes distributed autonomy's reality, requiring a 3.9% firmware-line reach cap.
  4. The server-side aggregate cap holds even under backend compromise; A+ is then the cheapest solution.
  5. A real attack in which a distributed, floor-equipped fleet caused larger grid impact than a centralized one.

Limitations:

24. Conclusion

Since the probability of a breach cannot be zero, the question moves from "how do we defend" to "what structure do we permit in which N MW move when the defence fails". This article defined that quantity as Cyber Blast Radius and compared centralized control and distributed autonomy against the same attack on the same grid, fleet-wide, across all channels, per year.

The result is more cautious than the draft. For the same control-plane compromise, distributed autonomy turns 35 GW into 117 MW. But once attack surfaces are counted and the shared vendor firmware-line channel is added, software-only distributed autonomy is no safer than centralization. What lowered the total was not the architecture but a cap on the MW behind one credential and one firmware line — on the device side an output floor held by a monitor circuit independent of the main firmware, on the server side an aggregate cap held by an independent safety monitor. The cap can sit on either side. On the device side, B6/B7 are one-half to one-seventh of certified A; on the server side, A+ is one-third; with assumptions varied, either is below A in over nine draws out of ten.

Centralization keeps its advantages where the fleet is small, where the server-side cap can be strongly protected, where the vendor channel is relatively weak, and where flexibility value must be maximised. On economics, the floor's lost value and the expected blackout cost are the same order; nothing is settled there.

Still, for the real structure — GW-scale fleets behind one credential and a few firmware lines, with no cap on server or device — the conclusion does not move. However far probability is lowered, the consequence of a structure that can move 35 GW does not change. Only a cap on authority changes it.

The secure grid is not the grid that can never be hacked. It is the grid that continues to function when something is hacked.

One question remains: in Japan, how many GW can one attacker move at once? Measuring it, publishing it, and capping it is the next step.

Citation and reuse

Of the attachments, the CSV, JSON, Python and JavaScript files (data/, model/, simulator/, results/*.csv, *.json) are released under CC BY 4.0. Cite I-S3 Co., Ltd. and this article's URL and you may reuse, modify and redistribute them, including commercially. If you publish results recomputed with different assumptions, please keep the attribution to the original model.

Text and figures (PNG) may be quoted with attribution. For wholesale reproduction of figures, or if you hold different data or assumptions, use the contact form; we will recompute and append how the results change.

When citing "the I-S3 Cyber Blast Radius model", please give the version (1.0) and the base values in data/assumptions.csv; other base values give other numbers.

Suggested citation

English

Masuda, S. (2026). From Cybersecurity to Cyber-Resilience: How to Connect Tens of Millions of DER in a World That Cannot Be Fully Defended. I-S3 Co., Ltd. https://www.i-s3.com/cyber-resilience-der-blast-radius-en.html

Japanese

益田周防海(2026)「CybersecurityからCyber-Resilienceへ——完全には守れない世界で、数千万台の分散電源をどう接続すべきか」株式会社I-S3、2026年10月7日公開、https://www.i-s3.com/cyber-resilience-der-blast-radius.html

BibTeX

@misc{masuda2026cbr,
  author = {Masuda, Suomi},
  title  = {From Cybersecurity to Cyber-Resilience: How to Connect Tens of Millions of DER in a World That Cannot Be Fully Defended},
  year   = {2026},
  publisher = {I-S3 Co., Ltd.},
  url    = {https://www.i-s3.com/cyber-resilience-der-blast-radius-en.html},
  note   = {Japanese original: https://www.i-s3.com/cyber-resilience-der-blast-radius.html. Model, data and code (CC BY 4.0): https://www.i-s3.com/cyber-resilience-der-blast-radius.html#assets}
}

Key files for recomputation

Reproduce: python3 model/cbr_model.py && python3 model/grid_sim.py && python3 model/correlated_failure.py && python3 model/monte_carlo.py && python3 model/make_figures.py && node simulator/parity_test.js (README).

Change log

Attachments

The files below are generated with the same assumptions and calculations as the text; use them to check numbers or recompute.

Assumptions and data

Model (Python 3, numpy, matplotlib)

Results (CSV / JSON)

Simulator

Research notes and review

Text and summaries

References

ID は本文・data/assumptions.csv・data/incidents.csv から参照する。I=事例、R=事故報告・研究、S=規格・規制、W=アーキテクチャ・理論の先行研究、J=日本のDERデータ、D=仮定表の根拠。アクセス日は特記なき限り 2026-10-06。

I. サイバー事例

R. 事故報告・系統研究

S. 規格・規制

W. アーキテクチャ・理論の先行研究

J. 日本の DER データ

D. 仮定表(assumptions.csv)の根拠

Suomi Masuda / I-S3 Co., Ltd. Based on public sources and an open model. Not advice for a specific system design or investment.