Issue #4 ·

Physical AI Safety Dispatch — July 2026

Fourth issue. This month went deep: fleet safety, standards bodies, diagnostic coverage, common cause failures, enterprise procurement. The audience that stayed through this is the audience that matters.

Month four. The theme: deep expertise. I deliberately pushed deeper into technical territory — IEC 61508 internals, diagnostic coverage ranges, common cause failure analysis, enterprise procurement gates. The question was whether the audience would follow. They did.


Three posts that went deepest

1. Fleet safety — one robot is not a thousand

Single-robot safety is a solved problem for industrial arms behind cages. ISO 10218 covers it. ISO 13849 grades it. TÜV certifies it. But fleet safety — 100+ heterogeneous robots sharing a workspace — has no standard methodology. No composable safety guarantees across different vendors. No fleet-level anomaly detection that doesn't drown operators in N× false positives. No root-cause attribution protocol for multi-robot incidents. Amazon's DeepFleet system (arXiv 2508.08574) manages approximately 5 million robot-hours and uses a vertex reservation safety bridge — but it's proprietary, not a standard. The gap between single-robot safety and fleet safety is where the next decade of Physical AI infrastructure gets built.

Source: arXiv 2508.08574; ISO 10218:2025; VDA 5050.

Read the full post →


2. Diagnostic coverage — the number that determines your SIL

DC measures the fraction of dangerous hardware failures that automatic diagnostics can detect. IEC 61508-2 Table A.2 defines three levels: Low (60–90%) = SIL 1 capable. Medium (90–99%) = SIL 2 capable. High (≥99%) = SIL 3 capable. A safety system with 60% DC detects 6 out of 10 dangerous faults. One with 99% detects 99 out of 100. That gap is the difference between a robot that can work near humans and one that must stay behind a cage. Most founders building Physical AI have never seen this table.

Source: IEC 61508-2:2010, Table A.2.

Read the full post →


3. The five hardest unsolved problems

I listed the five problems that, if any one is solved cleanly in the next three years, creates a decade-defining company. Real-time verification of learned policies against physical safety constraints. Fleet-level anomaly detection without false-positive explosion. Continuous certification that survives OTA model updates. Composable safety guarantees across heterogeneous fleets. Root-cause attribution after multi-robot coordinated failures. Each sits at the intersection of AI/ML, hardware safety, and scale. Single-robot solutions don't transfer.

I said I've picked which two I care about most. I didn't say which two publicly.

Read the full post →


What Month 4 was about

Month 4 was deliberately harder than Months 1-3. More equations. More standard references. More acronyms. The technical depth was a choice — fleet safety, robot procurement, diagnostic coverage, common cause failures, standards bodies, and the five hardest unsolved problems. Real audience numbers ship with the August retrospective; the Month 4 analytics window doesn't close until early August.


What I'm reading

IEC 61508-6:2010, Annex D — Common cause failure scoring. The β factor: the fraction of failures that defeat redundancy. β = 10% is the DEFAULT if you implement no CCF defenses — meaning one in ten dangerous failures affects both channels simultaneously. That's not redundancy. That's shared fragility. I wrote Post #45 about this and it got the most saves of any post in Month 4.

Enterprise robot procurement processes. MANTEC (2025), Robotic Systems Authority (2026), EVSINT (2026). I mapped the five gates that enterprise buyers actually use: pilot → technical evaluation → safety review → procurement + legal → phased rollout. Gate 3 (safety review) is where most vendors die silently. A vendor with CE marking enters Gate 3 pre-qualified. Without it, the safety review is the longest gate.


A note I won't post on LinkedIn

Common cause failures are the most underrated risk in Physical AI safety. Everyone builds dual-channel architectures and calls it redundancy. But if both channels share a power supply, the same firmware version, the same thermal environment, or the same engineer's design assumptions — a single cause can take out both. IEC 61508 quantifies this as β. I've calculated β for several commercial robot safety architectures using publicly available documentation. In every case where I could find enough information, the β estimate was above 5% — meaning the redundancy claim was aspirational, not engineered. The companies don't know this because they've never done the analysis. The standard requires it. Nobody checks.

— Mati


Physical AI Safety Dispatch is a monthly newsletter by Mati Melchior. Published on the 1st of every month.

We use cookies

This site uses essential cookies to function and, with your consent, analytics cookies (Google Analytics) to understand how the site is used. Learn more.