How to Calculate the EN 50159 Safety Code Length (Annex C.4)

Annex C.4 of EN 50159 gives a small set of formulas for how long the safety code must be in closed (Category 1) transmission systems. It is only three equations, but each parameter carries assumptions that are easy to get wrong, and the text itself is brief in places. This post walks through the model one parameter at a time. It also collects the questions and mistakes that come up most often in real projects.

The model in one paragraph

Every safety message is protected by two layers. The transmission code (a CAN CRC, an Ethernet FCS, and so on) belongs to the non-trusted transmission system. The safety code is generated and checked end to end by the safety-related equipment. The annex identifies three ways a corrupted message can cause a hazard:

  1. A hardware fault in the transmission system corrupts messages.
  2. EMI causes bit errors that neither code detects.
  3. The transmission code checker fails silently and stops filtering corrupted messages.

Each path gets its own equation, and the sum must stay within the target:

(1)  R_HW × p_US × k1      = R_H1      with k1 ≥ n × m,  m ≥ 5
(2)  p_UT × p_US × f_W     = R_H2
(3)  k2 × p_US × (1/T)     = R_H3

R_H1 + R_H2 + R_H3 ≤ R_H

p_UT ≈ 2^-b   (b = redundancy bits of a proper transmission CRC)
p_US ≈ 2^-c   (c = bits of the safety code)

The key to reading all three is that they share one structure:

hazard rate = (attempts per hour) × (probability that an attempt slips through the safety code)

An “attempt” is a corrupted message reaching a code check — the safety code check in equations 1 and 3, and both the transmission code check and the safety code check in equation 2. Once you see each equation this way, most of the confusion goes away. Keep the units consistent: rates and frequencies in 1/h, probabilities and k-factors dimensionless, and T in hours.

The parameters

ParameterMeaningUnit
R_HTarget hazardous failure rate of the complete transmission system1/h
R_H1Hazardous failure rate of hardware faults without transmission code checker fault1/h
R_H2Hazardous failure rate of EMI1/h
R_H3Hazardous failure rate due to faults of the transmission code checker1/h
R_HWHardware failure rate of the non-trusted transmission system1/h
p_UTProbability the transmission code misses an error–
p_USProbability the safety code misses an error–
f_MMaximum message frequency per receiver1/h
f_WFrequency of corrupted messages1/h
nConsecutive corrupted messages before safe fall-back–
mSafety factor included within k1 (m ≥ 5)–
k1Factor for hardware faults including safety margin (k1 ≥ n × m)–
k2Percentage of hardware faults that result in undetected disabling of transmission decoding–
TTime span in which more than a defined number of corrupted messages triggers the safe fall-back. Equation 3 gives the minimum interval in which only one detected error is allowedh

Equation 1: hardware faults

The attempts. A hardware fault produces a stream of corrupted messages. The system receives up to n of them before it enters the safe state, and each one has probability p_US of slipping through. The chance that at least one slips through is about n × p_US. This is the usual reading of the factor k1 ≥ n × m.

Why the margin m. Hardware faults are not random, so the n attempts are not independent draws with probability 2^-c. Two realistic examples show why:

  • An intermittent fault defeats a consecutive counter. A loose connector corrupts roughly one message in ten. Good messages keep resetting the “n consecutive errors” counter, so the safety code faces thousands of attempts instead of n.
  • An address-line fault delivers the wrong message. A gateway reads the neighbouring buffer and forwards a different, perfectly valid message. A CRC alone cannot catch this. Source identifiers and sequence numbers inside the protected data can.

The annex explains the margin briefly: because hardware failures cannot be assumed to be random, a safety margin m ≥ 5 is built into k1 = n × m. In practice, a margin does not replace design measures, so make sure the mechanisms behind the model really exist: error counters that are not reset by every good message, and source identifiers and sequence numbers inside the safety code’s coverage.

Why p_UT does not appear. Equation 1 has no p_UT factor, and the annex does not spell out why. The usual reading is that hardware faults can corrupt data outside the span the transmission code protects: before encoding, after decoding, or inside gateways and switches that decode and re-encode frames and issue a fresh, valid CRC. In effect, no credit is taken for the transmission code on this path.

Where R_HW comes from. The annex defines R_HW only as the hardware failure rate of the non-trusted transmission system. What follows is practical guidance, not annex text. In practice, R_HW is the sum of the failure rates of every element between the points where the safety code is generated and where it is checked: gateways, switches, repeaters, transceivers, communication modules, cabling and connectors. Use manufacturer FIT or MTBF data (1 FIT = 10⁻⁹/h, λ ≈ 1/MTBF), or reliability handbooks such as SN 29500, IEC 61709 or MIL-HDBK-217F. Three points deserve attention:

  • Safety equipment at the ends is excluded. It belongs to the safety equipment’s own SIL analysis.
  • Redundant paths do not reduce R_HW, at least not for a simple failover architecture where either path’s message is accepted on its own: a fault on either one can still deliver a corrupted message, so for integrity you sum both. A design that requires both paths to agree, and falls back to the safe state on a mismatch, is a different case, covered later in this post.
  • Without data, 10⁻⁴/h is a defensible starting point. It corresponds to an MTBF of about 10,000 hours, a pessimistic figure rather than a typical one for transmission hardware.

Equation 2: EMI

Why there is no failure rate. Nothing fails here. EMI is an environmental condition acting on a healthy system. f_W is simply the rate of attempts: each corrupted message tests both codes once.

Choosing f_W. The annex gives two ways to set it:

  • Take f_W = f_M. This is the worst-case estimate, and it is self-justifying, since there cannot be more corrupted messages than messages. For cyclic transmission f_M is well defined; for non-cyclic transmission the maximum possible frequency must be taken. It already covers bursts in terms of attempt count, because a burst cannot exceed “every message corrupted, always”. The cost is more safety-code bits.
  • Limit f_W with safe counters or safe timers. If more than one wrong message is received within a defined time interval, safe communication is aborted and the safe fall-back state is entered. For example, abort on the second detected error within one minute, so f_W ≤ 60/h instead of 36,000/h. This holds whatever the burst statistics are. The price is availability: bursts cause trips.

A third route sometimes seen in practice, estimating f_W from a measured bit error rate, is not one of the annex’s options. I would avoid it as a primary argument: bit error rates are not guaranteed in a railway environment, and bursts break the random-error model behind the estimate.

The real risk of bursts is independence. Multiplying p_UT × p_US assumes the two codes fail independently, but both see the same error pattern. The classic mistake is using the same CRC polynomial in both layers. An error pattern that is a multiple of that polynomial then defeats both codes, and the product becomes fiction. Only claim credit for p_UT if you can argue independence. Otherwise take p_UT = 1, as the annex’s own footnote suggests.

Where the window lives. The window must live in the safety layer. A counter in the non-trusted transmission layer cannot be credited.

Equation 3: silent failure of the transmission checker

Which “decoder” fails. It is the transmission code checker, such as the CAN controller or the Ethernet MAC. It is not the safety code checker, which keeps working and is the last remaining barrier. That is why p_US appears in the equation. If the safety checker itself failed, the answer lies in the safety equipment’s SIL design, not in this annex.

Where the failure rate went. The annex defines k2 as the factor describing the percentage of hardware faults that result in undetected disabling of transmission decoding. Its informative derivation assumes that in only 1 of 10,000 hardware faults the transmission code checker fails undetected, and that the average duration of this state (without EMI) is T = MTBF = 1/R_HW. R_HW then cancels out:

k2 = 10⁻⁴ × R_HW × (1 / R_HW) = 10⁻⁴

Note that the annex reuses the symbol T for this duration. It has nothing to do with the window T. The annex says that if periodic checking of the transmission encoding mechanism is possible, k2 can be neglected. (By analogy with proof-test intervals in IEC 61508, k2 then behaves like λ_silent × τ / 2 for a test interval τ. That is my extension, not annex text.) The annex also remarks that its estimate is very pessimistic, since a small degradation of transmission quality would usually lead to the safe fall-back state. Without any justification, take k2 = 1, which means assuming the checker is always broken.

Why 1/T is the attempt rate. With the checker broken, every corrupted message reaches the safety layer. The window mechanism caps the attempts at about one per T. Without a window, the attempts would be f_W, and equation 3 would collapse into equation 2 with p_UT = 1. If no such mechanism is implemented, the annex requires the safe fall-back state to be entered immediately after the first detected error, otherwise other measures against possible error conditions must be introduced.

This 1/T only covers tolerating a single detected error before fall-back: trip on the second, not the sixth. If your design instead tolerates N errors within a wider window W (say, five errors in an hour rather than one in a minute), equation 3 itself doesn’t hand you that directly. The condition becomes W ≥ N × T_req, where T_req is the minimum T that equation 3 gives for the one-error case with your R_H3 budget. So “five in an hour” is only valid once you’ve checked it against that inequality, it isn’t automatically equivalent to “one per minute” just because the two sound similarly strict.

If you take p_UT = 1, equation 3 becomes redundant. You are already assuming the checker filters nothing, so losing it changes nothing. You can set R_H3 = 0 and reassign its share of the budget. This is my reading, not annex text: the annex still lists all three terms, so document the argument explicitly and agree it with your assessor.

Last remarks

Shortening c is also possible, within limits. The annex allows halving c (at least) to reach the same target by repeating each message and checking the consistency of two mutually independent copies — two independent witnesses do the work of extra code bits. It even hints that some further improvement is possible beyond half, but recommends stopping there rather than doing the finer maths. This is not free: it only counts as a real gain once you have shown that common-cause failures — a single fault corrupting both copies the same way — are negligible.

This independence has to be real, not assumed. Sending both copies over the same network is not enough: hardware faults there tend to be persistent, so the same defect can corrupt both copies the same way and still pass the consistency check, one of the reasons equation 1 gives no credit for the transmission code either. The trick reliably helps against equation 2’s random bit errors; getting a genuine benefit against R_HW or a shared transmission code checker needs two physically separate paths.

A common legacy mistake is computing T = N_max × Δt, the maximum number of bad packets multiplied by the packet period. That is the system’s reaction time to a burst, not the window of equation 3. It does not contain p_US or R_H3, so it cannot guarantee the target. With a consecutive counter, it does not even bound the attempt rate. With N = 3 and Δt = 100 ms, it yields T = 0.3 s and implies about 12,000 attempts per hour, against a requirement of about 1.4 per hour. If your mechanism really tolerates N errors in a window W, the correct condition is W ≥ N × T_req.

A worked example

This example checks a design where the choices are already made: T, n and the window mechanism are already implemented, not still open. Channel: 16-bit transmission CRC, 32-bit safety code, R_HW = 1 × 10⁻⁴/h (no vendor data), n = 3, m = 5 (k1 = 15), f_W ≤ 60/h via a one-minute window, independence between the two codes credited, and a safety layer that already uses a 10-minute timer for T. Target: R_H = 1 × 10⁻¹⁰/h.

Results:

  • Eq. 1: R_HW × p_US × k1 = 10⁻⁴ × 2⁻³² × 15 ≈ 3.5 × 10⁻¹³/h.
  • Eq. 2: p_UT × p_US × f_W = 2⁻¹⁶ × 2⁻³² × 60 ≈ 2.1 × 10⁻¹³/h.
  • Eq. 3, with no periodic checking of the transmission code checker (k2 = 1): R_H3 = k2 × p_US / T = 1 × 2⁻³² / (1/6) ≈ 1.4 × 10⁻⁹/h. Total ≈ 1.4 × 10⁻⁹/h — over the 1 × 10⁻¹⁰/h target. This design fails.
  • Eq. 3, with periodic checking in place (k2 = 10⁻⁴): R_H3 = k2 × p_US / T = 10⁻⁴ × 2⁻³² / (1/6) ≈ 1.4 × 10⁻¹³/h. Total ≈ 7.0 × 10⁻¹³/h — comfortably inside the target. This design passes.

No solving needed here, just substitution: plug in what the design already has and check the sum against the target. The result also shows how much a single assumption, whether the checker is periodically tested, can decide whether a fixed design passes or fails.

If T, n or the window aren’t decided yet, the direction reverses: instead of checking a chosen T, you solve the equations for the minimum T (or c, or n) the design must have, and you can also choose not to split the target budget equally across the three equations. That derive-first approach, and how to use the freed-up budget well, is covered in full in our EN 50159 course.

The most common mistakes

  1. Claiming p_UT credit while using the same CRC polynomial in both layers, or without any argument for independence.
  2. Crediting error counters in the transmission layer. Only safe counters and timers in the safety equipment count.
  3. Double-counting p_UT when a safety-layer window already sees only frames that passed the transmission CRC.
  4. Consecutive-error counters that fully reset on each good message. Intermittent faults defeat them.
  5. Computing T as N_max × packet period instead of deriving it from equation 3.
  6. Confusing the two meanings of T in the annex: the window, and the duration of the silent-fault state in the k2 note.
  7. Thinking the eq. 3 “decoder” is the safety checker, and concluding that p_US is irrelevant there.
  8. Mixing units, such as messages per second against hazard rates per hour.
  9. Treating network redundancy as reducing R_HW. For integrity it adds.
  10. Over-investing in precise R_HW figures when a sensitivity argument shows they do not matter, or under-investing when the safety code is short and they do.
  11. Checking the safety code in non-safety hardware, such as a standard gateway. A single fault could then defeat both barriers, and the model no longer applies.
  12. Using p_US = 2^-c without demonstrating the code’s properness for the actual message length and polynomial.

Closing thought

Annex C.4 looks like arithmetic, but it is really a set of architectural requirements in disguise. Every favourable number you use must be backed by a design feature: an independent polynomial, a safe window, a non-resetting counter, a periodic checker test. When in doubt, take the conservative value (p_UT = 1, k2 = 1, f_W = f_M) and let a few extra safety-code bits buy you a simpler, more defensible safety case.

This article is guidance on interpreting the standard and does not replace the standard itself or your assessor’s judgement.


Go Deeper on Safety-Related Communication

If you want to build your expertise in EN 50159 and the safety-related communication behind interlocking and remote control systems, our online courses at RAMSRail.com are designed for engineers and project managers who need practical, working knowledge of these standards.

Explore our RAMS training courses at RAMSRail.com