Practices - CAVOK Services AB https://cavokservices.se Safety Engineering and Aviation Thu, 05 Dec 2024 16:35:02 +0000 en-US hourly 1 https://cavokservices.se/wp-content/uploads/2024/10/cropped-icononly_transparent-1-32x32.png Practices - CAVOK Services AB https://cavokservices.se 32 32 239595418 5 ways to mess up your FTA https://cavokservices.se/5-ways-to-mess-up-your-fta/ https://cavokservices.se/5-ways-to-mess-up-your-fta/#respond Thu, 05 Dec 2024 16:34:13 +0000 https://cavokservices.se/?p=341 FTA is an analysis method that breaks down an undesired behavior of a system to different root causes and thus reveals systematic failures and weak points of the system. Just like FMEA, FTA also originates from aerospace engineering, hence is it highly utilized in safety critical engineering. One of the most significant differences between FMEA […]

The post 5 ways to mess up your FTA first appeared on CAVOK Services AB.

]]>
FTA is an analysis method that breaks down an undesired behavior of a system to different root causes and thus reveals systematic failures and weak points of the system. Just like FMEA, FTA also originates from aerospace engineering, hence is it highly utilized in safety critical engineering. One of the most significant differences between FMEA and FTA is that FTA takes simultaneous errors into account while FMEA does not.

In regards of ISO26262 FTA is most commonly used as the mandatory deductive analysis tool for ASIL C and D safety artifacts. The most beneficial aspect of FTA however is the effective and structured breakdown of a safety artifact. If you are already experienced with FTAs here are 5 mistakes to avoid:

I.  Seeing FTA as a requirement breakdown method

It’s a common mistake, when safety engineers set a requirement as the top element of the FTA. For example, “Blockage of steering shall be avoided”, then as the root failures they define the requirements that shall be fulfilled in order to fulfill the top requirement. This is a way of skipping a few steps between analysis and the outcome, and therefore it can jinx the whole point of the FTA. Remember, the top of the fault tree is the failure that is caused by not fulfilling a functional safety requirement or safety goal. And on the branches of the fault tree there should be root errors. Only and the end of a branch shall we define the preventing requirement.

II.  Skipping levels

If we return to the previous example an error of skipping levels would be to flatten the fault tree and skipping interim failures. This usually happens when the product is almost ready for production and the functional safety work was not carried out parallel to the product development. An example of this mistake would be listing a root cause to a top failure of “Picture freeze of the rear-view camera on the display” as “ISP buffer freeze”. This mistake not only costs us higher ASIL level requirements (as the common cause errors are not discovered, thus decomposition is not possible) but poses a higher risk on missing a possible fault on lower levels.

III.  Listing unplausible errors

Similar to the FMEA this is where domain knowledge and industry experience plays a large role in the completeness of the analysis. And again, important to mention that FTA is not a design tool either. It’s an analysis tool for a system that has been designed and the limitations are well known. For example, we can’t expect an ISP buffer freeze causing a complete image freeze if the buffer cache memory exceeds the size of a full frame.

IV.   Not minding the independence

A very handy feature of the FTA is that if the precondition to a failure is two other failures occurring simultaneously, we are allowed to decompose the ASIL of the top failure. This method is supported and described in the ISO26262 standard and comes it useful when we want to decrease the integrity level of our safety requirements. But very often the requirement of the independence is forgotten. For example, if out top failure is “communication loss due to undervoltage” we can’t decompose it to “undervoltage in the MCU” and “undervoltage in the CAN phy” since they are not independent errors.

V.   Not breaking down the timing

Usually forgotten and described in later stages of the development, but FTTI is just as important as the requirement itself. In the concept phase the FTTI is usually omitted or just added as <tbd>, but it should be handles as an important part of the requirement. The same applies when breaking them down, The timing of the resulting requirements and even the safety mechanisms depend largely on the needed timing. For example, if an MCU can not fulfill the fault detection time requirement an complete redesign might be necessary, and with the progression of the development it gets exponentially more expensive. Therefore, the timings shall be handled early and with the same focus as the requirements.

The post 5 ways to mess up your FTA first appeared on CAVOK Services AB.

]]>
https://cavokservices.se/5-ways-to-mess-up-your-fta/feed/ 0 341
The simplest introduction to FMEA https://cavokservices.se/the-simplest-introduction-to-fmea/ https://cavokservices.se/the-simplest-introduction-to-fmea/#respond Wed, 13 Nov 2024 08:54:24 +0000 https://cavokservices.se/?p=315 FMEA, Failure Mode and Effect Analysis originates from the 1940s and still today it’s widely used in all branches of product design and development. Many standards require an FMEA analysis to help eliminating potential hazards of a system/component or process. FMEA is basically a structured answer to the question: “What could go wrong?”. It is […]

The post The simplest introduction to FMEA first appeared on CAVOK Services AB.

]]>
FMEA, Failure Mode and Effect Analysis originates from the 1940s and still today it’s widely used in all branches of product design and development. Many standards require an FMEA analysis to help eliminating potential hazards of a system/component or process. FMEA is basically a structured answer to the question: “What could go wrong?”. It is important that FMEA is NOT a design tool, it aims to analyse something that is at the end of the design lifecycle, and it considers only one failure a time.

I.

Let’s say that our company plans to release a saltshaker with a metal lid, which we have to analyze to ensure product safety and quality.

In the first step let’s list what could go wrong:

  1. The holes get blocked, and the salt won’t come out or only a few small pieces will fall through with each shake
  2. The cap will corrode and contaminate the salt
  3. When the salt gets stuck at the bottom due to humidity and the user smashes it to the table, the glass can break and cuts the user’s hand

Here we can see the first catch of FMEA: to find sensible failure modes. For example, a lightning striking the cap of the saltshaker and melting it thus closing the holes is technically a failure mode, but no one would consider this during the design. Therefore, it is important to include someone in the analysis with domain experience to advise which faults are sensible for consideration.

II.

As we have 3 failure modes let’s have a deeper look at them and see what is the most severe. Using common sense anyone would say that failure nr. 3 is the most severe possibly seeing the bleeding hand on front of our eyes. But FMEA is smarter than common sense: it uses 3 aspects to rate the severity of a failure: severity, probability and controllability. Let’s see an example for each:

  • Severity: the magnitude of the harm it could cause(to people, infrastructure, business, legal matters etc.):
    • High severity: a phone battery exploding
    • Low severity: camera flash malfunction
  • Probability/Exposure: how likely is it that the failure will occur during the design lifetime
    • High probability: discoloration of a shoe sole
    • Low probability: discoloration of a glass bottle
  • Controllability: if we discover the problem occurring can we intervene (move the people away from the threat)
    • High controllability: incorrect set temperature of an air conditioner
    • Low controllability: sudden explosion of a car tire

It is very important that when quantifying the above aspects, the scale should always spread from most favorable (1) to least favorable (10). The most common is the 1-10 but other scales are also allowed, provided that same scale is used for all of the 3 aspects!

Let’s grade now our failures from above on a 10-level scale:

  1. The holes get blocked, and the salt won’t come out or only a few small pieces will fall through with each shake
    1. Severity: 3, since it only effects the product usability
    2. Probability: it will most probably happen during the lifetime of the saltshaker, so we can grade it as an 8
    3. Controllability: 1, since the user can just buy a more fine-grained salt
  2. The cap will corrode and contaminate the salt:
    1. Severity: 7, since it’s not going to be lethal, but could pose a high health risk on prolonged use
    2. Probability: we are talking about a metal lid, it’s very probable that normal humidity and the corrosive nature of salt will corrode the lid during the product lifetime, so it’s 8
    3. Controllability: 5 
  3. When the salt gets stuck at the bottom due to humidity and the user smashes it to the table, the glass can break and cut the user’s hand
    1. Severity: 8, since it’s not lethal but poses a high risk of human injury
    2. Probability: the thick glass can not break when hit against a usual table with human force. It most likely breaks when hit against an uneven concrete/stone surface. We can state that it’s very unlikely that the glass breaks during expected use: 2
    3. Controllability: thin glass breaks suddenly, but our saltshaker is thick, so it’ll chip first. Also, we can control the force and surface we use. Let’s grade it to a 4.

It’s apparent that the gradings are not fully objective. There might be scales with examples and explanations for each grade, but overall, this is a subjective grading where FMEA shows one of its weaknesses: without domain experienced members in the analysis team, it’s very easy to over/underestimate the gradings.

III.

Now as we graded our failures let’s calculate the RPN aka. the Risk Priority Number. This dimensionless number will guide us to weight the risks of out failures.  The RPN is the product of the grades of the three aspects of each failure mode:

  1. 3*8*1 = 24
  2. 7*8*5 = 280
  3. 8*2*4 = 64

We can see from the above results, that the highest risk is contaminating the salt, and not cutting the user with broken glass.

IV.

Now that we have a single number for each failure, it’s time to decide which are the ones we need to focus on aka. define corrective actions. Deciding this is the consideration of the team or the company. Sometimes high severity failures are always considered regardless of the RPN, sometimes the top few RPNs are focused on and sometimes there is a limit, for example RPN 250, which is subject to corrective actions. These actions usually aim to decrease the number of one aspect. Let’s see now some examples for corrective actions on our highest RPN, failure nr.2:

  • Replace the cap with plastic
  • Apply protective layer on the cap
  • Do not sell the product in high humidity areas
  • Etc.

What is left now is to decide on the corrective action(s) and re-evaluate the RPN value of the fault ensuring that the new RPN will stat under our limit.

The post The simplest introduction to FMEA first appeared on CAVOK Services AB.

]]>
https://cavokservices.se/the-simplest-introduction-to-fmea/feed/ 0 315