CAVOK Services AB https://cavokservices.se Safety Engineering and Aviation Tue, 10 Dec 2024 18:37:38 +0000 en-US hourly 1 https://cavokservices.se/wp-content/uploads/2024/10/cropped-icononly_transparent-1-32x32.png CAVOK Services AB https://cavokservices.se 32 32 239595418 Safety vs reliability https://cavokservices.se/safety-vs-reliability/ https://cavokservices.se/safety-vs-reliability/#respond Tue, 10 Dec 2024 18:37:36 +0000 https://cavokservices.se/?p=344 During my career I have seen several instances, where safety and reliability are not properly understood and often interchanged with each other. Sometimes on a concept level or vehicle function design level if the engineer is not experienced with ISO26262 (like non-EU OEM for example) but knows the basics of the ASIL concept tends to […]

The post Safety vs reliability first appeared on CAVOK Services AB.

]]>
During my career I have seen several instances, where safety and reliability are not properly understood and often interchanged with each other. Sometimes on a concept level or vehicle function design level if the engineer is not experienced with ISO26262 (like non-EU OEM for example) but knows the basics of the ASIL concept tends to classify all of the supplier requirements on the highest ASIL, just to be sure that the vehicle would be safe. This is how I met reliability requirements mistaken to safety requirement. To clean up the confusion I prepared a very simple example to demonstrate how these two concepts are entirely different.

  • Let’s look at the picture above. A safe and reliable product is a Casio watch. It works all the time and poses no danger on the user whatsoever.
  • An unsafe and reliable product is for example a skateboard. It’s a simple design with reliable parts, and we can’t really think of a skateboard that is non-functioning for some reason other than being broken or destroyed. However, it is a very dangerous vehicle: anyone without the proper training AND protective gear is subject to risk of physical harm.
  • A safe but unreliable product is basically any smartphone app. An app can not possibly malfunction in such a way that could harm the user. But they are usually very unreliable, they can freeze, lag and crash any time.
  • An unsafe and unreliable product is the BMW N57D30 engine, that has a tendency of sucking in the engine oil through the bearings of the turbocharger and thus making it impossible to turn it off by cutting the fuel supply. This usually ends in the fire of the engine and the whole vehicle.

The main takeaway from the above is, that a product can be safe without being reliable.

The post Safety vs reliability first appeared on CAVOK Services AB.

]]>
https://cavokservices.se/safety-vs-reliability/feed/ 0 344
5 ways to mess up your FTA https://cavokservices.se/5-ways-to-mess-up-your-fta/ https://cavokservices.se/5-ways-to-mess-up-your-fta/#respond Thu, 05 Dec 2024 16:34:13 +0000 https://cavokservices.se/?p=341 FTA is an analysis method that breaks down an undesired behavior of a system to different root causes and thus reveals systematic failures and weak points of the system. Just like FMEA, FTA also originates from aerospace engineering, hence is it highly utilized in safety critical engineering. One of the most significant differences between FMEA […]

The post 5 ways to mess up your FTA first appeared on CAVOK Services AB.

]]>
FTA is an analysis method that breaks down an undesired behavior of a system to different root causes and thus reveals systematic failures and weak points of the system. Just like FMEA, FTA also originates from aerospace engineering, hence is it highly utilized in safety critical engineering. One of the most significant differences between FMEA and FTA is that FTA takes simultaneous errors into account while FMEA does not.

In regards of ISO26262 FTA is most commonly used as the mandatory deductive analysis tool for ASIL C and D safety artifacts. The most beneficial aspect of FTA however is the effective and structured breakdown of a safety artifact. If you are already experienced with FTAs here are 5 mistakes to avoid:

I.  Seeing FTA as a requirement breakdown method

It’s a common mistake, when safety engineers set a requirement as the top element of the FTA. For example, “Blockage of steering shall be avoided”, then as the root failures they define the requirements that shall be fulfilled in order to fulfill the top requirement. This is a way of skipping a few steps between analysis and the outcome, and therefore it can jinx the whole point of the FTA. Remember, the top of the fault tree is the failure that is caused by not fulfilling a functional safety requirement or safety goal. And on the branches of the fault tree there should be root errors. Only and the end of a branch shall we define the preventing requirement.

II.  Skipping levels

If we return to the previous example an error of skipping levels would be to flatten the fault tree and skipping interim failures. This usually happens when the product is almost ready for production and the functional safety work was not carried out parallel to the product development. An example of this mistake would be listing a root cause to a top failure of “Picture freeze of the rear-view camera on the display” as “ISP buffer freeze”. This mistake not only costs us higher ASIL level requirements (as the common cause errors are not discovered, thus decomposition is not possible) but poses a higher risk on missing a possible fault on lower levels.

III.  Listing unplausible errors

Similar to the FMEA this is where domain knowledge and industry experience plays a large role in the completeness of the analysis. And again, important to mention that FTA is not a design tool either. It’s an analysis tool for a system that has been designed and the limitations are well known. For example, we can’t expect an ISP buffer freeze causing a complete image freeze if the buffer cache memory exceeds the size of a full frame.

IV.   Not minding the independence

A very handy feature of the FTA is that if the precondition to a failure is two other failures occurring simultaneously, we are allowed to decompose the ASIL of the top failure. This method is supported and described in the ISO26262 standard and comes it useful when we want to decrease the integrity level of our safety requirements. But very often the requirement of the independence is forgotten. For example, if out top failure is “communication loss due to undervoltage” we can’t decompose it to “undervoltage in the MCU” and “undervoltage in the CAN phy” since they are not independent errors.

V.   Not breaking down the timing

Usually forgotten and described in later stages of the development, but FTTI is just as important as the requirement itself. In the concept phase the FTTI is usually omitted or just added as <tbd>, but it should be handles as an important part of the requirement. The same applies when breaking them down, The timing of the resulting requirements and even the safety mechanisms depend largely on the needed timing. For example, if an MCU can not fulfill the fault detection time requirement an complete redesign might be necessary, and with the progression of the development it gets exponentially more expensive. Therefore, the timings shall be handled early and with the same focus as the requirements.

The post 5 ways to mess up your FTA first appeared on CAVOK Services AB.

]]>
https://cavokservices.se/5-ways-to-mess-up-your-fta/feed/ 0 341
When the agile craze (almost) took over the car industry https://cavokservices.se/when-the-agile-craze-almost-took-over-the-car-industry/ https://cavokservices.se/when-the-agile-craze-almost-took-over-the-car-industry/#respond Tue, 19 Nov 2024 21:45:59 +0000 https://cavokservices.se/?p=324 The wind of change About a decade ago the whole automotive industry got stirred up by the “next big thing” that would increase work effectivity, streamline processes, while making the company future proof. Not too long after we recovered from the LEAN craze of the 90s, the new world-changing way of working was already knocking […]

The post When the agile craze (almost) took over the car industry first appeared on CAVOK Services AB.

]]>
The wind of change

About a decade ago the whole automotive industry got stirred up by the “next big thing” that would increase work effectivity, streamline processes, while making the company future proof. Not too long after we recovered from the LEAN craze of the 90s, the new world-changing way of working was already knocking on the industry’s doors. Management gurus, snake oil salesmen and money-making MBAs, all without engineering background took no time to jump on the new wave and milk it recklessly. Hyping up the old-fashioned profit-oriented CEOs with the promise of something as innovative as the first iPhone: the verdicts from the studies were clear, the numbers didn’t lie. And as some early adopters gave in to the promises, the rest didn’t want to turn into the next Nokia who totally missed the signs of the times, so they listened to the over-enthusiastic consultants.

And just like that “agile” became the buzzword of the early 2010s all around the automotive industry.

The fact that agile way of working was meant for smaller, fast paced SW companies with short feedback loops (both in the sense of organization and time), low iteration costs, and basically no regulations usually didn’t bother the advocates. Some people saw the future of agile within automotive pretty dark, but these rational voices usually were too few or just suppressed by management hubris.

The practice

Several Tier-1s and OEMs started shoehorning the agile practices into the V model and took the challenge to revolutionize the industry. PI plannings, sprint reviews, daily standups became part of our lives; we attended funky trainings where we have gotten familiar with MVPs, user stories and much more. We found ourselves following the ways of working of a startup while working for monstrous enterprises with rigid processes, conservative culture, legal certifications and a hierarchy chain that could reach the sky.

We can see now, what the critics saw in the beginning: HW, safety critical SW, and mechanics development, where the dependence on suppliers is enormous and thanks to certifications an iteration can take up more than a year is not exactly the environment that the writers of the agile manifesto had in mind.


Many managers set up to this new trend without properly understanding how agile could actually improve (or in many cases hinder) their effectivity, some even played Frankenstein and came up with a mashup of the V-model and agile and called it “Vagile” (this is not a joke, I personally heard this from one of my directors). The verdicts from studies focused on the software centered companies were loud and clear: this is the future; you are either in or out for good. In many cases this ended up in a cargo cult philosophy: we do standups and PI plannings, use scrum or Kanban boards and then magically we’ll turn into a highly profitable company with spotless customer satisfaction.

What is left

Most of the players on the automotive market already accepted that the next big thing wasn’t meant for them, but some great things stuck and without a doubt significantly improved our everyday life at work. Without the aim of completeness just a few examples: the Atlassian toolchain became the gold standard at every size, daily standups and scrum boards allow teams to track their progress accurately, and even team sizes and competences were more rationalised. The winners are definitely the SW only Tier-1s and Tier-2s who could adapt the best agile by adding some extra steps for compliance.

The post When the agile craze (almost) took over the car industry first appeared on CAVOK Services AB.

]]>
https://cavokservices.se/when-the-agile-craze-almost-took-over-the-car-industry/feed/ 0 324
What is systems engineering? https://cavokservices.se/what-is-systems-engineering/ https://cavokservices.se/what-is-systems-engineering/#respond Sat, 16 Nov 2024 20:21:10 +0000 https://cavokservices.se/?p=319 Every time someone outside of the automotive industry asks me what do I work with, I find myself in a quite funny situation. Most of the people know what SW development is, they know that the people that create the apps on their phones are the SW developers. Everyone has seen a PCB probably, so […]

The post What is systems engineering? first appeared on CAVOK Services AB.

]]>
Every time someone outside of the automotive industry asks me what do I work with, I find myself in a quite funny situation. Most of the people know what SW development is, they know that the people that create the apps on their phones are the SW developers. Everyone has seen a PCB probably, so they would understand what HW development is about. Mechanical guys also have an easy way to explain their daily work, everyone knows what a cogwheel is. But what about systems engineering?

Not so long ago my organist friend of 10 years asked me what do I do for a living actually? Other than I work on XX product I couldn’t really translate my work into everyday language. Because how would telling that “I negotiate and break down the customer specification, design the functional behaviour and carry out safety work needed for ISO26262 compliance” sound to my musician friend? Encrypted probably. And even if I say “systems engineering” that also doesn’t help a lot, since it could mean anything.

At the start of my career even as an engineer from a prestigious university it took more than a year to wrap my head around requirement management, traceability, coverage and other metrics, and most importantly the significance of requirements. If I say “I write concise paragraphs about how the product shall be” most of the people would think about a law related position.

So recently I came up with this example:

Imagine that a customer wants to build a billiard table, from a carpenter who doesn’t know what billiard is. So the customer tells the carpenter that he wants a table covered in green fleece with 6 holes. Then the carpenter delivers a table to the customer that is fully wrapped in green fleece and has 6 holes equally spaced in the middle. Now, I’m working for the carpenter and my job is to ensure that this doesn’t happen and even to make sure that the table will not collapse or harm anyone.

This still doesn’t describe all the aspects of this “technical janitor” kind of role, where you get involved in a huge variety of tasks sometimes feeling like an octopus with each leg in a different area, but maybe gives a hint about the importance of systems engineering.

The post What is systems engineering? first appeared on CAVOK Services AB.

]]>
https://cavokservices.se/what-is-systems-engineering/feed/ 0 319
The simplest introduction to FMEA https://cavokservices.se/the-simplest-introduction-to-fmea/ https://cavokservices.se/the-simplest-introduction-to-fmea/#respond Wed, 13 Nov 2024 08:54:24 +0000 https://cavokservices.se/?p=315 FMEA, Failure Mode and Effect Analysis originates from the 1940s and still today it’s widely used in all branches of product design and development. Many standards require an FMEA analysis to help eliminating potential hazards of a system/component or process. FMEA is basically a structured answer to the question: “What could go wrong?”. It is […]

The post The simplest introduction to FMEA first appeared on CAVOK Services AB.

]]>
FMEA, Failure Mode and Effect Analysis originates from the 1940s and still today it’s widely used in all branches of product design and development. Many standards require an FMEA analysis to help eliminating potential hazards of a system/component or process. FMEA is basically a structured answer to the question: “What could go wrong?”. It is important that FMEA is NOT a design tool, it aims to analyse something that is at the end of the design lifecycle, and it considers only one failure a time.

I.

Let’s say that our company plans to release a saltshaker with a metal lid, which we have to analyze to ensure product safety and quality.

In the first step let’s list what could go wrong:

  1. The holes get blocked, and the salt won’t come out or only a few small pieces will fall through with each shake
  2. The cap will corrode and contaminate the salt
  3. When the salt gets stuck at the bottom due to humidity and the user smashes it to the table, the glass can break and cuts the user’s hand

Here we can see the first catch of FMEA: to find sensible failure modes. For example, a lightning striking the cap of the saltshaker and melting it thus closing the holes is technically a failure mode, but no one would consider this during the design. Therefore, it is important to include someone in the analysis with domain experience to advise which faults are sensible for consideration.

II.

As we have 3 failure modes let’s have a deeper look at them and see what is the most severe. Using common sense anyone would say that failure nr. 3 is the most severe possibly seeing the bleeding hand on front of our eyes. But FMEA is smarter than common sense: it uses 3 aspects to rate the severity of a failure: severity, probability and controllability. Let’s see an example for each:

  • Severity: the magnitude of the harm it could cause(to people, infrastructure, business, legal matters etc.):
    • High severity: a phone battery exploding
    • Low severity: camera flash malfunction
  • Probability/Exposure: how likely is it that the failure will occur during the design lifetime
    • High probability: discoloration of a shoe sole
    • Low probability: discoloration of a glass bottle
  • Controllability: if we discover the problem occurring can we intervene (move the people away from the threat)
    • High controllability: incorrect set temperature of an air conditioner
    • Low controllability: sudden explosion of a car tire

It is very important that when quantifying the above aspects, the scale should always spread from most favorable (1) to least favorable (10). The most common is the 1-10 but other scales are also allowed, provided that same scale is used for all of the 3 aspects!

Let’s grade now our failures from above on a 10-level scale:

  1. The holes get blocked, and the salt won’t come out or only a few small pieces will fall through with each shake
    1. Severity: 3, since it only effects the product usability
    2. Probability: it will most probably happen during the lifetime of the saltshaker, so we can grade it as an 8
    3. Controllability: 1, since the user can just buy a more fine-grained salt
  2. The cap will corrode and contaminate the salt:
    1. Severity: 7, since it’s not going to be lethal, but could pose a high health risk on prolonged use
    2. Probability: we are talking about a metal lid, it’s very probable that normal humidity and the corrosive nature of salt will corrode the lid during the product lifetime, so it’s 8
    3. Controllability: 5 
  3. When the salt gets stuck at the bottom due to humidity and the user smashes it to the table, the glass can break and cut the user’s hand
    1. Severity: 8, since it’s not lethal but poses a high risk of human injury
    2. Probability: the thick glass can not break when hit against a usual table with human force. It most likely breaks when hit against an uneven concrete/stone surface. We can state that it’s very unlikely that the glass breaks during expected use: 2
    3. Controllability: thin glass breaks suddenly, but our saltshaker is thick, so it’ll chip first. Also, we can control the force and surface we use. Let’s grade it to a 4.

It’s apparent that the gradings are not fully objective. There might be scales with examples and explanations for each grade, but overall, this is a subjective grading where FMEA shows one of its weaknesses: without domain experienced members in the analysis team, it’s very easy to over/underestimate the gradings.

III.

Now as we graded our failures let’s calculate the RPN aka. the Risk Priority Number. This dimensionless number will guide us to weight the risks of out failures.  The RPN is the product of the grades of the three aspects of each failure mode:

  1. 3*8*1 = 24
  2. 7*8*5 = 280
  3. 8*2*4 = 64

We can see from the above results, that the highest risk is contaminating the salt, and not cutting the user with broken glass.

IV.

Now that we have a single number for each failure, it’s time to decide which are the ones we need to focus on aka. define corrective actions. Deciding this is the consideration of the team or the company. Sometimes high severity failures are always considered regardless of the RPN, sometimes the top few RPNs are focused on and sometimes there is a limit, for example RPN 250, which is subject to corrective actions. These actions usually aim to decrease the number of one aspect. Let’s see now some examples for corrective actions on our highest RPN, failure nr.2:

  • Replace the cap with plastic
  • Apply protective layer on the cap
  • Do not sell the product in high humidity areas
  • Etc.

What is left now is to decide on the corrective action(s) and re-evaluate the RPN value of the fault ensuring that the new RPN will stat under our limit.

The post The simplest introduction to FMEA first appeared on CAVOK Services AB.

]]>
https://cavokservices.se/the-simplest-introduction-to-fmea/feed/ 0 315