Back to the blog

Quality system & steering

Is your quality system reactive, or under control?

In brief

Everyone asks you where you stand. The customer, before trusting you with a critical part. The auditor, on day three. The banker sizing up the risk. And you answer what everyone answers: "we're ISO 9001 certified." True — and it says almost nothing.

Serious maturity scales do exist to answer that question. ISO 9004 offers a five-level self-assessment; the aerospace industry has its own. The catch is that they were designed for organizations of several thousand people, with a full quality department to run them. Here's the version that fits a shop of thirty.


"We're certified" doesn't answer the question

Certification is binary. You either conform or you don't. That's useful — essential, even, to sell to certain customers — but it's a threshold, not a measurement.

Take two equally certified machine shops, comparable in size, serving neighbouring markets. In the first, nonconformities are raised by the same three people, handled when there's time left over, and closed with the note "operator made aware." In the second, anyone on the floor can raise an issue from a tablet, every nonconformity is tied to the process that produced it, and the effectiveness of corrective actions is verified three months later before the file is closed.

Both have the same certificate on the wall. They are not in the same world. The first endures its problems; the second sees them coming. And no logo on a piece of paper captures that difference.

Certification says you have a system. It doesn't say whether that system works for you, or whether you work for it.

That's precisely the gap maturity scales fill. They don't ask "do you have a procedure?" but "what actually happens when something goes wrong?" Lived from the inside, the difference is enormous.

The six questions that really describe a quality system

After fourteen years looking at quality systems from the inside, in shops of every size, I've come to the view that six questions are enough to describe any of them. Not six clauses of a standard: six questions an owner can ask himself driving home.

Do we see problems, and quickly? That's detection. How many different people raised an issue at your shop this year? If the answer is "three, always the same three," it isn't that your shop has few problems. It's that most of them never surface. A system that sees nothing isn't a healthy system — it's a blind one.

Do we know where each thing came from? That's traceability. Is every nonconformity tied to the process that produced it, to the part, the lot, the measuring instrument? Without that thread, you accumulate incidents. With it, you accumulate knowledge.

Do we handle things on time? That's responsiveness. Not "did you open a corrective action," but how long passes between opening and closing — and above all, whether anyone checks, later, that the action actually worked.

Do we act before things break? That's prevention. How many of your processes have been through a serious risk analysis? How many actions did you launch this year on a problem that hadn't happened yet? This is almost always the weakest axis, and I'll explain further down why that's no small detail.

Do we decide on numbers? That's steering. Are your indicators current at the moment you make a decision, or reconstructed the night before the management review? A dashboard you fill in for the auditor is not a dashboard.

Does the loop turn on its own? That's improvement. Do management review decisions turn into real actions, with an owner and a date? And does the same problem come back every six months under a slightly different label?

Taken separately, these six questions are unremarkable. Taken together, on a single picture, they draw a shape — and that shape tells you more about a shop than a forty-page audit report.

Figure 1 — One shop's profile, six axes

Detection Traceability Responsiveness Prevention Steering Improvement

Profile of a machine shop of about thirty people. Solid line: today. Dashed line: the same shop twelve months earlier. Real progress on five axes — and prevention has barely moved. Illustrative data.

Five levels, described by what you see on the floor

On each of these six axes, you can place yourself on a five-step scale. The labels come from recognized frameworks — ISO 9004 for the normative base, the IAQG aerospace maturity model for the industrial version. But I'll describe them differently: by what you observe walking the shop floor.

Level 1 — Reactive. You act when the customer complains. Problems get solved, often solved well, but under pressure and without leaving any usable trace. Everyone has their own method. The know-how lives in the heads of two or three experienced people, and walks out the door with them.

Level 2 — Planned. Processes are written down and broadly followed. Nonconformities are recorded, corrective actions opened. But follow-up is uneven, deadlines slip, and you discover at management review that a dozen files have been dragging for months. This is where the vast majority of certified small manufacturers sit — and it's a perfectly respectable place to be.

Level 3 — Controlled. The data is reliable and continuously current. Every issue is tied to its process and its cause. Deadlines hold because they're visible, not because someone chases people. At this level, you can answer an auditor's question in minutes rather than two days of digging.

Level 4 — Optimized. You no longer just process: you analyze. Trends are tracked, recurring causes identified and attacked at the root. Action effectiveness is verified before files are closed. Management decisions rest on numbers nobody disputes, because they come out of the system rather than a spreadsheet rebuilt for the occasion.

Level 5 — Innovating. A meaningful share of problems is detected by the system itself, with no human needing to report them: an instrument with a lapsed calibration blocks release, an out-of-tolerance dimension automatically raises a nonconformity. The system no longer merely records what it's given — it watches.

Figure 2 — The scale

Reactive Planned Controlled Optimized Innovating you endure you record you control you anticipate the system sees

The same shop as Figure 1, overall level: Planned. Why 2, when four of its six axes sit higher? That's the point of the next section.

The rule that stings: your level is your weakest link

Here's the point I won't budge on, and it's the one that makes everyone grumble when I present it: the overall level of a quality system is not the average of its axes. It's the lowest of them.

The shop in Figure 1 is excellent at steering, solid on detection, traceability and improvement. It's weak on prevention. Take the average and you get something like 3.2 — a nice number to show. Take the minimum and you get 2. And 2 is what describes reality.

Why? Because a quality system isn't a portfolio where good holdings offset bad ones. It's a chain. You may detect admirably, trace perfectly and steer to the millimetre: if you never analyze your risks upstream, you'll keep discovering the same problems after the fact, with increasing elegance. You'll have become excellent at reacting. You won't have become good at preventing.

It's the criticism I've heard most often from experienced auditors, though never in those words: "your system runs well, but it always runs downstream."

Four excellent axes don't buy back a missing one. A quality system is worth its weakest link — and that's good news, because it tells you exactly where to work.

Because that's the real value of the rule. It isn't there to give you a bad grade. It's there to stop you wasting effort. A shop still investing in steering when it's already at level 4 on that axis and level 2 on prevention is spending its money in the wrong place. The average would have hidden that; the minimum shouts it.

The numbers to place yourself

Levels are fine. Numbers are better — especially in front of an owner who wants to know whether he's good or not. Here are the orders of magnitude commonly used in manufacturing, with the caveat that has to come with them: these are industry reference points, not normative thresholds. No auditor will withhold your certificate because you're running at 180 defective parts per million. These figures are there to place you against your peers, not to judge you.

What gets measuredIndustry orders of magnitude
Defective parts per million (DPPM)
precision machining
Under 100: the level of the best. From 100 to 500: competitive. Above 500: there's ground to gain.
Cost of poor quality
as a % of revenue
Manufacturers typically fall between 5 and 35%. Under 5%, you're among the best. Most owners don't know their own figure.
On-time delivery (OTD) 98% and above among tier-one aerospace and automotive suppliers. Often a contractual requirement rather than an internal goal.
Time to close a corrective action Under 30 days for a routine case, under 10 for a critical one. Beyond that, it isn't a corrective action any more — it's an open file.
Recurrence rate
same cause coming back
Under 10% in a mature system. Between 30 and 50% where action effectiveness is never verified — which is the most common case.

That last figure deserves a pause, because it's the most revealing of them all. A high recurrence rate doesn't mean your people work badly. It means you close files without checking they're actually fixed. That's a difference of method, not of competence — and it's one of the few things you can correct without hiring anyone.

As for the cost of poor quality, that's the one I'd recommend calculating first if you only calculate one. Not because it's the most precise, but because it's the only one that speaks to everyone in the company. Scrap, rework, recovery hours, customer returns, express freight to catch up a late shipment: add up a year, divide by your revenue. The result surprises people, every time.

Why most self-assessments are worthless

Now the bad news. These scales have existed for a long time. ISO 9004 has offered its self-assessment questionnaire for years. And in practice, almost nobody uses it — or rather, people use it badly.

The scenario is always the same, documented by practitioners as much as I've seen it in the field. The quality manager receives the grid. He fills it in alone, on a Friday afternoon, before the audit. He consults neither the process owners nor senior management, because that would take three meetings he won't get. He ticks boxes based on what he believes he knows and on what would be awkward to admit. The result goes in the binder.

This isn't bad faith. It's the logical consequence of a declarative tool: when the score depends on what you say about yourself, it measures your optimism, not your system.

And there's something worse than optimism: age. A self-assessment done in March says nothing about what you are in October. Your quality system moves every day — nonconformities open, actions close, deadlines slip. An annual snapshot of something that moves continuously is a souvenir, not an instrument.

The conclusion is direct: the level shouldn't be declared, it should be calculated. How many distinct people raised an issue over twelve months? The system knows. What share of nonconformities is tied to a process? The system knows. What's the median time to close a corrective action? The system knows that too — and it has no reason to round it in the flattering direction.

It's a complete reversal. You no longer ask the quality manager to rate himself: you show him what his own data says about him. And because the data changes every day, the score changes every day.

Figure 3 — What's missing, axis by axis

Axis and level reachedWhat the data says
Detection — Controlled Six distinct people raised an issue over twelve months. Eighteen nonconformities are waiting to be validated, five of them for more than a week.
Traceability — Controlled 91% of nonconformities are tied to a process. On the other hand, only 12% of corrective actions trace back to a risk analysis.
Responsiveness — Planned Median time to close: 34 days. Effectiveness verified before closing: 58% of files.
Prevention — Planned Three processes out of eleven are covered by a risk analysis. Two preventive actions opened this year.
Steering — Optimized Indicators current, management review held and its findings locked. Nothing is blocking at this level.
Improvement — Controlled 64% of review decisions are converted into actions with an owner and a date. Four families of nonconformities recur more than three times.

Not one of these lines is an opinion: they're counts. That's what makes them arguable in front of an auditor — in the good sense, since you can verify where every figure comes from. Illustrative data.

Where AI helps here — and where it has no business

I'm raising this because the question is coming, and because the answer is less obvious than it looks.

Let's start with what matters most: AI must not hand you your maturity level. Ever. That isn't a matter of principle, it's a matter of soundness. A maturity level is a deterministic calculation — fixed rules applied to real data. If AI "estimates" you're at level 3, you inherit two serious problems. First, you can't prove anything: faced with an auditor asking how that number was obtained, "the model estimated it" is not an acceptable answer. Second, the number isn't reproducible: run the analysis again next month with no data changed, and it may well have moved.

An indicator that serves as evidence has to be reproducible and explainable line by line. It's the same principle as product release: what carries your liability isn't entrusted to an intuition, however statistical.

That said — and this is the interesting part — there's real work left for AI, downstream of the calculation.

Grouping what the human eye no longer sees. Forty nonconformities spread over twelve months, entered by eight different people each with their own phrasing, can describe exactly the same problem without anyone noticing. Nobody rereads twelve months of history looking for similarities of vocabulary. AI does. When Figure 3 reports "four families of recurring nonconformities," that's precisely this matching work — and a human then validates the grouping, because similarity of words is not similarity of causes.

Drafting. The analysis behind a corrective action, the email explaining the situation to a customer, the minutes of the management review: these are long texts, repetitive in structure, that nobody enjoys writing. AI produces a first version in thirty seconds. The human corrects it, owns it and signs it. Liability doesn't change hands — only the blank page disappears.

Turning a gap into actions. "Prevention: level 2" helps nobody. "Here are the three processes covered by no risk analysis, two of which involve critical parts" helps immediately. Again, the finding comes from the calculation; AI does the work of putting it into plain language.

AI helps you understand your score and act on it. It doesn't hand it to you.

It's the same boundary as everywhere else in well-built trade software: the machine prepares, the human decides, and the software enforces the frame that keeps the two from being confused.

The real brake isn't the tool

Now for something articles on quality maturity carefully avoid, because no software fixes it: the main obstacle to reaching levels 3 and 4 isn't technical. It's cultural.

Go back to the first question, the one about how many people raise an issue. When the answer is "six out of thirty," it's almost never a problem of access to the tool. It's that speaking up costs something. Sometimes a remark in front of the others. Sometimes the belief, grounded in experience, that it will land on someone — often on the person who spoke.

I've seen this mechanism in nearly every shop I've worked in, and it's rarely the doing of a harsh owner. It settles in by small touches. A "you again?" said without thinking. A meeting that looks for who made the mistake before looking at why it was possible. A corrective action closed with "operator made aware," which names a culprit and corrects nothing. Nobody decided to install fear. It installed itself, and it holds up very well.

The result is paradoxical, and that's what makes it dangerous: the shop that reports few issues looks healthier than the one that reports many. The numbers are prettier, the dashboards greener. Except the first one doesn't know what's happening inside it, and the second does. A quality system moving up a level first sees its nonconformities rise. That's the sign it's starting to work, not that it's deteriorating — and it's a message an owner needs to hear before launching the effort, not six months in.

A quiet shop isn't a shop without problems. It's a shop where problems don't travel upward.

What unlocks this comes down to a few things, all of them demanding. The first is a simple rule, stated by management and above all held to over time: nobody is ever punished for reporting an issue — including their own. It's worth nothing until it has been publicly tested, the first time an employee reports their own mistake. That day, everyone watches what the boss does. The answer he gives sets the culture for the next two years.

The second is a shift in vocabulary that looks cosmetic and isn't. A nonconformity isn't attached to a person: it's attached to a process. That's not politeness, it's the only way to get a usable root-cause analysis — and it's exactly what the traceability axis measures. A system that looks for who is responsible produces culprits; a system that looks for the cause produces improvements.

The third, and most concrete: make visible what becomes of reports. Nothing kills reporting faster than an issue raised that vanishes into a form with no reply. If the person who raised it sees the action, its owner and its date — then, three months later, that it was verified — they'll do it again. If not, they'll draw the obvious conclusion and stop.

Software creates none of these three things. But it makes them measurable, which is no small matter: the number of distinct reporters, the share of nonconformities tied to a process rather than to a name, the proportion of actions whose effectiveness was verified. Those are three culture indicators dressed as quality indicators. When they rise, it isn't the tool improving — it's people starting to speak again.

And if you're at level 2?

You probably are. That's where the vast majority of certified small manufacturers sit, and I'd rather say so plainly than let you believe otherwise.

Three things to know before getting discouraged.

Progress isn't linear. Research on maturity models in smaller companies is clear on this: you don't climb one level a year across every axis. You move sharply on two, stall on three, and slide back on one when a key person leaves. That's normal. A model promising a steady ascent is describing a company that doesn't exist.

The step from 1 to 2 is the most profitable of them all. Going from "we endure" to "we record and follow up" requires no sophisticated tool and no new hire. It requires that issues be logged as they happen, and that someone look at them every week. The return on that effort far exceeds anything the later refinements will bring.

A shop of thirty doesn't have to play in the same league as an aircraft maker. The frameworks I've cited — ISO 9004, the aerospace models, the great industrial excellence prizes — were built for organizations with a full quality department. Researchers who have studied their use in smaller companies all reach the same conclusion: applied as-is, they're out of reach and demoralizing. A scale is only useful if it tells you what to do on Monday morning.


In closing: three questions, your numbers

I won't ask you to rate yourself — I've spent a whole article explaining that it's worthless. Instead I'll leave you three questions whose answers already sit somewhere in your files.

How many different people raised an issue at your shop this year? If it's fewer than five in a shop of thirty, your problem isn't the quality of your production. It's that you only see a fraction of what happens.

Among your nonconformities of the last twelve months, how many come back from a cause you had already dealt with? If you can't answer, that's an answer in itself: nobody is checking that corrective actions work.

How many of your processes went through a risk analysis before a problem occurred? That's the axis where almost everyone plateaus, and the one that sets your overall level.

If those three answers take you more than half a day to reconstruct, you've already learned something about your maturity level — and that's no disgrace. It's everyone's starting point. What matters is that next time, the answer sits one screen away rather than at the bottom of a filing cabinet.

The principles described in this article are the ones that guided the development of Asterion Solutions, a suite of trade-specific software built for small manufacturers who want to structure their quality without multiplying administrative tasks.

Free resource

Checklist: passing your ISO 9001 audit as an SME

Clause by clause, what an auditor will actually ask — plus the 3 questions they almost always ask.