Proof

How we validate

Every figure on this site comes from data the client already had, tested on periods the model never saw. This page is the method — including the machines where it did not work.

The method

Calibrate once, then replay month by month.

Calibration against replayROLLING VALIDATION
CALIBRATIONseen oncenever seen by the modelSCOREDEvery figure we publish is measured on the right-hand side of this line,never on the calibration period.
Calibration
Calibrated on one full year, on data never seen during calibration.
Replay
Replayed over 12 rolling windows — retrained each month, tested on the next, each month evaluated once out of sample.
Baseline
Every configuration compared with random alerting at the same alert budget, so a result has to beat chance at equal cost.
Causality
Only the past is ever used to produce a score, so what is validated offline is what runs in real time.
Where it pays, and where it does not

Two of four machines paid off. We publish the other two.

From the aerospace and defence line. Neither V1 nor V3 is priced into the case.

MachineDowntime anticipatedBreak-even per false alertVerdict
V4217 h · 56%5.9 hpays off under any assumption
V247 h · 50%37 minpays off if a simple inspection
V16 h · 16%~5 minnot cost-effective
V33 h · 4%~6 minno usable signal in the log

On V3 the major breakdowns are preceded by no build-up of micro-stops. There is nothing to see in the log, and it is not a tuning problem. We state that in the deliverable rather than tuning until something appears.

Alert load across the line runs at two to four alerts per machine per week, and roughly one alert in six is followed by a serious stop within the hour. That ratio is in the deliverable too, because it is what decides whether the alerts get acted on.

On the optimisation side

The zone comes first. Then we count the runs inside it.

Not a prediction

The recommended zone was delimited first. Only then did we look at which past production runs fell inside it, and what they actually yielded.

A statistical floor per zone

Each zone carries its own floor. Across the full 336-run history, no run inside a zone ever fell below that zone’s floor.

Within reach of the line

On a representative run, 8 of the 10 retained levers were already inside the recommended range — the zone is not a different factory.

Extrapolation is labelled

Any range that goes beyond what the line has already produced is flagged as extrapolation rather than presented as a finding.

Not every lever is a setpoint

Ambient temperature is observed, not set. It guides when to produce rather than what to adjust, and we say which is which.

Confirmable in a week

One run takes about five hours. Running the next productions inside the zone and comparing yield to history gives a first verdict within a week.

What the test can return

A validated case, or a documented no.

Both are real outcomes. The second one you keep too.

Next step

Name one line and one target variable.

We read the export you already produce and tell you within a week whether it carries the levers — before you commit to anything.

Request a demo

We use this only to reply to you. No mailing list, no resale.