Scale was not the hardest part. Context was.
Two machines of the same model behaved differently because load, maintenance history and calibration varied. A global threshold produced alert storms, while connectivity gaps erased the lead-up to failures.
Normal operating envelopes changed by asset, shift and product line.
Events had to survive hours without reliable WAN connectivity.
Engineers ignored a system that could not explain why a warning fired.
Compress locally; learn globally.
Protocol normalization and local buffer
Rolling vibration, thermal and load features
Fleet joins and ordered processing
Hierarchy-aware anomaly inference
Evidence, severity and recommended check
Warnings needed an operator contract.
Fleet patterns supplied a prior, then asset-specific residuals captured local behavior.
Each warning included changed features, a comparison window and the similar historical fault.
Work-order outcomes became labels instead of leaving model truth inside a separate system.
Fewer alerts, earlier interventions.
| Measure | Before | After |
|---|---|---|
| Telemetry capacity | Regional silos | 1B+ events/day |
| False alert rate | High / untracked | 63% lower |
| Connectivity loss | Data gaps | Buffered + replayed |
| Unplanned downtime | Baseline | 38% lower |
Move from telemetry storage to operational intelligence.
We can map the asset hierarchy, signal path and edge failure modes before choosing a model.