Industrial IoTStreaming MLAWS

One billion daily signals, reduced to the next useful action.

The client did not need more sensor charts. Plant engineers needed a small number of explainable warnings that survived poor connectivity and respected the differences between machines.

1B+Events per day
11Plants
38%Less downtime
63%Fewer false alerts
01 · Constraint

Scale was not the hardest part. Context was.

Two machines of the same model behaved differently because load, maintenance history and calibration varied. A global threshold produced alert storms, while connectivity gaps erased the lead-up to failures.

Plant variance

Normal operating envelopes changed by asset, shift and product line.

Disconnected edges

Events had to survive hours without reliable WAN connectivity.

Alert fatigue

Engineers ignored a system that could not explain why a warning fired.

02 · Signal path

Compress locally; learn globally.

SenseEdge gateway

Protocol normalization and local buffer

ShapeFeature windows

Rolling vibration, thermal and load features

StreamKinesis + Flink

Fleet joins and ordered processing

ScoreAsset models

Hierarchy-aware anomaly inference

ActMaintenance queue

Evidence, severity and recommended check

Failure path: gateways kept a bounded local buffer and continued a conservative ruleset when cloud scoring was unavailable; replay restored the fleet timeline after reconnection.
03 · Decisions

Warnings needed an operator contract.

D-01Hierarchical models

Fleet patterns supplied a prior, then asset-specific residuals captured local behavior.

D-02Evidence with every alert

Each warning included changed features, a comparison window and the similar historical fault.

D-03Feedback from maintenance

Work-order outcomes became labels instead of leaving model truth inside a separate system.

04 · Outcome

Fewer alerts, earlier interventions.

MeasureBeforeAfter
Telemetry capacityRegional silos1B+ events/day
False alert rateHigh / untracked63% lower
Connectivity lossData gapsBuffered + replayed
Unplanned downtimeBaseline38% lower
AWS IoT CoreKinesisApache FlinkTimestreamSageMakerONNXGrafana
Adoption lesson: engineers trusted the model only after the alert showed its evidence and maintenance feedback visibly improved the next prediction.

Move from telemetry storage to operational intelligence.

We can map the asset hierarchy, signal path and edge failure modes before choosing a model.

Design my IoT intelligence path →