
Edge AI Platform for Predictive Maintenance
Klyff
Industrial & Manufacturing
Edge AI
Enterprise Grade
Engineering resilience through Klyff's transformation.
We developed Klyff's edge AI predictive maintenance platform from concept to production deployment across 10+ factories. The platform collects machine vibration, thermal, and acoustic sensor data and runs ONNX-optimized inference models in under 8ms on ARM edge hardware — eliminating the need for cloud round-trips. A federated learning architecture aggregates model improvements across factory nodes without exposing sensitive operational data. A zero-touch OTA pipeline rolls model updates to the entire fleet with no downtime.
The Manufacturing Challenge
- Collect and analyze real-time machine vibration, thermal, and acoustic sensor data in factory environments with sub-10ms response requirements.
- Predict equipment failures before they cause costly production line shutdowns, replacing reactive maintenance.
- Minimize bandwidth and cloud costs by processing data locally at the edge using ONNX-optimized models on ARM hardware.
- Ensure model updates are securely aggregated across all factory nodes without exposing sensitive operational data to the cloud.
The Edge-First Solution
- Designed a federated learning architecture using PyTorch that trains models across distributed ARM edge nodes — model weights aggregate centrally on AWS, never raw sensor data.
- Developed SIMD-accelerated vibration signal processing algorithms in Rust, achieving under 8ms end-to-end inference latency on ARM Cortex-A72 edge hardware.
- Exported trained PyTorch models to ONNX Runtime with INT8 quantization, reducing model size by 65% and inference time by 3x compared to the original float32 models.
- Implemented a zero-touch OTA model deployment pipeline using AWS IoT Core that rolls updated models to 10+ factory device fleets with automatic rollback and health checks.
Operational Success
Reduction in unplanned factory downtime through early anomaly warnings — machines flagged 48–72 hours before failure.
End-to-end edge inference latency on ARM hardware using ONNX Runtime with INT8 quantization.
Model size reduction via ONNX INT8 quantization, enabling deployment on resource-constrained edge hardware.
Factories running the federated edge AI platform in production with zero-touch OTA model updates.
Sub-8ms Edge Inference Across a Federated Factory Fleet
By keeping vibration and thermal signal analysis local using ONNX Runtime on ARM hardware, the system eliminates expensive cloud data ingestion entirely. Federated aggregation continuously improves model accuracy without exposing factory operational secrets. Zero-touch OTA pipelines ensure every factory always runs the latest anomaly detection model with no manual intervention or downtime.


