Hardware/SoftwareUnited States2023
Crash observability and OTA for a Cortex-M fleet
Coredump capture, symbolication, and staged OTA that turn field failures into fixed builds.
- Client
- Memfault
- Region
- United States
- Sector
- Embedded observability + OTA
- Engagement
- 2023 · 15 weeks
- Team
- 2 firmware · 2 backend · 1 platform
- Status
- In production
The challenge
A connected-hardware maker was flying blind on field failures — crashes surfaced only as returns, with no root cause and no safe way to push a fix. They needed on-device fault capture plus a backend for symbolication, metrics, and staged OTA.
What we built
The full pipeline, end to end — not a blurb.
- 01Firmware
A Cortex-M fault handler capturing coredumps and metrics with compact upload and watchdog integration.
- 02Backend
A symbolication service, issue de-duplication, device metrics, and OTA release orchestration.
- 03Product
A crash explorer and fleet-health view with staged rollout and automatic halt.
- 04Operations
Kubernetes, object storage, and alerting.
Results
20 min
crash root-cause time via symbolicated traces, from 3 days
4.2 KB
median coredump upload overhead per event
0
bricked units across staged OTA rollouts
61%
lower field crash rate over three release cycles
Now the release gate for every firmware ship.
Have a constraint like Memfault’s? Bring us yours.
Next engagement
PriceOye
E-commerce web + operations
