AIUnited States2024
INT4/INT8 quantization pipeline for inference edge
A reproducible post-training quantization flow taking vision and sequence models to Jetson-class hardware.
- Client
- Latent AI
- Region
- United States
- Sector
- Edge model quantization
- Engagement
- 2024 · 12 weeks
- Team
- 3 ML · 1 systems
- Status
- Shipped
The challenge
Customer models were too large and too slow at FP16 for the edge targets they had to run on, and hand-tuned quantization didn’t transfer across a model family. Latent AI needed a repeatable INT4/INT8 flow that held accuracy without retraining.
What we built
The full pipeline, end to end — not a blurb.
- 01Calibration
Calibration profiling with per-channel INT4 weights, INT8 activations, and error compensation.
- 02Tooling
PyTorch → ONNX → TensorRT with custom low-bit kernels and an engine cache.
- 03Evaluation
A zero-leakage split harness, per-layer sensitivity sweeps, and an A/B against FP16.
- 04Deployment
Reproducible builds and target packaging for Jetson.
Results
4.1×
model compression, −0.7 pt accuracy vs. FP16
2.9×
higher throughput on Jetson Orin
63%
lower peak memory, 3.4 GB down to 1.25 GB
7
model families the flow was reused across
Adopted into the customer’s edge deployment pipeline.
Have a constraint like Latent AI’s? Bring us yours.
Next engagement
Memfault
Embedded observability + OTA
