Skip to content
All work

AIUnited States2024

INT4/INT8 quantization pipeline for inference edge

A reproducible post-training quantization flow taking vision and sequence models to Jetson-class hardware.

Client
Latent AI
Region
United States
Sector
Edge model quantization
Engagement
2024 · 12 weeks
Team
3 ML · 1 systems
Status
Shipped

The challenge

Customer models were too large and too slow at FP16 for the edge targets they had to run on, and hand-tuned quantization didn’t transfer across a model family. Latent AI needed a repeatable INT4/INT8 flow that held accuracy without retraining.

What we built

The full pipeline, end to end — not a blurb.

  1. 01Calibration

    Calibration profiling with per-channel INT4 weights, INT8 activations, and error compensation.

  2. 02Tooling

    PyTorch → ONNX → TensorRT with custom low-bit kernels and an engine cache.

  3. 03Evaluation

    A zero-leakage split harness, per-layer sensitivity sweeps, and an A/B against FP16.

  4. 04Deployment

    Reproducible builds and target packaging for Jetson.

Results

4.1×

model compression, −0.7 pt accuracy vs. FP16

2.9×

higher throughput on Jetson Orin

63%

lower peak memory, 3.4 GB down to 1.25 GB

7

model families the flow was reused across

Adopted into the customer’s edge deployment pipeline.

Have a constraint like Latent AI’s? Bring us yours.

Next engagement

Memfault

Embedded observability + OTA