My Inferentia port matched its reference token-for-token — and still output garbage
7.6 relevance
Score Breakdown
technical depth 9
novelty 7
actionability 7
community 5
strategic 7
personal 9
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Another Inferentia porting debugging story with a validation twist.
Summary
Porting Google's Gemma-4 31B dense model (60 layers, ~60GB bf16) to AWS Inferentia2 (inf2.24xlarge, TP=8) produced token-for-token identical output to CPU but was gibberish. Manual parallel_model_trace failed due to OOM from 8 simultaneous fp32 compiles and deadlock on serialized ranks, forcing a pivot to ModelBuilder which compiles one rank and loads weights per rank. Attention layout issues emerged: 50 sliding layers (head_dim 256, 16 KV heads) and 10 global layers (head_dim 512, 4 KV heads) — the 4 KV heads cannot be sharded across 8 ranks, requiring replication.