Field notes
What we are learning, in public.
Engineering deep-dives, design essays, and customer stories from the team building the GPU-native AI runtime.
Engineering
Disaggregated serving: how NVIDIA Dynamo changes inference economics
How prefill/decode disaggregation on NVLink clusters cuts first-token latency by 40%.
Jonas Vetter · May 18, 2026
DesignThe flat UI thesis: less affordance, more authority
Why we removed shadows, gradients, and animation from the GPU dashboard.
Priya Subramanian · May 11, 2026
ProductPolicy-as-code for inference: a DSL for GPU budget governance
Three failure modes we saw in customer rollouts — and how a typed policy layer fixed runaway GPU spend.
Hannah Schulze · May 4, 2026
CustomerNorthBank's 62% inference cost reduction on H100s
A case study on NVIDIA GPU fleet governance in regulated finance.
Oliver Reed · April 28, 2026
AITensorRT-LLM compilation in production: a 90-day field report
How quantization and fused kernels improved throughput 4x on Llama 3.1-70B.
Daniel Park · April 21, 2026
SecurityPrompt injection against NIM endpoints: a 90-day field report
What actually happens when inference APIs face hostile inputs.
Renata Lima · April 14, 2026