Autoscaling GPU Inference on Kubernetes: Cold Starts, Scale-to-Zero and Cost per Million Tokens
September 14, 2026 · IntelliSensei Team
Why GPU utilisation is the wrong scaling signal, how to cut a multi-minute GPU cold start to seconds, when scale-to-zero actually pays, and how to track cost per million tokens for a PyTorch inference fleet.
Read more