AWS Launches 13 SageMaker AI Inference Capabilities in 2026, Introduces GPU-Aware Routing Gateway

Amazon announced 13 new SageMaker AI inference features in 2026, including the HyperPod Inference Gateway for GPU-aware routing.

AWS Launches 13 SageMaker AI Inference Capabilities in 2026, Introduces GPU-Aware Routing Gateway

According to aws.amazon.com, Amazon SageMaker AI delivered 13 new capabilities in 2026 year-to-date across two deployment paths: managed endpoints and Amazon SageMaker HyperPod Inference. The announcement comes as generative AI inference presents unique challenges, with models spanning “tens to hundreds of gigabytes” and latency requirements “measured in tokens per second,” according to the same source.

The two deployment paths serve different use cases, according to aws.amazon.com. Managed endpoints are “fully managed by AWS” and best for “fast and fully managed deployment with minimal ops overhead,” while HyperPod offers “Kubernetes-native control over dedicated GPU clusters” suited for “train-to-serve multi-cloud/hybrid-cloud deployments.”

A key launch is the Amazon SageMaker HyperPod Inference Gateway, announced September 18, 2026, according to aws.amazon.com. The gateway addresses GPU waste caused by “round-robin and least-connections algorithms” that lack “visibility into what’s happening inside your GPUs.” The solution “deploys as a single EKS managed addon” and uses “real-time GPU signals to place every inference request on the best-suited pod,” according to the source.

According to aws.amazon.com, the gateway can reduce “first-token latency by up to 82%,” with one example showing chatbot response time dropping from “4.4 seconds for the first token” to “under 800 ms.”