Amazon SageMaker AI has delivered 13 inference launches year-to-date. According to an official review published by AWS, these technical rollouts are organized across two distinct operational deployment paths. Those supported paths consist of fully managed endpoints and Amazon SageMaker HyperPod Inference.
The updates encompass several specialized tools for model execution. AWS notes that the launches extend from inference recommendations to capacity-aware instance pools. The additions also incorporate architectural capabilities including tiered KV caching alongside disaggregated prefill and decode.

