AWS published a technical guide for deploying Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on SageMaker HyperPod using vLLM, featuring quantization and reasoning capabilities.
AWS released Ray Serve Deep Learning Containers to replace TorchServe, providing a supported, pre-assembled container for GPU inference workloads on Kubernetes clusters.