2026-09-14

9/14/2026, 12:00:00 AM ~ 9/15/2026, 12:00:00 AM (UTC) Recent Announcements Qwen3.6-35B-A3B-NVFP4 and Wan2.1-T2V-1.3B-Diffusers models now available on Amazon SageMaker JumpStart NVIDIA’s Qwen3.6-35B-A3B-NVFP4 and Alibaba’s Wan2.1-T2V-1.3B-Diffusers models are now available on Amazon SageMaker JumpStart, expanding the portfolio of foundation models available to AWS customers. These two models bring specialized capabilities spanning agentic coding with long-context reasoning and lightweight text-to-video generation, enabling customers to deploy high-performance, scalable AI solutions on AWS infrastructure.\n These models address different enterprise AI challenges with specialized capabilities:...

September 14, 2026

2026-08-11

8/11/2026, 12:00:00 AM ~ 8/12/2026, 12:00:00 AM (UTC) Recent Announcements Amazon EC2 R8a instances are now available in Canada (Central) region Starting today, Amazon EC2 R8a instances are now available in Canada (Central) Region. These instances, feature 5th Gen AMD EPYC processors (formerly code named Turin) with a maximum frequency of 4.5 GHz, deliver up to 30% higher performance, and up to 19% better price-performance compared to R7a instances.\n R8a instances deliver 45% more memory bandwidth compared to R7a instances, making these instances ideal for latency sensitive workloads....

August 11, 2026

2026-07-13

7/13/2026, 12:00:00 AM ~ 7/14/2026, 12:00:00 AM (UTC) Recent Announcements OpenAI GPT-5.6 Sol, Terra, and Luna now generally available on Amazon Bedrock GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock, bringing the smartest family of models from OpenAI yet to Bedrock’s next-generation inference engine built for high-performance, security and reliability. GPT-5.6 sets a new standard for intelligence and efficiency, allowing you to solve harder problems in less time and with more intelligence per token....

July 13, 2026

2026-07-06

7/6/2026, 12:00:00 AM ~ 7/7/2026, 12:00:00 AM (UTC) Recent Announcements Amazon SageMaker HyperPod now supports disaggregated prefill and decode Amazon SageMaker HyperPod now supports Disaggregated Prefill and Decode (DPD), an inference optimization that separates the two phases of large language model (LLM) inference — prefill and decode — onto dedicated GPU pools and transfers the key-value (KV) cache between them over Elastic Fabric Adapter (EFA) using GPU-Direct RDMA. Customers running LLMs in production for chat assistants, agentic pipelines, retrieval-augmented generation, and long-document analysis need consistent per-token latency and predictable throughput under mixed traffic, but when prefill and decode share the same GPU, a single long-context request can stall token generation for every concurrent request and force customers to over-provision one phase to protect the other....

July 6, 2026

2026-06-18

6/18/2026, 12:00:00 AM ~ 6/19/2026, 12:00:00 AM (UTC) Recent Announcements Amazon ECS announces faster service auto scaling Amazon ECS service auto scaling now detects and responds to load changes faster with support for high resolution (20-second) metrics and metric publishing optimizations. In AWS benchmarking tests, time to trigger scale-out improved from 363 seconds to 86 seconds (76% faster, 4.2x), and total time to scale and provision new tasks improved from 386 seconds to 109 seconds (72% faster, 3....

June 18, 2026