Dev15 min reading time

Disaggregated prefill and decode for LLM inference on SageMaker HyperPod

AWS Blog
Read full post
Amazon SageMaker HyperPod now supports Disaggregated Prefill and Decode (DPD) for large language model inference using vLLM, separating prefill and decode phases across GPU pools to reduce latency and improve concurrency for long-context workloads.

More in Dev

Introducing the Agents API

Covered by 3 sources
Dev1 min read

Native is now the future of mobile at Shopify

Simon Willison's Weblog
Dev5 min read

AWS open-sources Pizza Bot: email-style inbox for background AI agents

The New Stack (AI)