Machine LearningAI Research16 min reading time

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

AWS Blog
Read full post
Heidi Health, AWS, and NVIDIA collaborated to reduce automatic speech recognition (ASR) inference costs by 75% on Amazon EC2 using NVIDIA CUDA Multi-Process Service (MPS) and Triton Inference Server. This approach increased GPU utilization from 15-20% per request to supporting 92.1 requests per second per GPU, cutting required GPU instances from 16 to 4 while maintaining sub-second latency.

More in Machine Learning

Machine Learning4 min read

DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro

SiliconANGLE
Machine Learning4 min read

Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

The New Stack (AI)
Machine Learning3 min read

OpenAI Releases GPT-6 Astra for Coding and Computer Use

InfoQ (AI, ML & Data)