Products & tools · first seen 27 Aug, updated 27 Aug
Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2
Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts GPU…
Summary from AWS Machine Learning Blog.
Coverage 1 article · 1 outlet
-
AWS Machine Learning BlogReduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2