Tensorwire
Products & tools · first seen 27 Aug, updated 27 Aug

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

1 outlet Nvidia Amazon

Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts GPU…

Summary from AWS Machine Learning Blog.

Coverage 1 article · 1 outlet

  1. AWS Machine Learning Blog
    Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2