Tensorwire
Products & tools · first seen 15 Sep, updated 15 Sep

Optimizing cost and latency with Amazon Bedrock prompt caching

1 outlet Amazon

Prompt caching in Amazon Bedrock can cut input token costs by up to 90% when you repeatedly send the same context to foundation models. This post walks through six practical prompt caching scenarios using the Converse API: message content,…

Summary from AWS Machine Learning Blog.

Coverage 1 article · 1 outlet

  1. AWS Machine Learning Blog
    Optimizing cost and latency with Amazon Bedrock prompt caching