Tensorwire
Products & tools · first seen 11h ago, updated 11h ago

GPT-5.6 Quietly Broke Our Prompt Cache. One Message Boundary Fixed It.

1 outlet GPT & ChatGPT

How a model upgrade took our cache hit rate from ~90% to almost nothing, and why the fix was about where our prompt text sits across messages, not what it says. TL;DR GPT-4o caches any identical prefix in 128-token steps. It doesn’t care wh…

Summary from Towards AI.

Coverage 1 article · 1 outlet

  1. Towards AI
    GPT-5.6 Quietly Broke Our Prompt Cache. One Message Boundary Fixed It.