Products & tools · first seen 11h ago, updated 11h ago
GPT-5.6 Quietly Broke Our Prompt Cache. One Message Boundary Fixed It.
How a model upgrade took our cache hit rate from ~90% to almost nothing, and why the fix was about where our prompt text sits across messages, not what it says. TL;DR GPT-4o caches any identical prefix in 128-token steps. It doesn’t care wh…
Summary from Towards AI.