Release date 2026/07/02
Bugfix
Emitted the documented
X-AI-RateLimit-*headers on429responses from policies that partition byproviderormodel.Fixed an issue with token counters being decremented on AI Semantic Cache hits.
Fixed an issue where
modelorprovidermatchers withoutpartition_bywere skipped, causing rate limits to be missed.Fixed policy integration with
ai-proxyfor single-target and multi-target AI proxy routes.