The Response Rate Limiting Policy counts arbitrary units of usage that the upstream service reports back in a response header. This lets an upstream with variable per-request costs report its own usage instead of AI Gateway inferring cost implicitly. You can define as many named limits as you want, and instruct the upstream to increment them by any number of units.
Response Rate Limiting Policy
How it works
Define one or more limits in config.limits, each with its own second, minute, hour, day, month, or year threshold. From your upstream service, report usage against those names using a response header in the form:
<header-name>: <limit-name>=<value>[,<limit-name>=<value>]By default, the header is named x-kong-limit (config.header_name). For example, to report 1 unit of usage against a limit named video:
x-kong-limit: video=1AI Gateway removes this header before returning the response to the original client, and increments the named counter by the reported value.
This Policy doesn’t prevent the upstream from being called once a limit is reached. Every request, including the one that trips the limit, is still proxied to the upstream.
Example: Rate limit on usage reported by a self-hosted model backend
You can use this Policy in AI Gateway with a self-hosted, OpenAI-API-compatible backend, such as a vllm-type AI Model Provider, that reports its own usage on every response. The following configuration allows 2 units of a test limit per minute, aggregated by client IP address:
ai_gateway_policies:
- ref: my-response-ratelimiting
ai_gateway: !lookup {id: !env AI_GATEWAY_ID}
display_name: my-response-ratelimiting
name: my-response-ratelimiting
type: response-ratelimiting
enabled: true
global: false
config:
limit_by: ip
policy: local
limits:
test:
minute: 2Make sure to replace the following placeholders with your own values:
AI_GATEWAY_ID: Theidof your AI Gateway.
Once the upstream has reported 2 units of the test limit within a minute (for example, by returning x-kong-limit: test=1 on each of its first two responses), the third request still reaches the upstream, but AI Gateway discards that response and rejects the request with 429 Too Many Requests instead.
Strategies
Use config.policy to choose how counters are stored:
local: Counters are stored in-memory on the node. Minimal performance impact, but not shared across AI Gateway nodes.redis: Counters are stored on a Redis server and shared across nodes.
Using cloud authentication with Redis
If your AI Policy uses a Redis datastore, you can authenticate to it with a cloud Redis provider. This allows you to rotate credentials without relying on static passwords.
The following providers are supported:
- AWS ElastiCache
- Azure Managed Redis
- Google Cloud Memorystore (with or without Valkey)
Each provider also supports an instance and cluster configuration.
Fallback from Redis
When the redis strategy is used and an AI Gateway node is disconnected from Redis, the plugin will fall back to local rate limiting.
This can happen when the Redis server is down or the connection to Redis is broken.
AI Gateway keeps the local counters for rate limiting and syncs with Redis once the connection is re-established.
AI Gateway will still rate limit, but the AI Gateway nodes can’t sync the counters. As a result, users will be able
to perform more requests than the limit, but there will still be a limit per node.
Limit by
Use config.limit_by to choose what the Policy aggregates counters against: consumer, credential, or ip. If the AI Consumer or credential can’t be determined, AI Gateway falls back to ip.
Limit by IP address
If limiting by IP address, it’s important to understand how AI Gateway determines the IP address of an incoming request.
The IP address is extracted from the request headers sent to AI Gateway by downstream clients. Typically, these headers are named X-Real-IP or X-Forwarded-For.
By default, AI Gateway uses the header name X-Real-IP to identify the client’s IP address. If your environment requires a different header, you can specify this by setting the real_ip_header Nginx property. Depending on your network setup, you may also need to configure the trusted_ips Nginx property to include the load balancer IP address. This ensures that AI Gateway correctly interprets the client’s IP address, even when the request passes through multiple network layers.
Headers sent to the client
AI Gateway sends additional headers back to the client, indicating how many units are still available and how many are allowed in total. For example, for a limit named video with a per-minute threshold:
X-RateLimit-Limit-Video-Minute: 10
X-RateLimit-Remaining-Video-Minute: 9If more than one limit or time window is configured, AI Gateway returns a combination of all of them. You can hide these headers with config.hide_client_headers.
Headers sent to the upstream
AI Gateway also appends usage headers for each limit before proxying the request to the upstream service, so the upstream can decide whether to process the request at all when no limits remain. These headers are in the form X-RateLimit-Remaining-<LIMIT_NAME>, for example:
X-RateLimit-Remaining-Video: 3
X-RateLimit-Remaining-Image: 0