Response Rate Limiting Policy

The Response Rate Limiting Policy counts arbitrary units of usage that the upstream service reports back in a response header. This lets an upstream with variable per-request costs report its own usage instead of AI Gateway inferring cost implicitly. You can define as many named limits as you want, and instruct the upstream to increment them by any number of units.

How it works

Define one or more limits in config.limits, each with its own second, minute, hour, day, month, or year threshold. From your upstream service, report usage against those names using a response header in the form:

<header-name>: <limit-name>=<value>[,<limit-name>=<value>]

By default, the header is named x-kong-limit (config.header_name). For example, to report 1 unit of usage against a limit named video:

x-kong-limit: video=1

AI Gateway removes this header before returning the response to the original client, and increments the named counter by the reported value.

This Policy doesn’t prevent the upstream from being called once a limit is reached. Every request, including the one that trips the limit, is still proxied to the upstream.

Example: Rate limit on usage reported by a self-hosted model backend

You can use this Policy in AI Gateway with a self-hosted, OpenAI-API-compatible backend, such as a vllm-type AI Model Provider, that reports its own usage on every response. The following configuration allows 2 units of a test limit per minute, aggregated by client IP address:

kongctl
policy.yaml
ai_gateway_policies:
  - ref: my-response-ratelimiting
    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}
    display_name: my-response-ratelimiting
    name: my-response-ratelimiting
    type: response-ratelimiting
    enabled: true
    global: false
    config:
      limit_by: ip
      policy: local
      limits:
        test:
          minute: 2

Make sure to replace the following placeholders with your own values:

  • AI_GATEWAY_ID: The id of your AI Gateway.

Once the upstream has reported 2 units of the test limit within a minute (for example, by returning x-kong-limit: test=1 on each of its first two responses), the third request still reaches the upstream, but AI Gateway discards that response and rejects the request with 429 Too Many Requests instead.

Strategies

Use config.policy to choose how counters are stored:

  • local: Counters are stored in-memory on the node. Minimal performance impact, but not shared across AI Gateway nodes.
  • redis: Counters are stored on a Redis server and shared across nodes.

Using cloud authentication with Redis

If your AI Policy uses a Redis datastore, you can authenticate to it with a cloud Redis provider. This allows you to rotate credentials without relying on static passwords.

The following providers are supported:

  • AWS ElastiCache
  • Azure Managed Redis
  • Google Cloud Memorystore (with or without Valkey)

Each provider also supports an instance and cluster configuration.

Fallback from Redis

When the redis strategy is used and an AI Gateway node is disconnected from Redis, the plugin will fall back to local rate limiting. This can happen when the Redis server is down or the connection to Redis is broken. AI Gateway keeps the local counters for rate limiting and syncs with Redis once the connection is re-established. AI Gateway will still rate limit, but the AI Gateway nodes can’t sync the counters. As a result, users will be able to perform more requests than the limit, but there will still be a limit per node.

Limit by

Use config.limit_by to choose what the Policy aggregates counters against: consumer, credential, or ip. If the AI Consumer or credential can’t be determined, AI Gateway falls back to ip.

Limit by IP address

If limiting by IP address, it’s important to understand how AI Gateway determines the IP address of an incoming request.

The IP address is extracted from the request headers sent to AI Gateway by downstream clients. Typically, these headers are named X-Real-IP or X-Forwarded-For.

By default, AI Gateway uses the header name X-Real-IP to identify the client’s IP address. If your environment requires a different header, you can specify this by setting the real_ip_header Nginx property. Depending on your network setup, you may also need to configure the trusted_ips Nginx property to include the load balancer IP address. This ensures that AI Gateway correctly interprets the client’s IP address, even when the request passes through multiple network layers.

Headers sent to the client

AI Gateway sends additional headers back to the client, indicating how many units are still available and how many are allowed in total. For example, for a limit named video with a per-minute threshold:

X-RateLimit-Limit-Video-Minute: 10
X-RateLimit-Remaining-Video-Minute: 9

If more than one limit or time window is configured, AI Gateway returns a combination of all of them. You can hide these headers with config.hide_client_headers.

Headers sent to the upstream

AI Gateway also appends usage headers for each limit before proxying the request to the upstream service, so the upstream can decide whether to process the request at all when no limits remain. These headers are in the form X-RateLimit-Remaining-<LIMIT_NAME>, for example:

X-RateLimit-Remaining-Video: 3
X-RateLimit-Remaining-Image: 0

Help us make these docs great!

Kong Developer docs are open source. If you find these useful and want to make them better, contribute today!