This feature is currently in Tech Preview and should not be used in a production environment.
Before you use Headroom with the AI Prompt Compressor Policy, you need a Headroom instance accessible to your AI Gateway.
You can do this with one of the following:
To configure an AI Prompt Compressor Policy with Headroom as the compressor service, do the following:
curl -X POST https://{region}.api.konghq.com/v1/ai-gateways/{AIGatewayId}/policies \
--header "accept: application/json" \
--header "Content-Type: application/json" \
--header "Authorization: Bearer $KONNECT_TOKEN" \
--data '
{
"display_name": "AI Prompt Compressor with Headroom",
"name": "my-ai-prompt-compressor",
"type": "ai-prompt-compressor",
"config": {
"compression_ranges": [
{
"min_tokens": 20,
"max_tokens": 100,
"value": 0.8
}
],
"provider": "headroom",
"compressor_url": "http://headroom-service:8787/v1/compress",
"timeout": 45000,
"keepalive_timeout": 60000,
"log_text_data": false,
"stop_on_error": true,
"headroom": {
"proxy_token": "'$HEADROOM_PROXY_TOKEN'",
"ssl_verify": true,
"session_id_headers": [
"x-claude-code-session-id",
"x-claude-code-agent-id",
"thread-id",
"session-id",
"x-session-id"
]
}
}
}
'
ai_gateway_policies:
- ref: my-ai-prompt-compressor
ai_gateway: !lookup {id: !env AI_GATEWAY_ID}
display_name: AI Prompt Compressor with Headroom
name: my-ai-prompt-compressor
type: ai-prompt-compressor
config:
compression_ranges:
- min_tokens: 20
max_tokens: 100
value: 0.8
provider: headroom
compressor_url: http://headroom-service:8787/v1/compress
timeout: 45000
keepalive_timeout: 60000
log_text_data: false
stop_on_error: true
headroom:
proxy_token: !env HEADROOM_PROXY_TOKEN
ssl_verify: true
session_id_headers:
- x-claude-code-session-id
- x-claude-code-agent-id
- thread-id
- session-id
- x-session-id
Make sure to replace the following placeholders with your own values:
For more details, see the configuration reference.
The compressor service exposes a /v1/compress endpoint that compresses messages and returns them. This endpoint accepts OpenAI and Anthropic’s message formats. Requests in unsupported formats are forwarded unchanged. You can use this interface to compress prompts, check the current status, or integrate the service with the AI Prompt Compressor Policy.
Headroom uses loopback-trust by default. It answers unauthenticated calls on 127.0.0.1, and returns 404 to non-loopback callers unless it’s started with HEADROOM_COMPRESS_ALLOW_REMOTE=1. You can run Headroom as a co-located sidecar reachable from the AI Gateway data plane without authentication. Alternatively, you can point the Policy at a remote instance and configure a bearer token, sent as both the X-Headroom-Proxy-Token and Authorization: Bearer headers, set with the headroom.proxy_token configuration field.
The AI Prompt Compressor Policy uses Headroom in a stateful mode and derives a session identifier for each conversation, based on the configured config.headroom.session_id_headers, an AI Consumer, or a credential. Headroom can recognize the turns it has already compressed for that conversation. Recognized turns are replayed unchanged instead of compressed again. This ensures the upstream LLM provider’s cache hits on repeated turns.
If a session identifier isn’t present, then every client that opens with the same prompt shares one session state on the Headroom side. If sessions collide, then only one will hit the upstream cache.
- AI Gateway sends the user or agent’s request to the AI Prompt Compressor.
- The AI Prompt Compressor builds a messages array from the whole conversation.
- The AI Prompt Compressor sends a
POST request to Headroom’s /v1/compress endpoint with the messages array and a session identifier for the conversation.
- If Headroom recognizes the session identifier, it replays the turns it already compressed unchanged and compresses only the new messages. Otherwise, it compresses the whole conversation.
- Headroom returns
200 with the compressed messages and metadata.
If the call to Headroom fails, the AI Prompt Compressor rejects the request by default. For details, see Failure behavior.
The following diagram illustrates how the AI Prompt Compressor Policy processes and compresses incoming prompts using Headroom:
sequenceDiagram
actor User as User/Agent
participant KongAICompressor as AI Prompt Compressor Policy
participant Headroom
participant LLM as LLM Provider
User->>KongAICompressor: Sends request
activate KongAICompressor
KongAICompressor->>Headroom: POST /v1/compress with messages and session ID
alt Session recognized
Headroom->>Headroom: Replay cached turns, compress only new messages
else New session
Headroom->>Headroom: Compress whole conversation
end
alt Compression succeeds
Headroom-->>KongAICompressor: Return 200 with compressed messages and metadata
KongAICompressor->>LLM: Send compressed messages to upstream provider
LLM-->>KongAICompressor: Return response
KongAICompressor-->>User: Return response
else Compression fails
Headroom-->>KongAICompressor: Return failure response
KongAICompressor--xUser: Return HTTP 500, request not forwarded
end
deactivate KongAICompressor
By default, if the call to Headroom fails, times out, or returns an error, the Policy returns an HTTP 500 to the client rather than forwarding the request uncompressed.
A 503 response indicating a compression timeout is retried automatically before the Policy gives up; every other non-200 response (for example a 400 from a bad configuration, or a 401 from a missing or incorrect bearer token) fails immediately.
- Deterministic compression requires sessions which is only available in Headroom v0.37.0 or newer. Older images silently ignore the session id and run stateless.
- Headroom sessions are held in memory on a single Headroom process. You must point every data plane node at its own sidecar, or all of them at one shared instance. Never point an AI Gateway data plane at a load-balanced set of Headroom instances, since a session’s turns must all reach the same process.
- No MCP tool-response compression: only the LLM request path is supported.
- Only requests with a
messages or input conversation array are sent to Headroom. Other formats are forwarded uncompressed.