AI Prompt Compressor Policy

Related Documentation
Made by
Kong Inc.
Incompatible with
on-prem
Minimum Version
AI Gateway - 2.0
Previous Versions of this page

The AI Prompt Compressor Policy compresses messages before sending them to a Large Language Model (LLM), reducing text length while preserving meaning. It supports multiple cache-aware compression providers.

The following compression providers are available:

  • kong: Use the LLMLingua 2 library to compress prose in user messages.
  • headroom: Use Headroom to compress agent traffic such as tool results, large JSON payloads, search output, and logs returned by tools.

Why use prompt compression

Efficient prompt compression helps you manage token limits, cut costs, and speed up LLM requests, all while keeping sensitive data safe and your prompts focused.

The following table outlines common use cases for the AI Prompt Compressor Policy and the configuration options available to tailor its behavior.

Use case

Description

Cost reduction Reducing token count in prompts decreases API costs when calling large language models, especially for high-volume use cases.
Token limit management Compress verbose inputs like chat history or documents to stay within the LLM’s context window. Prevents truncation of important content.
Latency reduction Smaller prompts result in faster request/response cycles, improving performance for real-time applications like voice assistants.
Dynamic prompt optimization Automatically strip verbose or low-value content before sending to the LLM, keeping the focus on what’s most relevant.

Deterministic compression

Prompt compression is a powerful lever to reduce costs by shrinking the number of input tokens sent to the LLM provider. However, standard compression techniques can alter the text slightly on each run. Because LLM providers offer native prompt caching that relies on exact prefix matches, any variation in the compressed payload triggers a cache miss, forcing you to pay full price for those tokens. Deterministic compression solves this by ensuring that identical inputs always produce identical compressed outputs. This avoids cache misses and guarantees a cache hit on the provider’s side, allowing the AI Gateway to stack the cost-saving benefits of both input token reduction and provider-level caching.

LLMLingua compressor service

Kong provides a Docker image for a compressor service, which compresses LLM prompts before sending them upstream. It uses LLMLingua 2 to reduce prompt size, which helps you manage token limits and maintain context fidelity. The compressor service supports both HTTP and JSON-RPC APIs and is designed to work with the AI Prompt Compressor Policy in AI Gateway.

Kong provides the AI Prompt Compressor Service as a private Docker image in a Cloudsmith repository. Contact Kong Support to get access to it.

Once you’ve received your Cloudsmith access token, run the following commands in Docker to pull the image:

  1. To pull images, you must authenticate first with the token provided by Kong Support:

     docker login docker.cloudsmith.io
  2. Docker will then prompt you to enter username and password:

     Username: kong/ai-compress
     Password: <YOUR_TOKEN>

    This is a token-based login with read-only access. You can pull images but not push them. Contact Kong Support for your token.

  3. To pull an image, run docker pull with the appropriate image and version tag, for example:

     docker pull docker.cloudsmith.io/kong/ai-compress/service:v0.0.3
  4. You can now run the image by pasting the following command in Docker:

     docker run --rm -p 8080:8080 docker.cloudsmith.io/kong/ai-compress/service:v0.0.3

Image configuration options

You can configure the compressor service image using environment variables. These affect model selection, hardware usage, logging, and worker behavior.

Configuration option

Description

LLMLINGUA_MODEL_NAME Specifies the LLMLingua 2 model to use for compression. Defaults to microsoft/llmlingua-2-xlm-roberta-large-meetingbank.
LLMLINGUA_DEVICE_MAP Device on which to run the model. Supported values include cpu, cuda, auto, or mps.
LLMLINGUA_LOG_LEVEL Log level for the LLMLingua compression logic. Set to info, debug, or warning based on your needs.
GUNICORN_WORKERS Number of Gunicorn worker processes (for Docker deployments only). Defaults to 2.
GUNICORN_LOG_LEVEL Log level for Gunicorn server output (for Docker deployments only). Defaults to info.

Compressor Service endpoints

The compressor service exposes both REST and JSON-RPC endpoints. You can use these interfaces to compress prompts, check the current status, or integrate the service with the AI Prompt Compressor Policy and other upstream services.

  • POST /llm/v1/compressPrompt: Compresses a prompt using either a compression ratio or a target token count. Supports selective compression via <LLMLINGUA> tags.

  • GET /status: Returns information about the currently loaded LLMLingua model and device settings (for example, CPU or GPU).

  • POST /: JSON-RPC endpoint that supports the llm.v1.compressPrompt method. Use this to invoke compression programmatically over JSON-RPC.

LLMLINGUA prompt flow

  1. The user sends the final prompt to the AI Prompt Compressor Policy.
  2. The AI Prompt Compressor Policy checks the prompt for <LLMLINGUA>…</LLMLINGUA> tags.
    • If tags are found, only the tagged sections are sent to LLMLingua 2 for compression.
    • If no tags are found, the entire prompt is sent to LLMLingua 2 for compression.
  3. LLMLingua 2 applies the compression using the rule that matches the prompt’s configuration you set with the policy: by ratio, target token count, or conditional length-based rules.
  4. The compressed prompt is returned to the AI Prompt Compressor Policy.
  5. The AI Prompt Compressor Policy sends the compressed prompt to the Large Language Model (LLM).
  6. The LLM processes the prompt and returns the response to the user.

The following diagram illustrates how the AI Prompt Compressor Policy processes and compresses incoming prompts based on tagging and configured rules.

 
sequenceDiagram
    actor User
    participant KongAICompressor as AI Prompt Compressor Policy
    participant LLMLingua2 as LLMLingua 2 Compressor
    participant LLM as Large Language Model

    User->>KongAICompressor: Sends final prompt
    activate KongAICompressor
    KongAICompressor->>KongAICompressor: Check for LLMLINGUA tags

    alt If tagged content found
        KongAICompressor->>LLMLingua2: Compress tagged sections
        activate LLMLingua2
        LLMLingua2-->>KongAICompressor: Return compressed sections
        deactivate LLMLingua2
    else If no LLMlingua tags
        KongAICompressor->>LLMLingua2: Compress entire prompt
        activate LLMLingua2
        LLMLingua2-->>KongAICompressor: Return compressed prompt
        deactivate LLMLingua2
    end

    KongAICompressor->>LLM: Send compressed prompt
    deactivate KongAICompressor
    activate LLM
    LLM-->>User: Return response
    deactivate LLM
  

The AI Prompt Compressor Policy applies structured compression to preserve essential context of prompts sent by users, rather than trimming prompts arbitrarily or risking token overflows. This ensures the LLM receives a well-formed, focused prompt keeping token usage under control.

Prompt compression options

The AI Prompt Compressor Policy offers flexible compression controls to fit different use cases. You can choose between full-prompt compression, conditional strategies, or selectively compressing only parts of the prompt:

Configuration Option

Description

Compression by ratio Compress the prompt to a percentage of its original length (for example, reduce to 80%). This allows for consistent shrinkage regardless of the initial size.
Compression by token count Compress the prompt to a specific token target (for example, 150 tokens). Useful when working close to LLM context window limits.
Conditional rules Apply different compression strategies based on prompt length. For example, compress prompts under 100 tokens using a 0.8 ratio, and compress longer prompts to a fixed token count.
Selective compression with tags Wrap sections of the prompt in <LLMLINGUA>...</LLMLINGUA> to target only specific parts for compression, preserving untagged content as-is.

Headroom compressor service v2.2+

This feature is currently in Tech Preview and should not be used in a production environment.

Before you use Headroom with the AI Prompt Compressor Policy, you need a Headroom instance accessible to your AI Gateway.

You can do this with one of the following:

Configure Headroom connection

To configure an AI Prompt Compressor Policy with Headroom as the compressor service, do the following:

For more details, see the configuration reference.

Compressor Service endpoint

The compressor service exposes a /v1/compress endpoint that compresses messages and returns them. This endpoint accepts OpenAI and Anthropic’s message formats. Requests in unsupported formats are forwarded unchanged. You can use this interface to compress prompts, check the current status, or integrate the service with the AI Prompt Compressor Policy.

Headroom uses loopback-trust by default. It answers unauthenticated calls on 127.0.0.1, and returns 404 to non-loopback callers unless it’s started with HEADROOM_COMPRESS_ALLOW_REMOTE=1. You can run Headroom as a co-located sidecar reachable from the AI Gateway data plane without authentication. Alternatively, you can point the Policy at a remote instance and configure a bearer token, sent as both the X-Headroom-Proxy-Token and Authorization: Bearer headers, set with the headroom.proxy_token configuration field.

Headroom prompt flow

The AI Prompt Compressor Policy uses Headroom in a stateful mode and derives a session identifier for each conversation, based on the configured config.headroom.session_id_headers, an AI Consumer, or a credential. Headroom can recognize the turns it has already compressed for that conversation. Recognized turns are replayed unchanged instead of compressed again. This ensures the upstream LLM provider’s cache hits on repeated turns.

If a session identifier isn’t present, then every client that opens with the same prompt shares one session state on the Headroom side. If sessions collide, then only one will hit the upstream cache.

  1. AI Gateway sends the user or agent’s request to the AI Prompt Compressor.
  2. The AI Prompt Compressor builds a messages array from the whole conversation.
  3. The AI Prompt Compressor sends a POST request to Headroom’s /v1/compress endpoint with the messages array and a session identifier for the conversation.
  4. If Headroom recognizes the session identifier, it replays the turns it already compressed unchanged and compresses only the new messages. Otherwise, it compresses the whole conversation.
  5. Headroom returns 200 with the compressed messages and metadata.

If the call to Headroom fails, the AI Prompt Compressor rejects the request by default. For details, see Failure behavior.

The following diagram illustrates how the AI Prompt Compressor Policy processes and compresses incoming prompts using Headroom:

 
sequenceDiagram
    actor User as User/Agent
    participant KongAICompressor as AI Prompt Compressor Policy
    participant Headroom
    participant LLM as LLM Provider

    User->>KongAICompressor: Sends request
    activate KongAICompressor
    KongAICompressor->>Headroom: POST /v1/compress with messages and session ID

    alt Session recognized
        Headroom->>Headroom: Replay cached turns, compress only new messages
    else New session
        Headroom->>Headroom: Compress whole conversation
    end

    alt Compression succeeds
        Headroom-->>KongAICompressor: Return 200 with compressed messages and metadata
        KongAICompressor->>LLM: Send compressed messages to upstream provider
        LLM-->>KongAICompressor: Return response
        KongAICompressor-->>User: Return response
    else Compression fails
        Headroom-->>KongAICompressor: Return failure response
        KongAICompressor--xUser: Return HTTP 500, request not forwarded
    end
    deactivate KongAICompressor
  

Failure handling

By default, if the call to Headroom fails, times out, or returns an error, the Policy returns an HTTP 500 to the client rather than forwarding the request uncompressed.

A 503 response indicating a compression timeout is retried automatically before the Policy gives up; every other non-200 response (for example a 400 from a bad configuration, or a 401 from a missing or incorrect bearer token) fails immediately.

Limitations

  • Deterministic compression requires sessions which is only available in Headroom v0.37.0 or newer. Older images silently ignore the session id and run stateless.
  • Headroom sessions are held in memory on a single Headroom process. You must point every data plane node at its own sidecar, or all of them at one shared instance. Never point an AI Gateway data plane at a load-balanced set of Headroom instances, since a session’s turns must all reach the same process.
  • No MCP tool-response compression: only the LLM request path is supported.
  • Only requests with a messages or input conversation array are sent to Headroom. Other formats are forwarded uncompressed.

Forward proxy support

Set config.proxy on this entity to route its outbound requests through an HTTP forward proxy. Use this in network-isolated deployments where AI Gateway cannot open direct connections to LLM providers or auxiliary services.

The proxy record is identical for AI Model, AI MCP Server, and supported AI Policy entities. Existing capabilities such as load balancing, health checking, streaming, WebSocket, and HTTP/2 continue to work when the proxy is active.

For the full field reference, traffic flow, and limitations, see Forward proxy support.

Help us make these docs great!

Kong Developer docs are open source. If you find these useful and want to make them better, contribute today!