AI Gateway Enterprise: This plugin is only available as part of our AI Gateway Enterprise offering.
The AI Semantic Response Guard plugin extends the AI Prompt Guard plugin by filtering LLM responses based on semantic similarity to predefined rules. It helps prevent unwanted or unsafe responses when serving llm/v1/chat, llm/v1/completions, or llm/v1/embeddings requests through AI Gateway.
You can use a combination of allow and deny response rules to maintain integrity and compliance when returning responses from an LLM service.
The plugin analyzes the semantic content of the full LLM response before it is returned to the client. The matching behavior is as follows:
If any deny_responses are set and the response matches a pattern in the deny list, the response is blocked with a 403 Forbidden.
If any allow_responses are set, but the response matches none of the allowed patterns, the response is also blocked with a 403 Forbidden.
If any allow_responses are set and the response matches one of the allowed patterns, the response is permitted.
If both deny_responses and allow_responses are set, the deny condition takes precedence. A response that matches a deny pattern will be blocked, even if it also matches an allow pattern. If the response does not match any deny pattern, it must still match an allow pattern to be permitted.
Disables streaming (stream=false) to ensure the full response body is buffered before analysis.
Intercepts the response body using the guard-response filter.
Extracts response text, supporting JSON parsing of multiple LLM formats and gzipped content.
Generates embeddings for the extracted text.
Searches the vector database (Redis, Pgvector, or other) against configured allow_responses or deny_responses.
Applies the decision rules described above.
If a response is blocked or if a system error occurs during evaluation, the plugin returns a 403 Forbidden to the client without exposing that the Semantic Response Guard blocked it.
A vector database can be used to store vector embeddings, or numerical representations, of data items. For example, a response would be converted to a numerical representation and stored in the vector database so that it can compare new requests against the stored vectors to find relevant cached items.
The AI Semantic Response Guard plugin supports the following vector databases:
Using config.vectordb.strategy: redis and parameters in config.vectordb.redis:
Valkeyv3.14+: When you configure vectordb.strategy: redis, Kong Gateway queries the server and checks the server name field. If it detects Valkey request, it automatically uses the Valkey-specific driver.
Kong’s semantic features use the Redis Query Engine (RediSearch) to create and search vector indexes, and the Redis JSON data type to store indexed documents. Your Redis or Valkey deployment must support both, or vector index creation fails.
Self-hosted Redis: Use Redis Open Source 8.0 or later, which bundles the Redis Query Engine and JSON data type by default, or a distribution that adds them as modules, such as Redis Stack, Redis Enterprise, or Redis Cloud.
Self-hosted Valkey: Use Valkey 8.1 or later with the valkey-search and valkey-json modules enabled. A vanilla Valkey build without these modules can’t be used for semantic features.
AWS ElastiCache: ElastiCache doesn’t support custom Redis modules, so ElastiCache for Redis OSS can’t be used for semantic features at any version. Use ElastiCache for Valkey 8.2 or later, which has vector search built in natively.
Google Cloud Memorystore: Use Memorystore for Redis Cluster 8.0 or later or Memorystore for Valkey 8.0 or later, both of which include the required modules automatically. Standard (non-cluster) Memorystore for Redis instances, including the Basic tier, are frozen at Redis 7.2 and don’t support RediSearch or the JSON data type, so they can’t be used for semantic features.
If your plugin uses a Redis datastore, you can authenticate to it with a cloud Redis provider, or with an OAuth 2.0 token endpoint.
This allows you to seamlessly rotate credentials without relying on static passwords.
The following providers are supported:
AWS ElastiCache
Azure Managed Redis
Google Cloud Memorystore (with or without Valkey)
OAuth 2.0 v3.16+, using the client_credentials or password grant type
You need:
A running Redis instance on an AWS ElastiCache instance for Valkey 7.2 or later or ElastiCache for Redis OSS version 7.0 or later
$CLUSTER_ADDRESS: The Memorystore cluster address.
$GCP_SERVICE_ACCOUNT: The GCP service account JSON.
You need:
An OAuth 2.0 token endpoint that issues access tokens for the client_credentials or password grant type
A Redis deployment that accepts the issued access token as a bearer credential, either natively or through an OAuth-aware proxy (such as Envoy) placed in front of it
To configure OAuth 2.0 authentication with Redis, add the following parameters to your plugin configuration:
$INSTANCE_ADDRESS: The Redis instance or proxy address.
$OAUTH_TOKEN_ENDPOINT: The OAuth 2.0 token endpoint URL used to request access tokens.
$OAUTH_CLIENT_ID: Your OAuth 2.0 client ID.
$OAUTH_CLIENT_SECRET: Your OAuth 2.0 client secret.
Kong Gateway caches the acquired token for the duration of its validity and refreshes it asynchronously before it expires.
To use the password grant type instead, set oauth.grant_type to password and also provide oauth.username and oauth.password.
If your Redis deployment uses ACL-based authentication and needs a username sent alongside the token, also set one of:
oauth.redis_username: a static username to send with AUTH <username> <token>.
oauth.redis_username_claim: the name of a claim in the access token (for example, oid for Microsoft Entra ID) to derive the username from.