You can proxy requests to Ollama AI models through AI Gateway by creating AI Model Provider and AI Model entities. This reference documents all supported AI capabilities, configuration requirements, and provider-specific details needed for proper integration.
Ollama provider
Upstream paths
AI Gateway automatically routes requests to the appropriate Ollama API endpoints. The following table shows the upstream paths used for each capability.
|
Capability |
Path template |
Description |
Upstream path or API |
|---|---|---|---|
| Generate |
/chat/completions, /completions, or /responses
|
Text generation for chat completions and responses | /api/chat |
| Embeddings |
/embeddings
|
Vector embeddings from text input | /api/embed |
Supported capabilities
The following tables show the AI capabilities supported by the Ollama provider when configuring AI Models.
By default, AI Gateway uses the path templates shown in the tables below (e.g.,
/chat/completions,/embeddings, etc.). To customize these paths, configure theconfig.pathsfield in your AI Model entity. Custom paths take the form{configured_path}/{template_path}— for example, if you set a custom path of/v2, requests to/embeddingswould be routed to/v2/embeddings.
Text generation
Support for Ollama text generation capabilities:
|
Capability |
Streaming |
Model example |
Path template |
Min version |
|---|---|---|---|---|
| generate | Supported | llama3.2:1b |
/chat/completions, /completions, or /responses
|
2.0 |
Embeddings
Support for Ollama embeddings generation:
|
Capability |
Model example |
Path template |
Min version |
|---|---|---|---|
| embeddings | qwen3-embedding:8b |
/embeddings
|
2.0 |
Ollama base URL
The base URL is $UPSTREAM_URL.
AI Gateway uses this URL automatically. You only need to configure a URL if you’re using a self-hosted or Ollama-compatible endpoint, in which case set the upstream_url option in your AI Model configuration.
Configure Ollama
To use Ollama with AI Gateway, configure a new AI Model Provider. You can then access supported AI Models from Ollama.
Here’s a minimal configuration for chat completions: