Models¶
A LanguageModel configures LLM access for a LanguageCluster. The operator reads all LanguageModel resources in the namespace and registers them with the cluster's shared LiteLLM gateway. Agents on gateway-backed runtimes never hold provider credentials or connect to providers directly: they call the gateway with their own gateway key.
The exception is vendor-key runtimes — Claude Code and Cursor — which talk to their vendor directly with the vendor's own key and do not use the gateway or LanguageModels at all.
How It Works¶
One LiteLLM proxy (gateway) runs per LanguageCluster. When you add or remove a LanguageModel, the gateway restarts with the updated model list — no agent redeploy required.
Credential Management¶
API keys are never injected into agent pods. Store them in a Secret:
Reference the Secret from the model spec:
apiVersion: langop.io/v1alpha1
kind: LanguageModel
metadata:
name: claude-sonnet
spec:
provider: anthropic
modelName: claude-sonnet-4-5
apiKeySecretRef:
name: anthropic-credentials
key: api-key
The gateway pod mounts the Secret and presents a single OpenAI-compatible endpoint to agents. Rotating a key is a kubectl create secret operation — the gateway restarts, agents are unaffected.
Agent Integration¶
The operator injects two environment variables into every agent container:
| Variable | Value |
|---|---|
MODEL_ENDPOINT |
http://gateway.<namespace>.svc.cluster.local:8000 |
LLM_MODEL |
Comma-separated list of model names from spec.models[].name |
Both are also available through /etc/agent/config.yaml under the models: key:
models:
claude-sonnet:
role: primary
provider: anthropic
model: claude-sonnet-4-5
endpoint: http://gateway.my-cluster.svc.cluster.local:8000
Agents call the gateway with the model name they want. The gateway routes to the correct upstream provider.
Gateway Authentication¶
The gateway only serves requests carrying a valid key. Every agent gets its own:
- The operator injects
MODEL_API_KEYinto each agent's containers, next toMODEL_ENDPOINT. Agents send it as the bearer key (Authorization: Bearer $MODEL_API_KEY); the bundled gateway-backed runtimes do this for you. - A key is
sk-langop-<agent UUID>.<signature>, where the signature is an HMAC of the agent UUID under the cluster's secret. The gateway checks it without a key store, and records the agent UUID as the request's LiteLLMuser_id, so usage can be attributed per agent. - A request without a valid key gets 401. Custom agent images must send
MODEL_API_KEY; one that sends a fixed placeholder is rejected.
The HMAC secret lives in the Secret gateway-auth (key hmac-secret) in the cluster's namespace. The operator generates it on first reconcile; create it yourself beforehand to choose the value. Rotating it (edit the value, or delete the Secret to have a new one generated) restarts the gateway and replaces every agent's key, and running agents are restarted with their new key.
For access outside an agent — an admin, a script, a client using the gateway's external ingress — set a master key on the gateway, which the gateway also accepts:
apiVersion: langop.io/v1alpha1
kind: LanguageCluster
spec:
gateway:
deployment:
env:
- name: LITELLM_MASTER_KEY
valueFrom:
secretKeyRef:
name: gateway-master-key
key: key
Supported Providers¶
| Provider | Value |
|---|---|
| Anthropic | provider: anthropic |
| OpenAI | provider: openai |
| Google AI Studio (Gemini API) | provider: gemini |
| Azure OpenAI | provider: azure (+ endpoint, apiVersion) |
| AWS Bedrock | provider: bedrock (+ region) |
| Google Vertex AI | provider: vertex (+ project, location) |
| Any OpenAI chat-completions API (Ollama, vLLM, LM Studio…) | provider: openai-compatible (+ endpoint) |
| Any other LiteLLM provider (DeepSeek, Qwen, xAI, Mistral, Groq, OpenRouter…) | litellmProvider: <LiteLLM prefix> |
custom is deprecated and behaves exactly like openai-compatible. Providers that need more than one credential (Bedrock access keys, a Vertex service account, an Azure AD app) take a whole Secret through credentialsSecretRef, applied to that model only. See LanguageModel for each provider's fields and examples.
Other providers¶
Anything LiteLLM supports is reachable with litellmProvider, its LiteLLM provider prefix, plus the provider's API key. LiteLLM knows each one's base URL, so endpoint is only needed to override it:
| Provider | litellmProvider |
|---|---|
| DeepSeek | deepseek |
| Z.ai (GLM) | zai |
| Moonshot (Kimi) | moonshot |
| MiniMax | minimax |
| xAI (Grok) | xai |
| Mistral | mistral |
| Groq | groq |
| Together AI | together_ai |
| Fireworks AI | fireworks_ai |
| OpenRouter | openrouter |
| Alibaba Qwen (DashScope) | dashscope |
| Cerebras | cerebras |
| Perplexity | perplexity |
spec:
litellmProvider: deepseek
modelName: deepseek-chat
apiKeySecretRef:
name: deepseek-credentials
See LiteLLM's provider list for the rest. Every example in the repository is a complete, CI-validated spec.
Self-hosted models (Ollama, vLLM)¶
spec:
provider: openai-compatible
modelName: llama3.2
endpoint: http://ollama.default.svc.cluster.local:11434/v1
No apiKeySecretRef needed for unauthenticated endpoints. The same shape works for any server with an OpenAI-compatible API; their usual defaults:
| Server | endpoint (default port) |
|---|---|
| Ollama | http://<host>:11434/v1 |
| vLLM | http://<host>:8000/v1 |
llama.cpp (llama-server) |
http://<host>:8080/v1 |
| SGLang | http://<host>:30000/v1 |
| LM Studio | http://<host>:1234/v1 |
The backend only needs /v1/chat/completions. Clients that speak the Responses API (/v1/responses, e.g. Codex) or the Anthropic Messages API (/v1/messages) still work: the gateway translates both to chat completions for these models.
Multiple models¶
Agents can reference multiple models. Each model is registered with the same gateway; the agent chooses which to call at runtime:
Wildcard models: a whole vendor catalogue¶
A LanguageModel with modelName: "*" stands for every model the provider offers, so one resource covers a vendor like OpenRouter:
apiVersion: langop.io/v1alpha1
kind: LanguageModel
metadata:
name: openrouter
spec:
provider: openai-compatible
modelName: "*"
endpoint: https://openrouter.ai/api/v1
apiKeySecretRef:
name: openrouter-credentials
The gateway registers it as openrouter/*. An agent picks a model by the vendor's own name with model, and calls the gateway with <LanguageModel name>/<model>; the gateway strips the prefix before calling the provider:
# LanguageAgent
spec:
models:
- name: openrouter
model: anthropic/claude-sonnet-4.5 # the agent's LLM_MODEL is openrouter/anthropic/claude-sonnet-4.5
Each wildcard is addressed by its own resource name, so several vendors can coexist, and a name no LanguageModel serves is still rejected. Rate limits and the timeout apply to all of the vendor's models together. * is only valid on its own: partial patterns such as gpt-* are rejected.
Listing models
For a wildcard, the gateway's /v1/models does not return the vendor's catalogue. It lists
LiteLLM's built-in table of known models for the provider type — for an OpenAI-compatible
vendor like OpenRouter, that is OpenAI's model names. To discover what a vendor offers, use
the vendor's own model list API.
Load balancing: models that share a modelName¶
Agents call the gateway by modelName, not by the LanguageModel's own name. When several LanguageModels in a cluster have the same modelName, the gateway serves them as one model and spreads requests across all of them. Use this to run one model on several endpoints:
apiVersion: langop.io/v1alpha1
kind: LanguageModel
metadata:
name: llama-gpu-a
spec:
provider: openai-compatible
modelName: llama3.2
endpoint: http://ollama-a.inference.svc.cluster.local:11434/v1
---
apiVersion: langop.io/v1alpha1
kind: LanguageModel
metadata:
name: llama-gpu-b
spec:
provider: openai-compatible
modelName: llama3.2
endpoint: http://ollama-b.inference.svc.cluster.local:11434/v1
Because the same thing happens by accident, for example two models with the same modelName but different providers or keys, creating or updating a LanguageModel whose modelName is already taken returns a warning naming the other models, and the gateway logs one on start. Give each model a distinct modelName if you want them addressed separately.
Advanced: Extra Gateway Configuration¶
For LiteLLM settings with no LanguageModel field (retries, fallbacks, caching, spend callbacks, or a hand-written model entry) set LANGOP_GATEWAY_EXTRA_CONFIG on the gateway. It is a YAML mapping deep-merged into the generated LiteLLM config: mappings merge recursively, lists and scalars replace.
apiVersion: langop.io/v1alpha1
kind: LanguageCluster
spec:
gateway:
deployment:
env:
- name: LANGOP_GATEWAY_EXTRA_CONFIG
value: |
litellm_settings:
num_retries: 2
request_timeout: 600
A list such as model_list replaces the generated one entirely, so prefer LanguageModels for models. A value that is not a mapping stops the gateway from starting.
Rate Limiting¶
Limits are enforced by the shared gateway across all agents. Per-agent limits are not currently supported.
Related¶
- LanguageCluster — owns the shared gateway
- LanguageAgent — references models via
spec.models - LanguageModel API Reference — full field documentation