LanguageModel¶
The LanguageModel CRD configures LLM access through the cluster's shared LiteLLM proxy.
Overview¶
A LanguageModel defines: - Provider (Anthropic, OpenAI, Azure, etc.) - Model name and version - API credentials (via Secret references) - Provider-specific settings (region, project, API version, extra LiteLLM params) - Rate limits and timeout
Quick Example¶
apiVersion: langop.io/v1alpha1
kind: LanguageModel
metadata:
name: claude-sonnet
namespace: my-cluster
spec:
provider: anthropic
modelName: claude-sonnet-4-5
apiKeySecretRef:
name: anthropic-credentials
key: api-key
Complete API Reference¶
See the Complete API Reference for full field documentation including:
- LanguageModel - Top-level resource
- LanguageModelSpec - Specification fields
- LanguageModelStatus - Status and endpoint information
Providers¶
Set exactly one of provider or litellmProvider.
provider |
Reaches | Also needs |
|---|---|---|
openai |
OpenAI | API key |
anthropic |
Anthropic | API key |
gemini |
Google AI Studio (Gemini API) | API key |
azure |
Azure OpenAI; modelName is the deployment name |
endpoint, apiVersion, an API key or Azure AD app credentials |
bedrock |
AWS Bedrock | region, AWS access keys or a Bedrock bearer token |
vertex |
Google Vertex AI | project, location, a service-account JSON |
openai-compatible |
Any OpenAI chat-completions API: Ollama, vLLM, LM Studio… | endpoint |
custom |
Deprecated, identical to openai-compatible |
For anything else LiteLLM supports, set litellmProvider to its LiteLLM prefix (e.g. deepseek, dashscope, xai, mistral, groq, openrouter, hosted_vllm). The gateway calls the model as <litellmProvider>/<modelName>, so a new vendor needs no operator change:
spec:
litellmProvider: deepseek
modelName: deepseek-chat
apiKeySecretRef:
name: deepseek-credentials
The admission webhook rejects a spec that is missing what its provider needs (the "Also needs" column).
Wildcard models¶
modelName: "*" makes the LanguageModel stand for the provider's whole catalogue. Agents choose a model with spec.models[].model and call it as <LanguageModel name>/<model>. See Wildcard models.
Typed fields and params¶
| Field | LiteLLM param | Used by |
|---|---|---|
region |
aws_region_name |
bedrock |
project |
vertex_project |
vertex |
location |
vertex_location |
vertex |
apiVersion |
api_version |
azure |
params passes anything else straight into the model's LiteLLM params, keeping JSON types, and overrides the fields above:
spec:
params:
aws_bedrock_runtime_endpoint: https://vpce-0123.bedrock-runtime.us-east-1.vpce.amazonaws.com
extra_headers:
X-Team: platform
Credentials never go in params: keys such as api_key, aws_secret_access_key or client_secret are rejected. Use a Secret.
Key Concepts¶
Shared Proxy Registration¶
When you create a LanguageModel:
- The LanguageModel controller validates the spec and sets
status.phase: Ready - The LanguageCluster controller (which watches
LanguageModelresources) detects the new CR - The shared
gateway-configConfigMap is regenerated with all models in the namespace - The gateway Deployment rolls over with the updated configuration
All agents immediately have access to the new model via MODEL_ENDPOINT.
Credential Management¶
Credentials live in Secrets and are read only by the gateway, never by agent pods.
One key: apiKeySecretRef names a Secret and a key (default api-key):
Several values: credentialsSecretRef names a whole Secret. Its keys are recognised by name and applied to this model only, so two models can use different credentials for the same provider:
| Secret key | Becomes |
|---|---|
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN |
Bedrock SigV4 credentials |
AWS_BEARER_TOKEN_BEDROCK |
Bedrock bearer token |
VERTEX_CREDENTIALS, GOOGLE_APPLICATION_CREDENTIALS, service-account.json, credentials.json |
Vertex service-account JSON (read from the mounted file) |
AZURE_API_KEY, GEMINI_API_KEY, api-key |
the API key |
AZURE_TENANT_ID, AZURE_CLIENT_ID, AZURE_CLIENT_SECRET, AZURE_AD_TOKEN |
Azure AD credentials |
LiteLLM's own param names (aws_access_key_id, api_key, tenant_id…) also work as keys. Other keys are ignored and named in the gateway log. When both references supply an API key, apiKeySecretRef wins.
Rate Limiting¶
Configure per-model rate limits:
The shared proxy enforces these limits across all agents.
Provider-Specific Examples¶
AWS Bedrock (access keys)¶
kubectl create secret generic aws-bedrock \
--from-literal=AWS_ACCESS_KEY_ID=AKIA... \
--from-literal=AWS_SECRET_ACCESS_KEY=...
spec:
provider: bedrock
modelName: anthropic.claude-sonnet-4-5-20250929-v1:0
region: us-east-1
credentialsSecretRef:
name: aws-bedrock
AWS Bedrock (bearer token)¶
spec:
provider: bedrock
modelName: amazon.nova-pro-v1:0
region: us-west-2
credentialsSecretRef:
name: bedrock-token
Google Vertex AI¶
spec:
provider: vertex
modelName: gemini-2.5-pro
project: my-gcp-project
location: us-central1
credentialsSecretRef:
name: vertex-sa
Gemini (Google AI Studio)¶
Azure OpenAI¶
modelName is your Azure deployment name.
spec:
provider: azure
modelName: my-gpt4o-deployment
endpoint: https://my-resource.openai.azure.com
apiVersion: "2025-01-01-preview"
apiKeySecretRef:
name: azure-credentials
With an Azure AD app instead of a key, put AZURE_TENANT_ID, AZURE_CLIENT_ID and AZURE_CLIENT_SECRET in a Secret and reference it with credentialsSecretRef.
Self-Hosted (Ollama, vLLM)¶
spec:
provider: openai-compatible
modelName: llama3.2
endpoint: http://ollama.default.svc.cluster.local:11434/v1
The backend only needs /v1/chat/completions: the gateway also serves /v1/responses and /v1/messages for these models by translating to it.
Related Resources¶
- LanguageAgent - Reference models in agents
- LanguageCluster - Shared proxy architecture