Skip to content

Agents

A LanguageAgent is a container image running as an Argo Workflow. The operator handles everything around it — lifecycle, configuration, networking, persistent storage — so the image itself can focus on doing work.

An agent is either always on (spec.execution.mode: service, the default) or invoked — on a cron schedule, or by hand against the WorkflowTemplate the operator renders for it (mode: task). See Execution Modes.

This page covers what the operator injects into every agent pod, what the agent is expected to do with it, and how to wire everything together.

Configuration

Every agent pod gets one file mounted at /etc/agent/config.yaml. The operator assembles it from the agent's spec.instructions, referenced personas, resolved tool endpoints, and model configuration. It's reconciled on every change to the LanguageAgent or any resource it references.

# Agent identity
agent:
  name: data-analyst
  namespace: default

# From spec.instructions — omitted when empty
instructions: |
  You are a data analyst. Analyze CSV files and generate insights.

# Resolved from referenced LanguagePersona CRDs
personas:
  - name: analytical-persona
    tone: professional
    personality: "precise, data-driven, always cites sources before drawing conclusions"
    expertise: "data analyst specialising in statistical reasoning and business intelligence"

# Resolved tool endpoints, keyed by tool name. An external server (spec.tools[].url)
# keeps its URL and may carry headers; $(NAME) in a value is an agent environment
# variable the runtime substitutes when it connects.
tools:
  mem0-memory:
    endpoint: http://mem0-memory.tools.svc.cluster.local:8080/mcp
    protocol: mcp
  python-executor:
    endpoint: http://python-executor.tools.svc.cluster.local:8080/mcp
    protocol: mcp
  control-plane:
    endpoint: https://cloud.example.com/mcp
    protocol: mcp
    headers:
      Authorization: Bearer $(CONTROL_PLANE_TOKEN)

# Model configuration — all LLM traffic routes through the shared gateway
models:
  claude-sonnet:
    role: primary
    provider: anthropic
    model: claude-sonnet-4-5
    endpoint: http://gateway.default.svc.cluster.local:8000

The file is read-only. Agents should read it on startup to configure their runtime.

Environment Variables

The operator injects these into the agent container and all init containers:

Variable Value
AGENT_NAME metadata.name of the LanguageAgent
AGENT_NAMESPACE metadata.namespace of the LanguageAgent
AGENT_UUID Stable UUID assigned to this agent
AGENT_CLUSTER_NAME Name of the LanguageCluster this agent belongs to
AGENT_CLUSTER_UUID Kubernetes UID of the LanguageCluster
AGENT_EXECUTION_MODE service or task — the agent's spec.execution.mode. A runtime keeps running in service and exits when its work is done in task
AGENT_EVENT The run's event parameter, verbatim; also written to /etc/agent/event.json. Empty unless the run was started with one. See Execution Modes.
AGENT_TRIGGER The run's trigger parameter: what started it. schedule for scheduled runs; otherwise empty unless given.
MODEL_ENDPOINT Shared LiteLLM gateway URL — the same URL regardless of how many models are referenced
MODEL_API_KEY This agent's gateway key; send it as the bearer key to MODEL_ENDPOINT. Always set; the gateway rejects requests without a valid key. See Gateway authentication.
LLM_MODEL Comma-separated list of model names for all referenced models
MCP_SERVERS Comma-separated tool endpoint URLs — service-mode tools use in-cluster DNS; sidecar-mode tools use http://localhost:<port>/mcp
AGENT_INSTRUCTIONS Content of spec.instructions; only set when non-empty
AGENT_REPO_DIR Absolute path to the cloned repository; also the agent's working directory. Only injected when spec.repository is set
GIT_CONFIG_*, GIT_SSH_COMMAND, GIT_TERMINAL_PROMPT Git identity and credentials for spec.repository; agent and repository init containers only
GH_TOKEN / GITLAB_TOKEN The repository Secret's token for the vendor CLI (gh, glab); agent container only
OTEL_EXPORTER_OTLP_ENDPOINT Propagated from the operator when configured
OTEL_SERVICE_NAME Set to agent-<name> when OTEL is configured

Additional variables from spec.deployment.env and spec.deployment.envFrom are passed through unchanged.

Workspace

When spec.workspace.enabled is true (the default), the operator provisions a PersistentVolumeClaim named <agent-name>-workspace and mounts it into the agent container and all init containers. The workspace survives pod restarts, Workflow replacement, and the boundary between task-mode runs — each scheduled run gets a fresh pod attached to the same volume. By default the PVC is deleted when the LanguageAgent is deleted; set spec.workspace.retain: true to preserve it.

Field Default Description
spec.workspace.enabled true Create and mount a workspace PVC
spec.workspace.size 10Gi PVC storage request
spec.workspace.mountPath /workspace Mount path in the container
spec.workspace.storageClassName cluster default StorageClass for the PVC
spec.workspace.accessMode ReadWriteOnce PVC access mode (ReadWriteOnce or ReadWriteMany)
spec.workspace.retain false When true, the PVC is preserved after agent deletion; the orphaned PVC name is recorded in status.workspacePVCName
spec.workspace.initialFiles — Files seeded into the workspace on first boot (keys = filenames, values = file contents; not overwritten if already present)
spec.workspace.seedConfigMapRef — External ConfigMap whose keys/values are seeded as files; merged with initialFiles (initialFiles wins on collision)

The volume is named workspace in the pod spec. Init containers that need to pre-seed it should mount it by that name.

The container's root filesystem is read-only. /tmp is a memory-backed tmpfs capped at 1Gi: it counts against the container's memory limit, and a write past the cap fails with ENOSPC rather than OOM-killing the pod. Anything large — clones, builds, downloads — belongs on the workspace.

Repository

When spec.repository is set, the operator adds a repository init container that runs after the workspace seeder and clones a git repository into the workspace. Both init containers use the operator's git client image (ghcr.io/language-operator/git, pinned to the chart appVersion; override with config.git.repository/config.git.tag or --git-image). The agent container's working directory is set to the clone path, and AGENT_REPO_DIR is injected into every container.

The clone is clone-once: it is skipped when the target directory already contains a .git directory, so the agent's edits and commits survive pod restarts. Declaring spec.repository defaults spec.workspace.enabled to true, since the clone needs the workspace PVC to land in.

Field Default Description
spec.repository.url — Git repository to clone — HTTPS (https://...) or SSH (git@host:org/repo.git). Required. Use the https:// URL for public repositories: GitHub does not accept anonymous SSH, so an SSH URL always needs an ssh-privatekey
spec.repository.ref default branch Branch, tag, or commit SHA to check out
spec.repository.path repo name from URL Subdirectory under the workspace mountPath to clone into (relative path only)
spec.repository.depth 0 When > 0, shallow-clone to this history depth
spec.repository.secretRef — Secret with git credentials for private repos. Keys: token or username+password (HTTPS), ssh-privatekey (SSH)
spec.repository.vendor from the host github, gitlab or git; selects the CLI that receives the token (gh, glab, none)

When secretRef is set, the Secret is mounted read-only at /var/run/secrets/langop.io/git into the repository init container and the agent container; the operator never reads its contents. Git is configured through GIT_CONFIG_* environment variables: a default commit identity (<name>@<namespace>.langop.io, overridable with GIT_AUTHOR_*/GIT_COMMITTER_*), a host-scoped credential helper for HTTPS remotes or GIT_SSH_COMMAND for SSH remotes, so fetch and push work after the clone. The Secret's token is also exported as GH_TOKEN or GITLAB_TOKEN by vendor. See the LanguageAgent API reference for details.

Networking

Every agent gets a NetworkPolicy allowing inbound traffic from other agents in the same cluster namespace — add egress for public APIs via spec.networkPolicies (see the Network Policies guide).

A service-mode agent is addressable, and additionally gets:

  • A ClusterIP Service named after the agent, with one port entry per spec.ports
  • An Ingress at <agent-name>.<cluster domain>, when the LanguageCluster has a domain

It listens on the port(s) defined in spec.ports. What it serves there is up to the image — HTTP, WebSocket, OpenAI-compatible API, anything.

A task-mode agent gets neither Service nor Ingress: its pods exist only for the duration of a run, so there would be nothing to route to in between. spec.ports is rejected for task agents.

Init Containers

For runtimes that need configuration in a format other than config.yaml, use an init container to translate before the agent starts. The operator automatically mounts /etc/agent/config.yaml into every init container, so adapters can read the operator config and write whatever format the runtime expects.

spec:
  image: ghcr.io/myorg/my-agent:latest
  ports:
    - name: http
      port: 18789

  workspace:
    size: 10Gi

  deployment:
    initContainers:
      - name: seed-config
        image: myregistry/config-adapter:latest
        env:
          - name: STATE_DIR
            value: /workspace/.config
        volumeMounts:
          - name: workspace
            mountPath: /workspace

The init container runs to completion before the agent container starts. On subsequent restarts it can check for existing state and skip re-seeding user data while still applying any operator-managed config changes.

Example

apiVersion: langop.io/v1alpha1
kind: LanguageAgent
metadata:
  name: data-analyst
  namespace: default
spec:
  image: myregistry/agent-runtime:python-v1.0.0
  ports:
    - name: http
      port: 8080

  instructions: |
    You are a data analyst. Analyze CSV files and generate insights.
    Focus on trends, anomalies, and actionable recommendations.

  persona: analytical-persona

  tools:
    - name: mem0-memory
    - name: python-executor

  models:
    - name: claude-sonnet

  workspace:
    size: 10Gi
    mountPath: /workspace

  deployment:
    resources:
      limits:
        memory: 1Gi
        cpu: 500m
    livenessProbe:
      httpGet:
        path: /health
        port: 8080
      initialDelaySeconds: 10
      periodSeconds: 30
    readinessProbe:
      httpGet:
        path: /health
        port: 8080
      initialDelaySeconds: 5
      periodSeconds: 10

Checklist

A well-behaved agent image should:

  • [ ] Listen on the port(s) defined in spec.ports (default: one port named http on 8080)
  • [ ] Read AGENT_EXECUTION_MODE: keep running when it is service (or unset), exit when the work is done when it is task
  • [ ] Read /etc/agent/config.yaml on startup for instructions, personas, tools, and models
  • [ ] Respect the AGENT_* identity env vars
  • [ ] Route all LLM traffic through MODEL_ENDPOINT — agents never hold real API credentials
  • [ ] Use spec.workspace.mountPath (default /workspace) for persistent state