Share feedback
Answers are generated based on the documentation.

Use local and hosted models

Use sbx run --model to choose the model and service that answer your agent's requests. The agent runs inside a local sandbox. The model can run on your host, at a hosted provider, or at an inference endpoint you configure.

This page covers the built-in claude, codex, and opencode agents. Choose a model that supports the tool calls and context length your agent needs. --model isn't supported with cloud sandboxes or v3 kits.

Note

Model selection is experimental. Enable it before following these examples.

Bundled model service

Docker Sandboxes includes llmman, a model management tool installed alongside sbx. It serves local models and forwards requests to hosted providers or custom endpoints. You don't need to install it separately.

Installing Docker Sandboxes or enabling model selection doesn't start llmman. Docker Sandboxes starts it on your host as a background process when you first use sbx run --model, unless you select --provider ollama. Runs without --model don't start it.

Once started, llmman keeps running after the sandbox or CLI exits. Later model-enabled runs reuse the service, so sandboxes share its model store and loaded models.

On Linux, starting the service requires Docker Engine on the host to pull the inference server image.

Enable model selection

Run these commands on your host:

$ sbx settings set platform.allowExperimentalFeatures true
$ sbx settings set feature.model true

The --provider flag selects where the model runs:

ProviderModel destination
Omitted, or llmmanA local model managed by llmman
ollamaAn existing Ollama installation on your host
A hosted provider IDA provider supported by llmman, such as openai or anthropic
An ID from model.providersAn endpoint you configure

Run a local model

Pass a GGUF model reference or short name to --model:

$ sbx run --model gemma4 claude

Docker Sandboxes downloads the model if needed. The model runs on the host, so its memory and compute requirements are separate from the sandbox's resource limits. Replace claude with codex or opencode to use another agent with the same model.

Use Ollama

Install and start Ollama on your host, then select it with --provider:

$ sbx run --model gemma4 --provider ollama claude

Docker Sandboxes connects to Ollama at localhost:11434. It doesn't install, start, or manage the Ollama process.

For Docker Model Runner, see Run Claude Code in a Docker Sandbox with Docker Model Runner.

Use a hosted provider

Select a provider supported by llmman and a model available from that provider. For example, to run Codex with an OpenAI model, export OPENAI_API_KEY in your host shell, then run:

$ sbx run --provider openai --model gpt-5-nano codex

The host's llmman service forwards requests to the provider. It doesn't download or run the hosted model. For available providers and their API-key variable names, see the llmman provider documentation.

Provider authentication

Make the provider's API key available in the host shell before the first sbx run --model command starts llmman. For a custom endpoint, choose the variable name with apiKeyEnv.

llmman inherits the environment of the process that starts it. Changing a variable in another shell doesn't update an already-running service. After changing a key, stop the host's llmman serve process, then run sbx run --model from the shell containing the updated variable. This interrupts model requests from other sandboxes using that service.

Provider authentication for this route is handled by llmman on the host. Credentials stored with sbx secret set aren't automatically supplied to it. For the agents' default authentication flows, see Manage credentials.

Connect a custom endpoint

Use model.providers to connect to an OpenAI- or Anthropic-compatible inference endpoint, such as an internal GPU server. The endpoint must be reachable from your host.

The setting is a JSON object keyed by provider ID. Check its existing value before changing it:

$ sbx settings get model.providers

For an OpenAI-compatible endpoint, define a provider named company:

$ sbx settings set model.providers '{"company":{"url":"https://inference.example.com/v1","wire":"openai","apiKeyEnv":"COMPANY_API_KEY"}}'

Replace the URL with your endpoint's base URL. Setting model.providers replaces the whole object, so include any existing providers you want to keep.

FieldDescription
urlRequired HTTP or HTTPS base URL, usually ending in /v1. Use the base URL, without /chat/completions or /messages.
wireThe endpoint's API format: openai (default) or anthropic. This describes the endpoint, regardless of which agent you run.
apiKeyEnvName of the host environment variable containing the API key. Omit it for an endpoint that doesn't require a key.
nameOptional display name. Defaults to the provider ID.

Export COMPANY_API_KEY in your host shell as described in Provider authentication, then select the provider and a model served by that endpoint:

$ sbx run --provider company --model <MODEL_NAME> claude

Docker Sandboxes applies the provider configuration when you run with --model. You can use the same provider with codex or opencode.

Change an existing sandbox's model

Pass the sandbox name and your model selection:

$ sbx run --name <SANDBOX_NAME> --model <MODEL_NAME> --provider <PROVIDER_ID>

Changing the model recreates the sandbox container. The workspace and kit-owned volumes persist. Omit --provider to select a local model managed by llmman.

Use another provider for larger requests

Pair a local model with another provider to handle requests that exceed the local model's context capacity:

$ sbx run --model gemma4 \
    --overflow-provider openai --overflow-model gpt-5-nano claude

Configure provider authentication before starting the model service. You can also use a provider defined in model.providers. Both overflow flags are required, and the local model must use the default llmman provider. This option can't be combined with --provider ollama or a hosted provider selected with --provider.

Requests that fit the local model stay local. Requests routed to the overflow provider send their contents to that endpoint and can incur provider charges.