Skip to content

Run config

A run config describes how to execute a run: the task, the model, its API key, and options such as limits. karotte run --config accepts a path to a JSON file or the JSON itself:

uv run karotte run --config run_config.json

The JSON gets parsed as an EvaluationRunConfig.

Flags

Without --config, karotte run builds the config from its flags:

uv run karotte run --task find-python --model anthropic/claude-opus-5-5

--task and --model are then required. The run gets a random run_id, and transcript_file defaults to out/transcript.json.

Flag Field
--task task_id
--model model
--model-api-key model_api_key
--reasoning-effort reasoning_effort
--transcript-file transcript_file

With --config, the flags override the config's fields, so one config can be run against several models:

uv run karotte run --config run_config.json --model openai/gpt-5.5

The other fields have no flag; set them in a config.

Create a run config

uv run karotte create-run-config --model anthropic/claude-opus-5-5 --task find-python

It loads your tasks, so run it from the environment repo. It writes a random run_id, sets transcript_file to out/transcript.json, and fills in every other field with its default:

{
    "run_id": "3f9c1a2b",
    "task_id": "find-python",
    "model": "anthropic/claude-opus-5-5",
    "transcript_file": "out/transcript.json",
    ...
}

EvaluationRunConfig lists every field with its default.

If you rename a task or want to run another one, change task_id.

Extra config

extra_config is a free-form dict passed to your task unchanged; see Tasks and steps for how a task reads it.

{
    "extra_config": {
        "difficulty": "hard",
        "extra_system_prompt": "Explain each command before you run it."
    }
}

These keys have a built-in meaning:

Key Value Effect
extra_system_prompt string Appended to the system prompt after a blank line.
system_prompt_override string Replaces the system prompt. Wins over extra_system_prompt.
extra_task_instructions string or list of strings Appended to the step's instructions after a blank line. A string applies to every step; a list has one entry per step, and steps past its end are unchanged.
task_instructions_override string or list of strings Replaces the step's instructions. Same string-or-list rules. Wins over extra_task_instructions.
extra_artifact_paths string or list of strings Absolute paths saved as artifacts after each step, before pre_scoring_hook. Missing paths are skipped with a warning.

The first two are read by the default template's get_system_prompt(), so they only work in tasks that use it. Karotte itself applies the latter three.

Models

model is a litellm model id with its provider prefix, such as anthropic/claude-opus-5-5, openai/gpt-5 or together_ai/moonshotai/Kimi-K3.

karotte models list shows the models Karotte knows by name, each with its output-token ceiling and the reasoning effort levels it accepts. --json prints the catalog as JSON.

karotte models list

It also lists prefixes under which any model id is accepted, such as together_ai/ and fireworks_ai/. A model Karotte doesn't know still runs, with a 64k output-token ceiling and no reasoning effort parameter.

API keys

Without model_api_key, Karotte reads the key from the model provider's variable in your shell when the run starts:

Model Variable
anthropic/..., claude-* ANTHROPIC_API_KEY
openai/... OPENAI_API_KEY
gemini/... GEMINI_API_KEY
mistral/... MISTRAL_API_KEY
together_ai/... TOGETHERAI_API_KEY
xai/... XAI_API_KEY
other <provider>/... <PROVIDER>_API_KEY, e.g. DEEPSEEK_API_KEY, with a warning

To use another key, set model_api_key to the key itself or to a $VAR reference to another variable; karotte run --model-api-key and create-run-config --model-api-key set it for you. If the variable is unset, the run stops with an error, unless it uses the fake model or --prepare-only.

Rubric judge

By default, RubricJudge grades with the run's own model and API key. To grade with a different model, set rubric_judge_model. Karotte then reads its key from rubric_judge_api_key, or from the provider's variable in the table above. If that key is missing, the run still starts.

A model passed to RubricJudge(model=...) in the task overrides both.

The fake model and pt/ models can't grade. For those runs, set rubric_judge_model or pass a model to RubricJudge in the task.

Reasoning effort

reasoning_effort takes:

  • "min" or "max": the model's lowest or highest level, whatever the provider calls it
  • one of the provider's own level names, such as "medium", if the model accepts it
  • null: send nothing and use the provider's default, except for grok models, which get high

Providers don't agree on what the levels are called or how many there are, so Karotte automatically selects the right level if you pass min or max. karotte models list shows available levels per model.

Grok models (xai/) are also always sent temperature 0.7, by both the builtin agent and grok-build. Karotte sets no temperature for any other model.

Limits

Use these fields to limit how much work the student can do:

Field Type Meaning
turn_limit int or null Maximum number of model responses in the run. When it's reached, the run ends with an error.
step_time_limit_seconds float, list of float, or null Maximum wall-clock time per step in seconds, counted from when the step's instructions are sent.
on_step_time_limit "error" or "score" What happens when a step runs out of time. "error" (the default) ends the run with an error. "score" ends the step and scores the student's work so far.
inject_time_remaining_counter bool Tell the student how much time is left. Defaults to true.
step_context_window_limit int, list of int, or null Maximum input tokens per turn in a step, based on what the provider reported for the previous turn.
on_step_context_window_limit "error" or "score" Same as on_step_time_limit, but for the context limit.
inject_context_remaining_counter bool Tell the student how much context is left. Defaults to true.

The time and context limits apply to each step separately. Pass a single value to use it for every step, or a list with one value per step. Steps beyond the end of the list have no limit. For example, step_time_limit_seconds: 600 gives the student ten minutes for each step of a two-step task.

Karotte checks both step limits between turns. It doesn't interrupt a turn in progress, so a step can go over its limit by up to one turn. The context limit uses the input-token count the provider reports for each turn. It measures how large the student's context is, not how many tokens were billed over the whole step. If the provider doesn't report usage, the context limit never triggers.

When a counter is on, Karotte adds a note like Time remaining: 540 seconds or Context remaining: 12000 to the last tool result of each turn, so the model can pace itself. When it's off, the student hits the limit without warning. A counter has no effect unless its limit is set.

Limit support by agent

Agents that run a step in their own process can only enforce the limits their CLI supports.

Limit builtin external CLI agents
turn_limit Yes Yes Per step, through the CLI's --max-turns
step_time_limit_seconds Yes Yes Yes, by killing the CLI's processes
inject_time_remaining_counter Yes Yes No
step_context_window_limit Yes, if the provider reports usage No No
inject_context_remaining_counter Yes, if the provider reports usage No1 No

Agents

The agent produces the student's messages and tool calls for each step. Don't confuse it with the model: model picks which LLM answers, agent picks the harness that calls it. Whichever agent you use, Karotte still runs the task's hooks, judges and artifacts.

Agent What it does
builtin The default. Karotte calls the model through litellm, streams its responses, retries failed calls and runs the tool calls.
external A backend sends each of the model's messages, and Karotte runs the tool calls. Needs backend_uri; see Backend.
CLI agents A coding agent's own CLI, installed in the image, works on each step as the student and calls the tools itself: mistral-vibe or grok-build.

Set agent in the run config:

{
    "agent": "mistral-vibe",
    "model": "mistral/<model-id>",
    "model_api_key": "$MISTRAL_API_KEY"
}

karotte agents list shows the agents your Karotte version supports, and the version it pins for each CLI agent. --json prints the list as JSON.

karotte agents list

CLI agents

builtin and external work with any image. A CLI agent also has to be installed in the image at build time, so setting agent in the run config isn't enough. Add it when you create the environment:

karotte create-env my_env --agent mistral-vibe

or add it to an existing environment and rebuild:

uv run karotte agents add mistral-vibe
uv run karotte build

Both commands record the agent in the environment's .manifest.json. During the image build, karotte agents install installs the pinned version; it only works inside a build. karotte agents remove <name> removes an agent from the manifest, and the next build leaves it out.

CLI agents differ from builtin in a few ways:

  • They run as the student and call the tool server directly. If they bring their own tools, those replace the task's tools that do the same job; their shell replaces bash, for example.
  • They send their own system prompt instead of the task's.
  • reasoning_effort isn't applied, except by grok-build.
  • The student's firewall only lets them reach localhost and the model proxy. Without a proxy (no --proxy and no plugin default), Karotte listens on a local port and forwards the agent's model calls from there to the provider, adding the API key itself. The agent gets a placeholder key, so the real one never reaches the student.
  • mistral-vibe sends its calls to <proxy>/v1 as an OpenAI-compatible endpoint, or through that forwarder to Mistral's API when there's no proxy. grok-build does the same with the proxy, or goes through the forwarder to xAI's API without one.

grok-build specifics:

  • It only runs grok models (xai/); the run config is rejected for any other model.
  • All steps of a run share one Grok session, so later steps see the earlier conversation.
  • Grok doesn't list MCP tools to the model directly. The model finds the task's tools with search_tool and calls them through use_tool (as karotte__<tool>), and the transcript records those calls as use_tool.
  • Grok's web search, image editing, feedback and ask-the-user tools are turned off, along with its telemetry and update checks.
  • It sends temperature 0.7 and reasoning_effort (see Reasoning effort).

Fake model

The fake model replays messages you write in advance, so you can test your tools, hooks and judges without waiting for or paying for a real model. It needs no API key. To use it, set use_fake_model to true and write a get_messages function in src/environment/fake_model.py. The default template includes one:

from karotte import EvaluationRunConfig
from karotte.schemas import Message


def get_messages(config: EvaluationRunConfig) -> list[Message]:
    return [
        Message(role="assistant", content="path: /workdir/.venv/bin/python"),
    ]

Each turn uses the next message in the list. A message with text and no tool calls ends the step, so finish each step with one. If the list runs out, the run fails.

use_fake_model takes the place of whatever agent the config names.

Model endpoints and proxies

karotte run --proxy <url> sends model calls to <url> instead of to the provider. Pass the base URL without /v1:

  • Claude models go to <url>/v1/messages. Karotte sets ANTHROPIC_BASE_URL to <url> inside the container, and litellm adds /v1/messages.
  • Other models go to <url>/v1, so the endpoint needs an OpenAI-compatible API there.

You can use the same flag to point Karotte at any endpoint that serves these paths. The URL is used from inside the container, so localhost means the container, not your machine. Variables like ANTHROPIC_BASE_URL set in your shell don't reach the container.

The endpoint may not require a key and instead inject its own key. Karotte always sends a key, though, so if the key's variable isn't set, --proxy fills in a placeholder for both the run's model and the judge model.

If you don't pass --proxy, Karotte uses the URL a plugin registers under karotte.default_proxy_url (see Plugins). Without such a plugin, calls go directly to the provider, and a CLI agent's calls go through the forwarder described in CLI agents. --no-proxy ignores the plugin's URL.


  1. The external agent doesn't report token usage, so its context counter, if on, always shows the full limit. ↩