Run lifecycle¶
This page walks through one karotte run, from the command line to the finished transcript.
The first part happens on your machine, the rest inside the sandbox that the runtime starts.
On your machine¶
- Read the config.
--configcan be JSON or a path to a JSON file. Karotte looks up the API keys in your shell (see API keys). Then the run config preprocessors from plugins can rewrite the config, in order of their entry point names. - Pick the runtime. This is
--runtimeif you passed it, otherwise your platform's default for the task's hardware. Karotte checks that the runtime is installed and set up before it goes any further. An explicit--runtimeis checked before step 1, since it doesn't depend on the task. - Split parallel runs. With
-n N, runigets the run id<run_id>-<i>, the transcript file<name>_<i>.json, and the websocket port plusi. - Build the image and tag it
karotte.--devskips this step; see Development loop. - Remove leftover containers from earlier runs.
Karotte only removes containers whose names start with
karotte_run_followed by the longest prefix all the run ids share, so commands with different run ids don't touch each other's containers. -
Start one sandbox per run, named
karotte_run_<run_id>. Inside it, Karotte runs itself again:The directory that contains
transcript_fileis mounted into the sandbox, so the transcript and the artifacts end up next to each other on your machine. Onfirecracker, that directory goes into the VM on a drive and comes back out after the VM shuts down. The websocket port is published to your machine. -
Show progress. The terminal UI connects to each run's websocket and shows its events as they happen. With
--no-ui, Karotte prints every run's output straight to the terminal instead. - Report. Karotte prints where each transcript and artifact directory is.
With
--no-ui, it exits with a non-zero code if any run ended with an error.
Inside the sandbox¶
- Harden the process. Karotte removes every directory the student can write to from
PATH,LD_LIBRARY_PATH,LD_PRELOADandLD_AUDIT, and restarts itself. It also removes group and world write permission from mount points the runtime left writable, except the workdir and the temp directories. - Load the task from the
environmentpackage. - Start the tool server. An HTTP MCP server starts in a subprocess, on the host and port from
mcp_server_config. It runs as root; see Tools. - Create the agent and set up the student's firewall. The student can reach localhost and the sandbox's own addresses (see
KAROTTE_STUDENT_NETWORKin Runtimes). It can never reach the websocket port. It can only reach the tool server and the model proxy when a CLI agent needs them. If the firewall rules were applied, Karotte then tries to reach a few outside addresses as the student, and refuses to run if any of them answers. Withuse_fake_model, the fake model takes the agent's place here. - Start the event streams: the terminal output, the websocket, and the backend if
backend_uriis set. - Set up the task:
- Karotte calls
task.configure_tools()and registers the task's tools with the tool server. Tools that a CLI agent brings itself are skipped. - If a file already exists at
transcript_file, it's deleted. - Karotte records a
TaskStartedEvent. - The default student resource limits are applied, and Karotte logs which kind of confinement is in effect.
- Karotte calls
task.pre_hook(). Whatever it returns goes into the transcript as the metadata of aTaskPreHookCompletedEvent. - Karotte adds the system message from
task.system_prompt. It skips this when the property returnsNone, and for CLI agents, which send their own.
- Karotte calls
- Run the steps. For each step:
- Karotte records a
StepStartedEvent. - The step's instructions go to the agent as a user message, with any
extra_configchanges applied; see Extra config. - The agent works on the step (see The model loop).
- Files listed in
extra_config["extra_artifact_paths"]are saved as artifacts. - Karotte calls
step.pre_scoring_hook(), thenstep.judge.evaluate(transcript), which produces aScoringEvent. If either raises aStudentMisbehaviorError, the step scores 0; see Scoring. - Karotte calls
step.post_hook(). - Karotte records a
StepCompletedEvent. If the judge said not to continue, the remaining steps are skipped and the run counts as failed.
- Karotte records a
- Finish. Karotte records a
TaskCompletedEventwith statuspassedorfailed. If an exception happens anywhere in the run, Karotte records anErrorEventinstead and the run ends with statuserror. - Write the transcript to
transcript_file, even after an error. Then Karotte gives the transcript and the artifacts to the owner of the directory they were written to, so a container running as root doesn't leave root-owned files on your machine. Finally, the tool server stops.
The hooks are described in Tasks and steps.
sequenceDiagram
participant CLI as karotte run (sandbox)
participant MCP as Tool server<br/>(subprocess)
participant Runner as Runner
participant Agent as Agent
CLI->>CLI: Harden, load task
CLI->>MCP: Start tool server
CLI->>CLI: Firewall the student
CLI->>Runner: run()
Runner->>Runner: task.configure_tools()
Runner->>MCP: Register task tools
Runner->>Runner: Default resource limits
Runner->>Runner: task.pre_hook()
Runner->>Runner: System message
loop For each step
Runner->>Agent: Step instructions
loop Model loop
Agent->>Agent: Next model message
opt Tool calls
Agent->>MCP: Call tools
MCP-->>Agent: Results
end
end
Runner->>Runner: step.pre_scoring_hook()
Runner->>Runner: step.judge.evaluate(transcript)
Runner->>Runner: step.post_hook()
alt continue_task is false
Runner->>Runner: Stop (failed)
end
end
Runner->>Runner: TaskCompletedEvent
Runner->>Runner: Write transcript file
The model loop¶
The builtin agent calls the model through litellm with the transcript's messages and the task's tools, and streams the response.
If the response contains tool calls, Karotte runs them one at a time on the tool server, adds each result to the transcript, and calls the model again.
A turn without tool calls ends the step.
Some turns have neither text nor tool calls, because they were cut off at the output-token limit or only contain reasoning. These don't end the step. Karotte asks the model to continue instead, up to three times in a row, and stops with an error on the fourth.
The external agent runs the same loop, but gets each message from the backend instead of calling a model.
It waits up to two hours for each message.
A CLI agent runs its own program once per step and calls the tools itself.
Karotte checks the limits on turns, time and context between turns.
What the run produces¶
| Output | Where |
|---|---|
| Transcript file | transcript_file, which is out/transcript.json by default. Written once, when the run ends. |
| Artifacts | out/<run_id>_artifacts/, next to the transcript. If that directory already exists, Karotte adds _2, _3 and so on. With backend_uri set, they're uploaded to the backend instead. See Artifacts and transcripts. |
| Terminal output | Every event, printed as it happens. The terminal UI shows it per run; with --no-ui, it's printed directly. |
| Websocket | Every event, on the port from websocket_config, for the terminal UI. The student can't reach it. |
| Backend | Every event, sent over HTTP when backend_uri is set, plus the run's status and score. See Backend. |
To look at the transcripts in a directory later, run karotte dashboard out/.
Stop before the steps¶
--prepare-only sets up the sandbox the same way a real run does, up to and including the pre-hook and the system message.
Then it keeps the sandbox running without starting any steps or writing a transcript.
Use it to look around in a prepared sandbox:
The sandbox is ready once /tmp/karotte_prepared_<run_id> exists inside it.
Press Ctrl-C to stop it.
It doesn't work with -n greater than 1, and it doesn't need a model API key.