Skip to content

Concepts

Environments and the Karotte library

An environment is the isolated world a model (we call it the student) works in: the tasks, the tools it can use, the data and dependencies it needs, and the judges that score it. You define it as a Python project, created with karotte create-env, and Karotte builds it into a container image that runs in a sandbox. In the Python project it is the environment package, which depends on the karotte library.

Karotte provides everything an environment needs:

  • the harness that runs a task, talks to the student and scores the result
  • tools for the student, such as bash, view_lines_in_file and replace_in_file
  • judges, such as ExecutableJudge and RubricJudge
  • transcripts, streamed to the terminal UI or a backend
  • the karotte CLI, which builds images, runs tasks and shows transcripts

The environment adds everything else on top:

  • its tasks, in src/environment/tasks/
  • a Containerfile that builds it into a container image
  • data and dependencies for the student and for scoring

Since Karotte is a dependency, you get its fixes and features by updating it; see Updating. The files create-env generates come from templates.

Terminology

Term Meaning
Student The model under test. Its commands run as an unprivileged student user with resource limits, a firewall and a disk quota. See Student resources.
Task A verifiable unit of work that the student must complete. See Tasks and steps.
Step One part of a task. A step consists of instructions for the student and a judge that scores the result.
Judge Decides how well the student did in a step. It returns a score and whether the task should continue. See Scoring.
Agent What drives the student through the task. The builtin agent calls the model itself through litellm; others get the model's messages from a backend or run a CLI coding agent baked into the image. See Agents.
Image The environment built into a container image by karotte build or karotte run.
Sandbox One copy of the image executing one run: a VM or a container, depending on the runtime.
Runtime What runs the image: a VM (apple-container, firecracker) or a container engine (docker, podman, docker:gvisor, nerdctl). See Runtimes.
Run config A JSON file that names the task, the model, its API key and options such as hints and limits. See Run config.
Transcript The record of a run: every message, tool call and score. See Artifacts and transcripts.