Artifacts and transcripts¶
Every run produces a transcript. A task can also save files from the sandbox as artifacts, so you can inspect them after the run.
Artifacts¶
Call save_artifact() from a hook to keep a file or directory:
from pathlib import Path
from karotte import Step, save_artifact
class MyStep(Step):
def post_hook(self) -> None:
save_artifact(self.config, Path("/some/path/report.md"))
A directory is copied with everything in it.
A path that doesn't exist is skipped with a warning.
With save_artifacts: false in the run config, save_artifact() does nothing.
post_hook runs after the step is scored.
If your pre_scoring_hook or judge deletes student files, save them before that.
The default template's collect_submission() already saves the submissions it collects as artifacts.
Warning
save_artifact doesn't validate what it copies.
Only call it on files you have "taken into custody" via something like collect_submission().
See Scoring.
Where artifacts go¶
Without a backend, artifacts are copied into a <run_id>_artifacts/ directory next to the transcript, so out/<run_id>_artifacts/ with the default run config.
If that directory already exists, Karotte appends _2, _3 and so on instead of overwriting it.
karotte run prints where the transcript and artifacts are when it finishes.
With backend_uri set in the run config, artifacts are uploaded to the backend instead.
See Connecting a backend.
Artifacts without code changes¶
To capture extra files without touching the environment, list their absolute container paths under extra_artifact_paths in the run config's extra_config.
Karotte saves them after each step, before the pre_scoring_hook runs.
See Run config.
Transcripts¶
The transcript records every event of a run: messages, tool calls and their results, scores, token usage and errors.
Karotte writes it as JSON to the run config's transcript_file, which karotte create-run-config sets to out/transcript.json.
The file is written when the run ends.
A new run with the same config replaces it.
With -n 3, the runs write transcript_0.json to transcript_2.json.
The transcript holds the run id and a list of events.
Each event has a timestamp and a type.
The full schema is in karotte/schemas/transcript.py.
Streamed message chunks go to live viewers only and aren't saved.
To read a transcript in Python, validate it with the Transcript model:
from pathlib import Path
from karotte.schemas.transcript import Transcript
transcript = Transcript.model_validate_json(Path("out/transcript.json").read_text())
for message in transcript.messages:
print(message.role, message.content)
Viewing transcripts¶
karotte dashboard opens past transcripts in a terminal UI: