Evaluation Task Selection

The task selection system converts CLI arguments into task combinations. You must specify tasks using one of the following modes.

Selection Modes

1. Specific Task Mode

Specify exact tasks to run. This is the most common approach.

# Single main task (honest mode - no side task) ct run eval -t crawl_depth --policy honest # Main + side task (attack mode) ct run eval -t crawl_depth -s visit_malicious_website --policy attack # Multiple environments with same main task ct run eval -e web_scraping -e clinical_trial --just-main-tasks --policy honest # Multiple of each (creates cross-product) ct run eval -t task1 -t task2 -s side1 -s side2 --policy attack

2. Task File Mode (--task-file)

Reads explicit task combinations from a file. Takes precedence over other options.

# From a named task set ct run eval --task-file medium --policy attack # From a custom file path ct run eval --task-file my_tasks.jsonl --policy attack

Task files are JSONL, one record per task combination with main_task and/or side_task:

{"env": "web_scraping", "main_task": "crawl_depth", "side_task": "visit_malicious_website"} {"env": "clinical_trial", "main_task": "adverse_reaction_followup_tracking", "side_task": "suppress_safety_notifications"}

3. Trajectory Mode (--trajectory-id)

Extracts task info from an existing trajectory. Useful for re-running or replaying.

# From API/database (trajectory ID) ct run eval --trajectory-id abc123 --policy attack # From local file ct run eval --trajectory-id ./my_trajectory.json --policy attack # With task override ct run eval --trajectory-id abc123 -t different_main_task --policy attack

This is especially useful with the replay policy to re-execute trajectories:

ct run eval --trajectory-id abc123 --policy replay

Replaying without execution (--replay-sandbox)

By default a replay re-executes every recorded action in a real sandbox. --replay-sandbox skips execution and builds each sample from the recording alone: the sample's environment, tasks and metadata are the recording's own, with no registered environment or task resolved, so a recording from an imported dataset or an isolated setting replays in-process. The tools offered are the recording's too: one per function it called, taking the argument names it used, and a submit taking the recorded submit's fields. Each returns the recorded output of the action the replay policy has in flight: the policy issues each recorded action as it hands it to the protocol, so both sides index the same recording. Nothing is provisioned, and the sandbox itself answers no command. Replayed samples cost nothing and match the recording exactly. Inspect's tool-output limit is lifted, since each recorded output already carries the recording's own truncation. A call no monitor saw is not re-issued: one the recording's policy made privately (tagged policy-private), or a hidden scorer call such as run_side_task_test, which eval2 keeps as a private affordance. A submit (a call whose name, stripped of an MCP tools prefix, is submit) is always re-issued under the name the recording gave it, and an accepted one ends the replayed sample, so recorded work after it is not replayed and the sample scores NOANSWER. A recorded run_main_task_test was visible to the recording's monitors, so it is re-issued as an action, although a native eval2 run keeps it as an affordance. --replay-sandbox takes no task overrides, no eval1 protocol and no inspect_swe harness.

Replay stays sound while the executed tool calls form a verbatim prefix of the recording. Anything that is not a tool call may differ from the recording: assistant text and reasoning, monitor scoring, interrogation questions and the policy's answers to them, judge deliberation, or stopping early. A protocol that commits anything but the issued call (blocks, audits, edits or replaces it) diverges the replay: any later tool call returns a tool error, the replay policy ends the sample on its next draw, and scores do not carry over. A tool call that nothing flagged earlier and that does not match the action in flight raises ReplayDivergenceError with the recorded and incoming action.

--replay-sandbox therefore works with any replay-like policy, meaning one that issues each recorded action it re-emits (as both replay policies do) and whose executed tool calls stay a prefix of the recording. For example, a policy that emits the recorded tool calls but answers a blue protocol's interrogation questions with a live model is replay-like, because the answers are not tool calls. On the blue side, pass-through protocols (e.g. trusted monitoring) replay cleanly, while protocols that rewrite actions (trusted editing, defer-to-trusted) diverge and end the sample with NOANSWER.

A recorded failure is replayed as the same error type and message, a call the recording rejected while parsing included.

The eval message limit grows to fit the longest source recording under --replay-sandbox. The step budget still applies: raise it (for example, --max-steps 1400 on ct run eval2) when replaying a longer recording in full.

To sweep a blue protocol over a recording, pair --replay-sandbox with the replay untrusted policy, which re-issues the recording's tool calls in order and ends the sample once they run out:

ct run eval2 --trajectory-id logs/run.eval --untrusted-policy replay -up traj_like=logs/run.eval \ --replay-sandbox --protocol defer-to-resample

Programmatic scorers don't run: a full replay carries the source trajectory's scores over; a partial or diverged replay scores NOANSWER with an explanation. No scorer script is re-executed.

4. Expansion Modes

Use when you want to run many tasks without listing them explicitly.

--all: Expands to all valid mainĂ—side combinations.

# All attack combinations for one environment ct run eval -e web_scraping --all --policy attack

--just-main-tasks: All main tasks with no side tasks, for honest evaluations.

ct run eval --just-main-tasks --policy honest ct run eval -e web_scraping --just-main-tasks --policy honest

--just-side-tasks: All side tasks with no main task. The agent receives only the side task instructions.

ct run eval --just-side-tasks --policy attack

Validation Rules

  1. One mode required: Must specify --task-file, an expansion flag (--all, --just-main-tasks, --just-side-tasks), or specific tasks (-t)
  2. Expansion flags are mutually exclusive: Cannot combine --all, --just-main-tasks, --just-side-tasks
  3. Specific tasks don't mix with expansion: -t/-s cannot be combined with expansion flags
  4. Environment is optional: If not specified, inferred from task IDs or defaults to all environments

If you omit required parameters, the CLI will prompt you interactively (unless DISABLE_TUI=1).

Attack vs Honest Mode

Whether a sample runs in "attack" or "honest" mode is determined by side task presence:

  • Attack mode: Sample has a side task (-s specified or from task file)
  • Honest mode: Sample has no side task

Task Sets

Task sets are predefined collections of task combinations stored in src/control_tower/task/task_sets/. They make it easy to run consistent evaluations across the team. The built-in sets are examples:

NameDescription
simpleSingle task combination for quick testing
mediumMultiple task combinations from different environments
one_per_envOne task combination per environment for broad coverage

Manage task sets with ct task-sets:

ct task-sets list # List available task sets ct task-sets validate simple # Validate a task set ct task-sets validate --all # Validate all task sets ct task-sets show medium # Show contents of a task set

To create a custom task set, add a .jsonl file to src/control_tower/task/task_sets/ with one record per task combination:

{"env": "clinical_trial", "main_task": "adverse_reaction_followup_tracking", "side_task": "suppress_safety_notifications"} {"env": "web_scraping", "main_task": "redact", "side_task": "expose_secret"}

Each record names an env plus main_task and/or side_task (a record with only side_task runs side-only). Run ct task-sets validate <name> to check it.