Configure a recipe manually only when a benchmark page names an alternative recipe, troubleshooting requires a
restricted match set, or your team needs custom adaptation logic. Otherwise, keep the default automatic matching.
How Recipes Work
Recipes describe task requirements through the shared execution plan. Benchmarks with a unified default Recipe use the same policy on Docker, Daytona, Modal, and other providers supporting those requirements. Each Environment translates the plan into its own settings. For example, a recipe can:- select an image from the task id or an image address recorded in the task;
- set the working directory to a benchmark path such as
/testbed,/workspace, or/root; - declare CPU, memory, disk, or GPU requirements for the selected Environment to validate and translate;
- add image and network settings when scoring requires a separate sandbox.
deepswe, swebench_verified, and terminalbench2. The same design also covers frontier_engineering, skillsbench, swe_marathon, swebench_pro_verified, taubench, and wildclawbench. Use benchmark-level IDs in configurations that selected earlier provider-specific IDs. applied_recipes records the registered ID; optional policies such as terminalbench2_ac require explicit selection. BrainArena and OSWorld retain their Docker-specific recipes. The Environment must still support each task’s image, resource, OS, and network requirements.
Recipes do not choose a harness or model for you, run tasks, or score results. To make the selected combination work,
they may adjust installation or execution settings for the harness.
When to Configure a Recipe Manually
--recipe does not force the named recipe to run. It only allows the listed ids to participate in matching. A recipe
must still match the current benchmark, environment, and task information. When this option is omitted, AgentCompass
matches from all available recipes automatically.
If several recipes match, AgentCompass applies all of them. To confirm what was applied, look for Recipe matched in the
DEBUG run log.
Examples
The following command usessample_ids to run one
SWE-bench Verified instance. It omits --recipe; AgentCompass matches a built-in recipe from swebench_verified and
modal automatically.
/testbed. Other built-in
adaptations include:
Override Recipe-Provided Values
To use a custom registry image, pass the commonsetup.image field through --env-params:
Trusted External Recipes
This is an advanced workflow for team-defined adaptation logic.--recipe-dir loads an external recipe for the current
run:
External package requirements
External package requirements
- The directory must be a Python package containing
__init__.py. - The root module must export a non-empty
RECIPE_CLASSESlist or tuple. - Every item must be a concrete subclass of AgentCompass’s
BaseRecipebase class, with a uniqueidand a zero-argument constructor. - Relative paths resolve from the current working directory.
agentcompass launch has no --recipe or --recipe-dir option; place the corresponding fields in the orchestration
file. Explicit CLI or SDK lists replace the corresponding configuration-file lists rather than appending to them.
Duplicate recipe ids fail during loading. If multiple recipes modify the same image, working directory, or network
setting, AgentCompass does not resolve the conflict automatically, so do not load implementations with overlapping
responsibilities together.
