How to Run Benchmarks Concurrently#
CorralRunner schedules trials with asyncio and bounds task execution and
evaluation with the configured concurrency limits.
result = await runner.run(
"benchmark-2026-08-14",
task_ids=["task_1", "task_2", "task_3"],
trials_per_task=4,
k_values=[1, 2, 4],
max_parallel=16,
max_parallel_per_task=4,
max_parallel_evaluations=2,
max_parallel_total=18,
max_parallel_by_model={"gpt-5": 8},
max_parallel_by_environment={"wetlab": 1},
max_parallel_evaluations_by_environment={"wetlab": 1},
)
The following limits apply throughout the benchmark:
| Argument | What it bounds |
|---|---|
max_parallel |
All running task attempts in the benchmark. Defaults to 4. |
max_parallel_per_task |
Attempts of the same task. |
max_parallel_evaluations |
Concurrent evaluations. Defaults to max_parallel. |
max_parallel_total |
Executions and evaluations combined. Defaults to max_parallel + max_parallel_evaluations, or 2 * max_parallel when both phase limits use their defaults. |
max_parallel_by_model |
Attempts using each model ID. |
max_parallel_by_environment |
Attempts using each environment ID. |
max_parallel_evaluations_by_environment |
Evaluations using each configured environment name (or its runtime ID when no name is available). Environments without an entry use only the global evaluation limit. |
Evaluation runs independently after each task attempt finishes, so scoring does not consume execution capacity or delay dependency-ready tasks. The runner still waits for all evaluations before returning the report. The total limit provides a shared ceiling while retaining separate execution and evaluation limits. No environment-specific evaluation limits are enabled by default.
With the default max_parallel_per_task=1, a task's next trial starts only
after its preceding trial has been evaluated. This guarantees that reflective
agents receive that trial's score and state as previous_evaluation and
previous_state. Setting max_parallel_per_task above 1 explicitly allows
same-task trials to overlap; those trials can only observe evaluations that
finished before they started, so the immediately preceding trial is not
guaranteed to be available.
These benchmark limits are the concurrency controls. Corral does
not serialize requests that share an agent_id; a registered agent instance
must be reentrant and keep task-specific scratch in AgentSession or local
variables.
Per-task model and environment IDs belong to task metadata:
from corral import BenchmarkTaskMetadata
tasks = {
"wetlab_1": BenchmarkTaskMetadata(
agent_id="react-gpt5",
environment_id="wetlab",
model="gpt-5",
max_iterations=20,
)
}
Dependencies are metadata too. The runner makes a selected task set
dependency-closed automatically before starting trials. The runner waits
for same-round parents, propagates their runtime outputs, and marks descendants
unreachable when a parent produces no valid output. Pass
include_dependencies=False to request strict validation instead.
tasks = {
"prepare": BenchmarkTaskMetadata(
agent_id="react-gpt5",
environment_id="preparation",
),
"analyze": BenchmarkTaskMetadata(
agent_id="react-gpt5",
environment_id="analysis",
dependencies=("prepare",),
),
}
Task retries and cancellation are handled in the runner. Each task's commit history contains typed setup, action, tool-effect, and terminal commits. The execution head advances atomically, so a retry restores the latest projection and finishes an already-proposed Action before asking the agent for a new decision. Benchmark scheduling lives in the invoking process; task state is persisted in the commit store.