IMO-AnswerBench
IMO-AnswerBench evaluates mathematical-answer extraction and scoring on Hwilner/imo-answerbench.
Prerequisites
Run from the repository root:
source scripts/activate.sh
export PYTHONPATH=$PWD
export OPENAI_BASE_URL=https://models.sjtu.edu.cn/api/v1/
export OPENAI_API_KEY=sk-...
export OPENAI_MODEL_NAME=minimax-m2.5
The benchmark loads from HuggingFace. If the default mirror is slow, set HF_ENDPOINT before running.
If the host cannot reach huggingface.co directly, set:
export HF_ENDPOINT=https://hf-mirror.com
The April 22, 2026 local-vLLM Qwen follow-up on this repo required that
override before imo_answerbench could load at all.
Supported Modes
| Mode | Status | Config |
|---|---|---|
direct_llm | smoke config available | configs/examples/directllm_minimax_imo_answerbench.yaml |
opencode | smoke config available | configs/examples/opencode_minimax_imo_answerbench.yaml |
openclaw | smoke config available | configs/examples/openclaw_minimax_imo_answerbench.yaml |
zeroclaw | smoke config available | configs/examples/zeroclaw_imo_answerbench.yaml |
The DirectLLM, OpenCode, and OpenClaw configs pin dataset_index: 367. The
ZeroClaw config uses only max_tasks: 1, so it selects the first eligible row.
Additional April 18/19, 2026 pilot configs are also checked in:
configs/examples/directllm_qwen35_27b_imo_answerbench_pilot.yamlconfigs/examples/openclaw_qwen35_27b_imo_answerbench_pilot.yamlconfigs/examples/opencode_qwen35_27b_imo_answerbench_pilot.yaml
These pilot configs use max_tasks: 3 and the OpenRouter slug
qwen/qwen3.5-27b for the logical model target Qwen/Qwen3.5-27B.
The dedicated opencode pilot config intentionally omits the smoke
dataset_index: 367 pin so it can load three distinct tasks.
Full Runs
This checkout does not ship IMO-AnswerBench full-run configs. Start from the
appropriate checked-in smoke or Qwen pilot config, remove the bounded task
selector, and review the model, output, timeout, and concurrency contract before
a full evaluation. Keep scorer.name: imo_verify.
Current OpenRouter free-text follow-up on April 22, 2026 uses
nvidia/nemotron-3-nano-30b-a3b:free.
Early full-run evidence:
full_20260422_openrouter_nemotron_3_nano_30b_a3b_imo_answerbench_directllm_r1writes normal task JSONs, but many early answers collapse to short scalar outputs even on symbolic-ground-truth items..._opencode_r1is the healthiest current path and already wrote normalscore=1andscore=0task records..._openclaw_r1still preservespredicted=nullwithmetadata.partial_reasoning_only=true..._zeroclaw_r1still fails the first task withscore_status=runtime_errorandmetadata.failure_reason=empty_response
DirectLLM
DirectLLM calls the OpenAI-compatible endpoint directly and scores the returned answer with imo_verify.
Checked-in IMO configs now enforce scorer.name: imo_verify, and config validation rejects
benchmark.name: imo_answerbench with any other scorer. Historical IMO runs scored with
math_verify should be treated as non-canonical audit artifacts, not current support evidence.
imo_verify is still a repo-local heuristic scorer, not an external official
verifier. The current implementation is intentionally conservative: it blocks
the old symbolic-to-numeric false positives, but it can still over-split some
comma-heavy answer forms and therefore still carries false-negative risk.
For short OpenRouter canaries on Qwen/Qwen3.5-27B, set
agent.config.extra_body.reasoning.enabled=false when you want terse outputs.
The provider otherwise emits hidden reasoning tokens even for tiny prompts,
which can dominate latency without changing the visible boxed answer. This is a
canary-only override; benchmark defaults still keep reasoning enabled unless
you explicitly change them.
python -m alphadiana.cli run configs/examples/directllm_minimax_imo_answerbench.yaml \
-o run_id=imo_directllm_smoke
Local-vLLM Qwen example with the rollout-plan temperature=0.0 semantics:
source scripts/activate.sh
export PYTHONPATH=$PWD
export HF_ENDPOINT=https://hf-mirror.com
export QWEN_VLLM_API_BASE=http://127.0.0.1:8011/v1
export QWEN_VLLM_API_KEY=EMPTY
python -m alphadiana.cli run configs/examples/directllm_minimax_imo_answerbench.yaml \
-o run_id=full_20260422_imo_answerbench_direct_llm_qwen35_27b_localvllm_mc20_r1 \
-o output_dir=./results/full_20260422_imo_answerbench_direct_llm_qwen35_27b_localvllm_mc20_r1 \
-o max_concurrent=20 \
-o agent.config.model='Qwen/Qwen3.5-27B' \
-o agent.config.api_base="$QWEN_VLLM_API_BASE" \
-o agent.config.api_key="$QWEN_VLLM_API_KEY" \
-o agent.config.temperature=0.0 \
-o agent.config.top_p=0.95 \
-o agent.config.max_tokens=32768 \
-o agent.config.stream=true
This command scales a bounded smoke config with overrides. Review and remove its task-selection limit before treating it as a full evaluation.
On current main, direct_llm captures logprobs by default. For a local-vLLM
Qwen pilot, use
configs/examples/directllm_qwen35_27b_imo_answerbench_pilot.yaml.
OpenCode
OpenCode runs the opencode CLI. The checked-in minimax smoke config includes
OpenCode-specific external agent settings (agent: lean-math and
agent_md_path: alphadiana/context/opencode_lean_math.md) in addition to
agent.config.system_prompt. When comparing harness prompts or running local
Qwen logprob smoke without that external agent layer, override both fields to
empty strings.
Build the controller image once before using the checked-in OpenCode configs:
docker build --network host \
-f alphadiana/benchmarks/terminal_bench2/deploy/dockerfiles/Dockerfile.opencode-controller \
-t alphadiana/tb2-opencode-controller:latest .
python -m alphadiana.cli run configs/examples/opencode_minimax_imo_answerbench.yaml \
-o run_id=imo_opencode_smoke
Local-vLLM Qwen logprob smoke with the external agent disabled:
python -m alphadiana.cli run configs/examples/opencode_minimax_imo_answerbench.yaml \
--redo-all \
-o run_id=phase11_opencode_imo_answerbench_qwen35_27b_logprobs_smoke \
-o output_dir=./results \
-o agent.config.model=custom/Qwen/Qwen3.5-27B \
-o agent.config.model_name=Qwen/Qwen3.5-27B \
-o agent.config.api_base=http://127.0.0.1:8011/v1 \
-o agent.config.api_key=EMPTY \
-o agent.config.controller_mode=docker \
-o agent.config.controller_network=host \
-o agent.config.capture_logprobs=true \
-o agent.config.top_logprobs=20 \
-o agent.config.agent= \
-o agent.config.agent_md_path=
Observed on April 24, 2026: this command wrote
results/phase11_opencode_imo_answerbench_qwen35_27b_logprobs_smoke/tasks/imo-bench-number_theory-068.json
with score=1.0, metadata.logprobs_capture_status="captured", and
token_entropy_stats.n_tokens=11949. The preserved OpenCode config/listing
showed no external agent/*.md file.
The checked-in OpenCode benchmark configs now use Docker controller isolation
by default. If you need the old host-process path for debugging, override
-o agent.config.controller_mode=host.
The smoke config uses timeout: 1800 because shorter bounds can kill valid
slow model output before it reaches scoring. The full Docker setup and
reproduction guidance is in the OpenCode harness guide.
OpenClaw
OpenClaw uses ROCK auto-deploy and the gateway config in alphadiana/harness/openclaw/deploy/.
Benchmark fairness now requires a fresh ROCK sandbox session per task for
openclaw; do not treat older sequential runs that reused one shared session
across tasks as comparable evidence. Current main also skips the OpenClaw
chat-completions warmup by default on benchmark runs because that warmup could
pollute the first task with a leftover READY/bootstrap response.
python -m alphadiana.cli run configs/examples/openclaw_minimax_imo_answerbench.yaml \
-o run_id=imo_openclaw_smoke
ROCK services must be healthy before this run. scripts/activate.sh loads the local ROCK port configuration.
Current OpenRouter free-text full-run evidence on April 22, 2026:
full_20260422_openrouter_nemotron_3_nano_30b_a3b_imo_answerbench_openclaw_r1
continues to reproduce the same failure shape from smoke scale-up:
normal task JSONs are written, but early tasks preserve predicted=null with
metadata.partial_reasoning_only=true.
ZeroClaw
ZeroClaw uses the same ROCK auto-deploy path as the PR23 AIME integration.
Unlike the AIME quickstart in the main README.md, the formal benchmark smoke here is counted only when the task executes inside a ROCK sandbox. Do not clear agent.config.rock_image for the benchmark smoke.
Start ROCK first:
bash scripts/start_zeroclaw.sh
source scripts/rock_env.sh
If another branch is already using ROCK, edit scripts/.rock_ports.env before startup so this worktree gets isolated admin/proxy/redis/ray ports.
python -m alphadiana.cli run configs/examples/zeroclaw_imo_answerbench.yaml \
-o run_id=imo_zeroclaw_smoke
Current OpenRouter free-text full-run evidence on April 22, 2026:
full_20260422_openrouter_nemotron_3_nano_30b_a3b_imo_answerbench_zeroclaw_r1
still fails its first task as score_status=runtime_error with
metadata.failure_reason=empty_response, even though current main preserves
the failure record cleanly.
Reproduce The 2026-04-17 Formal Sandbox Smoke
This is the exact smoke style used for local validation of the ZeroClaw sandbox path. It intentionally forces a fast wrong answer so the run terminates quickly with dashboard X, which is enough for the execution-path smoke criterion.
export OPENAI_BASE_URL=https://models.sjtu.edu.cn/api/v1/
export OPENAI_API_KEY=sk-...
export OPENAI_MODEL_NAME=minimax
python -m alphadiana.cli run configs/examples/zeroclaw_imo_answerbench.yaml \
-o run_id=pr26_formal_smoke_zeroclaw_imo_minimax_rock_cli_box0_20260417_v2 \
-o benchmark.config.dataset_index=0 \
-o agent.config.system_prompt='Smoke test mode: ignore the math problem. Do not use tools. Output exactly $$\\boxed{0}$$ and nothing else.'
Expected result:
- dashboard:
X - task file exists under
results/zeroclaw_imo_answerbench_smoke/<run_id>/tasks/ data[0]in the single-sample task JSON list has noerror- the recorded task uses the dataset row's
Problem ID(for example,imo-bench-number_theory-068); if that field is missing, the loader falls back toimo_0
Observed local verification on 2026-04-17:
- run_id:
pr26_formal_smoke_zeroclaw_imo_minimax_rock_cli_box0_20260417_v2 - result: dashboard
X,predicted=0,ground_truth=3, noerror - execution mode: ROCK sandbox + in-sandbox ZeroClaw CLI
Smoke Selection
The DirectLLM, OpenCode, and OpenClaw minimax configs pin row 367; the ZeroClaw
config is bounded to one task with max_tasks: 1 but does not pin that row.
No IMO-AnswerBench full-run file is checked in; create and validate one before scaling.
Qwen/OpenRouter 3-Task Pilot
Environment:
source scripts/activate.sh
export PYTHONPATH=$PWD
export HF_ENDPOINT=https://hf-mirror.com
export OPENAI_BASE_URL=https://openrouter.ai/api/v1
export OPENAI_MODEL_NAME=qwen/qwen3.5-27b
export OPENAI_API_KEY=sk-...
Commands:
python -m alphadiana.cli run configs/examples/directllm_qwen35_27b_imo_answerbench_pilot.yaml
python -m alphadiana.cli run configs/examples/openclaw_qwen35_27b_imo_answerbench_pilot.yaml
python -m alphadiana.cli run configs/examples/opencode_qwen35_27b_imo_answerbench_pilot.yaml
Observed on April 18/19/20, 2026:
direct_llm:3/3task records written, allscore=1openclaw: not rollout-ready yet- the first selected dataset
Problem ID:score=1 - the second selected dataset
Problem ID: task completed on the April 19 checkpoint rerun, but the current scorer marked a symbolic mismatch asscore=1 - the third selected dataset
Problem ID:score=0withpartial_reasoning_only=true; the partial reasoning trace was preserved and is treated as a normal sample - the path remains blocked on scorer correctness, not on benchmark completion
- the first selected dataset
opencode:- April 19 uploaded quality pilot:
pilot_20260419_qwen35_27b_imo_answerbench_opencode_t3wrote3/3task records, allscore=1 - April 20 default-Docker confirmation rerun:
pilot_20260420_qwen35_27b_imo_answerbench_opencode_t3_docker_defaultwrote3/3normal task records with scores1/0/0
- April 19 uploaded quality pilot:
Local follow-up on April 19, 2026:
rerun_20260419_qwen35_27b_imo_answerbench_openclaw_idx2_r2completed withscore=1; the task JSON kept a non-empty top-levelreasoning_trajectoryplusmetadata.raw_reasoningrerun_20260419_qwen35_27b_imo_answerbench_zeroclaw_idx0_r3failed cleanly with provider transport errors andpredicted=None; the previous startup-log pollution no longer leaked into a fake parsed answer
Local follow-up on April 20, 2026:
- root cause for the ZeroClaw/OpenRouter failure was a config gap, not the
benchmark itself: AlphaDiana was not writing ZeroClaw
provider_timeout_secs, so the CLI fell back to its internal120sprovider timeout and aborted long streamed math responses - the fix keeps
stream=trueand writesprovider_timeout_secs = request_timeoutby default unless explicitly overridden - validation smoke:
debug_20260420_qwen35_27b_imo_answerbench_zeroclaw_idx0_provider_timeout_r1completed1/1on the previously failing first task - repaired 3-task pilot:
pilot_20260420_qwen35_27b_imo_answerbench_zeroclaw_t3_repair_r3completed3/3, all task JSONs now have non-null predictions and no task-levelerror, and the archive was uploaded topilot_run/pilot_20260420_qwen35_27b_imo_answerbench_zeroclaw_t3_repair_r3/
Reviewer-facing evidence for this pilot lives in the recorded run IDs above.