Skip to main content

Quick Start

This page runs your first benchmark end to end with the direct_llm baseline: write (or point to) a config, validate it, run it, and read the report. The direct_llm agent is the simplest path because it needs no sandbox, no gateway, and no ROCK services. It is a single-turn system+user chat to an OpenAI-compatible endpoint with no tools and no multi-turn loop, which makes it the reference baseline against the harnesses.

If you have not installed AlphaDiana yet, start with Installation first.

The flow at a glance​

Every run follows the same pipeline. A YAML ExperimentConfig is loaded, benchmark / agent / scorer are resolved from string-keyed registries, tasks are expanded into (task, sample_index) work items, and each item runs agent.solve -> scorer.score -> ResultStore.append:

config.yaml
→ ExperimentConfig.from_yaml (alphadiana/engine/config/experiment_config.py)
→ ConfigValidator (alphadiana/engine/config/validator.py)
→ Runner.setup() resolves benchmark / agent / scorer from registries
→ Runner.run() load_tasks → per (task, sample): agent.solve → scorer.score → ResultStore.append
→ ReportGenerator.generate → RunSummary
→ Runner.teardown()
→ results/<run_id>.jsonl (+ results/<run_id>/...)

The CLI entry point is the Click group main in alphadiana/cli.py; the orchestrator is Runner in alphadiana/engine/runner.py; results are written by ResultStore in alphadiana/analysis/io/result_store.py.

1. Point the agent at a model​

direct_llm builds its own OpenAI client and resolves the model, endpoint, and key from config first, then from environment variables. For a local vLLM server set:

# These names match configs/examples/direct_llm.yaml and the loader fallbacks.
export OPENAI_MODEL_NAME="Qwen/Qwen3-235B-A22B"
export OPENAI_BASE_URL="http://127.0.0.1:8011/v1"
export OPENAI_API_KEY="sk-EMPTY"

Use sk-EMPTY (any non-EMPTY string) for a keyless local server. The literal string EMPTY is treated as blank by both the validator and the agent. For a hosted provider, set OPENAI_BASE_URL=https://openrouter.ai/api/v1 and a real key.

2. Write (or point to) a config​

A ready-made example ships at configs/examples/direct_llm.yaml. It is also present in the public repository's default branch. The shape is:

run_id: "" # blank → auto-generated uuid4().hex[:12]

agent:
name: direct_llm
version: "1.0"
config:
model: "" # falls back to OPENAI_MODEL_NAME
api_base: "" # falls back to OPENAI_BASE_URL
api_key: "" # falls back to OPENAI_API_KEY
temperature: 0.6
max_tokens:

benchmark:
name: aime
config:
dataset: "HuggingFaceH4/aime_2024"
split: "train"

scorer:
name: numeric
config:
tolerance: 1e-6

max_concurrent: 1
output_dir: "./results"

$VAR / ${VAR} placeholders are expanded from the environment at load time. An unresolved placeholder becomes a blank string, and for direct_llm a blank model / api_base / api_key then falls back to OPENAI_MODEL_NAME / OPENAI_BASE_URL / OPENAI_API_KEY.

Top-level keys​

KeyMeaningDefault
run_idNamespaces all output; blank → uuid4().hex[:12]; / is replaced with _""
agent{name, version, config}; name selects the registered agentrequired
benchmark{name, config}; name selects the dataset loaderrequired
scorer{name, config}; e.g. exact_match, numericrequired
sandboxnull (no sandbox) or {name, config}; not needed for direct_llmnull
max_concurrentWorker count; 1 is sequential, otherwise a ThreadPoolExecutor; integer in [1, 64]1
num_samplesSamples per task; drives Pass@k / Avg@k1
output_dirWhere <run_id>.jsonl and <run_id>/ are written./results
metadataFree-form notes embedded into each record{}

Useful direct_llm agent.config keys​

KeyMeaningDefault
model / api_base / api_keyOpenAI-compatible target; blank → env fallbackenv
temperatureSampling temperature0.7
top_pNucleus sampling; only sent when explicitly setunset (None)
max_tokensOutput cap; if unset, auto-resolved from the server's max_model_len - 8192auto
streamStream the completiontrue
max_retriesRetries on rate-limit / timeout / 5xx with exponential backoff3
system_promptSystem message; default asks for a \boxed{} answerbuilt-in
capture_logprobs / top_logprobsCapture token logprobs; direct_llm captures by defaulttrue / 20

For the full schema and the dotted -o override syntax, see Configuration. For agents other than the baseline, see Harnesses.

3. Validate​

validate runs the same ConfigValidator as run, but without executing anything. It blocks on hard errors (missing agent.name, an agent.version with no digit, max_concurrent out of range, a scorer mismatch, and so on):

alphadiana validate configs/examples/direct_llm.yaml \
-o benchmark.config.max_tasks=1

Expected output:

Config is valid.

You can layer overrides on either validate or run. Each -o a.b.c=value is parsed into a nested dict and auto-cast to bool/int/float:

alphadiana validate configs/examples/direct_llm.yaml \
-o benchmark.config.max_tasks=20

4. Run​

alphadiana run configs/examples/direct_llm.yaml \
-o run_id=quickstart_aime_directllm_t1_k1 \
-o benchmark.config.max_tasks=1 \
-o num_samples=1

The runner loads the tasks, expands them into (task, sample_index) work items, and for each one calls agent.solve then scorer.score then appends a record. Runs are checkpoint-resumable: rerunning this command with the same explicit run_id skips any task that already has a scorer-matching valid_scored record. Provider/runtime errors and no-answer records remain retryable. Timeout-classified outcomes from the current harnesses are scored zero with finish_reason: timeout, so they are checkpoint-complete rather than retried. To ignore the checkpoint and recompute everything:

alphadiana run configs/examples/direct_llm.yaml --redo-all \
-o run_id=quickstart_aime_directllm_t1_k1 \
-o benchmark.config.max_tasks=1 \
-o num_samples=1

When the run finishes, the CLI prints the headline metrics:

Run completed: quickstart_aime_directllm_t1_k1
Accuracy: <model-dependent value>
Mean Score: <model-dependent value>
Pass@1: <model-dependent value>
Avg@1: <model-dependent value>
Tasks: 1/1 completed

5. Read the report​

run prints a summary automatically. To regenerate the Markdown report for a finished run later, point report at the directory containing the .jsonl:

alphadiana report ./results

It loads <run_id>.jsonl, deduplicates by (task_id, sample_index), and prints the same RunSummary metrics broken out per category.

Results layout​

run_id namespaces everything under output_dir. The flat .jsonl is the source of truth for metrics and checkpointing; the <run_id>/ directory holds the manifest and per-task artifacts:

results/
<run_id>.jsonl # one JSON record per (task_id, sample_index)
<run_id>/
run_manifest.json # expected task / sample counts, config metadata
artifacts/<task_id>/... # raw runtime artifacts, logprob sidecars
tasks/<task_id>.json # JSON list of samples (even when n=1)
lifecycle/<task_id>.jsonl # per-item lifecycle events
status/ # status / dashboard files

Each line in <run_id>.jsonl carries the problem, the model's answer, the score, and observability fields:

{
"task_id": "aime_0",
"problem": "Competition math problem...",
"ground_truth": "42",
"predicted": "42",
"correct": true,
"score": 1.0,
"score_status": "valid_scored",
"rationale": "Numeric match within tolerance",
"trajectory": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "...", "thinking": "..."}
],
"token_usage": {"prompt_tokens": 150, "completion_tokens": 1200},
"finish_reason": "stop",
"wall_time_sec": 47.3,
"timestamp": "2026-06-24T..."
}

predicted is the agent's extracted answer, correct is the scorer's verdict, and score_status is what checkpointing keys off: only a scorer-matching valid_scored record counts as complete. Errors are recorded too, with correct: null and score: null, so nothing is silently dropped. Current timeout-classified harness outcomes are the exception to the intuitive "no-answer means retry" rule: they are normalized to score: 0, correct: false, and score_status: valid_scored.

Accuracy, Pass@k, and Avg@k​

The report metrics come from ReportGenerator.generate (alphadiana/analysis/report.py):

MetricDefinition
Accuracycorrect / scored — fraction correct among records that were actually scored
Accuracy (total)correct / expected_sample_count — denominator is the planned count, so unfinished or errored samples count against it
Mean ScoreMean of the score field over scored records (equals accuracy for binary scorers)
Pass@kFraction of unique tasks with at least one correct sample (k = num_samples)
Avg@kPer task, the correct-sample rate n_correct / num_samples, then averaged across tasks

With num_samples: 1 (the default), Pass@1 and Avg@1 both equal Accuracy because there is exactly one sample per task. They diverge only with multi-sample runs. With num_samples: 4, for example, Pass@4 rewards getting the answer at least once across four draws, while Avg@4 measures average per-draw reliability. Choose the sample count required by the study protocol; the validator only requires a positive integer.

Next steps​

  • Try a different benchmark: GPQA-Diamond or AIME. Set num_samples above 1 only when you want a multi-sample Pass@k / Avg@k protocol.
  • Swap the baseline for a real agent scaffold: see the harnesses (opencode, openclaw, zeroclaw), which add tools, sandboxes, and memory.
  • Tune the run YAML, overrides, and run-id conventions in Configuration.