Skip to main content

Harness-Aware Evaluation

Reasoning performance in agent systems depends on more than the base model alone. It is also shaped by the agent framework, tool interface, execution environment, and evaluation protocol. Run the same model through two different agent scaffolds and you will routinely see very different accuracy. The scaffold is a measurement instrument, and a leaderboard number that does not name its harness is under-specified.

AlphaDiana treats the harness as a first-class, swappable variable. The benchmark, the model endpoint, and the scaffold are decoupled, so you can hold any two fixed and vary the third. This page explains that decoupling, the direct_llm reference baseline, the "agent scaffold tax," and the macro/micro study split.

Decoupling model from harness

A run is the product of three independent choices:

  • the benchmark (e.g. AIME, GPQA, HLE) under benchmark in the config,
  • the model endpoint (an OpenAI-compatible base URL + model name),
  • the harness (an Agent) selected by agent.name.

The endpoint is a logical evaluation input, but its config key is not uniform across transports. direct_llm, OpenCode, and ZeroClaw use api_base as the provider endpoint. OpenClaw uses lowercase api_base for an already-running agent gateway; its runtime provider endpoint is configured separately as OPENAI_BASE_URL / openai_base_url inside agent.config. See Provider endpoints and agent gateways.

Every harness implements the same contract: the Agent abstract base class in alphadiana/harness/base.py. It defines setup(config: dict), solve(task: BenchmarkTask, sandbox=None) -> AgentResponse, and a no-op teardown(), plus class attributes name and version. The engine sets agent.version, then calls setup() once and solve() per task (alphadiana/engine/runner.py). The uniform response contract makes comparison possible, but swapping harnesses is not generally a one-line edit: endpoint keys, controller/sandbox runtime, prompts, retries, and timeout controls can differ. Report and match those fields before associating an accuracy gap with a harness condition.

Every harness returns the same shape, the AgentResponse dataclass (AgentResponse in alphadiana/harness/base.py). Beyond the core fields (answer, trajectory, raw_output, token_usage, token_entropy_stats, wall_time_sec, metadata) it carries extended observability fields such as reasoning_trajectory, request_messages, response_json, system_prompt, finish_reason, and artifact_manifest, so two harnesses are compared on an identical record schema. Raw harness runtime artifacts are normalized into these fields by alphadiana/harness/proxies/preservation.py, and the result store at alphadiana/analysis/io/result_store.py writes them to disk.

The registry

Harnesses are looked up by string key through AgentRegistry (alphadiana/harness/registry.py), a classmethod-only singleton over a class-level _registry: dict[str, Type[Agent]]. The keys are the values you put in agent.name:

agent.nameHarnessScaffold
direct_llmDirect LLMnone (single-turn chat)
opencodeOpenCodeopencode CLI, multi-turn + tools
openclawOpenClawOpenClaw gateway + ROCK sandbox
zeroclawZeroClawZeroClaw agent loop

AgentRegistry.list() returns the sorted available names; AgentRegistry.get(name) raises KeyError (with the available names embedded in the message) on a miss. Registration is import-triggered, not auto-discovered: runner.py explicitly imports each agent module so the module-level AgentRegistry.register(...) call at the bottom of the file fires. Adding a new harness therefore means both adding the register() call and an import line in the runner.

The direct_llm baseline

DirectLLMAgent (alphadiana/harness/direct_llm.py, name = "direct_llm") is the no-harness reference: a single-turn system+user chat to an OpenAI-compatible endpoint with no tools, no multi-turn loop, and no code execution. It is the floor every scaffold is measured against, useful for establishing a clean baseline before measuring the effect of an agent framework.

It deliberately builds its own client with httpx.Client(trust_env=False) so inherited SOCKS/HTTP proxy environment variables cannot break it. The default system prompt asks the model for a \boxed{} answer, and _extract_answer prefers the boxed value. Reasoning is recovered from reasoning_content / reasoning model-extra fields and from <think>...</think> tags.

By default direct_llm also captures logprobs (capture_logprobs=True) and quantizes them to int16, so the baseline produces token-entropy data for free.

Config

Config keys are resolved from agent.config, with environment fallbacks for the three connection fields (OPENAI_MODEL_NAME, OPENAI_BASE_URL, OPENAI_API_KEY).

KeyDefaultNotes
modelOPENAI_MODEL_NAMEmodel name on the endpoint
api_baseOPENAI_BASE_URLOpenAI-compatible base URL
api_keyOPENAI_API_KEY / EMPTYuse sk-EMPTY, not the literal EMPTY
temperature0.7
top_p
max_tokensautoif unset, GETs {api_base}/models, uses max_model_len - 8192 (fallback 131072)
request_timeout600per-request seconds
streamTrue
stream_total_timeoutdefaults to request_timeout (600 seconds); 0/None disableson hit, answer=None, finish_reason="timeout"
max_retries3exponential backoff with jitter
system_promptboxed-answer prompt
enable_thinking
extra_bodypassthrough to the request body
capture_logprobsTrue
top_logprobs20
logprobs_formatint16int16 or float
warning

For local vLLM, use any non-EMPTY string such as sk-EMPTY. The literal "EMPTY" (case-insensitive) is treated as unset and falls back to the environment.

# configs/examples/direct_llm_gpqa_diamond.yaml (excerpt)
agent:
name: direct_llm
config:
api_key: sk-EMPTY
temperature: 0.6
capture_logprobs: true
python -m alphadiana.cli validate configs/examples/direct_llm.yaml
python -m alphadiana.cli run configs/examples/direct_llm.yaml \
-o run_id=concept_aime_directllm_t1_k1 \
-o benchmark.config.max_tasks=1 -o num_samples=1

The agent scaffold tax

A harness is not free. The same multi-turn loop, tool documentation, and skill preamble that can lift accuracy also injects tokens, latency, parsing steps, and new failure modes (timeouts, truncation, tool-call malformation, session overflow). On easy or self-contained problems the scaffold often costs more than it adds, so the agent can score below its own direct_llm baseline on the identical model and benchmark. We call that gap the agent scaffold tax.

This is exactly why direct_llm is the reference line and not just another harness. The interesting quantity is the signed difference

harness_accuracy − direct_llm_accuracy

which can be positive (the scaffold pays for itself) or negative (the scaffold taxes the model). Reporting an agent number without the matching direct_llm number hides which side of zero you are on.

Macro and micro studies

AlphaDiana supports two complementary modes of harness-aware study:

  • Macro — compare whole harness conditions (direct_llm vs opencode vs openclaw vs zeroclaw) on a fixed model and benchmark, then compare end-to-end accuracy and cost. The harness-specific runtime and transport fields are part of each condition; keep shared budgets fixed and disclose the fields that cannot be identical.

  • Micro — compare a matched scaffold intervention (tool exposure, skill loading, memory behavior, and any associated prompt changes) against its baseline condition. Micro studies need surgical control over the request stream, which is what the proxies provide (below). The Tool / Skill / Memory axes page covers the ablation design.

Proxies for micro intervention

Two distinct proxies live under alphadiana/harness/proxies/, with opposite lifecycles. They are not wired together.

  • LogprobCaptureProxy (logprob_proxy.py) is an in-process threaded HTTP proxy that the harness agents spin up so the in-sandbox CLI agent's OpenAI-bound traffic is routed through it. It injects logprobs=True / top_logprobs, can normalize messages, and captures request/response summaries for observability. This is the mechanism that gives the agent harnesses the same logprob data direct_llm gets directly.

  • tool_filter_proxy.py is a standalone aiohttp CLI proxy (not imported by any agent) for experimental intervention on the request stream. It mutates POST /v1/chat/completions to filter tools by name (--allow / --block regex), strip or replace the system prompt (--harness-strip is section-aware per harness via harness_strip.py, used for the "no-tools" micro-cell), strip AlphaDiana's prepended user-intro block, and apply OpenRouter provider / reasoning overrides.

python -m alphadiana.harness.proxies.tool_filter_proxy \
--port 9100 --upstream http://127.0.0.1:8000/v1 --api-key sk-EMPTY \
--harness-strip zeroclaw
note

Unlike direct_llm (which forces trust_env=False), tool_filter_proxy uses trust_env=True so it honors HTTP(S)_PROXY to reach upstreams on locked-down hosts.

Skills

A related, prompt-level lever is skills: file bundles under alphadiana/harness/skills/<name>/, each with a top-level SKILL.md (YAML frontmatter with name + description). Shipped bundles include advanced-maths (a symbolic/numeric protocol) and anthropic-bundle.

Select one with agent.config.skill_folder, which accepts three forms: an empty value disables it; an absolute path is used as-is; a bare name resolves to alphadiana/harness/skills/<name>/.

HarnessHow the bundle is mounted
opencodeshutil.copytree into <workdir>/skills/<name> so the read tool can index it
zeroclawsandbox.upload() of each file into <workspace_dir>/skills/<name>/
warning

Skills are not auto-injected into context. The system prompt must instruct the model to read the mounted SKILL.md. Skill efficacy is therefore a prompt-level concern, which makes it a clean micro axis to ablate.

See also