Skip to main content

Evaluation Axes: Tool, Skill, Memory

AlphaDiana can compare matched harness conditions that change tool exposure, skill loading, or memory behavior. These are useful controlled comparisons, but they are not always literal one-variable interventions: disabling tools can also replace prompt text, and enabling a skill or memory mode can add instructions as well as runtime state. There are three axes:

  • Tool — whether the harness exposes its native tools (file, bash, sandbox).
  • Skill — whether a skill bundle is loaded into the harness.
  • Memory — whether the harness keeps a persistent memory store, and at what scope.

For a defensible comparison, keep the model, sampling parameters, benchmark, scorer, and unrelated harness settings fixed, then document the complete condition bundle that changed. Treat the score delta as an association with that bundle, not as an isolated causal contribution from a single capability.

A useful reference point is the Direct LLM baseline (direct_llm harness), the model without an agent scaffold. A harnessed cell below a matched DirectLLM cell may be described as an agent scaffold tax, provided both cells use the same model and evaluation protocol. This checkout does not contain the raw result artifacts needed to substantiate a particular headline delta; use the result store from the run being reported. See Harnesses for the four harnesses themselves.

What "off" means per axis

"Off" is defined per axis and the definitions are not interchangeable.

Axis"On""Off"
ToolHarness exposes its native tools and uses the corresponding prompt conditionTools are filtered and the system/user prompt condition may also be rewritten
SkillA skill bundle is mounted and named in the promptNo skill bundle is loaded and the skill-introduction prompt is absent
MemoryHarness's native memory backend is enabled (persistent_memory: true)Harness defaults with no memory store and no memory-encouraging prompt

Note that Tool-off and Memory-off are not the same configuration. Tool-off strips every tool and rewrites the system prompt — it is the floor on what the harness can do. Memory-off keeps the harness's native tools and is the correct baseline for measuring memory's contribution. Conversely Tool-on (harness default) and Memory-off are the same data point reported under two names for table self-containment.

For the Skill axis mechanics (skill bundles mounted via agent.config.skill_folder), see the Harnesses overview. The rest of this page covers the Memory axis, the only axis that adds a second dimension (scope) on top of on/off.

The Memory axis

Tool exposure and skill loading are one-shot on/off edits. Memory adds a scope dimension: how long the store survives before it is cleared. Widening scope makes memory less monotonic, not more, which is why it is the deliberate counterexample to the framework's monotonicity thesis.

Scopes

The experimental design discusses three possible persistence scopes. Only the checked-in Cross-Task ZeroClaw config is directly runnable from this checkout.

ScopeWhat persistsnBoundary
Cross-SampleIntended: memory accumulates across samples of the same problemmultipleRequires task-scoped namespacing or an explicit reset between problems; the runner does not provide this boundary
Cross-TaskMemory accumulates across different problems run sequentiallyconfigurableStore remains available within the run
TransferIntended: build a store on one dataset, then evaluate with recall-on / store-offconfigurableRequires separately controlled build and frozen evaluation stages

Static paper figures may illustrate proposed or previously reported comparisons, but they are not current support evidence. Do not quote exact memory deltas from this page without the corresponding run IDs and task-level artifacts.

Mechanisms are harness-native, not a shared framework

There is no runner-level snapshot framework. Each harness uses its own native memory backend, which is the point: the same "memory on" knob produces opposite signs because the harness, not the model, decides whether memory anchors behavior (OpenCode) or injects low-relevance noise (OpenClaw, ZeroClaw).

HarnessBackendMechanism
OpenCodeOpenCode session chain + per-task /compact; persistent HOME (opencode.db sqlite) bind-mounted at {workdir}/.controller-homeCompaction summaries carry the agent's own prior clean solves as a behavioral template — anchoring, not knowledge transfer
OpenClawmemory-lancedb vector plugin, backed by an OpenAI-compatible embedding endpoint (for example a local vLLM serving an embedding model)One distilled [fact] sentence per problem via a forced store-turn; recall injected under a <relevant-memories> "untrusted historical data" guard
ZeroClawsqlite + vector (embeddings via an OpenAI-compatible endpoint), memory_store / memory_searchKeyed [math] insights; self-pollutes by also storing its own system prompt as a [conversation] memory and recalling it as junk

Configuration

Memory is enabled per harness through agent.config. The master switch is the same across all three harnesses; the remaining keys are harness-specific.

KeyHarnessPurpose
persistent_memoryall threeMaster memory on/off
compact_after_taskOpenCodeRun /compact after each task
fresh_sessionOpenCodeFresh session per task; fills and injects the harness memory bank instead of chaining sessions
memory_freezeOpenCodeTransfer mode: frozen tasks fork from a post-build HOME snapshot
oracle_feedbackall threePost-solve reflection turn that reveals ground_truth (oracle-feedback v2)
context_limit / output_limitOpenCodeDeclare token budget so OpenCode's native autocompact fires before the provider's hard wall
memory_embedding.{base_url, model, dimensions, search_mode}ZeroClawVector recall; omit base_url to fall back to FTS-only
memory_lancedb.{api_key, model, base_url, dimensions, db_path, auto_capture, auto_recall}OpenClawFlat keys in one dict: embedding endpoint (api_key, model, base_url, dimensions) and the LanceDB store path. auto_capture / auto_recall are read on the gateway path only; the local-agent path hardcodes autoCapture: false and autoRecall: true

OpenCode's flags are parsed by OpenCodeAgent.setup() in alphadiana/harness/opencode/agent.py. ZeroClaw's memory store-turn is _memory_store_via_agent in alphadiana/harness/zeroclaw/agent.py, which skips the write when a task carries metadata['memory_mode'] == 'frozen'. OpenClaw runs memory through an embedded openclaw agent --local two-turn flow (alphadiana/harness/openclaw/agent.py) so the memory-lancedb plugin's autoRecall / autoCapture hooks fire; the chat/completions path never triggers them.

Transfer is data-driven: each task carries metadata.memory_mode, either build (write + recall) or frozen (recall only). The default is build.

Running a memory experiment

python -m alphadiana.cli run \
configs/memory_experiments/exp1_zw_aime_memory_seq.yaml \
--redo-all

The shipped config is a ZeroClaw Cross-Task run in sqlite FTS mode (no embedding endpoint needed). It reads five environment variables: OPENAI_MODEL_NAME, OPENAI_BASE_URL, OPENAI_API_KEY, ROCK_BASE_URL, and ROCK_PROXY_URL. An unset one is blanked rather than reported. The three OPENAI_* are named back to you when the run dies, but an unset ROCK_* fails opaquely: the runner quietly falls back to a default localhost ROCK port and you get a connection error.

The config also pins rock_image: zeroclaw-reasoning:0.6.9, which makes the runner auto-create a ROCK sandbox, so the ROCK mode prerequisites on the ZeroClaw page apply first: build that image locally (it is not published on Docker Hub), then start the host ROCK services with bash scripts/start_zeroclaw.sh and export their URLs with source scripts/rock_env.sh. Both scripts need the ref/ROCK checkout that the one-time setup in Installation creates, and abort without it, which is what leaves the two ROCK_* variables unset.

The remaining Cross-Sample and Transfer configs are not shipped in the repository. Increasing num_samples and changing a prompt is not sufficient to implement Cross-Sample isolation: the runner has no general store reset between task boundaries. A faithful experiment must use separate per-problem runs, task-namespaced stores, or an explicit verified reset. The ZeroClaw gate _has_memories counts successful store turns rather than inspecting the store, so it can fire even if the model never actually called memory_store.

Reading the results

Memory results land through the result store at alphadiana/analysis/io/result_store.py; the run loop lives under alphadiana/engine/ (alphadiana/engine/runner.py). The plots and the scope ladder are regenerated from the recorded results.

Report the configured sample count, run IDs, and uncertainty with every result. Do not infer precision or robustness from an intended design when the matching artifacts are not present in the checkout.