Read and export results¶
In the terminal¶
harnesslab experiment list
harnesslab experiment show <experiment-id> # matrix, aggregates, every run
harnesslab sweep report <experiment-id> # recompute a sweep's recommendation
harnesslab experiment compare <experiment-id> <a> <b> # paired evidence: B vs A per task
harnesslab ablate report <experiment-id> # component verdicts of an ablation
harnesslab grow list
harnesslab grow show <session-id> # version lineage and gate results
Ids accept a unique prefix or the exact name.
Export an experiment¶
harnesslab experiment export <experiment-id> > experiment.json
harnesslab experiment export <experiment-id> --no-events --no-artifacts -o summary.json
The document contains:
experiment: name, suite, status, timestamps, Harness Lab version and commit, environment and its hash, the spec that produced it;variants: id, runner, model, description, configuration andconfig_hash, sweep factors,harness_hashand the bundle's file list;tasks: id, name, version,task_hash,spec_hash,prompt_hash, base commit, tags;runs: one entry per run with status, outcome, timestamps, areproducibilityblock (base commit, hashes, runner, configuration, models requested and resolved, CLI version), the fullmetrics, the verifier result, and, unless disabled, the normalized events and the text artifacts (diff, verifier output) inline;aggregates: per-variant statistics (see Metrics);sweep_report: the recommendation report when the experiment was a sweep, including paired evidence against the runner-up and the baseline;ablation_report: per-component verdicts when the experiment was an ablation.
The same document is served at /api/experiments/{id}/export.json. It is plain JSON with no
NaN values, so it loads with any tool.
Export a grow session¶
writes the session's current accepted bundle into the directory plus lineage.json, which lists
every version with its status, reason, window tasks and fixes, gate pass rate, median llm_calls
and cost, optimizer model and cost, and rationale. The bundle is ready to be used as
harness: on any variant.
Analyse in Python¶
import json
import pandas as pd
doc = json.load(open("experiment.json"))
runs = pd.json_normalize(doc["runs"])
runs.groupby("variant")[["metrics.verified_pass", "metrics.llm_calls", "metrics.wall_time_seconds"]].mean()
aggregates already holds per-variant means, medians, standard deviations, minima and maxima,
and per-task pass rates, so most comparisons need no computation.
Statistics¶
Every decision Harness Lab supports (compare view, sweep recommendations, ablations, the optional
statistical grow gate) uses the same paired analysis, in harnesslab.experiments.stats:
- Pair by task. Both sides ran the same tasks from the same base commit. Within a task, repetitions are averaged first, so a task with five repetitions counts once. Tasks without a verified run on both sides are left out.
- Per-task differences of pass rate (B − A), of cost (reported, else estimated, else total
tokens) and of
llm_calls. - Cluster bootstrap over tasks: 2,000 seeded resamples of the tasks, percentile 95% interval of the mean difference, and P(> 0), the share of resampled means above zero.
- Exact sign test on tasks where one side did better (ties dropped), as a check that does not depend on resampling.
- Verdict:
betterorworseonly when the interval excludes zero and at least 5 tasks (configurable with--min-tasks) were paired; otherwiseno evidenceornot enough tasks.
The intervals are honest about small suites: with the three-task demo suite they are wide and no verdict is given. Use them to decide whether a difference is worth acting on, and add tasks before trusting a close call.
Reading a single run¶
Statuses and outcomes are independent (see Concepts). When a run looks
wrong, read in this order: the verifier's stderr (why it failed), the diff (what the agent
changed), the timeline (what it did), then the runner metadata in the reproducibility record
(permission denials, API retries, unknown records in the stream). The possible_suite_access
flag in the metadata marks runs whose shell commands mentioned the suite directory.