Changelog¶
The source of truth is CHANGELOG.md
in the repository; this page mirrors it.
0.2.0 (2026-10-06)¶
- Improvement tasks (
improve:in a task): the repository already passes and the agent gets several rounds to make a measured objective better. Baseline and every round are evaluated on a scratch copy with the hidden tests, so hidden tests never reach the agent's worktree; the best checkpoint is kept (anchored againstgit gc) and worse rounds are reverted; an optional in-loop evaluator lets the agent measure a limited number of times per round. A failed checkpoint or revert stops the protocol without losing any round's data. A run passes when its final state passes the gate and beats the baseline. Metrics: baseline, best, final, ratio, score (1 for minimizing to 0), progress to target, curve, per-round history, evaluator calls;improve_roundsandimprove_eval_budgetvariant options make rounds and budget sweep factors; a paired improvement-ratio interval. Bundleddemo-improvesuite andimprove-budgetsweep. - Safety as a measured axis: every run gets risky-action findings from its trace (credential
access, privilege, destructive, suite access, network, outside the worktree), each executed or
blocked, a
safeverdict, and per-variant safe and safe-pass rates. Tasks can plant canaries and lure URLs (safety:); reads, leaks and followed lures are findings, and canary tokens are detected before redaction but never recorded by a run (events, artifacts, verifier output, runner metadata, exports). The Claude runner records hook events (--include-hook-events). The bundledharnesses/sentinelis a rule-based pre-tool-use decider (python -I -m harnesslab.safety.decide, policy set with--denyand--allow) to A/B test; bundleddemo-safetysuite. harnesslab suite mine: builds a task suite from a repository's git history. Every commit that changes source and tests becomes a candidate (start at the parent, the commit's tests as hidden verifier, its source change as reference solution, its message as prompt with hidden test names scrubbed); kept only when the tests fail at the parent and pass at the commit. Writessuite.yaml, task files andmining_report.jsonwith the reason for every skipped commit.- Git fixture repositories are now copied with only the base commit and its ancestors (fetched by SHA; no later commits, branches, tags or remote), so an agent cannot read a mined task's answer out of the history. Existing internal clones are not reused.
- Paired statistics (
harnesslab.experiments.stats): task-level cluster bootstrap intervals and an exact sign test, with verdicts only above a minimum number of paired tasks. Shown on the compare view and byharnesslab experiment compare, in sweep reports (recommendation vs runner-up and vs an optionalbaselineconfiguration), and available as grow gate rules (gate.require: not_worse_ci | better_ci). harnesslab ablate run|report: tests every component of a harness bundle against its own absence and the whole bundle against a minimal one; verdictshelps,hurts,no evidence,not enough tasks; dashboard section andablation_reportin exports. The fake runner gainscomponent_solvesandcomponent_llm_callsfor offline ablation demos.- Harness bundles: a directory (system prompt, skills, hooks, agents, fake.yaml) any variant can
carry with
harness:; applied by the Claude Code (--append-system-prompt-file,--plugin-dir), Codex (prompt prefix), generic and fake runners;harness_hashrecorded on variants and runs and folded intoconfig_hash, so bundles work as sweep factors. llm_callsmetric (Claude Code, generic JSONL;nullfor Codex) in metrics, aggregates, the compare view, exports and as a sweep objective.- Growing Harness loop (
harnesslab grow run|resume|list|show|export): failure window → optimizer → window check → held-out gate → accept or roll back; budgets, retirement, resumable state, per-version audit files; grow sessions and harness versions in the database and dashboard. - Optimizer plugins:
claude-cli(Claude Code binary, no tools, structured output),manual,fake; leak controls (scrubbed optimizer view, candidate lint) so hidden test sources, names, quoted lines and, by default, assertion details never reach an optimizer (optimizer.verifier_detail: summary | full);harnesslab harness check. - Optimizers may not add or edit
hooks.jsonunlessoptimizer.allow_hooks: true: hook commands run on the host outside the agent's tool allowlist.harness checkwarns about bundles with hooks. - The
claude-clioptimizer charges every attempt, including failed ones, tomax_optimizer_cost_usd, and removes its temporary working directory. - Failure cases shown to the optimizer always come from a failing run, also with
repetitions > 1;split.fractionsno longer rejects valid fractions because of rounding. - Bundled
harnesses/baselineandgrow/demo-fake.yaml,grow/claude-grow.yamltemplates (copied byharnesslab init). - Documentation site (this site), generated CLI reference, docs consistency tests.
0.1.0 (2026-09-22)¶
First release.
- Experiment runner: same task, identical base commit, one isolated git worktree per run, independent verifier (exit code + optional partial score), normalized traces, metrics, aggregates, JSON export.
- Harness adapters: OpenAI Codex CLI (
codex exec --json), Claude Code CLI (claude -p --output-format stream-json), deterministic fake runner, generic command runner; plugin API (harnesslab.api,plugins:lists, entry points). - Configuration sweeps (
harnesslab sweep): model × reasoning effort × toolset × compaction × action policy grids with budgets, cheapest verified configuration per workload, factor effects, holdout tasks. - Local dashboard (FastAPI + HTMX): experiments, task × variant matrix, run detail with timeline/diff/verifier output, compare view, sweep recommendations.
- Bundled demo suite (
harnesslab run demo),harnesslab initscaffolding,harnesslab doctor.