Honest limits¶
What Harness Lab does not do, or does only partly, as of the current version.
Isolation¶
Runs are isolated with git worktrees, not containers. Agents run with your privileges and can reach the host filesystem. Hidden tests can be found on disk by a determined agent; such runs are flagged, not prevented. See Security model.
Platforms¶
No native Windows: process groups and worktree cleanup are POSIX-only. WSL works.
Statistics¶
The compare view shows raw deltas. With repetitions there are no confidence intervals or significance tests yet, and the grow gate compares plain pass rates. On small suites the gate can decide on noise. Task-level bootstrap intervals are planned.
Metrics that depend on the harness¶
reported_cost_usdexists only when the harness reports cost (Claude Code does, Codex does not). Give other harnesses a pricing table.llm_callsis exact for Claude Code and the generic JSONL protocol andnullfor Codex, whose stream does not expose model invocations.- Action granularity cannot be switched inside closed CLIs;
action_policysteers it through a prompt and the realized granularity is measured.
Growing the harness¶
- Only the outer harness can be grown for Claude Code and Codex: prompts, skills, hooks, agents. The control loop stays closed.
- The bundled demo suite has three tasks, enough to prove the mechanics, not to measure an effect.
- The optimizer sees scrubbed verifier output. By default unittest and pytest assertion details
are replaced by a placeholder; expected values printed in other ways (custom runners,
printcalls, exception messages) are not recognised. Write hidden tests accordingly. - Optimizers may not write
hooks.jsonunlessoptimizer.allow_hooks: true; hook commands run on the host outside the agent's tool allowlist, so enable it only on a machine you can throw away. - No real Claude Code grow session has been run by the maintainers at the time of writing; the
claude-clioptimizer's flags and result format were verified against the installed CLI, and the full loop is verified with the fake runner and a replayed CLI. - Runners read the bundle from its source directory, not from the per-experiment snapshot, so editing a bundle while an experiment runs desynchronises hash and content.
Sweeps¶
Grids explode; only random subsampling and budgets bound them. Successive halving and Bayesian search are not implemented. There are no per-runner concurrency limits and no rate-limit handling beyond counting Claude Code's rate-limit events.
Dashboard and data¶
Tables only, no charts. No live streaming while a run executes; refresh the page. No pagination
or filtering on the experiment list. No cleanup or delete commands: experiments, artifacts, kept
worktrees and fixture caches accumulate under .harnesslab until you remove the directory.
Schema migrations are forward-only column additions, not a migration tool.
Scope by design¶
No prompt optimization of task prompts, no reinforcement learning, no LLM judges, no accounts, teams, billing, distributed workers, Kubernetes, vector databases or model routing.