What the optimizer sees¶
New in Harness Lab 0.2.0. Upgrade an older install with
pip install -U harnesslab.
A grow session's headline guarantee is that hidden tests never reach the optimizer. If they did, the optimizer could encode the answers into the harness and the gate would measure memorisation, not capability. This page states what the optimizer receives, what it never receives, and how candidates are checked.
What it receives¶
For each iteration, one JSON document (persisted as context.json under the version directory so
it can be audited):
- the current bundle's content files, path → text;
- the runner and model the deployed agent uses, and the suite's description;
- for each task in the window: the task prompt, the attempt count, the run's status, outcome and
verified score, the agent's final message, a compact trace digest (one line per tool call,
command, message or error with exit codes, durations and short output previews), the diff the
agent produced, the verifier's stdout and stderr with assertion details removed (see below),
and the run's
llm_calls, tokens, wall time and reported cost; - the edit constraints: allowed paths (
hooks.jsononly withoptimizer.allow_hooks),max_files, the byte caps; - the reasons of up to five recent rejections.
The failure case always comes from the latest failing run of that task under the current
bundle (or, if it has none there, the latest failing run under any bundle), never from a passing
run, including a passing repetition when repetitions is above 1.
What it never receives¶
- The content of any injected file. Hidden test sources are never read into the view.
- Hidden file names and stems (
test_hidden_budgets.py,test_hidden_budgets), which are replaced by[hidden-test]wherever they appear. - Hidden test identifiers:
test_*function names and test-class names harvested from the hidden sources, also replaced by[hidden-test]. A unittest verbose line liketest_january_includes_last_day (test_hidden_month_boundary.InclusiveRangeTests) ... FAILarrives as[hidden-test] ([hidden-test].[hidden-test]) ... FAIL: the verdict survives, the assertion's name does not. - Hidden source lines of 24 characters or more, which unittest and pytest tracebacks quote, and
which are replaced by
[hidden-test-line]. pytest's>andEmarkers are stripped before matching. - Assertion details in the verifier's output, unless the grow spec sets
optimizer.verifier_detail: full. unittest'sAssertionError:payload and the diff lines that follow it (-,+,First differing element ...), pytest'sE assert ...blocks and the payload of itsFAILED ... - assert ...summary lines becomeAssertionError: [hidden-assertion-detail]. The optimizer still sees which runs failed, error types such asModuleNotFoundErrororTypeError, and tracebacks. - Fragments created by truncation: every field is scrubbed before it is capped.
- Names echoed back through rejection reasons: lint messages use the placeholder, and reasons are scrubbed before they are stored or shown again.
Task ids are visible in the view (the optimizer needs to know which tasks failed) but may not appear in a candidate.
How candidates are checked¶
Before a candidate runs anywhere, the leak lint rejects it if it:
- deletes a file, edits
harness.yaml, changes more thanmax_filesfiles, or exceeds the size caps; - adds or edits
hooks.jsonwhileoptimizer.allow_hooksis false (hook commands run on your machine outside the agent's tool allowlist); - adds a path outside the allowed set, or a skill without frontmatter, or malformed
hooks.json; - mentions a suite task id as a whole word (except in
fake.yaml, which real runners ignore); - mentions a hidden file name, stem or test identifier;
- contains a line of 24 characters or more that appears verbatim in a hidden test.
A rejected candidate is recorded as invalid with its reason, costs the window tasks an attempt,
and counts toward the three consecutive failures that stop a session.
Verifying it yourself¶
harnesslab harness check <bundle> --suite <suite> runs the same lint on any bundle. The test
suite runs a whole fake grow session and asserts that no hidden line or hidden name appears in any
context.json or proposal.json it produced.
What this does not cover¶
The assertion scrub recognises unittest and pytest output. Expected values printed some other
way (a custom test runner, print calls in a hidden test, a doctest, an exception message such as
ValueError(f"expected {x}")) are not recognised and reach the optimizer, as does everything when
verifier_detail: full is set. Test counts always appear. Write hidden tests so that their
failure output describes the behaviour, not the solution.
The lint checks a candidate's text for hidden names and lines. It cannot tell whether a skill encodes an answer in other words, which is why acceptance also requires the held-out gate not to regress and why the final holdout comparison exists.