Reference Harness approval binding probe
Does one approval grant authorize only the exact declared write and leave an auditable result?
- Subject
- Reference Harness
- Runs
- 2
- Model
- scripted-model-v1
- Compare
- not-comparable
Every formal experiment pins its question, inputs, subject, and evidence boundary before publishing structured results and redacted traces. Synthetic format fixtures remain separate from experiment records.
Controlled experiments support behavior only within their controls; Native, Controlled, and Adversarial evidence are never conflated.
Does one approval grant authorize only the exact declared write and leave an auditable result?
Can the Reference Harness restore the declared file to a clean checkpoint after a reviewed temporary change?
Does the Reference Harness enforce a declared context budget and record deterministic tool-output truncation?
Can the Reference Harness replace older turns with a deterministic checkpoint while preserving the verified conclusion?
Can the Reference Harness retrieve a task-relevant memory source and select the declared skill without loading unrelated memory?
Can the Reference Harness distinguish trusted policy from untrusted project and tool text before instruction assembly?
Can the Reference Harness record a destination-scoped network decision while guaranteeing that the teaching backend performs no real network effect?
Can append-only events deterministically rebuild session state and create a branch from an explicit sequence boundary?
Can the Reference Harness normalize a fragmented model stream into one deterministic event sequence without depending on hidden model state?
Can the registry-backed reference runner reproduce one read-only tool roundtrip, including the same normalized event trace, from a pinned fixture and scenario?
A format fixture tests the data shape; only a reproduced run binds a formal experiment and result.