LAB · CONTROLLED · REPRODUCED RUN

Reference Harness sandbox and network policy probe · reference-sandbox-network-v1-run-001

This is a controlled experiment artifact for Reference Harness. It demonstrates the public runner pipeline, not any vendor agent's native behavior.

reference-sandbox-network-v1-run-00113 eventsclean
TRACE 0.1 · CONTROLLED

A formal experiment run, event by event

SESSIONevt-000T+00:00:00Z

session.start

Actor reference-agent produced sequence 0. Redaction status is clean.

{
  "max_turns": 8,
  "tool_count": 1
}

1 events are visible; hidden chain-of-thought is not part of the trace.

Download JSONL
HARNESS MODEcontrolled
MODELscripted-model-v1
NETWORKnot recorded
COMPARABILITYnot-comparable
WHAT THIS PROVES

A fixed fixture and prompt can produce a structured result, valid trace, identical normalized fingerprint, and downloadable replay through the public runner.

WHAT IT DOES NOT PROVE

Real model capability, vendor agent behavior, OS sandbox strength, performance cost, or cross-agent ranking.

View full experiment record