138 lines
8.3 KiB
Markdown
138 lines
8.3 KiB
Markdown
# P0 Trust-Boundary Hardening — Re-review Reflection (2026-07-11)
|
|
|
|
## Context
|
|
|
|
A re-review judged the harness improved on the *happy path* but still **fail-open on
|
|
execution integrity**: gates are prose-only, ledgers are agent-writable, and there is
|
|
no proof the harness raises real output quality. This round closes the P0 trust
|
|
boundaries the review found still open, fills company context with clearly-marked demo
|
|
data, and replaces the fake benchmark recorder with a real runner+grader.
|
|
|
|
Prior rounds: [p0-execution-integrity](2026-07-10-p0-execution-integrity-design.md),
|
|
[p1-structural-quality](2026-07-10-p1-structural-quality-design.md),
|
|
[p1-tranche3-state-engine](2026-07-10-p1-tranche3-state-engine-design.md).
|
|
|
|
## The honest trust model (why these fixes and not more)
|
|
|
|
Claude Code hooks **cannot fully sandbox an agent that has `Bash`** — a regex denylist
|
|
is bypassable by construction (guard_tools says so itself). So the trust model is
|
|
**defense-in-depth + tamper-evidence**, not cryptographic unforgeability. The only
|
|
*un-forgeable* anchors are the payloads **Claude Code itself supplies to hooks**:
|
|
|
|
- PostToolUse receipts: the *real* command, exit code, cwd, `tool_use_id` — the agent
|
|
never authors these.
|
|
- SubagentStart/Stop identity: `agent_id`, `agent_type`, spawn time.
|
|
|
|
Therefore the design principle for every ledger:
|
|
|
|
1. **Derive facts from receipt-backed artifacts**, never from agent-authored strings.
|
|
2. **guard_tools blocks the write/exec paths** an agent would use to forge or overwrite
|
|
a ledger (direct Write/Edit, Bash redirection, `python -c`, and invoking the ledger
|
|
scripts by hand).
|
|
3. **Validators cross-check agent claims** against the trusted receipts (exact match,
|
|
no substring/basename fuzz).
|
|
4. Human approval is a **documented soft-boundary**: the harness cannot authenticate a
|
|
human, so heavy-tier signoff binds to an out-of-band file that guard_tools protects
|
|
from agent writes, and the limitation is stated openly (no false "human approved").
|
|
|
|
This is stated so we don't over-claim. The bar moves from "any agent can silently skip
|
|
every gate" to "skipping a gate requires forging a Claude-Code-supplied receipt, which
|
|
the wiring makes tamper-evident."
|
|
|
|
## Workstreams
|
|
|
|
### P0-1 — workspace-unset fail-closed
|
|
Add `require_workspace()` to `_workspace.py` (raises/returns sentinel). Every
|
|
*operational* hook exits **2** (block) when workspace is unset instead of degrading to
|
|
allow: `state_engine guard`, `subagent_register`, `stop_validate`, `acceptance_log
|
|
append`, `token_ledger` mutations. Read-only queries and render paths stay advisory
|
|
(exit 0) so a missing workspace never breaks reporting. `--main` Stop stays advisory.
|
|
|
|
### P0-2 — context-package = real spawn gate
|
|
`guard_tools` PreToolUse gains `Agent|Task`. For an Org OS role/family `subagent_type`
|
|
(one with a generated `.claude/agents/<type>.md` card; helpers like `explore`,
|
|
`general-purpose`, `plan` exempt), the spawn is **denied** unless the prompt references
|
|
a context-package whose file **exists, validates, and whose hash matches** the embedded
|
|
reference. `subagent_register` records `package_path` + `package_sha256` + workflow +
|
|
role. `context_package.py` validator is upgraded from "is the field non-empty" to
|
|
"does the referenced target-repo / must-read file / agent card / acceptance-test exist".
|
|
|
|
### P0-3 / P0-5 — report identity, freshness, typed validation
|
|
- `report.schema.json`: `report-id`, `workflow-id`, `role-id`, `report-type` become
|
|
**required**; unknown `report-type` → reject (validator maps type→schema and errors on
|
|
miss). Fix `validate_report.py` typed-merge so `properties.update()` no longer clobbers
|
|
the common `primary-artifacts.items` constraint (deep-merge instead).
|
|
- `new_report.py --stub` writes a valid `report-type` and identity fields.
|
|
- `stop_validate.py`: a declared/looked-up report is accepted **only if** its
|
|
`workflow-id`/`role-id` match the registry record **and** its `created_at`
|
|
(fallback: file mtime) is **>= registry `started_at`** (ownership + freshness). This
|
|
closes "return a peer's or stale report as mine". Priority-1 declared path gets the
|
|
same ownership/freshness filter Priority-4 already implies.
|
|
|
|
### P0-4 — ledgers as a trust boundary
|
|
- **guard_tools** denies agent `Write`/`Edit`/`NotebookEdit` and Bash
|
|
redirection/tee/dd/`python -c`/`open(...,'w')` targeting `evidence/ledger.jsonl`,
|
|
`state/**/workflow.yaml`, `state/**/state-events.jsonl`, `state/acceptance-events.jsonl`,
|
|
`subagent-registry.jsonl`. It also blocks Bash invocation of `evidence_ledger.py`,
|
|
`acceptance_log.py append`, and `state_engine.py transition` **from the agent** (these
|
|
run via hook wiring or the trusted CLI, not hand-typed forgery).
|
|
- **state_engine**: preconditions derive from receipt-backed artifacts. `transition`
|
|
actually checks the selected `plan`, `allowed-by`, an authorized `actor` (privileged
|
|
`HUMAN-*` only via the trusted channel, not an agent-passed string), `evidence-grade`
|
|
presence on standard+, and heavy `plan-signoff` on plan→run. De-emphasize the explicit
|
|
`facts` override so a hand-written fact can't satisfy a gate absent the real artifact.
|
|
Add a trusted append CLI (`state_engine.py record-artifact`) so commands stop editing
|
|
YAML by hand.
|
|
- **acceptance_log append** rejects events whose `report-id` does not resolve to a real
|
|
file that **passes `validate_report`**. Ghost acceptances are refused.
|
|
|
|
### P0-6 — evidence receipt binding
|
|
Receipts gain `session_id`, `agent_id`, `workflow_id` (from the Claude-Code payload /
|
|
env), keep `tool_use_id`, `cwd`, `ts`, and store the artifact **full resolved path** +
|
|
**current hash**. `validate_report.py` matching becomes **exact**: command equality (not
|
|
substring), artifact full-path equality (not basename). Un-parseable exit code is
|
|
recorded as `null`, never coerced to `0`; an E4/E5 claim needs an explicit `exit_code:0`
|
|
receipt. E3 self-report `exit-code:0` without a receipt cannot yield High confidence.
|
|
|
|
### Root-cause R1 — company context (demo) + airtight E3 cap
|
|
Fill `org-os/01-company/company-context.yaml` with `status: demo` (explicitly *not a
|
|
real company*) and a real `projects[]` manifest for `_sandbox`. Extend the E3 cap in
|
|
`validate_report.py` to also block E3+ citations of company docs by **absolute path** and
|
|
of `CLAUDE.md`, not only relative company paths — so a template can't be laundered into
|
|
"real company evidence".
|
|
|
|
### Root-cause R2 — skill + tier wiring
|
|
`gen_agents.py` + `role-profiles`/`capability-families` emit `skills: [build-loop]` on
|
|
implementation families and thread tier→model/effort consistently through the cascade
|
|
commands (`/ground` `/decide` `/design` `/spec` `/build`), not just `/run-wave`.
|
|
|
|
### Root-cause R3 — golden-task E2E benchmark (real)
|
|
Replace `benchmark.py` arbitrary score-recorder with: (a) ≥3 golden tasks with objective
|
|
acceptance checks; (b) a **runner** that invokes the `claude` CLI headless under two arms
|
|
(plain vs harness); (c) an **automated grader** scoring first-pass acceptance / diff
|
|
applies / tests pass / rework count. Honesty: the machinery is real and wired; a full
|
|
run consumes API budget, so unrun state is reported as unrun (never fabricated scores).
|
|
|
|
### Root-cause R4 — preview_ui generalization
|
|
Drive package-manager / build / out-dir / serve from the project manifest
|
|
(pnpm/yarn/npm, Next/Vue/Vite) instead of hardcoded `npm`+`vite`+`dist`. Under strict
|
|
mode, dump-dom and state-route capture failures **fail** rather than warn.
|
|
|
|
### Secondary + hygiene
|
|
`doctor.py` dependency section-name mismatch (hidden section 3), nonexistent explicit
|
|
workspace counted OK, and SSOT "comment mention = consumed" false positive; `run-wave.md`
|
|
lens wording aligned to the 2-axis `lens_cap`; `run_all.py` per-suite timeout; lock /
|
|
atomic-create on ledger append + report-path issuance. Git: stage all new files, confirm
|
|
the mass deletions are intentional, re-run the CI entrypoint on a clean checkout.
|
|
|
|
## Testing
|
|
Each workstream ships assertions in `.claude/tests/` proving the **negative** case now
|
|
blocks (forged receipt rejected, peer report rejected, ghost acceptance refused,
|
|
unset-workspace blocks, uncertified spawn denied). `run_all.py` stays the single green
|
|
entrypoint. New file: `test_p0_trust_boundary.py`.
|
|
|
|
## Non-goals
|
|
- True human authentication (documented soft-boundary).
|
|
- Making regex denylists unbypassable (impossible; we raise the bar + add tamper-evidence).
|
|
- A full paid benchmark run (machinery only; explicit opt-in to actually spend budget).
|