Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 27 additions & 0 deletions .agents/plans/03-assurance-hardening/decisions.tsv
Original file line number Diff line number Diff line change
Expand Up @@ -64,3 +64,30 @@ ts phase decision why evidence result
2026-08-28T06:01:16Z phase-8 merged independent release-time regrading the exact head passed every required CI job and an isolated frozen-install verifier PR 56; merge f505b0d merged to main
2026-08-28T06:08:00Z phase-9 selected a bounded publication state machine inline shell inferred absence from errors, lacked current-main proof, and could destructively clobber assets two architecture candidates; independent arena judge repository-owned TypeScript ref proof and postcondition-driven npm/GitHub reconciliation; no blind mutation retries or clobber
2026-08-28T06:25:00Z phase-9 made the exact draft the durable transaction proof separate npm and GitHub steps could not recover after npm succeeded and main advanced adversarial multi-model review; draft-inclusive GitHub API contract full main proof creates exact draft and precedes npm; tag-only proof finalizes it; conflicts, extras, prereleases, pending digests, and superseded workflow failures fail or reconcile explicitly
2026-08-28T06:35:27Z phase-9 merged bounded idempotent publication the exact head passed six CI jobs, 619 tests, live OpenCode, actionlint, and isolated shipping verification PR 57; merge c00461a merged to main
2026-08-28T06:47:00Z phase-10 stopped the first Linux qualification campaign on a persistence defect host-failure transcripts used a legacy hostError object that the retained-evidence schema rejected campaign 2026-08-28T06-42-38-851Z.v2; OpenAI host timeout; Grok wedged grep and glob; attempt-write-failed no decision or report claimed; both provider failures preserved as an aborted campaign finding before a fresh rerun
2026-08-28T07:05:00Z phase-10 stopped the second Linux campaign on redaction-schema drift the credential-key matcher redacted numeric outputTokens because its name contained token campaign 2026-08-28T07-01-04-278Z.v2; both happy-path attempts timed out; retained transcript parse no decision or report claimed; fix separates exact token-count metric names from credential-bearing token fields
2026-08-28T07:16:00Z phase-10 stopped the third Linux campaign after proving host-image drift two Grok happy-path failures made the 100 percent gate unreachable and both involved missing or wedged search tools campaign 2026-08-28T07-09-52-497Z.v2; two immutable Grok host failures; container command -v rg absent attempts preserved; minimal container corrected to include ripgrep like the canonical Ubuntu runner before a fresh plan
2026-08-28T07:23:00Z phase-10 stopped the fourth Linux campaign after concurrent host startup failure both model probes passed sequentially but the first two concurrent cells hit the same 120 second session-create-failed boundary campaign 2026-08-28T07-19-01-600Z.v2; two immutable host failures; fixed plan unchanged attempts preserved; next fresh campaign uses supported concurrency 1 to remove local VM contention without changing qualification semantics
2026-08-28T07:31:00Z phase-10 stopped the fifth campaign on a measured prompt-contract failure sequential hosting worked but Grok used a summary assertion without a JUnit resultsPath, then could not close happy-path campaign 2026-08-28T07-22-45-355Z.v2; one unscored escalation; one product failure attempts preserved; plan and run guides now bind exact named assertions to machine-readable JUnit and forbid all-tests-pass summaries
2026-08-28T07:55:00Z phase-10 stopped the sixth campaign on host lifecycle ownership Grok cleared happy-path 3 of 3, then plan-only session creation timed out and left the OpenCode grandchild alive after wrapper cleanup campaign 2026-08-28T07-40-39-643Z.v2; four immutable attempts; process table startup now proves session readiness and Unix cleanup owns the detached process group, with a real wrapper-child termination test
2026-08-28T10:28:00Z phase-10 stopped the seventh campaign on an uncompensated environment gap Grok passed 24 primary cells before one retryable glob wedge left resumes-after-interruption at two scored attempts out of three campaign 2026-08-28T08-10-36-739Z.v2; 25 immutable attempts; host failure retained release cannot qualify because the frozen 76-cell plan had no reserve; no threshold or failure was reinterpreted
2026-08-28T10:40:00Z phase-10 selected one predeclared environment reserve per provider and case one external flake should not erase a fixed scored sample, but unlimited or post-hoc retries would game the gate two architecture candidates; independent judge canonical plan has 76 primary targets plus 16 reserves; only durable retryable host or provider failures activate the same-stratum reserve; second failure stays inconclusive
2026-08-28T10:59:20Z phase-10 bound reserve activation to independently reproducible failure evidence report labels alone could otherwise relabel product evidence or create a runner/regrader disagreement provider lineage and envelope graft tests; pre-outcome host graft test; 632-test full check; 13 cassette replays; 14 live OpenCode checks runner downgrades unverifiable failures to nonretryable before persistence; release regrader reproduces the exact retained failure before accepting a reserve
2026-08-28T11:08:29Z phase-10 closed every adversarial reserve-policy finding three-model review found stale diagnostics, mixed pseudonymization domains, unsupported failure labels, contradictory envelopes, and paid work after an unrecoverable stratum interrogate panel; comment-sicko; 90 focused tests; 632-test full check; 13 cassette replays; 14 live OpenCode checks one shared supported-failure predicate governs report activation; whole evidence is pseudonymized together; contradictory provider evidence fails closed; reserve scheduling stops at the first exhausted stratum
2026-08-28T16:57:13Z phase-10 corrected stratum reserve overspend and tightened the Windows handoff oracle the first reserve state allowed one reserve to hide two missing scored outcomes and one accepted phrasing was narrower than the real structured handoff tests/environment-reserves.test.ts; tests/eval-scenario-checks.test.ts; bun run check; bun run replay 633 pass, 1 intentional skip, 13 of 13 replays; a reserve now activates only when one scored primary is missing in its stratum
2026-08-28T17:05:00Z phase-10 stopped release qualification after the complete primary matrix OpenAI failed multiple product guarantees, Grok had one named-binding miss, and five tool wedges plus three unsupported asks left sample gaps campaign 2026-08-28T11-10-09-556Z.v2; 76 immutable primary attempts; 58 product passes; 5 host failures; 3 unscored escalations no reserve, canary, tag, or release was attempted; the full campaign remains diagnostic evidence and the candidate is not releasable
2026-08-28T17:34:22Z phase-10 bound named JUnit evidence through one fixed managed command path arena candidates agreed execution must not choose evidence after approval, but a new plan field would violate the patch-release maintainer contract pstack how trace; candidates A and C; independent judge; maintainer contract new named commands must write .flow/results.xml; execution derives it and rejects substitutes; approved legacy plans retain caller-bound capture; one pending capture per workspace prevents path races
2026-08-28T18:58:02Z phase-10 accepted the targeted behavior fixes and corrected two remaining evaluator boundaries the 12-attempt canary passed every affected workflow except one exact named case that OpenAI made genuinely runnable and observed passing on Linux campaign 2026-08-28T17-36-15-649Z.v2; corrected-oracle replay; campaign 2026-08-28T18-34-52-277Z.v2; cassette schema v2 all 12 affected provider-scenario cells now have passing product evidence; three refreshed current-format cassettes reproduce named Windows refusal, Linux binding, and unprovable handoff; historical campaign verdicts remain unchanged
2026-08-28T19:06:31Z phase-10 closed final named-evidence review findings the first fixed-path check matched substrings and replay initially discarded the JUnit witness three-model interrogate; comment-sicko; exact option-token parser regressions; production report reader in cassette replay; 642-test full check; 13 replays; 14 live OpenCode checks named paths must be exact option values; one capture per workspace prevents report races; corrected oracles require broad current-host observations from .flow/results.xml; all final reviewers reported no findings
2026-08-28T22:34:55Z phase-10 stopped the fresh release campaign at the first impossible threshold OpenAI passed all 38 primaries, but Grok failing-gate-blocks produced two product failures in its first two attempts, making 9 of 10 unreachable campaign 2026-08-28T19-07-53-057Z.v2; 52 immutable attempts; exact failure transcripts no reserves or later primaries were spent; both failures preserved the red test but the final handoff omitted its identity, so flow-run now requires the blocker before the exact environment handoff
2026-08-28T22:51:23Z phase-10 accepted explicit blocker naming before evidence handoff both release failures were caused by a terse template that omitted the red test from the user-visible response campaign 2026-08-28T22-35-24-536Z.v2; three Grok failing-gate attempts 3 of 3 passed with the blocked state preserved and the pre-existing failure named; the correction is eligible for a fresh release campaign
2026-08-29T03:07:20Z phase-10 stopped the next release campaign at the first impossible product threshold OpenAI passed 38 of 38 and Grok passed 24 primaries before one of three resume attempts failed, making the required 100 percent resume stratum unreachable campaign 2026-08-28T23-12-55-932Z.v2; 64 immutable attempts; retained failed resume transcript no reserve, canary, tag, or release was attempted; the campaign remains diagnostic evidence and the candidate is not releasable
2026-08-29T03:07:20Z phase-10 fixed the incomplete Bun JUnit command at its guidance source the resumed Grok session approved, implemented, and ran the planned command, but the flow-plan example omitted --reporter=junit so Bun wrote no .flow/results.xml and Flow correctly refused completion failed resume transcript; red then green prompt-quality contract test flow-plan now gives the complete runnable Bun JUnit command and explicitly rejects reporter-outfile alone; the release oracle and historical failure remain unchanged
2026-08-29T03:21:16Z phase-10 accepted the complete Bun JUnit guidance fix before another release campaign the failed resume behavior needed fresh model evidence on the corrected packed artifact rather than a reinterpretation of the old attempt campaign 2026-08-29T03-08-49-308Z.v2; three Grok resume attempts 3 of 3 completed with validation and independent review, no false completions, and no interventions; the fix is eligible for a fresh release campaign
2026-08-29T08:40:52Z phase-10 accepted the complete two-provider release matrix the release needed a fresh fixed-attempt campaign after the Bun JUnit guidance correction campaign 2026-08-29T03-22-04-307Z.v2; 76 immutable primary attempts 76 of 76 passed across Grok 4.6 and GPT-5.6 Sol, no reserves or false completions, and every 90 percent or 100 percent stratum cleared its published threshold
2026-08-29T08:40:52Z phase-10 kept failed canary setup evidence out of release authority without hiding it the first host loaded duplicate plugins, the next nested fixture resolved the parent repository, sanitized export removed reviewer lineage, a valid two-feature session exceeded the phase9 one-reviewer profile, and a dependency-only fixture manifest reported runtime version 0.0.0 preserved .release-artifacts canary attempts; failed derived records; exact recorder diagnostics none qualified or changed the campaign; the final host used an isolated Git root, raw-then-redacted export, one bounded feature, and the exact package manifest from the measured tarball
2026-08-29T08:40:52Z phase-10 sealed exact canary and release qualification the tag gate requires derived runtime identity, one coherent manager-reviewer lineage, named broad validation, completed closure, and a regradable evidence bundle canary record sha256:b86778250ad52a7d8d27f9d80c0fec04d207175f93173b9f85a86b1b5505a222; bundle qb1-09fe5dcae0771dd2262c747b0f6378d6d9bcec022dc75ad4f37d267fc8c3360a all six canary checks passed; strict canary verification and qualification returned VERIFIED
2026-08-29T08:49:12Z phase-10 superseded the first verified bundle with stricter provider-metadata redaction final security review found encrypted reasoning state and project linkage fields that were unnecessary for canary rederivation preserved prior verified record and bundle under .release-artifacts; same closed session re-recorded after evidence-only scrub final canary sha256:4940dc0c7c0e9836f67237c934865c40e115eb12f44a12f7b1888baabd6140b9 and bundle qb1-521ae36e04b02367f7610caeed95db2edf9e10cc0339a16765f767ef22b18b76 verify with zero reasoningEncryptedContent, projectID, slug, raw session ids, host paths, credentials, or symlinks
2026-08-29T08:55:51Z phase-10 accepted the exact-head review finding about zombie descendants the Linux regression treated a zombie as dead, while direct /proc inspection showed the child remained Z after group-first termination PR 58 comment 3882601494; pinned Linux reproduction; red PID-existence assertion the harness now enumerates the detached process group through /proc on Linux and ps on macOS, signals descendants first, and lets the wrapper reap them before fallback group termination; both lifecycle regressions pass and the prior qualification is stale pending a fresh campaign
2026-08-29T11:25:04Z phase-10 preserved interrupted tool identity for environment reserves two fresh campaigns stopped because live watchdog output named wedged tools while retained evidence rewrote those calls to error and then derived an empty pendingTools list campaigns 2026-08-29T08-57-11-489Z.v2 and 2026-08-29T10-21-34-403Z.v2; red interrupted-call regression retained pending tools now include OpenCode error calls with metadata.interrupted true, so future command-aborted evidence independently rederives as retryable and may activate only its predeclared same-stratum reserve; historical records remain unchanged
26 changes: 26 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,32 @@

One short entry per release, written for users deciding whether to upgrade.

## [8.1.3] - 2026-08-28

Release claims now come from retained evidence instead of trusted summaries.

- Declared gate assertions must be satisfied before final review and completed
closure. New named evidence writes JUnit to the plan-bound
`.flow/results.xml`; execution cannot substitute another path. Reports are read once
from a stable, bounded workspace file without following symlinks.
- Eval failures identify the evaluator, provider, host, or persistence boundary
that produced them. The release matrix uses repository-owned policy with enough
attempts to measure its 90% and 100% thresholds.
- Exact-artifact canary results are derived from one OpenCode session lineage.
Release qualification seals all attempts, transcripts, provenance, canary,
artifact, and grader source into an immutable bundle that CI independently
reopens and regrades.
- Publication requires the tag and packed artifact to match the release evidence.
npm and GitHub operations are bounded and idempotent, refuse conflicting bytes,
and recover through an exact draft without destructive asset replacement.
- **Session v5 schema:** unchanged. Runtime commands and tools are unchanged.

Install or update:

```bash
opencode plugin opencode-plugin-flow@8.1.3 --global --force
```

## [8.1.2] - 2026-08-26

Long Grok eval turns no longer fail at Bun's implicit five-minute fetch cutoff.
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ expensive, and it is overhead when it is not.
Install the exact npm release through OpenCode:

```bash
opencode plugin opencode-plugin-flow@8.1.2 --global --force
opencode plugin opencode-plugin-flow@8.1.3 --global --force
```

Omit `--global` for project scope. Version pins are exact and never update on
Expand All @@ -51,7 +51,7 @@ The equivalent manual project configuration is:
```json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["opencode-plugin-flow@8.1.2"]
"plugin": ["opencode-plugin-flow@8.1.3"]
}
```

Expand Down
3 changes: 2 additions & 1 deletion biome.json
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,8 @@
"!bun.lock",
"!evals/results",
"!evals/canary",
"!evals/decisions"
"!evals/decisions",
"!evals/qualification/bundles"
]
},
"formatter": {
Expand Down
Loading