fix(status): serve menubar-json status from a disk-persisted snapshot to eliminate per-poll re-parse latency - #1
Open
dgabehar wants to merge 40 commits into
Open
Conversation
… cache String.prototype.slice returns a V8 SlicedString: a view that retains a reference to its ENTIRE parent string. The parsers store short previews of message text (userMessage.slice(0, 500/2000)) in the long-lived session cache. Session files routinely carry 100KB+ strings (agent- injected system prompts, tool results), so every cached preview pinned its full parent buffer for the life of the process. Measured on 3.2GB of kiro CLI session files (6,659 files, largest 40MB): cold parse, default heap, before: 4.33GB peak -> OOM crash cold parse, 8GB heap, before: 5.67GB peak (kiro provider alone) after kiro flatSlice: 0.91GB peak after parser.ts cache sites too: 0.64GB peak original failing command (cold, default heap, all providers): 0.89GB peak -> completes Warm runs were always fine (~0.29GB) because the cache's JSON round-trip flattens the strings on load — which made this bug appear intermittent: it only fired on a cold or invalidated cache. Fix: flatSlice() in content-utils.ts forces a flat copy via Buffer round-trip. Applied at the six kiro userMessage capture sites and the three shared cache-building sites in parser.ts (protects all providers). Regression test asserts the no-retention property via bounded heap growth over 1000 large-parent slices. AI-Origin: human
The kiro provider was reading the full working directory from session metadata (meta.cwd for CLI sessions, workspacePaths[0] for v2 IDE sessions, workspaceDirectory for workspace sessions) but discarding it via basename(), keeping only the leaf name for display. This meant computeAttributionRecords could never resolve kiro sessions to a git repo, so `codeburn sync push --attribution` produced 0 facts for all kiro-originated sessions. Now passes the full path as projectPath on emitted ParsedProviderCalls, which buildRepoGroups uses to resolve git identity and correlate commits with sessions via timestamp windows. The path stays local: only the normalized origin remote egresses in attribution spans. Behavior changes beyond attribution: - kiro calls now flow through canonicalizeProviderCallProject, so kiro sessions in LINKED GIT WORKTREES canonicalize to the main repository: their report project name changes from the worktree dir name to the main repo name (consistent with claude/codex behavior). - workingDirectory is now populated on kiro calls. - PROVIDER_PARSE_VERSIONS.kiro bumped (project-path-v1): cached entries predate projectPath and are served without re-invoking the parser, so without the bump this fix silently no-ops for every warm cache. The bump forces a one-time cold kiro re-parse on upgrade. ORDERING: this commit must land WITH (or after) the preceding SlicedString OOM fix. The forced cold re-parse it triggers is exactly the workload that OOM'd before that fix on multi-GB kiro stores. Perf: per-call canonicalization added a measured +5% to cold parse (.git-marker lstat walk per call). resolveCanonicalProjectPath is now memoized on cwd (cleared with the session cache), removing the redundant walks for all providers. Tests: projectPath emission fixtures for all three session formats (CLI, v2 IDE, workspace-session), fingerprint-change assertion, and a regression test seeding a pre-bump cache entry and proving the re-parse recovers projectPath. AI-Origin: human
… to eliminate per-poll re-parse latency codeburn status --format menubar-json took 25-90+ seconds per call because the menubar app spawns a fresh CLI process per poll, so every call re-JSON.parse'd the full session-cache blob and re-ran the full aggregation pipeline with no cross-process reuse. Adds a disk-persisted status snapshot keyed by a cheap corpus fingerprint (stat-only, no content read), with a settle-window debounce so rapid-fire source writes coalesce into one recompute instead of one per poll.
… to buildMenubarPayloadForRange
…th walks flatSlice returned strings within the bound unchanged, but provider adapters pre-truncate with .slice(0, 500) before the cache site, so those views still pinned their parent buffers. Always flatten; the round-trip is ~150ns per turn. Use utf16le so lone surrogates survive the copy. Cache the canonical-path Promise instead of the resolved value so calls in one Promise.all batch share a single walk. Document the one-time kiro re-parse and worktree regrouping.
fix: OOM on cold parse (V8 SlicedString retention) + kiro projectPath for repo attribution
Clear the per-directory Codex and Antigravity memo maps in the resident RSS guard; document the single cache-dir rule (XDG_CACHE_HOME no longer consulted, ledger migrated); stop output-overflow terminations from spending the resident's unexpected-death budget.
# Conflicts: # CHANGELOG.md
The separator regex retried its leading \s* from every offset, which is quadratic on the long whitespace-heavy commands agents emit; that one regex was ~30% of a warm status run on a multi-GB corpus. Match the separator alone and widen over whitespace by hand. Output is unchanged (differential check over 42k real commands).
…ared-cache perf(desktop): share cache state and eliminate duplicate cold hydration
…rator-regex perf(classifier): linear bash separator split
A warm launch rewrote the entire session cache whenever any provider appended a few KB: on a 6 GB corpus that is a 155 MB stringify + fsync every run. The on-disk cache is now a version-suffixed directory holding one shard per provider plus a small envelope, and a save rewrites only the providers marked dirty. - Dirtiness is tracked per provider (markCacheDirty) instead of one global flag, so an appended Claude session no longer republishes Codex, Copilot and the rest. - Shards carry a nonce in their filename and the envelope is renamed last, so a save is published at a single point: readers never see a half-updated set, and a writer that loses the refresh ownership fence leaves the canonical shards untouched. - A shard that fails validation is treated as an absent provider rather than rejecting the whole cache, so one malformed turn costs one provider's re-parse instead of every provider's history. - v7 migrates losslessly: the blob is re-laid-out into shards and removed only once that save publishes. Nothing re-parses. - Cold-parse progress saves now trigger every N files parsed rather than every 5s, so a slow cold parse no longer rewrites the growing cache on a wall clock.
Codex rollout files are append-only and the active ones run to hundreds of MB, but any growth re-read the file from byte 0 because the cache keyed only on mtime+size. The parser now records a restart point at every task_started boundary — the byte offset plus the state the single-pass decode carries across it — and a grown file with the same dev/ino picks up from there. The boundary sits at the task_started line itself, so the task it opens is re-decoded from the tail; the entry stores how many calls were decoded before that point so the resumed run starts from exactly those and cannot double-count the open task. An unusable or absent snapshot falls back to a full re-parse. CODEX_CACHE_VERSION is deliberately not bumped: the new fields are additive and absence-safe both ways, so a bump would discard a warm cache for nothing.
Two live processes share one cache directory routinely (a one-shot CLI beside the resident serve child, two menubar polls), and the shard layout had two ways to lose data there. - The atomic write used a FIXED temp name, so two writers publishing the envelope — every save does — shared one `envelope.json.tmp` and interleaved into a torn payload, or one deleted the shards the other's envelope named. 39 of 40 rounds ended in a total cache loss. The temp name carries a nonce again, as it did before the shard layout. - A save reused a shard filename from its own load snapshot without checking the file was still there. Another process republishing that provider unlinks the old shard, so the stale writer published an envelope naming a deleted file — read back as a corrupt provider and dropped whole, including PR-linked orphans no re-parse can recover. A reused shard is now existence-checked, and re-verified once more immediately before the envelope is published; a vanished one is rewritten from memory. Also: - Progress saves take a 30s floor beside the file counter. Only the claude scan reports per file; every other provider calls saveProgress once at its own boundary, so the counter alone never fired there. - The unreferenced-shard sweep waits an hour (temps still 5 minutes): an unreferenced shard may belong to a concurrent save whose envelope has not landed yet. - The sweep also retires the pre-v8 single-file temps in the parent directory, which nothing writes anymore. - The shard directory is created 0o700. - The claude and provider paths mark the cache dirty where they DELETE a stale entry, not only where they replace it: an unreadable file skips the replace, and the deletion would otherwise live only in memory. - Codex only treats a grown file as an append when the recorded boundary still lands just after a newline, so a same-inode rewrite that happens to end up larger re-parses instead of resuming mid-line.
perf(cache): per-provider session-cache shards and Codex incremental resume
scanProjectDirs ran every cached turn through cachedTurnToClassified — per-call reconstruction plus the turn classifier's category / retries / edit regexes — and only then applied the date slice, so a week view paid to classify all of history to keep a few percent of it. The keep/drop decision now runs on the raw CachedTurn (calls map 1:1 onto assistantCalls, so callsInRange sees the same survivors as the classified slicer) and only survivors are classified, still from their complete call list. The branch and PR-set carries still walk the full ordered turn list, and buildSpawnPrSets still reads the pre-slice turns. Real corpus, warm one-shot status --format menubar-json --no-optimize, identical payloads: today 1561ms -> 1318ms, week 1560ms -> 1410ms, month 1788ms -> 1675ms.
A memo hit must not re-render. The generated stamp is minted per render, so asserting it is unchanged across the hit states that contract directly instead of leaving it implied by byte equality.
…after-slice perf(parser): classify only the turns that survive the date slice
A provider's shard held its whole history, so one appended session rewrote 95 MB. Each provider's files are now split by the UTC month of their first turn - a bucket that is stable across appends, so a growing session never migrates shards - and every shard records the newest month it holds so a ranged load can skip the ones that cannot contribute. Dirty tracking is per bucket: markCacheDirty takes an optional file path and marks both the bucket the entry was last saved in and the one it is in now. A save writes only dirty buckets, carries the refs of months it never loaded, and merges the on-disk shard back in when a bucket is dirty but was never loaded. v8 and v7 caches re-lay-out losslessly.
…uery Every per-file markCacheDirty call site now names the file, so a parse, a re-parse, a failure marker, an orphan eviction and the durable age-out each dirty exactly the month they touched. The two section-level marks (a fingerprint reset, the durable stamp) stay provider-wide. parseAllSessions derives a month scope from its dateRange and threads it through every loadCache call, so a today/week query stops reading the months it cannot report on.
A file's shard span was read off turns[0]/turns[-1], but several providers emit turns non-chronologically (cursor by ROWID, goose/crush/copilot by a DESC ordering). That produced until < bucket - an empty span, so the shard was unreachable at every scope and its sessions re-parsed every run. cacheFileSpan now takes the min and max month over all turns. An entry re-bucketing out of a month the run never loaded (a re-parse that moved its oldest turn, or the getagentseal#441 failure marker that has no turns at all) left the old copy in the carried shard, so one path lived in two shards and a later load could resolve to the stale one. A save now prunes those paths from the shards it carries, and a load merges shards in envelope order, resolving any duplicate to the freshest fingerprint and dirtying both buckets so the next save retires the loser. A carried month whose shard another writer had republished was dropped from the envelope outright, losing expired-transcript PR orphans no re-parse can recover. The envelope is now re-read just before publishing and the current shard name adopted; a ref is dropped only when that envelope lacks it too. The same re-read moves every merge read after the ownership fence and gives the merge one optimistic retry, so the read-modify-write window shrinks to the publish itself. Also: retire an orphaned v8 directory / v7 file left by an interrupted re-layout, age-guarded, once a v9 envelope is published.
…th-shards perf(cache): shard the session cache by provider and month, load only the months a query needs
Reading, decoding and line-parsing a Claude session JSONL is per-file work that touches nothing shared, so it moves onto worker_threads for a large cold parse. parseClaudeFileFull() is the extracted unit both sides run; a worker runs it against an empty dedup set and returns the result as a JSON string, and the parent installs results in the order the serial loop would. Everything with cross-file state stays on the main thread, and a file whose message ids were already claimed (or whose worker failed) re-parses in-process, so the output is identical to the serial path. Thread count is decided per parse: never with <=2 cores, under 2 GB free memory, fewer than 200 pending whole-file re-parses or under 200 MB behind them, so warm and incremental runs spawn nothing. CODEBURN_PARSE_WORKERS overrides it. The pool is terminated when the parse ends, so the resident serve child accumulates no threads.
Pins the gates that keep threads off low-spec machines and warm runs, that a forced worker count bypasses them, that results come back in submission order, that a dead pool reports failure instead of throwing (and the serial fallback lands on the same result), that no thread outlives a parse across back-to-back parses, and that a cold CLI parse with and without workers produces the same payload and the same cache shards.
os.freemem() reports free pages on macOS, not available memory: on an idle 128 GB machine it reads a few hundred MB, so the 2 GB gate switched the worker pool on and off between runs on the platform the desktop app ships to. The gate and the budget now use process.availableMemory() (cgroup/rlimit-aware in a container), falling back to os.totalmem(): serial under 4 GB available, budget min(0.25 * available, 2 GB). An 8 GB box earns 8 threads, a 4 GB box none. The verbose line now carries every decision input — cores, available GB, pending files and bytes — on both the gate and the go path, so one support log explains itself.
The comment at the install site claimed only that an overlapping worker result 'is discarded'. State why the empty-set result is installable at all — an empty id intersection is proof a serial parse would have dropped nothing — and why the tempting shortcut is wrong: parsedTurnsToCachedTurns delta-encodes gitBranch across turns, so dropping one turn changes whether a LATER turn carries a gitBranch key. Overlap discards the whole file, never individual turns. Tests: the end-to-end determinism check now runs both parses over the SAME corpus, so cache shard BODIES are compared byte for byte instead of just their keys, and a new resumed-session fixture (a transcript restating another file's message ids, in both filename orders) makes install order decide the answer. Verified by mutation: removing the discard guard fails it, and yielding worker results out of order fails it. CODEBURN_VERBOSE now reports how many worker results were re-parsed in-process on id overlap, which is what the new test asserts on. The worker bundle's source map is excluded from the published package (-1.8 MB).
…cold-parse perf(parser): parallelize the cold Claude parse across worker threads, hardware-adaptively
110 resume splits take ~1.3s locally but exceeded vitest's 5s default on the CI runner once the parallel suite also hosts the parse-worker tests.
…ume-timeout test(codex): explicit timeout for the every-boundary resume differential
…rker pool Codex is the bigger half of a real cold parse (4 GB of rollouts against 1.8 GB of Claude sessions) and was still decoding one file at a time. A whole-file rollout decode now runs on the getagentseal#1008 pool. parseCodexFileFull is the serial decode with the codex cache switched off: no hit lookup, and the entry it would have written comes back to the parent instead. The worker runs it against an EMPTY dedup set and returns the calls, the keys it claimed, and that entry; the parent installs all three in the serial loop's order, so an empty key intersection is the proof that a serial parse would have dropped nothing either. On overlap the whole file is discarded and re-parsed in-process -- which is what makes a forked rollout safe, since it replays its parent's token_count history under the parent's key namespace and collides outright. Nothing cross-file moves off the main thread: the dedup set, canonical project paths and the codex cache's per-directory state all stay in the parent, and a file the cache can serve exactly or resume into from a byte offset never reaches a worker. The decision is per provider -- the Claude scan and the provider loop run one after the other, so at most one pool is alive -- and the pool is terminated when its scan ends. The workload gate is now files OR bytes rather than both, and the count takes max(files / 50, bytes / 200 MB): a corpus of a few hundred multi-hundred-MB rollouts is as parallelisable as a few thousand small transcripts, and would otherwise have earned one thread or none.
A mixed Claude + Codex corpus with forked rollouts in both creation orders, asserting an identical payload, byte-identical cache shards and a byte-identical codex-results.json between CODEBURN_PARSE_WORKERS=0 and =3, with the codex discard count pinned above zero so the overlap path is really exercised. Plus: a resumable rollout never reaching a worker (the decision line reports no full parses pending after an append), the off-thread decode matching parseCodexFileFull exactly including the cache entry it hands back, no cache file written by the decode itself, no leaked threads, and the files-OR-bytes gate.
…t per parse Three fixes from review, all measured on this box. The workload gate was files OR bytes. The files arm is wrong: 250 pending files holding 117 KB between them spawned 5 threads and ran ~5% SLOWER than serial, and a file count only starts paying for itself around 400. Gate on bytes alone; the count still takes max(files / 50, bytes / 200 MB), so a few hundred huge rollouts keep their threads. The flat 256 MB per-worker memory budget was contradicted by the Codex workload: a 260 MB rollout peaks near 430 MB in its worker, linearly across the pool. It is now derived per parse as clamp(256 MB, 2 x average pending file + 128 MB, 1 GB), which leaves a corpus of small Claude transcripts where it was and stops over-subscribing on rollouts. The parent's buffer of up to pool.size finished results is part of that peak and is named in the comment. The worker/file pairing at both install sites was positional, guarded only by position (Claude) or a path membership check (Codex). Each worker now echoes its path and the parent asserts it, outside the per-file try: a misalignment would install one session's turns under another's path -- a wrong number nobody would ever notice -- so it fails the run rather than being swallowed as a parse failure. On the Claude side that meant hoisting the whole worker-result block above the try, which is safe because an append never consumes a result in either its shortcut or its straddled-fallthrough case.
…codex-parse perf(parser): parallelize the cold Codex rollout parse on the same worker pool
…d month-shard cache re-layout Upstream restructured the session cache from a flat session-cache.v7.json to a month-sharded-by-provider layout (v8, now v9 via envelope re-layout) and split parseAllSessions into a sync wrapper plus an AsyncLocalStorage-scoped parseAllSessionsInCacheScope for correctness around concurrent cache-dir changes. - src/session-cache.ts: cleanupOrphanedTempFiles conflict resolved by keeping both sweeps upstream added (parent-dir versioned-cache-file sweep, and shard-dir blanket .tmp + unreferenced-shard sweep) but extending the parent-dir sweep to also catch status-snapshot.json.*.tmp, since that file lives in the parent cache dir (orthogonal to the month-shard subdirectory) and upstream's shard-dir sweep never sees it. Also repointed the two remaining getCacheDir() call sites in the status-snapshot section to getCodeburnCacheDir(), the renamed/relocated version of that helper. - src/parser.ts: kept CorpusFingerprint/computeCorpusFingerprint verbatim (untouched by upstream's diff), took upstream's parseAllSessions / parseAllSessionsInCacheScope split entirely. Verified: tsc --noEmit clean; tests/cli-status-menubar.test.ts (15/15), tests/session-cache*.test.ts + tests/session-cache-shards.test.ts (119/119), tests/parser*.test.ts (34/34) all pass. A full `vitest run` also surfaced 3 pre-existing failures confined to app/renderer/*.test.ts (missing jsdom in that nested workspace's own node_modules) — unrelated to session-cache.ts/ parser.ts and outside this merge's scope. Co-Authored-By: Claude <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
codeburn status --format menubar-jsontook 25-90+ seconds per call regardless of cache state, causing the polling CodeBurnMenubar.app to stall/stop refreshing. Root cause (confirmed via live profiling, not the original findings doc's hypothesis): the menubar spawns a fresh CLI process per poll, so every call re-parses the full session-cache blob and re-runs the full aggregation pipeline with zero cross-process reuse. This fix adds a disk-persisted status snapshot keyed by a cheap corpus fingerprint, with a settle-window debounce so rapid-fire source writes coalesce into one recompute instead of one per poll.Changes
src/session-cache.ts— disk-persisted status snapshot keyed by a cheap corpus fingerprint (stat-only, no content read)src/parser.ts— settle-window debounce so rapid-fire source writes coalesce into one recompute instead of one per pollsrc/main.ts— wires thestatus --format menubar-jsonpath to read from the persisted snapshot instead of re-parsing + re-aggregating on every polltests/cli-status-menubar.test.ts— new coverage for the snapshot/fingerprint/debounce behaviorSPEC-perf-cache-fix.md— fix design specRally Story
N/A
Test Plan
npm test(includes newtests/cli-status-menubar.test.tscoverage for the snapshot/fingerprint/debounce path)References
PERF-DEFECT-FINDINGS.md(repo root)Checklist
docs/release-notes.md+README.md)context-packs validatepasses (if directory files changed)npm testgreenGenerated by Vera (GitHub Backplane) — Session: b05c*** | Dispatch: 9*