Skip to content

Dashboard Agent V1 — chat, reports, Investigate, Watch - #4418

Draft
kathiekiwi wants to merge 222 commits into
mainfrom
feat/dashboard-agent-flows
Draft

Dashboard Agent V1 — chat, reports, Investigate, Watch#4418
kathiekiwi wants to merge 222 commits into
mainfrom
feat/dashboard-agent-flows

Conversation

@kathiekiwi

@kathiekiwi kathiekiwi commented Jul 29, 2026

Copy link
Copy Markdown
Collaborator

An AI assistant in a side panel on every dashboard page, behind the dashboard-agent feature flag. It reads runs, errors, queues, deploys and health through the public API (read-only, delegated user token), answers with rich cards, and can keep watching things after the conversation ends.

What's inside

  • Foundation@internal/dashboard-agent-contracts (trigger:// URI grammar, intents, watch specs, block envelope), investigations + watches tables, head-start reliability fix, eval sample-rate gate.
  • Reportsget_report renders the deterministic health report as a card (metric grid, sparklines, Next steps button row); stale telemetry is flagged and never trusted for advice.
  • Investigate — hypothesis-driven investigation on a live card with system-owned identity and revisions; entry buttons on failed runs, errors, backed-up queues and waiting runs; code-grounded when a repo is connected; server-generated follow-ups (Show code, View similar, Watch for a repeat).
  • Watch (Dashboard Agent: Watch (background condition watches + wake notifications + alerts) #4456) — one-shot background watches with a compact creation card on run/queue/error/health pages: resolution + observed outcome model, exactly-once wake delivery, expiry sweep, optional investigate-on-attention, standing email/Slack/webhook alerts with one-click unsubscribe. Queue conditions: drain, above/below N, stalled, oldest-age SLA.
  • UI — blank-state hero (Ask AI) with a Tab-to-accept placeholder and colored smart prompts, fullscreen mode, one persistent agent spinner across all phases.
  • Smart prompts — page-aware chips on every env page (37 routes, 24 page kinds); investigate/status chips appear only on loader-backed abnormal state.
  • Tooling — navigation, TRQL queries with live charts, deploy correlation, docs answers; golden eval suite; seeder for a live playground project (db:seed:agent-examples, with --heartbeat / --degrade / --recover for demos).

How to review

GUIDEBOOK.md — 10-minute local setup and a hands-on walkthrough of every case. Component gallery at /storybook/agent-ui.

Notes

  • Everything is gated by canAccessDashboardAgent; no behavior change with the flag off.
  • The seeder's --heartbeat mode is a review-stand crutch and will be removed before merge.

ericallam added 30 commits July 13, 2026 20:38
…signals

Gauges are read inside the enqueue/dequeue Lua and returned on the script reply
as a 2-tuple; counters are cumulative odometers. The run-queue Redis carries no
metrics stream of its own.
…counters

entryOrderKey returns a string built with BigInt math so ordering stays correct at real epoch magnitudes. Odometer keys are namespaced by definition name. The consumer reports null lag for a missing consumer group instead of 0, and empty gauge values parse as NaN rather than 0.
…ng order keys

The wait-time quantile materialized view now excludes wait_ms = 0 rows so it matches the count aggregation. order_key accepts a string or a number. Migration comments no longer contain semicolons that split the migration into invalid statements.
…rride

The queues list tolerates a metrics query failure by rendering without metrics and logging a warning. UsageSparkline renders its total override even when every bucket is zero. The queue detail page returns 404 and its loader skips the metrics query when the feature flag is off. The seed script validates bucket size and only writes ClickHouse against a local host.
A bucket-led ORDER BY DESC combined with fillGaps emitted an ascending WITH FILL (positive step, ascending bounds), which produces invalid or empty fills. Skip the gap-fill rewrite for descending orders and let the plain descending query stand. Adds a DESC fillGaps test.
Packs the stream sequence with a 1e6 factor (was 1e5) so up to 1M entries per millisecond per shard fit before a seq could spill into the next millisecond's range, far above what a single Redis stream can produce. ms*1e6 stays within UInt64. Also fixes the webapp mapping test that still expected a numeric order_key after the switch to a BigInt-derived string.
The queues list and queue detail pages now use the shared TimeFilter (any preset period or a custom date range) and everything on the page follows it: header tiles, per queue metric columns, charts, and stats. The custom period buttons, hand rolled chart cards, and duplicated metric fetch loops are replaced by the ChartCard and Chart primitives, UsageSparkline, and a shared useMetricResourceQuery hook. The ClickHouse list queries take an explicit end bound so fixed ranges query only their window.
Queries using deltaSumTimestampMerge failed with an unknown function error, which broke the queue detail stats and the started counts on the built in Queues dashboard.
The queues list header tiles now render the same line chart, grid, and tooltip as the rest of the metrics charts instead of a row sparkline, with the headline value in the tile header. The env saturation tile draws the environment concurrency limit and burst limit as labeled reference lines. Chart tooltips keep a gap between the series label and the value, and the shared line chart gains showDots and referenceLines options.
Adds an Allocation tab to the Queues page (behind the queue metrics UI flag): overview cards, a burst-aware capacity bar showing each queue allocation and its live usage in a distinct color, an inline-editable limits table with per-queue locks, load-weighted auto-balance, and a review dialog that bulk-applies limits as overrides through the existing concurrency system.

The queue list now defaults to Busiest ordering (with Backlog and Name options). ClickHouse ranks queues by activity over the last 15 minutes and returns just the requested page of names, so the cost per page is one small aggregate regardless of environment size; idle queues follow in name order and any failure falls back to name ordering. The classic page keeps plain name order.
The fallback WHERE injection only targeted the top-level SELECT, so a
query shaped as an outer aggregation over a FROM subquery failed to
compile: the time column only exists inside the subquery. Descend into
the subquery so the fallback lands next to the table reference.
Adds two rollups fed from the raw landing table: a per-queue 5-minute
tier and an environment-level 1-minute tier (gauges plus TDigest wait
quantiles). Ranking now reads the 5m tier and returns the page and the
ranked total in one windowed query instead of two scans.

The 5m materialized view reads raw rather than cascading off the 10s
table: deltaSumTimestamp states hold a single first/last segment, so
merging states in an MV's hash-ordered GROUP BY double-counts bridging
spans. For the same reason the env tier carries no counter columns, and
env-wide counter totals must group by queue before summing.
The built-in queues dashboard's enqueued vs started chart merged counter
states across queues, which mixes unrelated cumulative counters and
returns wrong totals; it now merges per queue and sums outside. Env
header tiles and saturation charts read the environment rollup, so their
cost no longer scales with queue count, and coarse-bucket ranges are
served from the 5m rollup automatically. Queue list ranking runs as one
query, time bounds are aligned to the bucket grid, and repeated
auto-refresh reads share ClickHouse query-cache entries.
… rollup

The env rollup's win comes from dropping the queue dimension, not from
coarser buckets: row count is queue-independent (~8640/day/env), so full
10-second granularity stays cheap at any range. Env header tiles and
saturation charts now resolve short-range detail exactly like the
per-queue charts, and the current-value tiles read the latest 10-second
bucket instead of a minute-wide one.
The simulator's --reset only cleared the raw and 10s tables, leaving
stale rows in the 5m and env rollups. It also force-merges the rollups
after seeding so current-value widgets read cleanly.
kathiekiwi and others added 14 commits August 3, 2026 07:59
…cess

The env layout loader queried the feature flag unconditionally; without
agent access the panel never mounts, so the read was wasted — and main's
new environment-ownership test (which stubs a minimal prisma) caught it.
- the empty chat centers a hero: sparkles icon, 'Ask Trigger' at blank-slate
  title size with the Beta badge, a one-line subtitle, a three-row composer
  with the send button inside, and the suggested prompts as a wrapping row
  of buttons colored by meaning (action indigo, status secondary, explain
  tertiary, docs the docs style) — one slot-to-variant mapping
- an Expand button next to Close takes the panel over everything right of
  the nav bar, like a page: no route, no modal, no remount — the chat keeps
  its transport and draft text; content stays mounted underneath; the
  transcript column gets a prose max-width; the preference persists
- storybook: hero states at panel and fullscreen widths
…op geometry

- the blank state's field shows the top suggested prompt as its placeholder;
  Tab drops it into the field as editable text, never sending it
- the send and stop buttons share identical square geometry
… land

- cmd+J is contextual: closed opens the panel, open starts a new chat;
  closing is Esc or the header's x — the New chat tooltip now shows cmd+J
  (displayed once, registered once)
- in-flight tool work renders as a bare spinner line, not a bordered pill —
  chips are for artifacts that stay, progress is transient
- error evidence and navigate targets normalize the API's friendly id to
  the raw fingerprint, so View similar failures opens the error page
  instead of 'Error not found'
The header button is the ask-ai Button variant (dot-matrix logo, 'Ask AI');
the hero title follows suit.
…er the agent works

One component in the spinner primitives; chat progress, pending tools, the
history thinking/watching markers, panel loading, chart loading and testing
hypotheses all use it, so agent activity reads as the agent rather than
generic loading.
- AgentSpinner rests on the playlist's first shape, so mounting shows no
  logo-head flash — a spinner is born spinning
- the pending indicator keeps one stable element across tool changes: the
  label swaps, the animation never restarts
…estarts

The pending tool line, the generic activity row and the investigation card's
own progress collapse into a single ChatProgress mounted once at the end of
the live turn: phases only swap its label (card phrase > tool phrase >
activity), decided in the pure progress-line module. ChatPendingTool is gone;
the card renders no spinner of its own; AgentSpinner has exactly one live
render site.
Only the runs list and run detail described themselves to the dashboard
agent, so every other page fell to "other" and offered generic chips.
Add handle mappers for the errors list, an error group, the queues list,
a queue, the deployments list and a deployment — loader data only, no
added queries — plus list page kinds in the contracts and an optional
deployment status. Investigate chips now appear for an unhealthy queue
and a deploy that didn't land.
Only the runs, errors, queues and deployments pages described themselves
to the dashboard agent; everything else fell to "other" and offered the
generic chips. Add handle mappers for the remaining 37 env-scoped routes
and 24 page kinds in the contracts, so each page offers an explain and a
docs question about what it actually shows.

Investigate and status chips stay gated on loader data: a scheduled task
with no schedule attached, all its schedules disabled, a paused queue, a
batch whose runs failed, a wait token past its timeout, a bulk action
still running, a spent quota, a prompt pinned to an override, a session
whose run failed. Loader data only, no added queries, no new signals.
…ions + alerts) (#4456)

Stacked on #4418 — the diff against that branch is the complete Watch
feature, extracted so the base agent PR can land without it.
@kathiekiwi kathiekiwi changed the title Dashboard Agent: UI & investigations Dashboard Agent V1 — chat, reports, Investigate, Watch Aug 3, 2026
The dot-matrix logo is drawn on canvas with a white-based ramp, so on the
light theme it was white ink on a white surface: the chat spinner, the
panel's hero logo and the Ask AI button's glyph all disappeared. Adds a
useThemeMode hook and an AgentMonoLogo wrapper that picks the ink from the
active theme, and routes every mono call site through it.
…st darkening it

Each monoLight stop is now the cool grey whose contrast against a white panel
matches that stop's contrast against a dark one, so the ramp keeps mono's
shape: the middle stop stays the extreme the head glows with and the third
stays the dim tail. With mode="light" handling the ghost grid and the opacity
pair untouched, the light logo is a mirror of the dark one rather than a
different-looking icon.
charcoal-100 ink on dark, charcoal-800 on light — pure white/black stays
reserved for the animated peak, so rest is calm and thinking visibly glows.
The trigger-green /25 border washes out on white; light uses the unified
success token at higher alpha.
- ask-ai hover fills with the brand green; ink flips to charcoal-800 (white
  on that green is ~1.4:1) with a hover-swapped dark logo — the canvas can't
  repaint on :hover
- docs buttons keep the docs-blue ink on the light theme (monochrome stays
  for the dark ones) — a text-bright label on white read as plain grey
- the send arrow and stop glyph are white; stop gets a theme-stable filled
  neutral so the glyph has a surface on light
bg-charcoal-700 rendered as black blobs in light-mode prose; the chip now
uses the theme-mapped --muted. Fenced blocks stay deliberately dark.
Raw charcoal read as black bars in light-mode prose; the header uses the
theme-mapped --muted and cells the grid token.
…ding the dashboard

- the seeder now stages the email-sends QUEUE counter (what the watch checks
  and queue pages read) alongside the env-level one — a drain watch on the
  stand no longer one-shots against an empty live counter
- vite ignores seed-*.mts: editing or running a seeder was full-reloading
  every open dashboard tab every few seconds
'Morning after the deploy': report, chart, a two-revision investigation
(latest-wins), two watch confirmations, a one-shot result, a real wake with
backing watch rows (one fired, one still active for the live chip), and
docs citations. Seeded idempotently; --showcase re-seeds just this chat
over a running stand.
…xit agent fullscreen on navigation

An already-late queue was recommended the age SLA — already true, so every
watch one-shot with "that already happened". Late queue now recommends the
drain (the recovery); a healthy queue recommends the age SLA.

Navigating to another page drops the fullscreen takeover back to the side
panel.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants