docs(tt4,nfr): record Tier-4 + NFR verification (ADR-0008 A1, ADR-0057)

Document the regression-hardening work from this session:

- ADR-0008 Amendment 1 — Tier 4 realized: the portable-pty + vt100 harness,
  expectrl dropped, serial/fail-fast design, the four flows, sidebar-region
  reads, command pacing (issue #39), CI default-target (advances TT5).
- ADR-0057 (new) — NFR verification strategy: contrast + ΔE2000 gates
  (NFR-5/7), startup/RSS measured (NFR-1/3: release ~29 ms / ~10 MB),
  NFR-2 by architecture, NFR-4/6 reviewer note. Records the two shipped
  light-theme contrast defects this work caught and fixed.
- requirements.md: TT4 [~]->[x], TT5 [/] (only a Windows exec runner left),
  NFR-1/2/3/5/7 ->[x], NFR-4/6 ->[/], each with evidence.
- CHANGELOG.md (new): Keep a Changelog + SemVer, [Unreleased] + [0.2.0].
- README ADR index updated; handoff-76; the working plan.

Docs-only (CI gate skips via paths-ignore).
This commit is contained in:
claude@clouddev1
2026-06-22 13:15:45 +00:00
parent fd63de3441
commit 88204f25c5
7 changed files with 590 additions and 29 deletions
+53
View File
@@ -0,0 +1,53 @@
# Changelog
All notable, user-facing changes to RDBMS Playground are documented here.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0/).
## [Unreleased]
### Added
- Install via **Scoop**, **Homebrew**, and **winget** in addition to the
existing channels.
- Tier-4 end-to-end test suite that exercises the real application in a
terminal (cold launch, save/reopen, export/import, undo) — guards against
regressions the unit and render tests can't see.
- Automated colour-accessibility checks: every theme now has its
foreground/background contrast and its syntax-colour distinctness verified
on each build.
### Fixed
- **Light theme:** string-literal and flag colours in syntax highlighting
were below the WCAG-AA contrast bar; both are now legible. Two dark-theme
token colours that were hard to tell apart have been separated.
- Windows `install.ps1` now works on the in-box Windows PowerShell 5.1 and
updates `PATH` for the current session.
## [0.2.0] - 2026-06-17
### Added
- **Version surfaces:** a `--version` / `-V` command-line flag and an in-app
`version` command, both reporting the build's exact version.
- **Installation options:** publish to crates.io (`cargo install
rdbms-playground`), `cargo binstall` support, a one-line `curl | sh`
installer, and a Windows `install.ps1` — so you no longer have to hand-pick
a release asset.
- **Documentation & landing site** at <https://relplay.org> — the canonical
user guide plus screencast demos.
- A demonstration-mode key alias (Ctrl-G acting as F1) so screencasts can
show the in-app help/hint key.
### Fixed
- Corrected several errors in the contextual-hint help text.
## [0.1.0] - 2026-06-15
First public release: the cross-platform terminal sandbox for learning
relational-database concepts — tables, keys, relationships, indexes, queries
and query plans — in a guided simple mode or a full advanced (SQL) mode, with
projects, undo, and export/import.
[Unreleased]: https://git.lazyeval.net/oli/rdbms-playground/compare/v0.2.0...HEAD
[0.2.0]: https://git.lazyeval.net/oli/rdbms-playground/compare/v0.1.0...v0.2.0
[0.1.0]: https://git.lazyeval.net/oli/rdbms-playground/releases/tag/v0.1.0
+53
View File
@@ -147,3 +147,56 @@ and on a nightly schedule for any extended coverage.
only slow tier and is kept narrow. only slow tier and is kept narrow.
- Adding a feature implies adding tests at the appropriate tier - Adding a feature implies adding tests at the appropriate tier
(or tiers); coverage is not retrofitted later. (or tiers); coverage is not retrofitted later.
## Amendment 1 — 2026-06-22: Tier 4 realized (TT4)
Tier 4 was a specification only (no PTY deps, no tests) until this
amendment. It is now implemented in `tests/e2e_pty.rs`, and the realized
shape refines the original decision in three ways.
**Tooling — `expectrl` dropped.** The decision named
`portable-pty` + `expectrl` + `vt100`. We kept **`portable-pty`** (0.9, to
spawn the real binary in a PTY at a fixed size) and **`vt100`** (0.16, to
parse the output stream into an inspectable cell grid), both current and
maintained. We **dropped `expectrl`**: it bundles its own PTY abstraction
(which conflicts with `portable-pty`) and matches line-by-line, a poor fit
for a full-screen TUI. A small hand-rolled `wait_for(predicate, timeout)`
that polls the vt100 screen replaces it — fewer dependencies, and assertions
read the actual rendered grid.
**Harness shape.** Each test spawns the binary
(`env!("CARGO_BIN_EXE_rdbms-playground")`) under a fresh PTY at **100×30**
(the wide three-region layout, ADR-0046), with its **own temp `--data-dir`**
so it never touches real projects or resume state, and `--theme dark` for
deterministic styling. Tests run **serially** (one PTY + child at a time)
via a poison-tolerant global lock, so timing stays predictable on the
low-parallelism self-hosted runner. Waits are **tight and fail-fast** (3 s):
a slow wait is a genuine hang, and a timeout panics with a full screen dump
for debuggability.
Two non-obvious robustness rules the harness encodes:
- **Read table presence from the Tables *sidebar* region, not the whole
screen** — the Output panel echoes every command, so a table name typed
in a command would pollute a whole-screen match.
- **Pace each command to completion before the next** (real-user / cast
cadence). A command submitted while the previous one's worker rebuild is
still in flight can be misread against a stale schema cache (issue #39,
discovered here — interactive use is unaffected).
**The four flows** mirror this ADR's Tier-4 scope: (1) cold launch → first
DDL → graceful quit (Ctrl-C); (2) create → quit → reopen via `--resume`,
asserting the *column* (not just the table name) round-trips; (3) export →
import into a fresh project → schema **and** data rebuilt; (4) `undo` after
`DROP TABLE`, through the Y/N confirm modal.
**CI.** The target runs by default in `cargo test`, so the existing Linux
gate exercises it on every push — this **advances TT5** (Tier 4 now runs in
CI on Linux). PTY-in-container was validated (openpty + child spawn work
under the Docker runtime the Gitea runner uses); the first real CI run is
the final confirmation. **Windows execution remains out of scope** (no
Windows runner), and the original Tier-4 note about a *nightly broader
coverage* schedule is **not** implemented — the focused four flows are the
committed TT4 scope; broader coverage is deferred.
**Piggybacked NFR measurement.** The same harness measures NFR-1 (startup
to first frame) and NFR-3 (idle RSS, Linux) — see ADR-0057.
+122
View File
@@ -0,0 +1,122 @@
# ADR-0057: Non-functional-requirement verification strategy
## Status
Accepted (2026-06-22).
## Context
The non-functional requirements (`requirements.md` NFR-1..7 — startup,
input latency, memory, distinctive design, colour use, cross-platform
parity, light/dark legibility) were quality bars that had **never been
formally verified**, even though v0.2.0 ships public binaries. With users
now pulling updates, we want each NFR either **gated by an automated test**
(so a regression fails CI) or **measured-and-documented** with a recorded
figure and a reviewer judgement — not left as an aspiration.
The decision below was taken with the user: contrast is a hard gate;
startup and memory are measured against generous bounds (not the literal
target — a tight timing/RSS gate would flake on a shared runner); the
qualitative bars are verified by argument and reviewer note.
## Decision
### NFR-5 / NFR-7 — colour contrast & legibility: **gated**
Unit tests in `src/theme.rs` (Tier 1, run by the existing gate):
- **`all_text_colours_meet_wcag_aa_contrast`** — every text-bearing
foreground (`fg`, `muted`, mode labels, `system`/`error`/`warning`,
`plan_efficient`, and all `tok_*`) must clear **WCAG-AA 4.5:1** against
`bg`, in **both** themes.
- **`advanced_mode_border_meets_ui_contrast`** — the advanced-mode border
carries a mode-warning signal, so it meets the **3:1** non-text /
UI-component threshold. The plain `border` is decorative structural
chrome and is **deliberately exempt** (it sits at 2.20:1 dark / 1.82:1
light by design, so panel borders recede behind content — recorded here
rather than silently tolerated).
- **`syntax_token_colours_are_perceptually_distinct`** — contrast against
the background is necessary but not sufficient: two token colours can
each clear 4.5:1 yet be hard to tell *apart* on a good screen. Every pair
of syntax-token classes that can share a line must clear **CIEDE2000
(ΔE2000) ≥ 15** (post-tuning minimum is ~18). The metric is itself
validated against the Sharma et al. reference vectors
(`delta_e_2000_matches_reference_vectors`).
Computing the real ratios while writing these tests **uncovered two
shipped defects** in the v0.2.0 light theme — `tok_string` at 4.42:1 and
`tok_flag` at 3.15:1, both below the 4.5:1 the scheme promises — plus two
dark-theme token pairs that were perceptually close despite distinct hex
(`tok_type``tok_keyword`, `tok_function``tok_identifier`, ΔE2000 ~14).
All four were fixed in the same change (commit `65eab71`); the gates now
hold the palette to both bars permanently. A dev tool,
`scripts/palette-preview.py`, renders the live palette with contrast +
ΔE2000 for future palette work.
### NFR-1 / NFR-3 — startup & memory: **measured, generously bounded**
The Tier-4 PTY harness (ADR-0008 Amendment 1) measures both:
- **`startup_to_first_frame_is_reasonable`** — spawn → first rendered frame.
- **`idle_memory_footprint_is_reasonable`** — child `/proc/<pid>` VmRSS
after the app is idle (Linux only).
These run against the **debug** binary, so they assert *generous* bounds
(startup < 2000 ms, RSS < 150 MB) for **gross-regression detection**, not
the literal NFR targets — a tight 500 ms / 50 MB gate on a debug binary
under a loaded shared runner would flake. The **real release figures**
(measured 2026-06-22, 5 trials each, this dev machine) are the documented
verification:
| NFR | Target | Release median | Margin |
| --- | --- | --- | --- |
| NFR-1 startup → first frame | < 500 ms | **~29 ms** | ~17× under |
| NFR-3 idle RSS | < 50 MB | **~10 MB** | ~5× under |
(Measurement granularity is the harness poll interval, ~20 ms — figures
are "tens of ms", not sub-ms precise.) Per the user decision these are
**not** wired as tight CI gates; the generous debug-bound tests are the
ongoing regression signal.
### NFR-2 — input responsiveness: **verified by architecture**
"Long-running queries execute off the UI thread so the interface stays
responsive." This is structural and already true: database work runs on a
dedicated **worker thread** (ADR-0010), and terminal input is read on a
**separate Tokio task** (`spawn_event_reader` in `src/runtime.rs`), so a
running query cannot block keystroke handling or rendering. There is no
automated keystroke-latency test — a deterministic sub-16 ms measurement
in a PTY harness would be flaky and low-value — so NFR-2 is verified by
this architectural argument rather than a gate.
### NFR-4 / NFR-6 — distinctive design & cross-platform parity: **reviewer note**
Qualitative bars. **NFR-4** (deliberate, identifiable palette/layout using
box-drawing) is satisfied by the bespoke two-theme palette, the
three-region framed layout, and the annotated plan/relationship rendering —
reviewer judgement, not automatable. **NFR-6** is partly evidenced by the
CI matrix (Linux + macOS **execute** the suite; Windows is **build-only**,
no runner) and is otherwise a reviewer judgement; one documented divergence:
on terminals/multiplexers without true-colour passthrough (e.g. a tmux
session missing `Tc`/`RGB`), the 24-bit palette is quantised to the
256-colour cube and near hues can collide — a terminal-capability issue,
not a palette one.
## Consequences
- Contrast and perceptual distinctness are now **permanently gated** — the
palette cannot regress below WCAG-AA or into look-alike token colours
without failing the build.
- Startup and memory have **recorded figures** and a generous
gross-regression test; the literal targets are met with large margins.
- The qualitative and architectural NFRs are **documented with evidence**
rather than asserted.
- A latent behaviour found during this work — fast DDL→insert misparsing
against a stale schema cache — is tracked as **issue #39** (deferred; no
interactive-user impact).
## Requirements bookkeeping
NFR-1, NFR-3, NFR-5, NFR-7 → `[x]` (gated test and/or measured with
evidence). NFR-2 → `[x]` (architectural). NFR-4, NFR-6 → `[/]` (reviewer
note; not fully automatable).
+2 -1
View File
File diff suppressed because one or more lines are too long
+98
View File
@@ -0,0 +1,98 @@
# Session handoff — 2026-06-22 (76)
Continues from handoff-75 (D3 package managers). This session took up the
**regression-hardening** trio identified as "what's next" now that public
binaries ship: **TT4** (Tier-4 PTY tests), **NFR verification**, and a
**CHANGELOG** — plus a real palette accessibility fix the work uncovered.
## §1. State
**Branch `main`.** Three commits this session:
- `65eab71` `fix(theme)` — WCAG-AA palette fix + contrast/ΔE2000 gates +
`scripts/palette-preview.py`.
- `fd63de3` `test(tt4)` — Tier-4 PTY end-to-end suite (`tests/e2e_pty.rs`).
- **(pending)** a docs commit — ADR-0008 Amendment 1, ADR-0057, README index,
`requirements.md`, `CHANGELOG.md`, this handoff, and the plan doc
`docs/plans/20260622-tt4-nfr-changelog.md`. *(Propose + confirm before
committing; not yet committed at time of writing.)*
**Test baseline: 2519 passed / 0 failed / 1 ignored** (was 2509; +4 theme
gates, +6 e2e_pty). `clippy --all-targets -D warnings` + `fmt --check` clean.
Full suite verified green 3× under parallel load.
## §2. What shipped
**Palette accessibility (ADR-0057, commit `65eab71`).** Computing real WCAG
ratios while writing the gate **found two shipped defects** in v0.2.0's light
theme — `tok_string` (4.42:1) and `tok_flag` (3.15:1), below the 4.5:1 the
scheme promises — and two dark token pairs that were perceptually close
despite distinct hex. All four fixed (only `theme.rs` colours changed, 4
values). New gated tests in `src/theme.rs`:
- `all_text_colours_meet_wcag_aa_contrast` (≥4.5:1, both themes),
- `advanced_mode_border_meets_ui_contrast` (≥3:1; plain border exempt),
- `delta_e_2000_matches_reference_vectors` (CIEDE2000 validated vs Sharma data),
- `syntax_token_colours_are_perceptually_distinct` (ΔE2000 ≥15).
Dev tool **`scripts/palette-preview.py`** renders the live palette (truecolor
swatch + contrast + ΔE2000), supports `theme:key=HEX` overrides for previewing
changes.
**TT4 — Tier-4 PTY tests (ADR-0008 Amendment 1, commit `fd63de3`).**
`tests/e2e_pty.rs` drives the **real built binary** in a pseudo-terminal:
`portable-pty` 0.9 + `vt100` 0.16 (**`expectrl` dropped** — conflicting PTY
layer; replaced by a hand-rolled `wait_for` on the vt100 grid). 100×30, temp
`--data-dir`, serial, tight 3 s fail-fast waits. The four ADR-0008 flows all
green. Runs by default in `cargo test` → the Linux gate now exercises Tier 4
(**advances TT5**). PTY-in-container validated (openpty+spawn under Docker).
- **Gotchas baked into the harness:** read table presence from the Tables
**sidebar** region (the Output panel echoes commands, polluting whole-screen
matches); **pace commands to completion** (a command sent mid-rebuild
misparses — issue #39); simple-mode insert needs the `values` keyword and a
single-value insert needs `add column` (a 2nd col in `with pk …` makes a
**compound PK**).
**NFR verification (ADR-0057).** All seven NFRs formally verified:
- NFR-1/3 **measured** via the PTY harness — release figures **~29 ms** startup
/ **~10 MB** idle RSS (well under 500 ms / 50 MB). Harness tests gate generous
*debug*-binary bounds for gross-regression detection (tight gates declined as
flaky — user decision).
- NFR-5/7 **gated** (the contrast/ΔE2000 tests above).
- NFR-2 by architecture (worker thread ADR-0010 + separate input task).
- NFR-4/6 reviewer note (`[/]`).
**CHANGELOG.md** — Keep a Changelog + SemVer; `[Unreleased]` + `[0.2.0]`
(0.1.0 = pre-history). v0.2.0 tag = `bd5be5e`; the palette fix + package
managers + TT4 are post-tag → Unreleased.
## §3. Decisions taken with the user
- Palette: fix the two light contrast bugs **and** the two dark near-duplicates
(4 hex changes, approved after a rendered preview).
- Borders <3:1 → **accept + document** (decorative chrome).
- Work directly on `main`; ADR-0057 numbered now (not draft).
- NFR perf: measured + generously bounded, **not** tightly CI-gated.
- CHANGELOG depth: 0.2.0 + Unreleased only.
- Schema-cache finding → **file a Gitea issue, defer the fix** (see §4).
## §4. Open / follow-ups
- **Issue #39** (filed this session): simple-mode **Form-B insert**
(`insert into T values (…)`) submitted faster than the post-DDL **schema-cache
refresh** validates against stale schema and is misparsed as SQL. No
interactive-user impact (cast driver + tests pace); deferred fix. Repro +
diagnosis in the issue.
- **TT5 remaining:** only a **Windows execution runner** now (macOS + Tier-4
gaps closed this session). First real CI run is the final confirmation that
the PTY tests pass in the Gitea container (validated locally; very low risk).
- Untouched larger items from the "what's next" survey: hint/help issues
**#36/#37/#38**; the big features (V4 session journal, TU1 tutorial). The
CHANGELOG can later feed `--release-notes-url` into the CI winget job
(handoff-75 §6).
## §5. Process pins
- Commits user-confirmed, no AI attribution, append-only, **push is the user's
step**.
- `/runda` + DA pass was run on both the plan and the implementation (it caught
the two light-theme contrast bugs and the #39 schema-cache behaviour — both
before finalizing).
- Consider a `cargo sweep` at this milestone (release build added ~ a GB).
+189
View File
@@ -0,0 +1,189 @@
# Plan — TT4 (Tier-4 PTY tests) + NFR verification + CHANGELOG
Status: **approved 2026-06-22**, implementation pending. Supersedes the
transient harness plan-mode copy. This is the persistent, version-controlled
plan per project convention (`docs/plans/`).
## Context
RDBMS Playground now ships public binaries (v0.2.0 on crates.io, plus
Scoop/Homebrew/winget). With real users pulling updates, the priority is
**regression protection before the next release**. Three tracked gaps are taken
up together:
- **TT4** (`requirements.md`, ADR-0008 Tier 4): the four critical end-to-end
flows are specified but **no PTY harness exists** — no deps, no tests. Tiers
13 never exercise the *real built binary* through a terminal.
- **NFR-1..7**: quality bars (startup, latency, memory, contrast, distinctive
design, parity) that were **never formally verified** despite shipping.
- **CHANGELOG**: none exists; useful now that releases are public (and lets the
CI `winget` job pass `--release-notes-url` later).
Baseline (Phase 1): `cargo test` = **2509 passed / 0 failed / 1 ignored**;
clippy + fmt clean. This is the number Phase 5 verifies against (new tests add
to it; nothing may regress).
User decisions (Phase 3):
- **TT4 in CI:** own `e2e_pty` test target, runs by default in `cargo test`
(so the Linux gate exercises it every push). **Tight, fail-fast timeouts**
the runner is self-hosted, low-parallelism, low expected flakiness (so a
failure surfaces fast rather than hanging on a generous timeout).
- **NFR:** WCAG contrast = hard gated test; startup + idle-memory = measured via
the PTY harness against *generous* bounds + numbers recorded; NFR-2/4/6 =
verify-by-argument + reviewer note. Each requirement → `[x]`/`[/]` with
evidence.
- **CHANGELOG:** Keep a Changelog + SemVer; `[Unreleased]` + a single
consolidated `[0.2.0]` only (0.1.0 treated as pre-history).
---
## Workstream 1 — TT4: Tier-4 PTY harness + 4 flows
### Tooling (refines ADR-0008's stated `portable-pty`+`expectrl`+`vt100`)
- **`portable-pty`** (wezterm, actively maintained) — spawn the real binary in a
PTY with a fixed window size.
- **`vt100`** — parse the PTY output stream into an inspectable cell grid.
- **Drop `expectrl`** — it bundles its *own* PTY abstraction (conflicts with
portable-pty) and is line-oriented, a poor fit for a full-screen TUI. Replace
it with a small hand-rolled `wait_for(predicate, timeout)` polling the vt100
screen. This deviation is recorded in the ADR-0008 amendment.
- **Checkpoint at implementation:** verify `vt100` is still maintained/current
(global currency rule). If stale, evaluate `avt` as the screen parser before
committing the dep. `cargo add` picks latest; pin sensibly.
### Harness (`tests/e2e_pty.rs` — new `[[test]]` target)
A `support` module providing:
- `struct PtyApp` — opens a PTY (e.g. **100×30**, the wide three-region
layout), spawns `env!("CARGO_BIN_EXE_rdbms-playground")` with a **per-test
temp `--data-dir`** (`tempfile::TempDir`), `--theme dark`, `TERM=xterm-256color`;
a reader thread feeds bytes into a `vt100::Parser`.
- `send(bytes)` (raw; `\r` = Enter, `\x03` = Ctrl-C), `send_line(s)` =
`s` + `\r`.
- `wait_for(&str, Duration)` — poll the vt100 screen text until the substring
appears or the (tight, ~**3 s**) timeout fires; on timeout **panic with a full
screen dump** (fail-fast + debuggable). A `screen_text()` helper for assertions.
- `quit()``send("\x03")`, wait for child exit, assert success.
- Run **serially** (`serial_test` dep, or a shared `Mutex`/single-threaded
target) — PTYs + temp dirs shouldn't race; also keeps timing stable.
- **Instrument** with `eprintln!`/screen dumps on every wait so a CI failure is
diagnosable from logs (per the project's logging discipline).
### The four flows (one `#[test]` each, mirroring ADR-0008 §Tier-4)
1. **Cold launch → DDL → quit.** Launch fresh; `wait_for("SIMPLE")` +
`wait_for("(none yet)")`; `send_line("create table Customers with pk
id(serial)")`; `wait_for("Customers")` in the Tables panel; `quit()`.
2. **Save → restart → reopen.** Launch project path P under the temp data-dir;
create a table; `quit()` (autosave already persisted per-command); relaunch
**same P**; `wait_for("Customers")` — schema restored from text/rebuild.
3. **Export → import → rebuild.** Project A: create table + insert a row;
`send_line("export A.zip")` (explicit path under the temp dir);
`wait_for` export-success note; `quit()`. Fresh project B:
`send_line("import <path>/A.zip")`; `wait_for` import note (rebuild auto-runs
on missing `.db`); assert the table + row are present (`show data` /
Tables panel).
4. **Undo after DROP.** Create table; `send_line("drop table Customers")`;
confirm gone; `send_line("undo")`; **`wait_for("Restore that earlier
state?")`** (the real modal); `send("y")`; `wait_for("Customers")` restored.
### CI
No workflow edit needed for execution: `e2e_pty` is a default target, so the
existing `gate` job's `cargo test --no-fail-fast` runs it on the Linux
container. **Validate locally first** that a PTY opens inside the
`rdbms-playground-ci` image (container `/dev/ptmx`); if not, add `--init`/pts
mount notes to the ADR. Advances **TT5** (Tier-4-in-CI on Linux); Windows
execution runner stays out of scope (no runner).
---
## Workstream 2 — NFR verification
### Gated: WCAG contrast (NFR-5, NFR-7) — `src/theme.rs` `#[cfg(test)]`
- Add a `relative_luminance(Color) -> f64` + `contrast_ratio(a,b) -> f64`
helper (WCAG 2.x sRGB formula) in the test module; a `match Color::Rgb` channel
extractor (panics on non-RGB — also a guard against a future named-colour
regression).
- For **both** `Theme::light()` and `Theme::dark()`, assert **≥ 4.5:1** for
normal-text foregrounds on `bg`: `fg`, `error`, `warning`, `system`,
`plan_efficient`, `mode_simple`, `mode_advanced`, and the info-carrying syntax
tokens (`tok_keyword/identifier/type/number/string/flag/function/error`).
- **`muted`/`tok_punct`** (deliberately dim secondary text) and **borders**
(`border`/`border_advanced`, non-text UI): assert the WCAG **non-text 3:1**
threshold, with a comment citing the WCAG large-text/non-text allowance. If
any *fails even 3:1*, that's a real finding → surface to the user, don't
silently relax. (Supersedes the existing inequality-only theme tests.)
- This runs in the existing gate — permanent regression protection on the palette.
### Measured + documented: startup (NFR-1) + idle memory (NFR-3)
Two measurement tests in `e2e_pty.rs`:
- **Startup:** time from spawn to `wait_for("SIMPLE")` (first rendered frame);
assert a **generous** bound (e.g. < 1500 ms locally) and `eprintln!` the actual.
- **Idle memory:** after the app is idle, read the child's `/proc/<pid>/status`
`VmRSS` (Linux); assert a generous bound and print actual. (Linux-gated via
`#[cfg(target_os = "linux")]`.)
The point is gross-regression detection, not a tight SLA gate (user decision).
### Verify-by-argument + reviewer note (NFR-2, NFR-4, NFR-6)
- **NFR-2** (input off-thread): already architecturally true — DB runs on the
worker thread (ADR-0010) and input on a separate Tokio task (confirmed in
`runtime.rs`). Record the evidence; optionally a test that a query doesn't
block a subsequent keystroke render.
- **NFR-4** (distinctive design) / **NFR-6** (cross-platform parity): qualitative
/ partly-CI. Written reviewer verification (DA-hat) in the NFR verification doc;
parity partly evidenced by the CI matrix (Linux+macOS execute, Windows builds).
### Requirements bookkeeping
Move NFR-5/7 → `[x]` (gated test), NFR-1/3 → `[x]` (measured + bounded test +
recorded numbers), NFR-2 → `[x]` (architectural evidence), NFR-4/6 → `[/]` with
the reviewer note (honest: not fully automatable). Each with a commit/test ref.
---
## Workstream 3 — CHANGELOG
- New **`CHANGELOG.md`** at repo root, **Keep a Changelog** + SemVer.
- `[Unreleased]` (empty/seeded) + one consolidated **`[0.2.0] - <release date>`**
with Added/Changed/Fixed of the notable user-facing changes (install methods,
packaging, `--version`/`version`, seed improvements, hint feature, readline
keys, cross-mode history) mined from ADRs 00530056 + handoffs 7375. 0.1.0 =
pre-history (a one-line note).
- Optional follow-up (flag, don't auto-do): wire `--release-notes-url` into the
CI `winget`/release path — out of scope unless you want it now.
---
## Docs / ADR updates (project discipline)
- **Amend ADR-0008** (Tier-4 section): record the realized tooling
(`portable-pty` + `vt100` + hand-rolled `wait_for`, `expectrl` dropped),
serial execution, tight-timeout/fail-fast stance, the default-target CI wiring,
and the four implemented flows. Update `docs/adr/README.md` in the same edit.
- **New ADR-0057 — NFR verification strategy** (next free number; assigned at
merge per ADR-0000): the contrast-gate + thresholds (incl. the 3:1
muted/border allowance), the measured-not-tightly-gated stance for
startup/memory, and the verify-by-argument treatment of NFR-2/4/6. Add a
short **NFR verification record** (the measured numbers + reviewer notes) —
either in the ADR or a `docs/` companion. Update the README index.
- Update `docs/requirements.md`: TT4 `[~]``[x]`, TT5 `[/]` note (Tier-4 now in
Linux CI; Windows-exec still pending), NFR rows as above.
- A handoff note (next sequence number) per project convention.
---
## Verification (Phase 5)
1. `cargo test` — full suite green; new e2e_pty + contrast tests pass; **2509 +
new, 0 failed, ≤1 ignored**; no regressions vs baseline.
2. `cargo clippy --all-targets -- -D warnings` + `cargo fmt --check` clean.
3. Run `e2e_pty` repeatedly (e.g. 510×) locally to confirm non-flaky under the
tight timeouts before relying on the gate.
4. Confirm a PTY opens inside the CI image (or document the fix).
5. DA-hat pass over each workstream + the NFR reviewer judgements, written down.
6. `cargo sweep` at the milestone (CLAUDE.md build hygiene).
## Risks / checkpoints
- **vt100 currency** — verify before adding; fall back to `avt` if stale.
- **PTY-in-container** — validate early; the whole TT4 CI value depends on it.
- **muted/border contrast** — may fail even 3:1; if so it's a finding to escalate,
not silently relax (could mean a small palette tweak — touches a decided area,
so escalate).
- **Commits** — confirm each message; no AI attribution; append-only; never push.
+72 -27
View File
@@ -918,28 +918,34 @@ since ADR-0027.)
for representative views. for representative views.
- [x] **TT3** Tier 3: synthetic event-loop integration tests - [x] **TT3** Tier 3: synthetic event-loop integration tests
covering the user-facing flows in this checklist. covering the user-facing flows in this checklist.
- [~] **TT4** Tier 4: PTY-based end-to-end for the four critical - [x] **TT4** Tier 4: PTY-based end-to-end for the four critical
flows named in ADR-0008 (cold launch → DDL → quit; save → flows named in ADR-0008 (cold launch → DDL → quit; save →
reopen; export → import → rebuild; undo after DROP). reopen; export → import → rebuild; undo after DROP).
*(Verified 2026-06-07: **nothing is wired** — no *(Implemented 2026-06-22 — ADR-0008 **Amendment 1**;
`portable-pty` / `expectrl` / `vt100` dependencies, no PTY test `tests/e2e_pty.rs` (commit `fd63de3`). Drives the actual built
files; ADR-0008 §Tier-4 is a specification only. The Tier-3 binary in a real pseudo-terminal: `portable-pty` + `vt100`
`tests/it/*_e2e.rs` files are synthetic event-loop tests, not (**`expectrl` dropped** — conflicting PTY layer, replaced by a
PTY. Correcting a stale `CLAUDE.md` line that read "Tier 4 is hand-rolled `wait_for` on the vt100 grid). Fixed 100×30, per-test
wired only for the listed critical flows" — it was not wired at temp `--data-dir`, serial, tight fail-fast 3 s waits; table
all. Genuinely deferred.)* presence read from the Tables sidebar region (Output echoes
commands); commands paced to completion (stale-schema-cache
misparse on faster-than-human input → issue #39). All four flows
green (flow 2 asserts a **column** round-trips, not just the
table name). Was previously a spec-only deferral; the earlier
2026-06-07 "nothing is wired" note is now resolved.)*
- [/] **TT5** CI runs all tiers on Linux, macOS, and Windows on - [/] **TT5** CI runs all tiers on Linux, macOS, and Windows on
stable Rust. stable Rust.
*(Partial, 2026-06-15. **CI is live** on the self-hosted Gitea *(Partial, updated 2026-06-22. **CI is live** on the self-hosted
Actions (`docs/ci/adr/`): the gate runs `clippy -D warnings` + Gitea Actions (`docs/ci/adr/`): the gate runs `clippy -D warnings`
`cargo test` (Tiers 13) on the **Linux** runner for every branch + `cargo test` (now Tiers **14** — the `e2e_pty` target runs by
push / PR, and `release-macos` runs the suite natively on the default) on the **Linux** runner for every branch push / PR, and
**macOS** runner. **Windows is build-only** — cross-compiled, not `release-macos` runs the suite natively on the **macOS** runner.
executed (no Windows runner). **Tier 4** (PTY, TT4) is still **Tier 4 in CI on Linux is now met** (TT4); PTY-in-container was
unwired, so "all tiers" is not yet fully met. "Stable Rust" is validated (openpty + spawn under Docker), first real CI run is the
satisfied by the flake's pinned `1.95.0` (a stable release, not final confirmation. **Windows is build-only** — cross-compiled,
nightly). Remaining for full TT5: a Windows execution runner and not executed (no Windows runner). "Stable Rust" is satisfied by
Tier-4 PTY in CI.)* the flake's pinned `1.95.0`. **Remaining for full TT5: a Windows
execution runner** (the macOS + Tier-4 gaps are now closed).)*
## Cross-cutting ## Cross-cutting
@@ -1015,41 +1021,80 @@ target is measurable, it is stated numerically; where it is
necessarily qualitative, the criterion is named and the bar is necessarily qualitative, the criterion is named and the bar is
"reviewer judgement against the criterion." "reviewer judgement against the criterion."
- [ ] **NFR-1 Performance — startup.** Cold launch to first All seven were formally verified 2026-06-22 (ADR-0057). Approach:
contrast/distinctness **gated** by tests, startup/memory **measured**
against the targets, the rest **verified by argument / reviewer note**.
- [x] **NFR-1 Performance — startup.** Cold launch to first
rendered frame under 500ms on commodity hardware (developer rendered frame under 500ms on commodity hardware (developer
laptop, mid-range desktop). Measured in CI on the Linux runner laptop, mid-range desktop).
as a regression gate. *(Verified 2026-06-22, ADR-0057 — measured via the Tier-4 PTY
- [ ] **NFR-2 Performance — input latency.** Keystroke-to-render harness; release-binary median **~29 ms** to first frame (5
trials, this dev machine), ~17× under target. The harness test
`startup_to_first_frame_is_reasonable` gates a generous
debug-binary bound for gross-regression detection; a tight 500 ms
CI gate was declined as flaky on a shared runner — user decision.)*
- [x] **NFR-2 Performance — input latency.** Keystroke-to-render
latency under 16ms during normal editing; long-running queries latency under 16ms during normal editing; long-running queries
must execute off the UI thread so the interface remains must execute off the UI thread so the interface remains
responsive (typing, scrolling, mode switching) while a query is responsive (typing, scrolling, mode switching) while a query is
running. running.
- [ ] **NFR-3 Performance — resource footprint.** Idle memory *(Verified 2026-06-22, ADR-0057 — by architecture: database work
runs on the dedicated worker thread (ADR-0010) and terminal input
on a separate Tokio task (`spawn_event_reader`, `runtime.rs`), so
a running query cannot block input/render. No automated sub-16 ms
latency test — it would be flaky and low-value.)*
- [x] **NFR-3 Performance — resource footprint.** Idle memory
under 50MB on the smallest target platform; no busy-loops; CPU under 50MB on the smallest target platform; no busy-loops; CPU
near zero when waiting for input. near zero when waiting for input.
- [ ] **NFR-4 Visual quality — distinctive design.** Colour *(Verified 2026-06-22, ADR-0057 — release-binary idle RSS
**~10 MB** (5 trials, Linux), ~5× under target. The event loop
blocks on `recv` (no busy-loop). The harness test
`idle_memory_footprint_is_reasonable` gates a generous
debug-binary RSS bound on Linux.)*
- [/] **NFR-4 Visual quality — distinctive design.** Colour
palette and typography are deliberate and consistent across palette and typography are deliberate and consistent across
views; layout uses Unicode box-drawing and symbols where they views; layout uses Unicode box-drawing and symbols where they
add clarity; rendering avoids the generic flat-default look add clarity; rendering avoids the generic flat-default look
that ships with most TUI frameworks. Criterion: a reviewer can that ships with most TUI frameworks. Criterion: a reviewer can
identify the app from a screenshot of any view. identify the app from a screenshot of any view.
- [ ] **NFR-5 Visual quality — colour use.** Colour conveys *(Reviewer note 2026-06-22, ADR-0057 — satisfied by the bespoke
two-theme palette, the three-region framed layout, and the
annotated plan/relationship rendering; qualitative, not
automatable.)*
- [x] **NFR-5 Visual quality — colour use.** Colour conveys
information rather than decoration: mode indication, query information rather than decoration: mode indication, query
result types (numeric vs text vs null), error severity, result types (numeric vs text vs null), error severity,
syntax highlighting categories. Foreground/background syntax highlighting categories. Foreground/background
combinations meet WCAG-AA contrast (4.5:1 for normal text) combinations meet WCAG-AA contrast (4.5:1 for normal text)
even though we have not committed to broader accessibility. even though we have not committed to broader accessibility.
- [ ] **NFR-6 Cross-platform parity.** Behaviour and visual *(Verified 2026-06-22, ADR-0057 — **gated** in `src/theme.rs`:
every text foreground clears 4.5:1 on both themes; the
advanced-mode border clears 3:1 (plain border decorative-exempt);
token pairs clear ΔE2000 ≥ 15. Writing the gate caught + fixed two
shipped light-theme defects (`tok_string` 4.42:1, `tok_flag`
3.15:1) and two near-duplicate dark pairs, `65eab71`.)*
- [/] **NFR-6 Cross-platform parity.** Behaviour and visual
quality are equivalent across Linux, macOS, and Windows on quality are equivalent across Linux, macOS, and Windows on
crossterm-supported terminals. Platform-specific divergence crossterm-supported terminals. Platform-specific divergence
(e.g. font fallbacks) is documented, not silently tolerated. (e.g. font fallbacks) is documented, not silently tolerated.
- [ ] **NFR-7 Light and dark background support.** The colour *(Reviewer note 2026-06-22, ADR-0057 — partly evidenced by the CI
matrix (Linux + macOS execute the suite incl. Tier 4; Windows is
build-only). Documented divergence: on terminals/multiplexers
without true-colour passthrough the 24-bit palette is quantised to
256 colours and near hues can collide (a terminal-capability
issue). Full parity awaits a Windows execution runner — see TT5.)*
- [x] **NFR-7 Light and dark background support.** The colour
scheme remains legible and visually coherent on both light and scheme remains legible and visually coherent on both light and
dark terminal backgrounds. The mechanism (auto-detect via dark terminal backgrounds. The mechanism (auto-detect via
terminal query, explicit user setting, or both) is an terminal query, explicit user setting, or both) is an
implementation choice, but the outcome is non-negotiable: no implementation choice, but the outcome is non-negotiable: no
dark-on-dark or light-on-light readability failures on either dark-on-dark or light-on-light readability failures on either
background. background.
*(Verified 2026-06-22, ADR-0057 — the WCAG-AA contrast gate
(NFR-5) runs over **both** themes, so the legibility outcome is
enforced on light and dark. Selection mechanism is `--theme` +
`COLORFGBG` auto-detect (OSC-11 querying deferred).)*
--- ---