perf: 7.5x cheaper terminal frames, 33x smaller scrollback, and no more git-status UI freeze - #44
Merged
Merged
Conversation
_row_to_strip built a rich.Text one character at a time: two dict.get, a Text.append, an eight-field style compare and a fresh rich.Style per cell, then Text.render(). Under a streaming agent that was 53% of the widget's CPU (4.39s of 8.2s in a cProfile of a 200x40 headless app). Replace it with a single-pass run-length encoder that emits rich Segments directly, plus a per-widget cache of rich.Style keyed by the Char's style fields — a real screen resolves to ~39 distinct styles, so essentially every Style construction leaves the render path. The style key is a plain tuple slice of the Char (one C-level op rather than eight attribute loads), which also keeps it usable as a GC-friendly storage key later. Cursor and selection are applied as overlays on top of the runs, in the same order the old code called Text.stylize() in, so the cursor's fg/bg swap still wins over character styles and the selection still wins over both. Measured on a 200x40 screen, 40 rows per pass: _row_to_strip 10.51 ms -> 1.38 ms (7.6x) profile_app 8 1 100% -> 71% of one core Behaviour note: the old loop's forced stylize at x == ncols-1 painted the final column with the *previous* cell's style, so a style change on the very last column was dropped. The run-length encoder gives that cell its own style. A differential harness over 389 rows (random SGR, sparse rows, unicode, cursor at every column, six selection spans, five widths) found no other difference.
ScrollbackScreen.index() kept every scrolled-off line as a dict[int, Char] — one dict plus one pyte Char per cell. pyte's Char is a NamedTuple *subclass*, and CPython's tuple-untracking optimisation (_PyTuple_MaybeUntrack) gates on PyTuple_CheckExact, so those cells are GC-tracked for as long as they are in scrollback and every gen2 collection walks all of them. Measured at 5000 lines x 200 cols: 91 MB per terminal, and a ~560 ms full-collection pause with six terminals — a visible freeze. Store each line as (text, runs) instead: a str plus plain tuples of (start, end, style_key), with style keys interned per screen. Plain exact tuples of primitives do get untracked, so full scrollback costs the GC nothing. Rendering a scrollback line now replays the runs directly, which is also far cheaper than the per-cell live-screen path. Row->text now lives on ScrollbackScreen (scrollback_text / screen_text / line_text) so the widget's selection code and ipc.read_agent_output share one implementation instead of each walking row dicts. Measured (200x40, 5000 lines): RSS per terminal 91.3 MB -> 2.8 MB gen2 pause, 6 terminals 561.9 ms -> 14.3 ms render 40 scrollback rows 9.59 ms -> 0.19 ms _DEFAULT_MAX_SCROLLBACK stays at 5000 — the compact form lands far under the memory target, so there is no reason to shorten history.
pyte's Screen.draw is the hottest thing lazyagent does under a streaming agent — after the render and scrollback fixes it was 84% of the remaining per-frame cost. For every single character it calls wcwidth, tests DECAWM and IRM, re-reads self.cursor and self.buffer, and builds the cell with Char._replace, which is ~10x the cost of the tuple construction it wraps (275k _replace calls per 60 repaints). Add a fast path for the overwhelmingly common case: a run of printable ASCII that fits before the right margin. Printable ASCII is exactly the set of characters that are unconditionally one cell wide, so no wcwidth call is needed at all. Everything the general path exists for — wrapping and the pending-wrap state, insert mode, a non-Latin-1 active charset (\e(0 remaps ASCII to box drawing, so the translate pass is not optional), wide characters, combining marks, unprintables — is excluded up front and delegated to the original implementation. Within one draw() call the cursor attributes cannot change, so cells differ only in their data. Build each distinct character once and share the Char; it is immutable, and sharing also cuts allocation in the live screen buffer. Measured (200x40, min of 4 interleaved runs): pyte parse per 5.6 KB repaint chunk 8.62 ms -> 2.52 ms frame: ingest + render 40 rows 13.36 ms -> 6.27 ms CPU scrolled-back view 13.40 ms -> 4.77 ms CPU Verified by a fuzz differential against the stock draw: 120 runs (60 random streams mixing wide chars, combining marks, wrapping, insert mode, charsets, scrolling and SGR, x use_utf8 on/off), comparing the full buffer, cursor, dirty set and scrollback. Zero mismatches.
Filed under [Unreleased] rather than appended to [0.6.0] — fold into the next release heading if you prefer the existing habit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
_refresh_git_statuses ran on the message pump. It shells out twice per worktree — `git status --porcelain` walks the whole working tree — and did so for every worktree, serially, every 30 seconds. Measured on a repo with 22 worktrees: 2.8-3.7 s per sweep, so roughly six seconds of every minute with the app unable to process a keystroke, a terminal chunk or a scroll. It also ran on startup and on every worktree create/remove, including the ones driven over MCP. Its two neighbours in the same file are already @work(thread=True); this one was not. Three changes: - Move both sweeps into worker threads and apply the result on the main thread via call_from_thread, mirroring _refresh_selected_diff. The full sweep keeps its place on the rescan paths, where all worktrees genuinely need refreshing for the sidebar. - Poll only the selected worktree on the timer, and drop the interval from 30 s to 60 s. Git subprocesses per minute on that repo: 88 -> 2. The selected worktree is also refreshed when selection moves, so arriving at a worktree does not show a stale badge. - Merge fetched statuses into the cache instead of replacing it, so a single-worktree refresh cannot drop the other sidebar badges. Alt+G refreshes the selected worktree's git status on demand — git status only, no worktree re-list. priority=True so it works while a terminal pane has focus, which is when you want it. test_sweep_does_not_block_the_message_pump is the regression guard: it makes the fake git calls sleep and asserts the call returns immediately. Removing the decorator fails it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Scrolling the sidebar with j/k fired a `git status` for every worktree passed through. `exclusive=True` on the worker cancels *waiting* on a superseded fetch, but a thread worker cannot interrupt a subprocess that has already started — so a quick scroll down a long sidebar still paid for every one. The throttle has to happen before the worker is spawned. The refresh now waits _GIT_STATUS_DEBOUNCE (0.4 s) for the selection to settle, and any move cancels a fetch queued for the worktree just left, including moving to the orchestrator. Alt+G is unaffected: an explicit request fires immediately and drops the pending debounce. The pass-through and cancellation tests assert on the timer rather than on wall-clock sleeps — a 50 ms window raced with the test's own awaits while panels mounted, which would have been flaky on a loaded CI box. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…crollback # Conflicts: # CHANGELOG.md
Diagnostic traces from a real session named this: rendering was 43s against 1.2s of parsing. 3147 stdout chunks averaged 76 characters, and 95% of them were under 100 — yet every one repainted all 36 visible rows, because recv() called a bare refresh(), which clears Textual's per-line strip cache. 29% of those repaints overran a 16ms frame; the worst was 200ms for a 77-character chunk. The event-loop watchdog caught 19 stalls of 102-222ms, all with render_line on the stack. pyte already records which rows changed in Screen.dirty and expects the consumer to read and clear it; we never did. Drive refresh_line() from that set instead, falling back to a whole-widget repaint when a scroll shifts every virtual row or when more than half the screen changed. The cursor needs handling on top: it is painted as a block on one cell and pyte does not mark a row dirty when only the cursor moved, so track (y, x, hidden) and repaint the row it left and the row it arrived on. Tracking x matters as much as y — sliding along a row otherwise leaves the old cell painted, which the differential harness caught. Also hoist what is constant for a render pass out of the per-line path. rich_style walks the DOM ancestor chain on every access (~50us, against ~35us to encode a whole row) and scrollable_content_region recomputes region arithmetic; render_lines now resolves both once and render_line falls back when called outside a pass. Measured on the logged pattern (small chunks, in-place updates): per chunk 9.25 ms -> 1.76 ms (5.2x) rows rendered 42.0 -> 1.7 (25x) Verified by rendering every chunk through Textual's strip cache and again with the cache cleared, comparing the two: 463 renders across in-place updates, bare cursor moves, hide/show, erase, scrolling, insert/delete lines, scroll regions, an active selection, scrolled-back history, unicode and 400 fuzz iterations — no stale rows. Full-widget output is byte-identical to the previous commit.
The logs pinned the "slow to open an agent" complaint: spawn.mount_pane went from 40ms to 1169ms over one session. It was not the terminal. get_diff inlined the full contents of every untracked file. It sniffed 8 KB to detect binaries and then read the whole file anyway, so a worktree holding a few 30 MB eval dumps returned 287,712,500 bytes of "diff". That string is handed to a Static on the main thread, and Textual measures a Static's height by word-wrapping all of its content on *every layout pass* — around 1 ms per KB of real diff. Mounting an agent pane triggers a layout, which is why spawning got slower as bigger diffs landed in the tab, and why the watchdog kept catching 100-300ms stalls in _apply_diff and in resolve_box_models. Three changes: - Cap get_diff. Untracked files are stat'd first and anything over 32 KB is listed by name and size instead of being read; total output is capped at 64 KB. On the affected worktree: 287 MB -> 65 KB. - Cap again at the widget, including per-line length, since one minified line wraps into thousands of visual rows and it is the wrapped count that costs. get_diff already bounds its output; this is the backstop for anything reaching the widget another way. Showing the worst real diff: ~60s of layout work -> 40ms. - Debounce the diff refresh on selection change, folding it into the timer git status already used. Scrolling the sidebar was launching three git subprocesses per worktree passed through, and parsing each result on the main thread. Caps were picked by measuring the real diff, not guessed: 256 KB still cost 257ms per layout, 128 KB 125ms, 64 KB 54ms. 64 KB is ~800 lines, more than anyone reads in a tab. Alt+G no longer cancels a pending selection refresh — that refresh now also covers the diff, which the shortcut deliberately does not.
Several files were committed with CRLF, and nothing in the repo normalized them: there was no .gitattributes and core.autocrlf is unset. That is a merge hazard, not just a cosmetic one. When two branches each convert such a file to LF independently, every line differs from the CRLF merge base on both sides, so git reports a conflict spanning the whole file for what may be a one-line change. Merging the terminal-performance work into the diagnostics branch hit exactly this: 16 conflict hunks covering all of app.py, of which 2 were real. `* text=auto eol=lf` stores every text file with LF and checks it out with LF everywhere; binary files are still detected and left alone. The 23 files that were CRLF are renormalized here in one commit, so the whitespace churn is isolated and easy to skip with `git log -w` or a blame ignore-revs entry. This lands on develop, the common ancestor of the branches in flight, so each of them inherits the normalization through a normal merge rather than converting the same files again. Branches crossing this commit should merge with `-Xrenormalize`, which now works reliably because the attributes exist to drive the conversion. Content is unchanged: every renormalized blob is byte-identical to its parent once CR is stripped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three pure-performance fixes to the terminal hot path. No feature or behaviour change intended; the one visible difference is a pre-existing bug the rewrite stops reproducing (documented below).
Profiled against the real
ScrollableTerminalin a headless Textual app, 200x40 screen, an Ink-style producer emitting ~20 full-screen repaints/sec (what Claude Code does)._row_to_stripwas 53% of the widget's CPUpyte.Screen.drawFix 1 — run-length encoded render path
_row_to_stripbuilt arich.Textone character at a time: twodict.get, aText.append, an eight-field style compare and a freshrich.Styleper cell, thentext.render(console). That was 53% of the widget's CPU (4.39s of 8.2s).Now a single pass emits
rich.Segments directly, with a per-widgetStylecache keyed by theChar's style fields — a real screen resolves to 39 distinct styles, so essentially everyStyleconstruction leaves the render path. The key is a plain tuple slice of theChar(char[1:]), which is one C-level op instead of eight attribute loads, and doubles as the storage key for Fix 2.Cursor and selection are applied as overlays on top of the runs, in the same order the old code called
Text.stylize()in, so the cursor's fg/bg swap still wins over character styles and the selection still wins over both. Cache is cleared innotify_style_updatealongside the other cached style state.Fix 2 — compact scrollback storage
ScrollbackScreen.index()kept every scrolled-off line as adict[int, Char]. pyte'sCharis a NamedTuple subclass, and CPython's tuple-untracking optimisation (_PyTuple_MaybeUntrack) gates onPyTuple_CheckExact— so a NamedTuple subclass is never untracked. Every cell in scrollback stays GC-tracked forever and is traversed by every gen2 collection. Measured: aCharcosts 12x a plain tuple in GC traversal, and six terminals with full scrollback produced ~560 ms full-collection pauses — a user-visible freeze.Lines are now stored as
(text, runs): astrplus plain tuples of(start, end, style_key), style keys interned per screen, trailing blank cells dropped (the renderer re-pads them). Plain exact tuples of primitives do get untracked, so full scrollback costs the GC essentially nothing. Rendering a scrollback line replays the runs directly, skipping the per-cell pass entirely.Row→text moved onto
ScrollbackScreen(scrollback_text/screen_text/line_text) so selection, double/triple-click andipc.read_agent_outputshare one implementation instead of each walking row dicts._DEFAULT_MAX_SCROLLBACKstays at 5000. The compact form lands ~7x under the memory target on its own, so there was no reason to shorten history.Fix 3 — plain-ASCII fast path in
pyte.Screen.drawWith 1 and 2 in place, parsing was 84% of the remaining per-frame cost, and
Screen.drawwas ~78% of that. For every single character pyte callswcwidth, tests DECAWM and IRM, re-readsself.cursorandself.buffer, and builds the cell withChar._replace— a NamedTuple_replaceis ~10x the cost of the tuple construction it wraps (275k_replacecalls per 60 repaints).pyte_patch.pynow adds a fast path for a run of printable ASCII that fits before the right margin. Printable ASCII is exactly the set of characters that are unconditionally one cell wide, so nowcwidthcall is needed at all. Everything the general path exists for is excluded up front and delegated to the original: wrapping and the pending-wrap state (cursor.x == columns), insert mode, a non-Latin-1 active charset (\e(0remaps ASCII to VT100 box drawing, so the translate pass is not optional), wide characters, combining marks, and unprintables.Within one
draw()call the cursor attributes cannot change, so cells differ only in theirdata. Each distinct character is built once and theCharis shared — it is immutable, and sharing also cuts allocation in the live screen buffer.Fix 4 — git status off the message pump (merged from #45)
Found while profiling the lag above, and it turned out to dominate everything else here by three orders of magnitude.
_refresh_git_statusesran on the message pump, shelling out twice per worktree, serially, for every worktree, every 30 seconds. On a repo with 22 worktrees that is 2.8–3.7 s per sweep — roughly six seconds of every minute with the app unable to process a keystroke. It also fired on startup, onr, on create/remove, and on every MCP worktree change.call_from_thread.j/kdoesn't fetch status for every worktree passed through.exclusive=Truecancels waiting on a superseded fetch but can't interrupt a subprocess already running, so the throttle happens before the worker spawns._git_statusesmerges rather than replaces, so a single-worktree refresh can't blank the other badges.Alt+Grefreshes the selected worktree's git status on demand — status only, no worktree re-list.priority=True, so it works while a terminal pane has focus.Regression guard:
test_sweep_does_not_block_the_message_pumpmakes the fake git calls sleep and asserts the call returns immediately. Verified it fails with the decorators removed, as do the debounce tests with the debounce bypassed.Measurements
Commands (benchmarks live outside the repo and are not committed):
Per-frame cost — the headline
Fixed workload, identical bytes for every configuration, CPU time, min of 4 interleaved runs (the box was busy; interleaving and taking minima keeps the comparison fair):
Components
_row_to_strip, 40 rows_StreamFilter.filter, same chunkScrollback memory + GC (5000 lines x 200 cols)
Per-terminal cost at full scrollback: 91.3 MB → 2.8 MB (target was <20 MB). Gen2 pause with six full terminals: 561.9 ms → 14.3 ms (target was <100 ms).
End-to-end (
profile_app.py 8 1, base and new run back to back)The CPU percentage understates this badly: the old code is saturated, so it drops frames and falls behind the producer — it got through 116 chunks while pinning a core, where the new code got through all 318 and still had headroom. On a quiet machine the same pair reads 100% / 197 chunks vs 93% / 318 chunks after fixes 1 and 2 alone.
bench_frame.pyis the load-independent number.Scaling with the number of worktrees
Modelling what lazyagent actually does — one pane visible via the
ContentSwitcher, the rest hidden. Hidden panes skip render/refresh/scroll but still parse every byte, so a background worktree with a busy agent is not free (min of 3):Worktrees with no output cost nothing per frame, before and after. But the GC pause was global and scaled with every worktree holding scrollback, whether or not its agent was still doing anything — a worktree whose agent finished hours ago still added ~130 ms to every full collection, freezing the whole UI:
That dimension is now flat.
Nothing scales with conversation length any more
bench_longconv.py, empty vs 5000-line scrollback: per-frame cost is flat. CPU now scales with output volume x number of live agents, not history size. The only remaining O(scrollback) operation is_extract_selection(0.8 → 3.4 ms at 5000 lines), which runs per click.Correctness
ColorParseErrorfallback, scrollback round-trip (text + styles + attributes), trailing-blank trim, gap filling, blank lines, FIFO eviction atmax_scrollback, custom-margins no-capture, key interning, scrollback-vs-live render equality, padding, selection over scrollback,line_text,ipc.read_agent_outputover real scrollback, and a 17-casedrawdifferential covering every branch the fast path bails out of.test_scrollback_entries_are_gc_untrackedassertsgc.is_tracked(...)is False for every stored entry, run and key after collection, with a comment explaining why. This is the invariant that silently rots if someone later swaps the tuples for a NamedTuple or dataclass — which would bring the pauses straight back.TestDrawFastPathfeeds identical input to a patched screen and a stock-drawscreen and compares the entire resulting state (buffer, cursor, dirty set). Cases cover wrapping, the pending-wrap state, DECAWM off, insert mode, SCS box drawing, SO/SI, wide characters, combining marks, Latin-1,DEL, and empty draws.render_lineat four scroll offsets with and without a selection, and a 120-run fuzz differential ofdraw(60 random streams mixing wide chars, combining marks, wrapping, insert mode, charsets, scrolling and SGR, xuse_utf8on/off) comparing full buffer + cursor + dirty + scrollback. Zero differences other than the one below.Behaviour note — latent bug found
The old loop's forced
stylizeatx == ncols - 1painted the final column with the previous cell's style, so a style change on the very last column was dropped. The full-widget dump shows it concretely: a red-background cell in the rightmost column renders with a default background today, and correctly red after this change. Any TUI that styles the right edge — a box border, a right-aligned badge, a full-width coloured bar — loses that cell's styling today.I did not go out of my way to fix this, and I could not reproduce it without deliberately writing wrong code in two places (the live and scrollback paths reach the final column differently, so bug-compatibility would have made them inconsistent with each other). Flagging rather than hiding it: it is the only visible change in the diff.
Also left alone, as pre-existing:
_char_rich_styleignoreschar.dimeven thoughpyte_patchtracks it. Behaviour preserved exactly —dimnow rides along in the style key and in scrollback, so honouring it later is a one-line change.recv()skips theRE_ANSI_SEQUENCEscan that maintainsself.mouse_tracking, and_flush_hidden_feeddoes not re-scan. An app that enables mouse tracking while its pane is hidden leavesmouse_trackingstale until the next visible chunk containing a DECSET sequence.pyte.Streamdefaults touse_utf8=True, under which pyte drops SCS (\e(0) and SO/SI entirely, so the active charset is always Latin-1 in practice. The fast path still checks it rather than depending on a flag set elsewhere;test_charset_selection_ignored_under_utf8documents this.Out of scope / follow-up
Fix 3 makes parsing cheaper but nothing rate-limits it — that half of the finding is deliberately not addressed here. Notes for whoever picks it up:
_HIDDEN_FEED_INTERVALdoes not help:_flush_hidden_feedjoins the buffered chunks and feeds the same bytes, so it saves the per-chunkrefresh()/scroll_end()/ regex pass but not the parse cost. Hidden agents are cheaper than they were, but still not free.index(), not by chunk boundaries, and the compact encoder is a pure function of the row — so batching or throttling feeds cannot change scrollback content. That stays a local change.on_resizeflushes buffered output beforeself._screen.resize(...), so output is applied at the geometry it was produced for;on_showflushes for immediate catch-up; the visible path readswas_at_bottombefore feeding and callsscroll_endafter._on_stdout(chars)runs per chunk regardless of visibility (hang detection needs it) and must not move behind a throttle;_after_stdout_processed()fires once per chunk when visible but once per flush when hidden, andMonitoredTerminal's sentinel scanning rides on it.feed, so it can stay outside any parse throttle — but see the latent bug above; it is currently skipped entirely while hidden._StreamFilter.filteris a per-character Python loop over every byte (0.48 ms/chunk) and could short-circuit chunks containing noESC;recv()'sRE_ANSI_SEQUENCEscan (0.05 ms/chunk) could be guarded on"\x1b[?" in chars.