Skip to content

Search and typeahead stall on full-text SQL, repeated file discovery, and slow mention providers #4413

Description

@brsbl

Summary

Thread search and composer typeahead can take seconds to respond. The investigation isolated three mechanisms: synchronous full-text SQL blocks the server, successive file queries repeat workspace discovery and sorting, and plugin mention results wait for the slowest provider. Users should receive useful suggestions promptly, and a search should not stall unrelated server requests.

This issue is the implementation handoff. PR #4402 is an unmerged prototype, now closed in favor of this issue; its code and measurements remain reference material. The prototype demonstrates that these mechanisms can be improved, but it does not solve total SQL cost and has a cold-search regression. A colleague should choose the smallest implementation that meets the acceptance criteria below rather than merge the prototype as-is.

Bottleneck Strongest controlled evidence Direction supported by the experiment
Full-text search blocks unrelated server work Concurrent health response p50/p95: 120.5/183.8 → 1.3/6.3 ms Move expensive matching off the server event loop; optimize SQL separately
Repeated file discovery/ranking Warm file request p50/p95: 56.2/75.0 → 37.0/46.3 ms Reuse bounded workspace listings and remove redundant sorting
Slow provider holds back fast suggestions First matching provider rendered p50/p95: 1765.7/1780.7 → 178.5/237.9 ms Publish providers independently and discard previous-query suggestions

These are controlled fixture comparisons, not claims about production-wide speedups. Original live-system observations and all final measurements are recorded below.

Versions and environment

  • Measured on 2026-09-26, using packaged source builds of bb-app 0.44.0.
  • Controlled baseline: a8cc3d740016b9e55cc80e222745525b4bbe6a27, the PR's merge base on main.
  • Controlled prototype: 3374bc3deb56c38e94e1c6ae9e2fcc8be305d327, the retained PR head.
  • API/CPU comparison: one blacksmith-4vcpu-ubuntu-2404 CI runner, four AMD EPYC vCPUs, about 15.4 GiB RAM, Node 22.22.0. Both revisions used the same lockfile hash and fixture.
  • Browser comparison: isolated branch web apps on a four-CPU Linux host, kernel 6.8.0-124-generic, Node 22.22.0, Chrome for Testing 154.0.8037.57. Viewport 1280 × 900 CSS pixels at 2× density.
  • Real packaged server and enrolled host daemon; synthetic data and deterministic mention providers. No agent inference was involved. Browser/API measurements used direct host loopback and excluded the Connect hop.
  • Initial diagnosis inspected source at a19b3ed4f7f834a1ecaa474161d031de219f5005 and a live setup with a heavily loaded ten-core Mac. Exact live app/OS versions were not retained with that first report.
  • At filing, main was 4354b88ce14457fbdf1fc1a8e0db2dd4b5ef90f3. Its only change since the measured baseline was plugins/thread-list/app/list/SidebarVisibilityControls.tsx; the implicated search paths were unchanged. The benchmarks were not rerun on this newer commit.

Steps to reproduce

1. Inspect the already-completed, reproducible comparison

Download the existing evidence; no build or private production database is needed to inspect it:

gh run download 36260015230 --repo get-bb/bb \
  --name search-profile-3374bc3deb56c38e94e1c6ae9e2fcc8be305d327 \
  --dir search-investigation-evidence

Start with summary.md, manifest.json, results/*.json, and results/cpu-summary.json. The archive also contains the exact harness, fixture, both packaged runtimes, logs, and 15 CPU profiles.

To repeat the complete API/CPU comparison in remote CI while the reference branch is retained:

gh workflow run ci.yml --repo get-bb/bb \
  --ref bb/debug-why-search-and-typeahead-is-so-slow-in-bb-thr_kz4aksn579 \
  -f search-profile-base=a8cc3d740016b9e55cc80e222745525b4bbe6a27

Verify that the dispatched head is 3374bc3deb56c38e94e1c6ae9e2fcc8be305d327. The harness switches revisions and installs/builds both packages inside a disposable CI checkout; do not run compare in an active development checkout. The workflow's continued presence in the closed PR is for reproduction, not a requirement for the eventual product fix.

2. Minimal scenarios exercised by that harness

  1. Thread search: a migrated SQLite database with 30 threads, 15 archived, and 54,000 message segments. Each thread has 1,800 segments containing repeated search needle workspace component … text. Search for search, needle, and workspace component with limitPerGroup=20. Start a /health request 15 ms after each search request. Observe both search latency and the unrelated request's delay. Each cold sample uses a fresh process and database connection.
  2. File suggestions: a Git workspace with 10,000 files in 100 directories, queried through POST /api/v1/files/paths and a real enrolled daemon. Each burst requests se, sea, sear, searc, search, with limit=16, files and directories included, hidden files excluded. Wait 2.1 seconds between bursts and 80 ms after each request within a burst. The first request is cold relative to the prototype's two-second listing cache; the next four are warm. These are serial API requests, not a simulated overlapping browser typing stream.
  3. Plugin suggestions: install the fixture plugin, which registers two # mention providers. One resolves in 20 ms, the other in 1,600 ms. The baseline calls the aggregate mention endpoint once; the prototype calls it independently with pluginId and providerId. Measure time to the first group and all groups.
  4. Rendered plugin suggestions: open an isolated thread composer with that plugin, type #arrival-000 through #arrival-009 at 80 ms per character, replacing the prior query each time. Time from the final input event to the first animation frame containing the matching provider row. Record first-provider and both-provider arrival separately. The fixture route was /projects/proj_khstsftpcs/threads/thr_uhqu2728zj; these IDs belong only to the synthetic UI fixture.

Fixture creation, exact requests, process setup, cleanup, and measurement code: retained harness. Reproduction notes.

Expected vs actual

Expected: fast plugin results appear while slow providers finish; replacing a query never presents stale matching suggestions as current. File discovery is reused across successive queries with bounded, explicit freshness. Full-text matching does not block unrelated server requests, and performance claims distinguish responsiveness from query latency.

Actual baseline: plugin groups arrive together after the slow provider, previous-query results can remain visible while waiting, settled file queries rediscover the same workspace, and synchronous SQL delays other server work.

Representative lines copied verbatim from the final benchmark's summary.md:

| warm-search healthMs | 120.5 / 183.8 | 1.3 / 6.3 | 40 / 40 |
| warm-files ms | 56.2 / 75.0 | 37.0 / 46.3 | 80 / 80 |
| mentions firstMs | 1601.6 / 1602.6 | 21.2 / 22.0 | 20 / 20 |
| cold-search searchMs | 142.7 / 167.2 | 245.0 / 299.9 | 10 / 10 |

All comparable result hashes matched. CPU-profile runs are excluded from timings.

Evidence

A. Original live-system symptoms — diagnostic observations, not a matched benchmark

Observation Recorded value Interpretation / limitation
Thread search for se 3.59 seconds Initial live API request
SQL ranking/grouping for that investigation 3.03 seconds, 66,260 matching segments Expensive matching precedes the result limit
File typeahead 0.59–10.7 seconds, 7,642 discovered paths Large variation under load; not a cache-only comparison
Plugin mentions 0.78–2.21 seconds Live providers and system load varied
Plugin catalog/contribution request 1.47 seconds Additional first-use dependency; its individual cause was not isolated
Mac load average 37 on 10 cores Contention can amplify latency; no measured attribution percentage
Swap used 8.3 GB Pressure indicator, not a direct measurement of paging during each request
Recorded daemon stalls Up to 2.55 seconds Further reason not to attribute the 10.7-second extreme solely to file scanning

These numbers were recovered from the investigation's original report. Raw scripts/fixtures from that early phase were not retained, and UI rendering was not timed then. They explain the reported severity but should not be used to calculate a before/after speedup against the later synthetic fixture.

An intermediate source-development comparison at 6548367b68 recorded search around 0.7 seconds before and after, concurrent health responses 716–912 → 11–23 ms, and a first worker startup of 11.9 seconds. That was not a packaged cold-start benchmark, and its temporary raw artifacts were removed. The retained packaged comparison below supersedes it for quantitative acceptance.

B. Controlled API timings

Same runner, fixture, and revisions; execution order before → after → after → before. Values are p50 / p95 in milliseconds. Cold means a fresh process/connection or expired listing cache, not a flushed OS page cache. These measurements exclude frontend debounce, DOM rendering, and Connect.

Measurement Baseline p50 / p95 Prototype p50 / p95 Samples per revision
Process startup to healthy 418.4 / 436.7 418.3 / 514.0 10
First search after startup 142.7 / 167.2 245.0 / 299.9 10
Warm full-text search 135.5 / 199.3 131.6 / 198.4 40
Health request during warm search 120.5 / 183.8 1.3 / 6.3 40
First file request in a burst 52.4 / 108.7 50.6 / 109.5 20
Subsequent file requests in a burst 56.2 / 75.0 37.0 / 46.3 80
First plugin group, API 1601.6 / 1602.6 21.2 / 22.0 20
All plugin groups, API 1601.6 / 1602.6 1601.7 / 1603.0 20

Trade-offs: the first search becomes 102.3 ms slower at the median; warm SQL/request latency and cold file latency barely change. The slow provider still takes about 1.6 seconds. The useful gains are server responsiveness, repeated file requests, and access to early provider results.

C. Controlled browser timings

Ten unique queries per revision, using the method above. Timing batches were unprofiled and unrecorded; traces and recordings came from separate runs. Both revisions used the same route, fixture, theme, viewport, and Chrome version. A recording command exceeded the browser plugin deadline, so the final after samples used a replacement isolated session with the same settings.

Measurement Baseline p50 / p95 Prototype p50 / p95
Final keystroke → first matching provider rendered 1765.7 / 1780.7 ms 178.5 / 237.9 ms
Final keystroke → both providers rendered 1765.7 / 1780.8 ms 1759.0 / 1790.3 ms
All ten browser samples, rounded to 0.1 ms
Query Before first After first Before both After both
arrival-000 1765.7 237.9 1765.7 1790.3
arrival-001 1758.4 190.9 1758.5 1788.1
arrival-002 1767.8 208.0 1767.9 1759.0
arrival-003 1763.2 173.9 1763.2 1748.8
arrival-004 1780.7 189.7 1780.8 1780.2
arrival-005 1770.2 178.5 1770.2 1788.4
arrival-006 1777.9 175.5 1778.0 1759.6
arrival-007 1768.7 177.5 1768.8 1749.1
arrival-008 1756.0 172.8 1756.1 1754.1
arrival-009 1765.1 180.0 1765.2 1754.6

At approximately 500 ms after replacing the query, the baseline still displays stale suggestions while the prototype displays the matching fast suggestion. Exact baseline/prototype revisions and 2× capture density are as listed above.

Before Prototype
Before: stale suggestions while providers finish Prototype: matching fast suggestion appears first

Before recording · Prototype recording

D. CPU profiles and causal mechanisms

1. Full-text matching runs synchronously on the server. The baseline HTTP handler calls the synchronous database search directly. The query groups FTS matches per token/thread, ranks and counts active/archived groups, applies the thread limit, then ranks matching segments for those threads. A small output limit does not eliminate the earlier ranking/counting work.

Across nine searches in each CPU-profile run, the SQLite-backed query call accounts for about 1,213 ms on the main server before and 1,201 ms on the search worker after. This is sampled time attributed to the synchronous query call, not an operator-by-operator SQLite execution profile. It corroborates moving the same cost away from the server event loop. It does not demonstrate a faster SQL plan. Startup, shutdown, logging, and GC are also present in the raw profiles; they must not be mistaken for search-only costs.

2. Completed workspace discovery is discarded. The baseline listing map coalesces only requests that overlap in flight and deletes the entry on completion. A later settled query repeats Git-ignore discovery and directory traversal. Plain fuzzy matching also sorts in Fzf and locally before the final merged comparator sort.

The prototype caches the underlying listing for two seconds, bounds it to 32 listings/100,000 paths, invalidates on watcher activity, and removes redundant sorts while retaining the final comparator. Across the matched daemon profiling workload, sampled node:path self-time on the daemon main thread falls 267.7 → 50.5 ms. Cache and sort changes were measured together; their individual contributions were not isolated. Warm requests still pay query-specific ranking, transport, and serialization costs.

3. Plugin calls run concurrently, but publication waits for all of them. The baseline service's Promise.all holds the response until all providers finish or time out. The UI uses one aggregate query and previous-query placeholder data. Therefore a 20 ms provider appears to take 1,600 ms when another provider is slow. The prototype's independent requests and query keys remove this barrier; slow-provider completion itself does not improve.

4. Debounce and discovery are additional stages, not substitutes for these fixes. File/plugin query debounce is 120 ms; thread search debounce is 150 ms. Browser timing includes this delay; API timing does not. Reducing debounce alone cannot explain or resolve seconds of backend waiting and could increase repeated work. The live 1.47-second catalog request and host contention need separate measurement; the deterministic provider comparison did not reproduce or solve their individual causes.

E. Evidence locations and reproducibility limits

  • Completed remote CI run and 77 MB raw evidence artifact. GitHub reports expiry 2026-12-25 at 17:43 UTC. Download it before expiry; closing the PR does not make retention indefinite.
  • manifest.json: immutable revisions, machine/runtime information, fixture identity, and runtime/archive hashes. Fixture archive SHA-256: 6bb9e601f527bfd2a9214de7ecc7c586a9426696ee11069bcd8765b1110d95b8.
  • results/: individual API samples, query shapes, comparable result hashes, and sampled CPU summary. Before/after results matched for search order, snippets, group totals, file suggestions, and provider groups.
  • profiles/: all 15 .cpuprofile files, including the search worker and daemon. For the main query comparison, start with before-cpu-profile-0/CPU.20260926.174826.6848.0.001.cpuprofile and after-cpu-profile-0/CPU.20260926.174849.7161.2.003.cpuprofile.
  • runner.mjs, fixture.tgz, before-runtime.tgz, after-runtime.tgz, and logs/: exact API reproduction inputs and outputs. CPU-profile runs are separate from latency samples.
  • A second copy, plus the browser traces, exact UI fixture, browser scripts, screenshots, and recordings, is retained in the originating managed worktree under .search-profile-evidence/. Browser raw traces/fixture are not in the CI artifact; browser timing samples are embedded above. Preserve/transfer those local files before deleting the environment. The report is BB thread thr_kz4aksn579, environment env_5pufnfjh5r.

The fixture intentionally stresses broad matching and slow providers. It does not establish production latency distributions, memory scaling across many workspaces, Connect latency, mobile behavior, or maximum throughput. The browser p95 is based on ten samples per revision. A separate final-head CI shard initially had two command-palette lazy-load timeouts and passed on retry without code/assertion changes; this is unrelated to the unprofiled timing results.

What was ruled out

Suggested priority and effort

High responsiveness impact: searches can stall unrelated server work, and common typing flows can wait seconds. No persisted-data loss was observed. Provider publication and bounded listing reuse are relatively contained; full-text SQL optimization and worker lifecycle/consistency need more care. Triage should set the actual priority and effort.

Proposed implementation

1. Publish mention providers independently

Use the prototype's per-provider query path as a starting point. Keep stable provider ordering, show available groups while others finish, and key results by query, trigger, project/thread context, plugin, and provider. Previous-query suggestions must not appear as current; abort obsolete HTTP work and ignore stale responses.

The prototype adds optional pluginId/providerId selectors together and preserves the existing unfiltered mention endpoint. Preserve shipped API/SDK/CLI compatibility. Keep provider timeout/failure handling independent so a failed provider does not hide successful results. Profile catalog availability separately before deciding whether prefetching or cache changes are needed.

2. Cache discovery, then rank the current query once

Reuse a bounded listing keyed by workspace and discovery options, not by each query string. Keep an explicit freshness bound and watcher invalidation for edits, renames, deletions, .gitignore, and nested roots. Unqueried file browsing must remain fresh. Retain in-flight coalescing during invalidation, and avoid persisting results from an invalidated scan.

Remove redundant fuzzy sorts only where the final comparator preserves score, tie-break ordering, limits, and highlight positions. Reference: listing cache and final ranking flow. Measure discovery and ranking separately before adding another cache or persistent index.

3. Separate server responsiveness from SQL optimization

The read-only search worker demonstrates the responsiveness fix. If retained, keep worker lifecycle/shutdown clear, skip cancelled queued queries, recover from failure, and verify source and packaged startup. Skipping queued requests does not interrupt an already-running SQLite query. Bound obsolete work under rapid typing rather than assuming cancellation of fetch cancels SQL.

Address the measured +102 ms median cold-search penalty explicitly. Eager worker initialization or warm-up is an option to measure, not a proven free improvement: it moves work to startup and adds a connection/worker footprint.

For actual query acceleration, collect query plans and timings for broad prefixes, selective tokens, and multi-token queries. Investigate reducing repeated FTS/ranking work and limiting snippet work to selected threads. Preserve exact active/archived counts, visibility/deletion rules, per-thread multi-token matching, ordering ties, and snippets. Do not approximate counts, change matching semantics, or silently move a shipped result contract to eventual consistency merely to improve a benchmark. SQL cost reduction remains open work; the reference PR does not provide a validated faster query.

4. Do not carry prototype complexity forward automatically

The PR contains 699 added lines of profiling script, CI workflow, and documentation. Retain the evidence and reuse the harness for validation, but permanent benchmark infrastructure is not required to ship these fixes. The last simplification request was interrupted by this handoff before edits: duplicate worker cleanup/error handling and the route's redundant result wrapper were identified as simplification opportunities, not implemented changes. Prefer clear local ownership over a new generic cache/worker framework.

Known prototype follow-ups

The one completed Slop Cop review covered 6548367b68 against a19b3ed4f7 and found no P0/P1 findings and three P2s. Later profiling/merge work did not receive a second code review. These are unimplemented prototype concerns, not evidence that they caused the original live incident:

  1. Search membership can change between worker matching and hydration. Concurrent archive/hide writes can produce mismatched groups or hidden hits; hydration rechecks deletion but not all membership rules. Reconcile visibility/archive state and define total-count consistency. Test interleaved persisted writes. Boundary.
  2. Invalidation drops pending discovery entries. Edits during a slow scan can let successive queries launch duplicate scans. Preserve in-flight coalescing while invalidating completed-cache eligibility. Test overlapping requests and refresh after settlement. Invalidation.
  3. The queue tail retains the last successful payload. result.catch(() => undefined) can retain rows/full message text until another search or shutdown. Make both barrier outcomes release payloads while preserving result delivery and failure recovery. Queue tail.

Acceptance criteria for the replacement fix

  • Reproduce on the current base and retain exact revisions, fixture identities, raw samples, profiles, and commands.
  • Fast-provider results render before the slow provider; query replacement has no stale rows, and final provider ordering/results remain correct. On the recorded fixture, first-result p95 below 300 ms is a demonstrated reference target, not a universal network SLA.
  • Successive file queries reuse discovery; expiration/invalidation restore fresh results; cache bounds and in-flight coalescing survive edits. Preserve ranking and limits. Aim to retain or improve the recorded 46.3 ms warm-file p95 on comparable hardware.
  • Full-text search does not block unrelated server work. The fixture's health p95 below 10 ms is a demonstrated reference target. Separately report first-search latency, warm query latency, startup, and CPU; do not call server responsiveness a SQL speedup.
  • Report and resolve or explicitly accept any cold-start trade-off. Verify packaged behavior rather than relying on a warm development worker.
  • Preserve full-text results, exact group totals, visibility/archive/deletion behavior, ranking ties, snippets, later writes, and existing unfiltered mention API behavior.
  • Exercise rapid query replacement, provider errors/timeouts, worker failures/shutdown, and concurrent database/file changes relevant to the chosen implementation.
  • Keep regression tests focused on these failure mechanisms. Run CI remotely and verify the changed real-product flow in an isolated branch app. Capture final evidence from the replacement head.
  • Keep raw profiling evidence accessible to the implementing colleague before CI retention expires or the original environment is removed.

Investigation and reference implementation: #4402. Originating BB thread: thr_kz4aksn579.

AGENT GENERATED

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    hostHost daemon, process lifecycle, memory, event looppartial-reproBug partially reproduced; some claims unverified; see linked reportperfpluginsPlugin SDK, runtime, marketplacethreadsTurns, timeline, messaging, forksuiApp shell, sidebar, composer, rendering

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions