You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Thread search and composer typeahead can take seconds to respond. The investigation isolated three mechanisms: synchronous full-text SQL blocks the server, successive file queries repeat workspace discovery and sorting, and plugin mention results wait for the slowest provider. Users should receive useful suggestions promptly, and a search should not stall unrelated server requests.
This issue is the implementation handoff. PR #4402 is an unmerged prototype, now closed in favor of this issue; its code and measurements remain reference material. The prototype demonstrates that these mechanisms can be improved, but it does not solve total SQL cost and has a cold-search regression. A colleague should choose the smallest implementation that meets the acceptance criteria below rather than merge the prototype as-is.
Bottleneck
Strongest controlled evidence
Direction supported by the experiment
Full-text search blocks unrelated server work
Concurrent health response p50/p95: 120.5/183.8 → 1.3/6.3 ms
Move expensive matching off the server event loop; optimize SQL separately
Repeated file discovery/ranking
Warm file request p50/p95: 56.2/75.0 → 37.0/46.3 ms
Reuse bounded workspace listings and remove redundant sorting
Slow provider holds back fast suggestions
First matching provider rendered p50/p95: 1765.7/1780.7 → 178.5/237.9 ms
Publish providers independently and discard previous-query suggestions
These are controlled fixture comparisons, not claims about production-wide speedups. Original live-system observations and all final measurements are recorded below.
Versions and environment
Measured on 2026-09-26, using packaged source builds of bb-app 0.44.0.
Controlled baseline: a8cc3d740016b9e55cc80e222745525b4bbe6a27, the PR's merge base on main.
Controlled prototype: 3374bc3deb56c38e94e1c6ae9e2fcc8be305d327, the retained PR head.
API/CPU comparison: one blacksmith-4vcpu-ubuntu-2404 CI runner, four AMD EPYC vCPUs, about 15.4 GiB RAM, Node 22.22.0. Both revisions used the same lockfile hash and fixture.
Browser comparison: isolated branch web apps on a four-CPU Linux host, kernel 6.8.0-124-generic, Node 22.22.0, Chrome for Testing 154.0.8037.57. Viewport 1280 × 900 CSS pixels at 2× density.
Real packaged server and enrolled host daemon; synthetic data and deterministic mention providers. No agent inference was involved. Browser/API measurements used direct host loopback and excluded the Connect hop.
Initial diagnosis inspected source at a19b3ed4f7f834a1ecaa474161d031de219f5005 and a live setup with a heavily loaded ten-core Mac. Exact live app/OS versions were not retained with that first report.
At filing, main was 4354b88ce14457fbdf1fc1a8e0db2dd4b5ef90f3. Its only change since the measured baseline was plugins/thread-list/app/list/SidebarVisibilityControls.tsx; the implicated search paths were unchanged. The benchmarks were not rerun on this newer commit.
Steps to reproduce
1. Inspect the already-completed, reproducible comparison
Download the existing evidence; no build or private production database is needed to inspect it:
Start with summary.md, manifest.json, results/*.json, and results/cpu-summary.json. The archive also contains the exact harness, fixture, both packaged runtimes, logs, and 15 CPU profiles.
To repeat the complete API/CPU comparison in remote CI while the reference branch is retained:
Verify that the dispatched head is 3374bc3deb56c38e94e1c6ae9e2fcc8be305d327. The harness switches revisions and installs/builds both packages inside a disposable CI checkout; do not run compare in an active development checkout. The workflow's continued presence in the closed PR is for reproduction, not a requirement for the eventual product fix.
2. Minimal scenarios exercised by that harness
Thread search: a migrated SQLite database with 30 threads, 15 archived, and 54,000 message segments. Each thread has 1,800 segments containing repeated search needle workspace component … text. Search for search, needle, and workspace component with limitPerGroup=20. Start a /health request 15 ms after each search request. Observe both search latency and the unrelated request's delay. Each cold sample uses a fresh process and database connection.
File suggestions: a Git workspace with 10,000 files in 100 directories, queried through POST /api/v1/files/paths and a real enrolled daemon. Each burst requests se, sea, sear, searc, search, with limit=16, files and directories included, hidden files excluded. Wait 2.1 seconds between bursts and 80 ms after each request within a burst. The first request is cold relative to the prototype's two-second listing cache; the next four are warm. These are serial API requests, not a simulated overlapping browser typing stream.
Plugin suggestions: install the fixture plugin, which registers two # mention providers. One resolves in 20 ms, the other in 1,600 ms. The baseline calls the aggregate mention endpoint once; the prototype calls it independently with pluginId and providerId. Measure time to the first group and all groups.
Rendered plugin suggestions: open an isolated thread composer with that plugin, type #arrival-000 through #arrival-009 at 80 ms per character, replacing the prior query each time. Time from the final input event to the first animation frame containing the matching provider row. Record first-provider and both-provider arrival separately. The fixture route was /projects/proj_khstsftpcs/threads/thr_uhqu2728zj; these IDs belong only to the synthetic UI fixture.
Expected: fast plugin results appear while slow providers finish; replacing a query never presents stale matching suggestions as current. File discovery is reused across successive queries with bounded, explicit freshness. Full-text matching does not block unrelated server requests, and performance claims distinguish responsiveness from query latency.
Actual baseline: plugin groups arrive together after the slow provider, previous-query results can remain visible while waiting, settled file queries rediscover the same workspace, and synchronous SQL delays other server work.
Representative lines copied verbatim from the final benchmark's summary.md:
A. Original live-system symptoms — diagnostic observations, not a matched benchmark
Observation
Recorded value
Interpretation / limitation
Thread search for se
3.59 seconds
Initial live API request
SQL ranking/grouping for that investigation
3.03 seconds, 66,260 matching segments
Expensive matching precedes the result limit
File typeahead
0.59–10.7 seconds, 7,642 discovered paths
Large variation under load; not a cache-only comparison
Plugin mentions
0.78–2.21 seconds
Live providers and system load varied
Plugin catalog/contribution request
1.47 seconds
Additional first-use dependency; its individual cause was not isolated
Mac load average
37 on 10 cores
Contention can amplify latency; no measured attribution percentage
Swap used
8.3 GB
Pressure indicator, not a direct measurement of paging during each request
Recorded daemon stalls
Up to 2.55 seconds
Further reason not to attribute the 10.7-second extreme solely to file scanning
These numbers were recovered from the investigation's original report. Raw scripts/fixtures from that early phase were not retained, and UI rendering was not timed then. They explain the reported severity but should not be used to calculate a before/after speedup against the later synthetic fixture.
An intermediate source-development comparison at 6548367b68 recorded search around 0.7 seconds before and after, concurrent health responses 716–912 → 11–23 ms, and a first worker startup of 11.9 seconds. That was not a packaged cold-start benchmark, and its temporary raw artifacts were removed. The retained packaged comparison below supersedes it for quantitative acceptance.
B. Controlled API timings
Same runner, fixture, and revisions; execution order before → after → after → before. Values are p50 / p95 in milliseconds. Cold means a fresh process/connection or expired listing cache, not a flushed OS page cache. These measurements exclude frontend debounce, DOM rendering, and Connect.
Measurement
Baseline p50 / p95
Prototype p50 / p95
Samples per revision
Process startup to healthy
418.4 / 436.7
418.3 / 514.0
10
First search after startup
142.7 / 167.2
245.0 / 299.9
10
Warm full-text search
135.5 / 199.3
131.6 / 198.4
40
Health request during warm search
120.5 / 183.8
1.3 / 6.3
40
First file request in a burst
52.4 / 108.7
50.6 / 109.5
20
Subsequent file requests in a burst
56.2 / 75.0
37.0 / 46.3
80
First plugin group, API
1601.6 / 1602.6
21.2 / 22.0
20
All plugin groups, API
1601.6 / 1602.6
1601.7 / 1603.0
20
Trade-offs: the first search becomes 102.3 ms slower at the median; warm SQL/request latency and cold file latency barely change. The slow provider still takes about 1.6 seconds. The useful gains are server responsiveness, repeated file requests, and access to early provider results.
C. Controlled browser timings
Ten unique queries per revision, using the method above. Timing batches were unprofiled and unrecorded; traces and recordings came from separate runs. Both revisions used the same route, fixture, theme, viewport, and Chrome version. A recording command exceeded the browser plugin deadline, so the final after samples used a replacement isolated session with the same settings.
Measurement
Baseline p50 / p95
Prototype p50 / p95
Final keystroke → first matching provider rendered
1765.7 / 1780.7 ms
178.5 / 237.9 ms
Final keystroke → both providers rendered
1765.7 / 1780.8 ms
1759.0 / 1790.3 ms
All ten browser samples, rounded to 0.1 ms
Query
Before first
After first
Before both
After both
arrival-000
1765.7
237.9
1765.7
1790.3
arrival-001
1758.4
190.9
1758.5
1788.1
arrival-002
1767.8
208.0
1767.9
1759.0
arrival-003
1763.2
173.9
1763.2
1748.8
arrival-004
1780.7
189.7
1780.8
1780.2
arrival-005
1770.2
178.5
1770.2
1788.4
arrival-006
1777.9
175.5
1778.0
1759.6
arrival-007
1768.7
177.5
1768.8
1749.1
arrival-008
1756.0
172.8
1756.1
1754.1
arrival-009
1765.1
180.0
1765.2
1754.6
At approximately 500 ms after replacing the query, the baseline still displays stale suggestions while the prototype displays the matching fast suggestion. Exact baseline/prototype revisions and 2× capture density are as listed above.
1. Full-text matching runs synchronously on the server. The baseline HTTP handler calls the synchronous database search directly. The query groups FTS matches per token/thread, ranks and counts active/archived groups, applies the thread limit, then ranks matching segments for those threads. A small output limit does not eliminate the earlier ranking/counting work.
Across nine searches in each CPU-profile run, the SQLite-backed query call accounts for about 1,213 ms on the main server before and 1,201 ms on the search worker after. This is sampled time attributed to the synchronous query call, not an operator-by-operator SQLite execution profile. It corroborates moving the same cost away from the server event loop. It does not demonstrate a faster SQL plan. Startup, shutdown, logging, and GC are also present in the raw profiles; they must not be mistaken for search-only costs.
2. Completed workspace discovery is discarded. The baseline listing map coalesces only requests that overlap in flight and deletes the entry on completion. A later settled query repeats Git-ignore discovery and directory traversal. Plain fuzzy matching also sorts in Fzf and locally before the final merged comparator sort.
The prototype caches the underlying listing for two seconds, bounds it to 32 listings/100,000 paths, invalidates on watcher activity, and removes redundant sorts while retaining the final comparator. Across the matched daemon profiling workload, sampled node:path self-time on the daemon main thread falls 267.7 → 50.5 ms. Cache and sort changes were measured together; their individual contributions were not isolated. Warm requests still pay query-specific ranking, transport, and serialization costs.
3. Plugin calls run concurrently, but publication waits for all of them. The baseline service's Promise.all holds the response until all providers finish or time out. The UI uses one aggregate query and previous-query placeholder data. Therefore a 20 ms provider appears to take 1,600 ms when another provider is slow. The prototype's independent requests and query keys remove this barrier; slow-provider completion itself does not improve.
4. Debounce and discovery are additional stages, not substitutes for these fixes.File/plugin query debounce is 120 ms; thread search debounce is 150 ms. Browser timing includes this delay; API timing does not. Reducing debounce alone cannot explain or resolve seconds of backend waiting and could increase repeated work. The live 1.47-second catalog request and host contention need separate measurement; the deterministic provider comparison did not reproduce or solve their individual causes.
results/: individual API samples, query shapes, comparable result hashes, and sampled CPU summary. Before/after results matched for search order, snippets, group totals, file suggestions, and provider groups.
profiles/: all 15 .cpuprofile files, including the search worker and daemon. For the main query comparison, start with before-cpu-profile-0/CPU.20260926.174826.6848.0.001.cpuprofile and after-cpu-profile-0/CPU.20260926.174849.7161.2.003.cpuprofile.
runner.mjs, fixture.tgz, before-runtime.tgz, after-runtime.tgz, and logs/: exact API reproduction inputs and outputs. CPU-profile runs are separate from latency samples.
A second copy, plus the browser traces, exact UI fixture, browser scripts, screenshots, and recordings, is retained in the originating managed worktree under .search-profile-evidence/. Browser raw traces/fixture are not in the CI artifact; browser timing samples are embedded above. Preserve/transfer those local files before deleting the environment. The report is BB thread thr_kz4aksn579, environment env_5pufnfjh5r.
The fixture intentionally stresses broad matching and slow providers. It does not establish production latency distributions, memory scaling across many workspaces, Connect latency, mobile behavior, or maximum throughput. The browser p95 is based on ten samples per revision. A separate final-head CI shard initially had two command-palette lazy-load timeouts and passed on retry without code/assertion changes; this is unrelated to the unprofiled timing results.
What was ruled out
Only debounce: 120–150 ms does not account for the observed seconds of server/provider waiting.
Only an overloaded Mac: the three mechanisms and prototype improvements were reproduced on controlled Linux fixtures. The Mac's load still plausibly amplified the original extremes.
Only rendering or Connect: baseline delays were present in direct API requests without either stage. The browser comparison separately confirmed the provider publication delay.
Serial execution of plugin providers: providers already execute concurrently; their aggregate response is the barrier.
A SQL speedup from the worker: warm request latency and sampled query cost remained essentially unchanged.
A demonstrated cache improvement for cold discovery: cold file p95 was approximately 109 ms in both revisions.
High responsiveness impact: searches can stall unrelated server work, and common typing flows can wait seconds. No persisted-data loss was observed. Provider publication and bounded listing reuse are relatively contained; full-text SQL optimization and worker lifecycle/consistency need more care. Triage should set the actual priority and effort.
Proposed implementation
1. Publish mention providers independently
Use the prototype's per-provider query path as a starting point. Keep stable provider ordering, show available groups while others finish, and key results by query, trigger, project/thread context, plugin, and provider. Previous-query suggestions must not appear as current; abort obsolete HTTP work and ignore stale responses.
The prototype adds optional pluginId/providerId selectors together and preserves the existing unfiltered mention endpoint. Preserve shipped API/SDK/CLI compatibility. Keep provider timeout/failure handling independent so a failed provider does not hide successful results. Profile catalog availability separately before deciding whether prefetching or cache changes are needed.
2. Cache discovery, then rank the current query once
Reuse a bounded listing keyed by workspace and discovery options, not by each query string. Keep an explicit freshness bound and watcher invalidation for edits, renames, deletions, .gitignore, and nested roots. Unqueried file browsing must remain fresh. Retain in-flight coalescing during invalidation, and avoid persisting results from an invalidated scan.
Remove redundant fuzzy sorts only where the final comparator preserves score, tie-break ordering, limits, and highlight positions. Reference: listing cache and final ranking flow. Measure discovery and ranking separately before adding another cache or persistent index.
3. Separate server responsiveness from SQL optimization
The read-only search worker demonstrates the responsiveness fix. If retained, keep worker lifecycle/shutdown clear, skip cancelled queued queries, recover from failure, and verify source and packaged startup. Skipping queued requests does not interrupt an already-running SQLite query. Bound obsolete work under rapid typing rather than assuming cancellation of fetch cancels SQL.
Address the measured +102 ms median cold-search penalty explicitly. Eager worker initialization or warm-up is an option to measure, not a proven free improvement: it moves work to startup and adds a connection/worker footprint.
For actual query acceleration, collect query plans and timings for broad prefixes, selective tokens, and multi-token queries. Investigate reducing repeated FTS/ranking work and limiting snippet work to selected threads. Preserve exact active/archived counts, visibility/deletion rules, per-thread multi-token matching, ordering ties, and snippets. Do not approximate counts, change matching semantics, or silently move a shipped result contract to eventual consistency merely to improve a benchmark. SQL cost reduction remains open work; the reference PR does not provide a validated faster query.
4. Do not carry prototype complexity forward automatically
The PR contains 699 added lines of profiling script, CI workflow, and documentation. Retain the evidence and reuse the harness for validation, but permanent benchmark infrastructure is not required to ship these fixes. The last simplification request was interrupted by this handoff before edits: duplicate worker cleanup/error handling and the route's redundant result wrapper were identified as simplification opportunities, not implemented changes. Prefer clear local ownership over a new generic cache/worker framework.
Known prototype follow-ups
The one completed Slop Cop review covered 6548367b68 against a19b3ed4f7 and found no P0/P1 findings and three P2s. Later profiling/merge work did not receive a second code review. These are unimplemented prototype concerns, not evidence that they caused the original live incident:
Search membership can change between worker matching and hydration. Concurrent archive/hide writes can produce mismatched groups or hidden hits; hydration rechecks deletion but not all membership rules. Reconcile visibility/archive state and define total-count consistency. Test interleaved persisted writes. Boundary.
Invalidation drops pending discovery entries. Edits during a slow scan can let successive queries launch duplicate scans. Preserve in-flight coalescing while invalidating completed-cache eligibility. Test overlapping requests and refresh after settlement. Invalidation.
The queue tail retains the last successful payload.result.catch(() => undefined) can retain rows/full message text until another search or shutdown. Make both barrier outcomes release payloads while preserving result delivery and failure recovery. Queue tail.
Acceptance criteria for the replacement fix
Reproduce on the current base and retain exact revisions, fixture identities, raw samples, profiles, and commands.
Fast-provider results render before the slow provider; query replacement has no stale rows, and final provider ordering/results remain correct. On the recorded fixture, first-result p95 below 300 ms is a demonstrated reference target, not a universal network SLA.
Successive file queries reuse discovery; expiration/invalidation restore fresh results; cache bounds and in-flight coalescing survive edits. Preserve ranking and limits. Aim to retain or improve the recorded 46.3 ms warm-file p95 on comparable hardware.
Full-text search does not block unrelated server work. The fixture's health p95 below 10 ms is a demonstrated reference target. Separately report first-search latency, warm query latency, startup, and CPU; do not call server responsiveness a SQL speedup.
Report and resolve or explicitly accept any cold-start trade-off. Verify packaged behavior rather than relying on a warm development worker.
Preserve full-text results, exact group totals, visibility/archive/deletion behavior, ranking ties, snippets, later writes, and existing unfiltered mention API behavior.
Exercise rapid query replacement, provider errors/timeouts, worker failures/shutdown, and concurrent database/file changes relevant to the chosen implementation.
Keep regression tests focused on these failure mechanisms. Run CI remotely and verify the changed real-product flow in an isolated branch app. Capture final evidence from the replacement head.
Keep raw profiling evidence accessible to the implementing colleague before CI retention expires or the original environment is removed.
Investigation and reference implementation: #4402. Originating BB thread: thr_kz4aksn579.
Summary
Thread search and composer typeahead can take seconds to respond. The investigation isolated three mechanisms: synchronous full-text SQL blocks the server, successive file queries repeat workspace discovery and sorting, and plugin mention results wait for the slowest provider. Users should receive useful suggestions promptly, and a search should not stall unrelated server requests.
This issue is the implementation handoff. PR #4402 is an unmerged prototype, now closed in favor of this issue; its code and measurements remain reference material. The prototype demonstrates that these mechanisms can be improved, but it does not solve total SQL cost and has a cold-search regression. A colleague should choose the smallest implementation that meets the acceptance criteria below rather than merge the prototype as-is.
These are controlled fixture comparisons, not claims about production-wide speedups. Original live-system observations and all final measurements are recorded below.
Versions and environment
a8cc3d740016b9e55cc80e222745525b4bbe6a27, the PR's merge base onmain.3374bc3deb56c38e94e1c6ae9e2fcc8be305d327, the retained PR head.blacksmith-4vcpu-ubuntu-2404CI runner, four AMD EPYC vCPUs, about 15.4 GiB RAM, Node 22.22.0. Both revisions used the same lockfile hash and fixture.6.8.0-124-generic, Node 22.22.0, Chrome for Testing 154.0.8037.57. Viewport 1280 × 900 CSS pixels at 2× density.a19b3ed4f7f834a1ecaa474161d031de219f5005and a live setup with a heavily loaded ten-core Mac. Exact live app/OS versions were not retained with that first report.mainwas4354b88ce14457fbdf1fc1a8e0db2dd4b5ef90f3. Its only change since the measured baseline wasplugins/thread-list/app/list/SidebarVisibilityControls.tsx; the implicated search paths were unchanged. The benchmarks were not rerun on this newer commit.Steps to reproduce
1. Inspect the already-completed, reproducible comparison
Download the existing evidence; no build or private production database is needed to inspect it:
Start with
summary.md,manifest.json,results/*.json, andresults/cpu-summary.json. The archive also contains the exact harness, fixture, both packaged runtimes, logs, and 15 CPU profiles.To repeat the complete API/CPU comparison in remote CI while the reference branch is retained:
Verify that the dispatched head is
3374bc3deb56c38e94e1c6ae9e2fcc8be305d327. The harness switches revisions and installs/builds both packages inside a disposable CI checkout; do not runcomparein an active development checkout. The workflow's continued presence in the closed PR is for reproduction, not a requirement for the eventual product fix.2. Minimal scenarios exercised by that harness
search needle workspace component …text. Search forsearch,needle, andworkspace componentwithlimitPerGroup=20. Start a/healthrequest 15 ms after each search request. Observe both search latency and the unrelated request's delay. Each cold sample uses a fresh process and database connection.POST /api/v1/files/pathsand a real enrolled daemon. Each burst requestsse,sea,sear,searc,search, withlimit=16, files and directories included, hidden files excluded. Wait 2.1 seconds between bursts and 80 ms after each request within a burst. The first request is cold relative to the prototype's two-second listing cache; the next four are warm. These are serial API requests, not a simulated overlapping browser typing stream.#mention providers. One resolves in 20 ms, the other in 1,600 ms. The baseline calls the aggregate mention endpoint once; the prototype calls it independently withpluginIdandproviderId. Measure time to the first group and all groups.#arrival-000through#arrival-009at 80 ms per character, replacing the prior query each time. Time from the final input event to the first animation frame containing the matching provider row. Record first-provider and both-provider arrival separately. The fixture route was/projects/proj_khstsftpcs/threads/thr_uhqu2728zj; these IDs belong only to the synthetic UI fixture.Fixture creation, exact requests, process setup, cleanup, and measurement code: retained harness. Reproduction notes.
Expected vs actual
Expected: fast plugin results appear while slow providers finish; replacing a query never presents stale matching suggestions as current. File discovery is reused across successive queries with bounded, explicit freshness. Full-text matching does not block unrelated server requests, and performance claims distinguish responsiveness from query latency.
Actual baseline: plugin groups arrive together after the slow provider, previous-query results can remain visible while waiting, settled file queries rediscover the same workspace, and synchronous SQL delays other server work.
Representative lines copied verbatim from the final benchmark's
summary.md:Evidence
A. Original live-system symptoms — diagnostic observations, not a matched benchmark
seThese numbers were recovered from the investigation's original report. Raw scripts/fixtures from that early phase were not retained, and UI rendering was not timed then. They explain the reported severity but should not be used to calculate a before/after speedup against the later synthetic fixture.
An intermediate source-development comparison at
6548367b68recorded search around 0.7 seconds before and after, concurrent health responses 716–912 → 11–23 ms, and a first worker startup of 11.9 seconds. That was not a packaged cold-start benchmark, and its temporary raw artifacts were removed. The retained packaged comparison below supersedes it for quantitative acceptance.B. Controlled API timings
Same runner, fixture, and revisions; execution order before → after → after → before. Values are p50 / p95 in milliseconds. Cold means a fresh process/connection or expired listing cache, not a flushed OS page cache. These measurements exclude frontend debounce, DOM rendering, and Connect.
Trade-offs: the first search becomes 102.3 ms slower at the median; warm SQL/request latency and cold file latency barely change. The slow provider still takes about 1.6 seconds. The useful gains are server responsiveness, repeated file requests, and access to early provider results.
C. Controlled browser timings
Ten unique queries per revision, using the method above. Timing batches were unprofiled and unrecorded; traces and recordings came from separate runs. Both revisions used the same route, fixture, theme, viewport, and Chrome version. A recording command exceeded the browser plugin deadline, so the final after samples used a replacement isolated session with the same settings.
All ten browser samples, rounded to 0.1 ms
At approximately 500 ms after replacing the query, the baseline still displays stale suggestions while the prototype displays the matching fast suggestion. Exact baseline/prototype revisions and 2× capture density are as listed above.
Before recording · Prototype recording
D. CPU profiles and causal mechanisms
1. Full-text matching runs synchronously on the server. The baseline HTTP handler calls the synchronous database search directly. The query groups FTS matches per token/thread, ranks and counts active/archived groups, applies the thread limit, then ranks matching segments for those threads. A small output limit does not eliminate the earlier ranking/counting work.
Across nine searches in each CPU-profile run, the SQLite-backed query call accounts for about 1,213 ms on the main server before and 1,201 ms on the search worker after. This is sampled time attributed to the synchronous query call, not an operator-by-operator SQLite execution profile. It corroborates moving the same cost away from the server event loop. It does not demonstrate a faster SQL plan. Startup, shutdown, logging, and GC are also present in the raw profiles; they must not be mistaken for search-only costs.
2. Completed workspace discovery is discarded. The baseline listing map coalesces only requests that overlap in flight and deletes the entry on completion. A later settled query repeats Git-ignore discovery and directory traversal. Plain fuzzy matching also sorts in Fzf and locally before the final merged comparator sort.
The prototype caches the underlying listing for two seconds, bounds it to 32 listings/100,000 paths, invalidates on watcher activity, and removes redundant sorts while retaining the final comparator. Across the matched daemon profiling workload, sampled
node:pathself-time on the daemon main thread falls 267.7 → 50.5 ms. Cache and sort changes were measured together; their individual contributions were not isolated. Warm requests still pay query-specific ranking, transport, and serialization costs.3. Plugin calls run concurrently, but publication waits for all of them. The baseline service's
Promise.allholds the response until all providers finish or time out. The UI uses one aggregate query and previous-query placeholder data. Therefore a 20 ms provider appears to take 1,600 ms when another provider is slow. The prototype's independent requests and query keys remove this barrier; slow-provider completion itself does not improve.4. Debounce and discovery are additional stages, not substitutes for these fixes. File/plugin query debounce is 120 ms; thread search debounce is 150 ms. Browser timing includes this delay; API timing does not. Reducing debounce alone cannot explain or resolve seconds of backend waiting and could increase repeated work. The live 1.47-second catalog request and host contention need separate measurement; the deterministic provider comparison did not reproduce or solve their individual causes.
E. Evidence locations and reproducibility limits
manifest.json: immutable revisions, machine/runtime information, fixture identity, and runtime/archive hashes. Fixture archive SHA-256:6bb9e601f527bfd2a9214de7ecc7c586a9426696ee11069bcd8765b1110d95b8.results/: individual API samples, query shapes, comparable result hashes, and sampled CPU summary. Before/after results matched for search order, snippets, group totals, file suggestions, and provider groups.profiles/: all 15.cpuprofilefiles, including the search worker and daemon. For the main query comparison, start withbefore-cpu-profile-0/CPU.20260926.174826.6848.0.001.cpuprofileandafter-cpu-profile-0/CPU.20260926.174849.7161.2.003.cpuprofile.runner.mjs,fixture.tgz,before-runtime.tgz,after-runtime.tgz, andlogs/: exact API reproduction inputs and outputs. CPU-profile runs are separate from latency samples..search-profile-evidence/. Browser raw traces/fixture are not in the CI artifact; browser timing samples are embedded above. Preserve/transfer those local files before deleting the environment. The report is BB threadthr_kz4aksn579, environmentenv_5pufnfjh5r.The fixture intentionally stresses broad matching and slow providers. It does not establish production latency distributions, memory scaling across many workspaces, Connect latency, mobile behavior, or maximum throughput. The browser p95 is based on ten samples per revision. A separate final-head CI shard initially had two command-palette lazy-load timeouts and passed on retry without code/assertion changes; this is unrelated to the unprofiled timing results.
What was ruled out
EventTimingProcessingEndwhen input events accumulate #2693 concerns a native Chromium renderer stall while the server remains responsive. None describes these measured mechanisms.Suggested priority and effort
High responsiveness impact: searches can stall unrelated server work, and common typing flows can wait seconds. No persisted-data loss was observed. Provider publication and bounded listing reuse are relatively contained; full-text SQL optimization and worker lifecycle/consistency need more care. Triage should set the actual priority and effort.
Proposed implementation
1. Publish mention providers independently
Use the prototype's per-provider query path as a starting point. Keep stable provider ordering, show available groups while others finish, and key results by query, trigger, project/thread context, plugin, and provider. Previous-query suggestions must not appear as current; abort obsolete HTTP work and ignore stale responses.
The prototype adds optional
pluginId/providerIdselectors together and preserves the existing unfiltered mention endpoint. Preserve shipped API/SDK/CLI compatibility. Keep provider timeout/failure handling independent so a failed provider does not hide successful results. Profile catalog availability separately before deciding whether prefetching or cache changes are needed.2. Cache discovery, then rank the current query once
Reuse a bounded listing keyed by workspace and discovery options, not by each query string. Keep an explicit freshness bound and watcher invalidation for edits, renames, deletions,
.gitignore, and nested roots. Unqueried file browsing must remain fresh. Retain in-flight coalescing during invalidation, and avoid persisting results from an invalidated scan.Remove redundant fuzzy sorts only where the final comparator preserves score, tie-break ordering, limits, and highlight positions. Reference: listing cache and final ranking flow. Measure discovery and ranking separately before adding another cache or persistent index.
3. Separate server responsiveness from SQL optimization
The read-only search worker demonstrates the responsiveness fix. If retained, keep worker lifecycle/shutdown clear, skip cancelled queued queries, recover from failure, and verify source and packaged startup. Skipping queued requests does not interrupt an already-running SQLite query. Bound obsolete work under rapid typing rather than assuming cancellation of fetch cancels SQL.
Address the measured +102 ms median cold-search penalty explicitly. Eager worker initialization or warm-up is an option to measure, not a proven free improvement: it moves work to startup and adds a connection/worker footprint.
For actual query acceleration, collect query plans and timings for broad prefixes, selective tokens, and multi-token queries. Investigate reducing repeated FTS/ranking work and limiting snippet work to selected threads. Preserve exact active/archived counts, visibility/deletion rules, per-thread multi-token matching, ordering ties, and snippets. Do not approximate counts, change matching semantics, or silently move a shipped result contract to eventual consistency merely to improve a benchmark. SQL cost reduction remains open work; the reference PR does not provide a validated faster query.
4. Do not carry prototype complexity forward automatically
The PR contains 699 added lines of profiling script, CI workflow, and documentation. Retain the evidence and reuse the harness for validation, but permanent benchmark infrastructure is not required to ship these fixes. The last simplification request was interrupted by this handoff before edits: duplicate worker cleanup/error handling and the route's redundant result wrapper were identified as simplification opportunities, not implemented changes. Prefer clear local ownership over a new generic cache/worker framework.
Known prototype follow-ups
The one completed Slop Cop review covered
6548367b68againsta19b3ed4f7and found no P0/P1 findings and three P2s. Later profiling/merge work did not receive a second code review. These are unimplemented prototype concerns, not evidence that they caused the original live incident:result.catch(() => undefined)can retain rows/full message text until another search or shutdown. Make both barrier outcomes release payloads while preserving result delivery and failure recovery. Queue tail.Acceptance criteria for the replacement fix
Investigation and reference implementation: #4402. Originating BB thread:
thr_kz4aksn579.