Skip to content

tool_search_tool reports "No tools found" for tools whose MCP server has not finished registering - indistinguishable from a genuine no-match #5069

Description

@doomslayer2k

Describe the bug

tool_search_tool returns No tools found matching pattern '<x>' — a clean, successful, well-formed negative — when the matching tools exist but their MCP server has not finished registering yet. There is no error, no warning, and no indication that the tool catalogue is still filling.

The result is indistinguishable from a real no-match, so the agent concludes the capability does not exist and proceeds on that false premise. In my case it spent a working session convinced an entire MCP namespace was unavailable, wrote that unavailability into a design document as an established constraint, and nearly filed an upstream issue based on it.

This is the agent-facing consequence of the startup-ordering defect in #4598, but it is a separate fix: even with perfect startup ordering, a search that lands during a reload or reconnect window should not report a clean negative.

Frequency: deterministic within roughly the first 1–3 minutes of a new session with a large MCP fleet. Also observed once as a 3 h 43 m window in a long-running session, cause unconfirmed.

Extensions/plugins: present (canvas extensions and plugins configured), but not required — the behaviour reproduces on a bare copilot session with only mcp-config.json.

Affected version

1.0.90-0

Steps to reproduce

  1. Configure a sizeable MCP fleet in ~/.copilot/mcp-config.json — enough that startup takes tens of seconds. In my environment 18 servers are merged and registration spans ~34 s.
  2. Start a new session.
  3. Within the first few seconds, call tool_search_tool with a pattern that certainly matches tools from one of the slower servers.
  4. Observe: No tools found matching pattern '<x>'.
  5. Wait ~3 minutes. Issue the identical call again.
  6. Observe: Found N matching tool(s): …

Measured, same session, same agent, same arguments (server names redacted):

03:32:48   pattern "<server-a>"  limit 3  ->  No tools found matching pattern '<server-a>'.
03:32:49   pattern "<prefix-b>"  limit 6  ->  No tools found matching pattern '<prefix-b>'.
03:32:49   pattern "<server-c>"  limit 5  ->  No tools found matching pattern '<server-c>'.

03:35:58   pattern "<server-a>"  limit 3  ->  Found 3 matching tool(s): …
03:35:58   pattern "<prefix-b>"  limit 6  ->  Found 6 matching tool(s): …
03:35:59   pattern "<server-c>"  limit 5  ->  Found 5 matching tool(s): …

All three patterns match tools that were present and callable in the same session minutes later.

Expected behavior

tool_search_tool should distinguish "not available yet" from "no such tool". Any of:

  • return a distinct status when MCP registration is still in flight, e.g. No matches yet — N of M MCP servers still registering;
  • block briefly until registration settles, or until a short timeout;
  • include the registration state in the result so the caller can tell a cold catalogue from an empty one.

What it must not do is return a bare negative that reads as authoritative absence. Agents treat deferred-tool search as the only way to discover whether a capability exists, so a false negative here silently removes capability for the rest of the session — and, because nothing is logged, it is invisible afterwards.

Additional context

Related, same subsystem, different faces:

Requested labels: area:tools, area:mcp if it exists, otherwise area:sessions.

App version         1.1.26
OS                  Windows 10 Enterprise 26H2 (build 26300.9457)
Architecture        AMD64
WebView2 runtime    154.0.4258.62
Copilot CLI         1.0.90-0
In-depth investigation

How the false negative was first hit. In a long-running session, seven consecutive tool_search_tool calls returned No tools found across a 3 h 43 m window — including patterns that had matched earlier the same day. Every call recorded success: true with a well-formed empty result, so nothing surfaced as an error anywhere.

Hypotheses tested and falsified. Compaction was the leading suspect, because the first empty result landed 3 minutes after a compaction completed. It was falsified three ways:

Harness Compaction Result after
Bare CLI 153 tokens works
Bare CLI 210,118 tokens works
App-hosted real auto-compaction works

Also ruled out: the regex itself (byte-identical calls returned opposite results hours apart), the model (identical throughout), and a custom agent's tools: allowlist (the agent's resolved list contains the MCP wildcard entries, and the same agent succeeds once warm).

What reproduced it. Issuing the search within seconds of session start. The confound that initially pointed at custom agents was that the first search in a different session returned tools from two fast-starting servers while the slow ones were still dialling — which looked like a filtering difference and was actually a timing difference.

Why it is invisible after the fact. The runtime's process log drops to a fixed ~118 lines/hour heartbeat after the first hour, so none of the failing window is traceable. --log-level debug does emit a useful per-turn turn tool surface resolved record — but it is a CLI flag with no environment-variable equivalent, so a session started by the desktop app cannot be put into debug mode. That diagnostic gap is worth fixing separately.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:mcpMCP server configuration, discovery, connectivity, OAuth, policy, and registryarea:toolsBuilt-in tools: file editing, shell, search, LSP, git, and tool call behavior

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions