Skip to content

stdio server never exits on stdin EOF while the client leaves stdout unread #3674

Description

@superbiche

Initial Checks

Release line

2.x (current stable)

Description

When a stdio client closes the server's stdin while a response is still being written to a full, unread stdout pipe, MCPServer.run() never returns. The process sits idle until it is signalled.

Steps: the client sends a request whose response is larger than the pipe buffer, does not read stdout, then closes stdin. The server keeps running indefinitely. If the client drains stdout, the same sequence exits 0 right away.

Thread dump of the hung server (faulthandler): the main thread is idle in the event loop's select, and one AnyIO worker thread is inside context.run(func, *args), blocked in the C-level write(2) behind stdout_writer's await stdout.write(...) / await stdout.flush() in mcp/server/stdio.py. AsyncFile.write calls to_thread.run_sync without abandon_on_cancel. So when stdin EOF ends the session and the task group cancels stdout_writer, the cancellation waits for a write that can only finish if the client reads. stdio.py already expects this case (closefd=False, because "a worker thread can still block on this descriptor after the transport exits"). The exit path still waits on that thread, though.

Impact is limited for clients that follow the spec's stdio shutdown, which escalates to SIGTERM and then SIGKILL. The TypeScript SDK client does this after 2 s per step. Any client that stops at "close stdin and wait" will leak the server process.

This is related to #2678 but is a different bug. #2678 is about responses lost on EOF, while this one is about the process never exiting on EOF. A fix for #2678 that drains pending responses before exiting needs a bound on that drain. Otherwise it would turn every undeliverable response into this hang.

Example Code

# server.py
from mcp.server.mcpserver import MCPServer

mcp = MCPServer("eof-hang-repro")


@mcp.tool()
def big() -> str:
    """Return a response larger than a pipe buffer."""
    return "x" * 1_000_000


if __name__ == "__main__":
    mcp.run()
# client.py: python client.py [--drain]
import json, subprocess, sys, threading, time
from pathlib import Path

drain = "--drain" in sys.argv
proc = subprocess.Popen([sys.executable, str(Path(__file__).with_name("server.py"))],
                        stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.DEVNULL)


def send(message):
    proc.stdin.write((json.dumps(message) + "\n").encode())
    proc.stdin.flush()


send({"jsonrpc": "2.0", "id": 0, "method": "initialize", "params": {
    "protocolVersion": "2025-06-18", "capabilities": {},
    "clientInfo": {"name": "repro", "version": "0"}}})
proc.stdout.readline()
send({"jsonrpc": "2.0", "method": "notifications/initialized"})
send({"jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {"name": "big", "arguments": {}}})
if drain:
    threading.Thread(target=proc.stdout.read, daemon=True).start()
time.sleep(1)
proc.stdin.close()
try:
    proc.wait(timeout=10)
    print(f"drain={drain}: exited {proc.returncode} after stdin EOF")
except subprocess.TimeoutExpired:
    print(f"drain={drain}: still running 10 s after stdin EOF")
    proc.terminate()
    proc.wait()

Output:

$ python client.py
drain=False: still running 10 s after stdin EOF
$ python client.py --drain
drain=True: exited 0 after stdin EOF

stdio.py is the same in 2.2.0, 2.3.0 and current main. In the downstream project where this came up (kwin-mcp), the hang still occurs after the server's own signal handling was hardened, because it depends only on run() returning after EOF.

Python & MCP Python SDK

Python 3.14.7 (Fedora 44, Linux 7.2)
mcp 2.3.0 (also reproduced on 2.2.0), anyio 4.15.1

@superbiche · user · drafted with Claude Opus 5.5, reviewed before posting; the voice and decisions are mine.

Activity

  1. added
    bugSomething isn't working
    v2Affects the v2 line (2.x on main)
    v1Affects the v1.x maintenance line
    on Oct 10, 2026
  2. nosmile99 commented on Oct 11, 2026

    @nosmile99

    Confirmed the reproduction on main at 91941ed (Linux, Python 3.12) with your exact server.py/client.py: drain=False leaves the server alive 10s after stdin EOF, drain=True exits 0.

    One finding on the suspected mechanism, verified empirically: abandon_on_cancel=True on the write is not sufficient on its own. I probed it in isolation — after abandoning, anyio.run() returns but the process still never exits, because anyio worker threads are non-daemon, so the abandoned thread blocked in write(2) keeps the interpreter alive in threading._shutdown. A complete fix has to get the blocked write off anyio's worker threads (e.g. a dedicated daemon writer thread) plus a bounded shutdown drain, which matches the bound you anticipate for #2678.

    The approach I'd take: move the stdout write+flush into a dedicated daemon thread fed by a FIFO queue (preserves ordering and per-line flush), and on transport shutdown wait up to a bounded timeout for the queue to drain. Write failures (e.g. EPIPE) are recorded by the thread and re-raised at shutdown, preserving the current error-surfacing behavior. I have this implemented locally with regression tests (full unread pipe exits; broken pipe still surfaces; existing tests/server/test_stdio.py green, ruff clean). Per CONTRIBUTING you have first call as reporter — leaving the fix to you unless a maintainer would rather I open the PR.

    Disclosure: the reproduction, analysis, and draft fix were produced with AI assistance and reviewed by me before posting.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingv1Affects the v1.x maintenance linev2Affects the v2 line (2.x on main)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions