Skip to content

Stdio client on Windows: dispose returns before the server tree exits, and a request sent during dispose never completes #1894

Description

@MarcelRoozekrans

Two related problems with stdio client teardown in ModelContextProtocol.Core 2.2.0, found while hosting per-workflow-run MCP servers on Windows. Both reproduce under load. #1751 (the cmd.exe /c interposition) and #1836 (stdin never closed on dispose) already cover the neighbouring issues, so this one only describes what they don't.

1. On Windows, dispose returns before the server process has exited

StdioClientTransport starts a Windows server as cmd.exe /c <command>. On dispose, StdioClientSessionTransport.CleanupAsync → DisposeProcess → KillTree does Kill(entireProcessTree: true), then WaitForExit on the cmd.exe process only. Termination of the children is asynchronous, and nothing waits for it. So DisposeAsync returns while the real server and its conhost.exe are still terminating, still holding the working directory and any open files.

Timestamps from one failing run: the client was disposed, and then a Directory.Delete of the server's working directory was attempted.

disposed 19:15:55.727  delete failed 19:15:55.730  directory free 19:15:55.781
  cmd.exe      exited 19:15:55.662
  dotnet.exe   exited 19:15:55.741   (after dispose returned, after the delete failed)
  conhost.exe  exited 19:15:55.736   (after dispose returned, after the delete failed)

Under concurrent load this failed in 53 of 96 test executions. A control program with the same cmd /c dotnet server tree never failed (0 of 160) once every process in the tree was waited on. So the cause is the descendants nobody waits for, not a Windows delay after exit.

Why it matters: a host that removes the server's working directory, or reuses its files, right after dispose gets "being used by another process".

Suggested fix: after the tree kill, wait on every process in the tree, not just the wrapper. Ideally, start the server inside a job object with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE, which would also remove the need for the cmd.exe wrapper (#1751).

2. A request sent during dispose is written but never completed

McpClient.DisposeAsync first disposes the session handler, which cancels the requests pending at that moment. It then disposes the transport, which, because of #1836, keeps the server's stdin open for ShutdownTimeout before killing it. A tools/call that another thread issues in that window is written to the server successfully. But nothing reads the reply any more, and nothing ever completes its pending entry, so the caller waits until its own timeout. In our case that was a 2-minute call timeout.

Suggested fix: once the session handler is disposed, refuse new requests with an ObjectDisposedException or OperationCanceledException, or complete them when the transport closes.

Workaround we use

We wait on the whole process tree ourselves, taking the wrapper's pid from StdioClientCompletionDetails.ProcessId, and we race each call against McpClient.Completion. Both would become unnecessary with the fixes above.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions