Two related problems with stdio client teardown in ModelContextProtocol.Core 2.2.0, found while hosting per-workflow-run MCP servers on Windows. Both reproduce under load. #1751 (the cmd.exe /c interposition) and #1836 (stdin never closed on dispose) already cover the neighbouring issues, so this one only describes what they don't.
1. On Windows, dispose returns before the server process has exited
StdioClientTransport starts a Windows server as cmd.exe /c <command>. On dispose, StdioClientSessionTransport.CleanupAsync → DisposeProcess → KillTree does Kill(entireProcessTree: true), then WaitForExit on the cmd.exe process only. Termination of the children is asynchronous, and nothing waits for it. So DisposeAsync returns while the real server and its conhost.exe are still terminating, still holding the working directory and any open files.
Timestamps from one failing run: the client was disposed, and then a Directory.Delete of the server's working directory was attempted.
disposed 19:15:55.727 delete failed 19:15:55.730 directory free 19:15:55.781
cmd.exe exited 19:15:55.662
dotnet.exe exited 19:15:55.741 (after dispose returned, after the delete failed)
conhost.exe exited 19:15:55.736 (after dispose returned, after the delete failed)
Under concurrent load this failed in 53 of 96 test executions. A control program with the same cmd /c dotnet server tree never failed (0 of 160) once every process in the tree was waited on. So the cause is the descendants nobody waits for, not a Windows delay after exit.
Why it matters: a host that removes the server's working directory, or reuses its files, right after dispose gets "being used by another process".
Suggested fix: after the tree kill, wait on every process in the tree, not just the wrapper. Ideally, start the server inside a job object with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE, which would also remove the need for the cmd.exe wrapper (#1751).
2. A request sent during dispose is written but never completed
McpClient.DisposeAsync first disposes the session handler, which cancels the requests pending at that moment. It then disposes the transport, which, because of #1836, keeps the server's stdin open for ShutdownTimeout before killing it. A tools/call that another thread issues in that window is written to the server successfully. But nothing reads the reply any more, and nothing ever completes its pending entry, so the caller waits until its own timeout. In our case that was a 2-minute call timeout.
Suggested fix: once the session handler is disposed, refuse new requests with an ObjectDisposedException or OperationCanceledException, or complete them when the transport closes.
Workaround we use
We wait on the whole process tree ourselves, taking the wrapper's pid from StdioClientCompletionDetails.ProcessId, and we race each call against McpClient.Completion. Both would become unnecessary with the fixes above.
Two related problems with stdio client teardown in ModelContextProtocol.Core 2.2.0, found while hosting per-workflow-run MCP servers on Windows. Both reproduce under load. #1751 (the
cmd.exe /cinterposition) and #1836 (stdin never closed on dispose) already cover the neighbouring issues, so this one only describes what they don't.1. On Windows, dispose returns before the server process has exited
StdioClientTransportstarts a Windows server ascmd.exe /c <command>. On dispose,StdioClientSessionTransport.CleanupAsync→DisposeProcess→KillTreedoesKill(entireProcessTree: true), thenWaitForExiton thecmd.exeprocess only. Termination of the children is asynchronous, and nothing waits for it. SoDisposeAsyncreturns while the real server and itsconhost.exeare still terminating, still holding the working directory and any open files.Timestamps from one failing run: the client was disposed, and then a
Directory.Deleteof the server's working directory was attempted.Under concurrent load this failed in 53 of 96 test executions. A control program with the same
cmd /c dotnet servertree never failed (0 of 160) once every process in the tree was waited on. So the cause is the descendants nobody waits for, not a Windows delay after exit.Why it matters: a host that removes the server's working directory, or reuses its files, right after dispose gets "being used by another process".
Suggested fix: after the tree kill, wait on every process in the tree, not just the wrapper. Ideally, start the server inside a job object with
JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE, which would also remove the need for thecmd.exewrapper (#1751).2. A request sent during dispose is written but never completed
McpClient.DisposeAsyncfirst disposes the session handler, which cancels the requests pending at that moment. It then disposes the transport, which, because of #1836, keeps the server's stdin open forShutdownTimeoutbefore killing it. Atools/callthat another thread issues in that window is written to the server successfully. But nothing reads the reply any more, and nothing ever completes its pending entry, so the caller waits until its own timeout. In our case that was a 2-minute call timeout.Suggested fix: once the session handler is disposed, refuse new requests with an
ObjectDisposedExceptionorOperationCanceledException, or complete them when the transport closes.Workaround we use
We wait on the whole process tree ourselves, taking the wrapper's pid from
StdioClientCompletionDetails.ProcessId, and we race each call againstMcpClient.Completion. Both would become unnecessary with the fixes above.