fix(mcp): a failed handshake must not strand the child process
Nightly Build / build (push) Successful in 4m5s
Nightly Build / build (push) Successful in 4m5s
An MCP server whose `initialize` answer is an error — a broken or version- mismatched connector — starts fine and then never exits. `McpServer::start` returned `Err` correctly, but the `Child` lives in the read-loop task rather than in the returned value, so `kill_on_drop` followed a task nothing ever drops. Every retry therefore left a live process holding three pipes and a pidfd, and the supervisor's retry ceiling is deliberately not permanent. The end state was not a dead connector but a dead instance: the process hit its 1024-descriptor limit, `accept()` began failing with EMFILE, and incoming connections queued on a socket nobody could accept from — while the process, the port and every other connector still looked healthy. Observed in production at ~5h from the first bad handshake to unreachable, with 229 orphaned interpreters. `stop_server`/`stop_all` had the same hole from the other side: they document the dropped handle as killing the process, but the task holds its end of stdin, so the child stayed blocked on a read that would never return. Both close with one seam. `McpServer` now owns a oneshot sender whose receiver the read-loop selects on; nothing ever sends, so the drop is the message. That covers a deliberate stop, the last `Arc` going away, and a `?` in `start()` unwinding past the local binding before it was ever returned — including the caller's `timeout`, which drops the same future. The loop then kills and, as importantly, reaps: an unreaped child trades the orphan for a zombie holding the same pipes. Both leaks are covered by tests that fail without the fix. Also raise LimitNOFILE to 65536: the installers write it, and update.sh heals an existing unit additively, leaving an admin's own value alone. The leak is the bug, but 1024 for a process sharing descriptors between the listener, every user's SQLite handles and three pipes per connector is thin regardless.
This commit is contained in:
@@ -62,9 +62,18 @@ release PR may merge — and a section is closed at the commit that bumps it.
|
||||
- Unencrypted users are unlocked and their runtimes started at boot, so Telegram, cron and
|
||||
the background agents work after a restart without anyone opening the web app first.
|
||||
- PDFs render through pdf.js instead of an iframe.
|
||||
- The service is allowed 65536 open files instead of the default 1024. New installs get it
|
||||
from the installer and existing ones from an ordinary update, unless you have set your
|
||||
own limit, in which case yours is left alone.
|
||||
|
||||
### Fixed
|
||||
|
||||
- A connector that fails to start no longer leaves its process behind. One that started
|
||||
but answered the handshake wrong — a broken or mismatched connector — was left running
|
||||
on every retry, and the accumulated processes eventually used up every file handle the
|
||||
server had: within hours the app stopped answering altogether, while the process, the
|
||||
port and every other connector still looked healthy. Stopping or deactivating a
|
||||
connector now genuinely ends its process too.
|
||||
- The server keeps running after you log out of the box; the install / update / uninstall
|
||||
scripts were hardened alongside it.
|
||||
- Skald survives a restart of the Docker daemon.
|
||||
|
||||
Reference in New Issue
Block a user