Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,6 +113,8 @@ The `install` command auto-detects installed coding agents and configures their

CBM automatically shares one per-account coordination daemon across Claude Code, Codex, OpenCode, and every other configured client. There is no opt-in setting for MCP servers or hook clients: the first daemon-backed CBM session starts it, each session registers its own work, and the final session shuts it down. The daemon owns long-lived background services such as watchers, shared indexing jobs, and the optional UI. Closing one session cancels work owned only by that session, while work still needed by another session continues.

The one opt-*out* is `CBM_IN_PROCESS=1`, for MCP hosts whose sandbox denies sockets so the daemon can never be reached. Such a session serves existing indexes read-only from its own process and shares nothing with other clients — no indexing jobs, watcher, UI, or coordination state; see [Running without the daemon](docs/CONFIGURATION.md#running-without-the-daemon).

The detached daemon does not depend on an MCP frontend's stderr. It keeps owner-only durable records under the canonical `${CBM_CACHE_DIR}/logs` directory (default `~/.cache/codebase-memory-mcp/logs`):

| File | Contents |
Expand Down
42 changes: 42 additions & 0 deletions docs/CONFIGURATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -159,6 +159,7 @@ These environment variables affect runtime behavior:
| `CBM_CACHE_DIR` | `~/.cache/codebase-memory-mcp` | Override the cache directory used for indexes, `_config.db`, and UI `config.json`. |
| `CBM_DIAGNOSTICS` | `false` | Enable periodic `snapshot.json` and retained `trajectory.ndjson` below a fresh owner-private directory in the system temp directory. The daemon records the randomized paths in the `diagnostics.start` discovery record (a single JSON line) in `${CBM_CACHE_DIR}/logs/cbm-daemon.log`; that one record is emitted even when `CBM_LOG_LEVEL` suppresses ordinary logging, so the paths always remain discoverable. |
| `CBM_DOWNLOAD_URL` | GitHub releases | Override the update download URL. |
| `CBM_IN_PROCESS` | *(unset)* | Set to `1` (or `true`) to serve MCP sessions read-only from their own process, without the coordination daemon. Only for hosts that cannot reach the daemon — see below. |
| `CBM_LOG_LEVEL` | role-aware | Set the log level to `debug`, `info`, `warn`, `error`, or `none` (or `0`-`4`). Thin MCP/CLI/hook frontends default to `warn`; the detached daemon and supervised index workers default to `info`. Physical workers retain INFO liveness records because their private logs drive the supervisor's no-progress timeout. Frontend messages use that session's stderr; detached daemon events use `${CBM_CACHE_DIR}/logs/cbm-daemon.log`. |
| `CBM_RUNTIME_DIR` | `%LOCALAPPDATA%` (Windows), `/private/tmp` (macOS), `/tmp` (other) | Parent directory for the daemon/CLI rendezvous directory, which CBM creates inside it as `cbm-daemon-<uid>` (`cbm-daemon-<key>` on Windows). Set it when the default ancestry cannot pass the private-directory check — see below. `CBM_CACHE_DIR` does **not** move the rendezvous. |
| `CBM_WORKERS` | auto-detected | Override the indexing worker count. |
Expand Down Expand Up @@ -202,6 +203,47 @@ it will coordinate through the default location instead.

Environment used by daemon-owned components—such as diagnostics, daemon logging, and process-wide indexing resource limits—is captured from the first daemon-backed session that starts the daemon. Later sessions join the existing process and cannot replace those values. To change them, close every daemon-backed session, update the relevant agent configurations consistently, and restart a session. `CBM_ALLOWED_ROOT` remains session-specific, a conflicting `CBM_CACHE_DIR` is rejected, and one-shot CLI commands use their own current environment without starting the daemon.

### Running without the daemon

`CBM_IN_PROCESS=1` is for MCP hosts whose sandbox denies socket system calls, so
the daemon can never be reached — for example a macOS seatbelt profile with
`(deny network*)`, under which even `AF_UNIX` `bind()` and `connect()` fail with
`EPERM`. Only `1` or `true` turn it on. It affects MCP server sessions only;
`cli`, hooks and `daemon` commands ignore it. Set it in the MCP server entry's
`env` rather than exporting it globally: some clients forward only selected
variables to MCP servers, and a global export would make every MCP session
read-only.

Such a session opts out of the shared daemon entirely:

- **Nothing is shared between clients.** Each client that sets it runs its own
standalone server process, which opens the indexes under `CBM_CACHE_DIR`
itself. There are no shared indexing jobs, no watcher or auto-indexing, no
UI, and no cross-session lease. It is also outside exact-build admission:
nothing checks that it runs the same build or uses the same `CBM_CACHE_DIR`
as the daemon, and `update` does not drain it — point it at the cache the
indexing sessions use, and restart it after updating.
- **Read-only.** It serves the `analysis` tool set (or `scout`, if that profile
was requested), so `index_repository`, `delete_project`, `manage_adr` and
`ingest_traces` are not offered. It never builds or refreshes an index: do
that from a daemon-backed session, or with
`codebase-memory-mcp cli index_repository` outside the sandbox.
- **No shared state is written.** It creates no rendezvous, takes no cohort,
lifetime or project lock, and writes neither `_config.db` nor any index page,
so it cannot disturb a daemon running alongside it. The one exception is
SQLite's own reader protocol: on an index in WAL mode, like any reader it
creates `<project>.db-shm` and an empty `<project>.db-wal` if they are absent.

It still needs a `CBM_CACHE_DIR` it can write: like every CBM process it sets
that directory to `0700` at startup, so a read-only mount or a sandbox that
denies writes there stops it before it starts. `search_code` and
`detect_changes` also stage short-lived scratch files in `${CBM_CACHE_DIR}/logs`
(creating it if needed), and `search_code` in `/tmp`.

Leave it unset whenever the daemon is reachable, when a session must index,
re-index, read or edit ADRs, or delete projects, or when you need the UI or
indexes that refresh themselves. It is not a fix for a rendezvous directory that
fails its checks — that is what `CBM_RUNTIME_DIR` is for.

### Roots that are always refused

Expand Down
7 changes: 7 additions & 0 deletions scripts/test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -384,6 +384,13 @@ CBM_TEST_BINARY="$WATCHDOG_BINARY" bash "$ROOT/tests/test_watcher_disabled.sh"
echo "=== Step 5f: worker request-scope regression ==="
CBM_TEST_BINARY="$WATCHDOG_BINARY" bash "$ROOT/tests/test_worker_session_scope.sh"

# Step 5g: CBM_IN_PROCESS (#2072) is a read-only, daemon-free stdio session.
# Next to a live private daemon it must write no shared state: no rendezvous,
# cohort or lifetime lock, and no byte of the daemon's index. Reuses the prod
# binary built in Step 5.
echo "=== Step 5g: CBM_IN_PROCESS writes no shared state (#2072) ==="
CBM_TEST_BINARY="$WATCHDOG_BINARY" bash "$ROOT/tests/test_in_process_mcp.sh"

# Step 6: security-strings URL allow-list regression. The MSYS2 CLANG64 toolchain
# bakes its package-tracker URL into the static Windows .exe; the binary string
# audit must allow-list it (Windows-only — Linux smoke never saw it).
Expand Down
114 changes: 114 additions & 0 deletions src/main.c
Original file line number Diff line number Diff line change
Expand Up @@ -2699,6 +2699,50 @@ static int main_run_daemon_ctl(int argc, char **argv, const cbm_daemon_ipc_endpo
return ui_result;
}

/* In-process sessions are read-only. The analysis/scout tool allowlist set in
* main() keeps every mutating tool out: dispatch_tool() refuses it before any
* handler runs. That allowlist is not a read-only contract (mcp.h: internal
* cache maintenance may still write), so these guards are the second line —
* a write that gets past it fails closed at the lease instead of proceeding
* uncoordinated. They cannot stop work a handler does before it reaches the
* lease, which is why the allowlist comes first.
*
* An unset guard would be no line at all: mcp_project_mutation_begin() is
* `!srv->mutation_begin || srv->mutation_begin(...)`, so an unset guard fails
* OPEN and every mutation proceeds. The daemon (daemon/application.c), the UI
* (ui/http_server.c) and the local CLI above all install one; an in-process
* session has no daemon to acquire a cross-session lease from, so it must
* refuse the mutation instead of taking it unguarded.
*
* Leaving the try-guard unset is deliberate: with mutation_begin set and
* mutation_try_begin NULL, mcp_project_mutation_try_begin() also returns false,
* so opportunistic writes during a read are refused too. */
static bool main_in_process_mutation_refused(void *context, const char *project) {
(void)context;
(void)project;
return false;
}

static void main_in_process_mutation_end(void *context, const char *project) {
(void)context;
(void)project;
}

/* Reject indexing with a message that says why, rather than letting the
* mutation guard report it as "blocked by an active index" — nothing is
* blocking, the mode simply cannot coordinate a write. Mirrors
* http_read_only_index_rejected in ui/http_server.c. */
static char *main_in_process_index_rejected(void *context, const char *repo_path,
const char *args_json) {
(void)context;
(void)repo_path;
(void)args_json;
return cbm_mcp_text_result("indexing is unavailable in CBM_IN_PROCESS mode: it needs the "
"coordination daemon's cross-session lease. Run "
"`codebase-memory-mcp cli index_repository` outside the sandbox.",
true);
}

int main(int argc, char **argv) {
/* Must remain the first statement: see allocator binding contract above. */
cbm_alloc_init();
Expand Down Expand Up @@ -3134,6 +3178,76 @@ int main(int argc, char **argv) {
return result;
}

/* In-process MCP: serve the stdio JSON-RPC loop directly, with no
* coordination daemon and therefore no AF_UNIX rendezvous at all.
*
* This exists for hosts whose sandbox denies socket syscalls outright —
* e.g. a macOS seatbelt profile that is `(allow default)` for the
* filesystem but `(deny network*)`, which covers network-bind and
* network-outbound and so makes both bind() and connect() EPERM. There the
* daemon handshake can never complete: presence is inferred from a file
* lock (permitted), so an EPERM connect is misread as "daemon still
* starting" and the whole MAIN_MCP_STARTUP_TIMEOUT_MS budget is spent
* retrying it before the client gives up.
*
* Opt-in via CBM_IN_PROCESS (1 or true, as for CBM_DIAGNOSTICS) so default
* behaviour is unchanged. The store is resolved from CBM_CACHE_DIR exactly
* as a daemon session does (cbm_mcp_server_new(NULL), as in
* daemon/application.c and ui/http_server.c), so this reads the same index
* the daemon builds.
*
* The session is READ-ONLY, and deliberately so. Without a daemon there is
* no cross-session mutation lease, no index executor and no exact-build
* admission, so a write here could not be coordinated against a daemon
* session touching the same project. It therefore serves only the
* analysis/scout tool allowlist, backed by the refusing guards above main()
* — see there for why an unset guard would have been worse than a refusing
* one. Indexing stays with the daemon, outside the sandbox. */
if (role == CBM_DAEMON_PROCESS_MCP_CLIENT) {
char inproc_buf[MAIN_PATH_CAP];
const char *inproc =
cbm_safe_getenv("CBM_IN_PROCESS", inproc_buf, sizeof(inproc_buf), NULL);
if (inproc && (strcmp(inproc, "1") == 0 || strcmp(inproc, "true") == 0)) {
cbm_mem_init(cbm_mem_ram_fraction_for_total(cbm_system_info().total_ram));
cbm_mcp_server_t *inproc_srv = cbm_mcp_server_new(NULL);
if (!inproc_srv) {
(void)fprintf(stderr, "codebase-memory-mcp: cannot create in-process MCP server\n");
return EXIT_FAILURE;
}
/* Serve only the analysis allowlist (scout stays scout; anything
* else becomes analysis). dispatch_tool() refuses a tool outside it
* before its handler runs, which the guards below cannot do: they
* sit inside the handlers, and handler work before them still
* reaches shared state — index_repository loads the index policy
* first, and with no server config that creates or opens
* ${CBM_CACHE_DIR}/_config.db before the executor could refuse. */
cbm_mcp_server_set_tool_profile(inproc_srv, tool_profile == CBM_MCP_TOOL_PROFILE_SCOUT
? CBM_MCP_TOOL_PROFILE_SCOUT
: CBM_MCP_TOOL_PROFILE_ANALYSIS);
/* The second line (see the guards above main()): no daemon means no
* cross-session lease, index executor or exact-build admission, so a
* write that got past the allowlist must fail closed rather than
* proceed uncoordinated. background_tasks off keeps
* maybe_auto_index() off the initialize path. */
cbm_mcp_server_set_background_tasks(inproc_srv, false);
cbm_mcp_server_set_index_executor(inproc_srv, main_in_process_index_rejected, NULL);
cbm_mcp_server_set_project_mutation_guard(inproc_srv, main_in_process_mutation_refused,
main_in_process_mutation_end, NULL);
char inproc_root[MAIN_PATH_CAP];
char inproc_allowed[MAIN_PATH_CAP];
const char *inproc_allowed_ptr = NULL;
if (main_session_context(NULL, inproc_root, inproc_allowed, &inproc_allowed_ptr)) {
(void)cbm_mcp_server_set_session_context(inproc_srv, inproc_root,
inproc_allowed_ptr);
}
setup_signal_handlers();
cbm_log_info("mcp.in_process", "reason", "CBM_IN_PROCESS", "daemon", "bypassed");
int inproc_rc = cbm_mcp_server_run(inproc_srv, stdin, stdout);
cbm_mcp_server_free(inproc_srv);
return inproc_rc < 0 ? EXIT_FAILURE : EXIT_SUCCESS;
}
}

cbm_daemon_ipc_endpoint_t *endpoint = main_daemon_endpoint_new();
if (!endpoint) {
/* #1582: this is where an ownership/ancestry refusal lands, and it was
Expand Down
Loading
Loading