Skip to content

fix(pipeline): spill before the post-extraction phases when they would cross the budget (#2184) - #2394

Open
DeusData wants to merge 1 commit into
mainfrom
fix/issue-2184-post-extraction-budget
Open

DeusData wants to merge 1 commit into
mainfrom
fix/issue-2184-post-extraction-budget

Conversation

@DeusData

Copy link
Copy Markdown
Owner

Extraction spills parsed results under the memory budget, but it can finish just below the spill point. The post-extraction phases (resolve, then semantic) then grow past the budget with nothing checking them. This is how #2184's openclaw run went over budget at lower worker counts.

This PR projects the post-extraction working set when extraction finishes. If resolve would cross the budget, results are spilled (parked) before resolve starts, not after. On a generous budget the decision is keep and nothing changes.

Test: parallel_post_extract_projection_spills_before_resolve fails before the fix (parked == -1, expected 8), passes after it, and fails again when it is reverted. parallel_spill_mode_builds_the_same_graph and the parallel (75), pipeline (288), mem (72), semantic (35), incremental (163) and complexity (5) suites pass. make -f Makefile.cbm lint-ci passes.

Proof (release build, 8 workers; peak = charged memory, which is what the budget is measured against):

run before after
Go, 2000 MB budget resolve 2660 MB (33 % over), no spill spill, 21,882 results parked; every phase under budget (semantic peak 1746 MB)
Go, 40000 MB 27.4 s keep, no spill, 27.5 s
django, 400 MB resolve 455 MB resolve 277 MB

Graph dumps are byte-identical before and after in every run (Go: 335,191 nodes / 1,813,857 edges).

Known remainder, not in this PR: on django at 400 MB the run still peaks at 476 MB inside semantic_edges. That is the graph floor plus about 198 MB of live semantic vectors and LSH, which batching cannot move. Making that retained set spillable is a separate follow-up.

Refs #2184

…d cross the budget (#2184)

Spill mode was entered only DURING extraction, when the charge crossed
budget - budget/16, and the final sweep at extraction end ran only if that
latch had already fired. A run that ended extraction just under the latch
kept every result in memory, and the phases that cannot spill -- registry
build, cross-LSP prepare (all_defs, per-language registries, surface rows,
module index) and resolve (edges into the graph buffer, cross-LSP appends
into the cached results) -- then grew the process well past the budget.
openclaw at 8 workers with a 4079 MB budget: 3678 MB charged at extraction
end, 4076 after registry build, 4933 after cross-LSP prepare, 5829 in
resolve (+43 %). The same repo at 18 workers latched during extraction and
stayed under budget.

Extraction end now projects the post-extraction growth from the result
counts (defs, calls + usages, files) and enters spill mode when charged +
growth would cross the latch, so the existing final sweep parks every
cached result before registry build. The per-unit costs (1280 B/def,
160 B/ref, 8 KiB/file) are fitted on the measured charge growth from the
parallel_extract to the parallel_resolve mark on five local corpora
(go, django, kotlin, rust, php) and cover openclaw; the model under-reads by
at most 7 % (inside the budget/16 margin) and over-reads by up to 46 %, the
safe side: a spill never changes graph content. The decision is a pure
function of counts (no RSS-window sampling) and is logged as
mem.post_extract.projection.

A charge test seam (cbm_mem_set_charged_for_tests) pins cbm_mem_charged(),
which otherwise reads the process footprint, so the "ended just under the
latch" case is testable deterministically.

Proof (release build, M5 Pro, CBM_WORKERS=8, peak charged = the budget
metric's high-water mark):
  go, CBM_MEM_BUDGET_MB=2000   before: extract 1348 -> registry 1497 ->
    lsp_prepare 2168 -> resolve 2660 MB (+33 % over budget, no spill)
    after: projection charged 1335 + growth 1909 > line 1875 -> spill,
    21,882 results parked; extract 413 -> registry 558 -> lsp_prepare 1151
    -> resolve 1297, semantic 1746 MB peak: every phase under budget.
    Node/edge dumps byte-identical (335,191 nodes / 1,813,857 edges); wall
    29.2 s vs 29.0 s.
  go, 40000 MB (generous): decision=keep, no spill, identical dumps, wall
    27.4 s vs 27.5 s.
  django, 400 MB: resolve mark 455 -> 277 MB after the spill, dumps
    identical; the remaining peak (476 MB) is the semantic pass's retained
    set on top of the graph floor, which this change does not address.

Refs #2184

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
@DeusData DeusData mentioned this pull request Sep 28, 2026
2 tasks

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant