cdc: massively speed replay with ordered concurrency and batching - #9
Draft
thesyncim wants to merge 10 commits into
Draft
cdc: massively speed replay with ordered concurrency and batching#9thesyncim wants to merge 10 commits into
thesyncim wants to merge 10 commits into
Conversation
Keep trigger suppression on the dedicated connection to remove a target round trip from every replayed transaction. Co-authored-by: Cursor <cursoragent@cursor.com>
Reuse validated target mappings until pgoutput reports a changed source definition, avoiding repeated catalog queries on hot tables. Co-authored-by: Cursor <cursoragent@cursor.com>
Queue source-ordered DML in bounded windows and validate every result before transactional progress or commit. Co-authored-by: Cursor <cursoragent@cursor.com>
Validate stream identity, mark first progress, and upsert the durable LSN in one guarded statement. Co-authored-by: Cursor <cursoragent@cursor.com>
Reuse exact SQL and parameter type parses across source transactions while bounding long-lived session state with protocol-level LRU eviction. Co-authored-by: Cursor <cursoragent@cursor.com>
Fold BEGIN into replay and guard progress plus COMMIT in one final window so mismatches abort server-side before data can commit. Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Reproduced deadlock and fix
The 2,048-transaction / 16-table reproduction filled the 512-transaction window and hung. The goroutine dump showed all workers idle and the scheduler unable to read: a later same-table job had been linked to a predecessor that had already emitted its completion event, so that dependency could never be released.
The scheduler now treats already-durable table tails as satisfied, workers never wait inside an open target transaction for coordinator permission, and WAL decoding runs through a one-item bounded read pump. Apply sessions also have a 2-minute
idle_in_transaction_session_timeoutfuse so any future stall rolls back and reconnects instead of pinning transactions indefinitely.Sustained performance
-race: 877 ms / 11.4k events/s / 6.6×; CI enforces at least 3×-race: 305 ms / 32.7k events/s / 16.2×Durability
synchronous_commitremains enabled for CDC. Each parallel target transaction atomically commits DML with one receipt range for its contiguous source batch. A coordinator atomically advances authoritative progress and removes only the contiguous receipt prefix. Therefore:The forced recovery test commits a later two-transaction batch first, leaves canonical progress absent, restarts with one worker, and verifies exact-once convergence plus receipt cleanup.
Validation
go test ./...make vetmake race-racemake e2ewith batch 128make crash-e2e(kills and resumes copy, indexes, catchup, and follow; exact final hashes)