Skip to content

Release idle shared-storage topics - #4016

Open
T4rk1n wants to merge 1 commit into
devfrom
fix/4010-free-shared-storage-topics
Open

T4rk1n wants to merge 1 commit into
devfrom
fix/4010-free-shared-storage-topics

Conversation

@T4rk1n

@T4rk1n T4rk1n commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Fixes #4010.

Shared-storage topics were never removed. Streaming uses one topic per page load, so every page that streamed kept its last 32 frames for the life of the process (and forever in Redis).

Topic lifetime. LocalSharedStorage, DiskcacheSharedStorage and RedisSharedStorage take topic_ttl (default 300s, None disables, <= 0 rejected). A topic nobody publishes to or reads for that long is dropped, buffer and sequence both.

  • Local: the engine counts calls holding each topic and sweeps idle, unheld topics at most every quarter ttl. A blocked poll keeps its topic, and a publish never lands in a topic that was just dropped.
  • Redis: the publish script PEXPIREs the counter and stream; each poll renews both (pipelined with the existing head/floor reads).
  • Diskcache: messages expire topic_ttl after publish, the counter after the last publish or poll.

Per-run downlink id. A released topic restarts at seq 1, but the browser kept its cursor for the whole page load. A page that idled past the ttl and streamed again would resume from its old cursor, get {reset: true} and fail the new stream. The renderer now picks a fresh downlinkId (and restarts its cursor at 0) each time it pins a connection from idle, and sends it on every stream request. get_stream_connection_id returns <end_id>:<downlinkId>, so each run has its own topic. The id only partitions the page's own signed space, so it is not signed, just validated (403 if malformed). The downlink lifecycle record stays keyed on the page. This also fixes a SharedWorker switching to another tab's endId with a stale cursor.

Proof. The issue's repro (20 page loads, ~50 KB frames), with a 2s ttl: 0 topics and 0 frames after idle, versus 20 topics / 640 frames before. test_stst012_idle_stream_topics_are_released drives this in a browser and fails with either half reverted (without the per-run id, the page's next run after idle never renders).

Contributor Checklist

  • I have run the tests locally and they passed (tests/shared_storage, tests/streaming incl. Redis, renderer karma suite)
  • I have added tests, or extended existing tests, to cover any new features or bugs fixed in this PR

optionals

Topics were never removed, so every page load that streamed kept its
last 32 frames for the life of the process (or forever in Redis).

All three backends now take topic_ttl (default 300s, None disables): a
topic nobody publishes to or reads for that long is dropped. The local
engine sweeps idle topics not held by any call; Redis PEXPIREs the
counter and stream on publish and poll; diskcache expires messages and
touches the counter.

A released topic restarts at seq 1, which would reset a page that
streams again after idling with its old cursor. The renderer now picks
a fresh downlinkId per run of streams and restarts its cursor at 0, so
each run reads its own topic (<end_id>:<downlinkId>). The downlink
lifecycle record stays keyed on the page.
@sonarqubecloud

Copy link
Copy Markdown

@github-actions

Copy link
Copy Markdown
Contributor

Dash performance benchmarks

✅ all within thresholds

scenario metric p90 (ms) median growth baseline p90 note
✅ callback_chain chain_ms 447.0 433.2 0.91x 499.1
✅ callback_chain graph_ms 3.2 3.2 1.0x 2.4
✅ callback_fanout fanout_ms 81.5 78.9 0.92x 92.5
✅ deep_nesting render_ms 55.6 53.1 1.02x 56.8
✅ full_children_replace replace_ms 5254.9 1973.2 17.8x 4697.1
✅ initial_render_large render_ms 595.1 570.5 1.0x 694.4
✅ initial_render_small render_ms 101.0 88.6 0.95x 104.0
✅ patch_append_nested append_ms 147.6 97.6 2.59x 192.3
✅ patch_append_toplevel append_ms 112.9 76.7 2.26x 140.2
✅ patch_scalar_update_large update_ms 156.1 144.6 1.05x 202.9
✅ wildcard_all_resolve wildcard_ms 297.3 283.9 0.96x 313.6
✅ wildcard_all_resolve graph_ms 1.1 1.1 1.0x 1.3

growth = late-third / early-third per-op time; ~1 is flat, a large value means the per-op cost scales with accumulated state.

machine scale vs baseline: 0.92x - divided out of the baseline ratios so they compare like for like (the absolute warn/fail ceilings are left un-scaled); calibrated on initial_render_small.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Shared-storage topics are never freed: streaming leaks ~32 frames per page load

1 participant