perf: batch fragmented CASE branch execution - #9820
Conversation
…ution Signed-off-by: Nicholas Gates <nick@nickgates.com>
Merging this PR will regress 19 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | case_when_nary_10_conditions[1000] |
285 µs | 404.6 µs | -29.55% |
| ❌ | Simulation | random_i8[0.5] |
67.9 µs | 93.3 µs | -27.25% |
| ❌ | Simulation | take_filter_primitive_nullable_slice_mask_random_indices[4096, 1000] |
100.9 µs | 131 µs | -22.99% |
| ❌ | Simulation | case_when_nary_equality_lookup[1000] |
246.4 µs | 317.3 µs | -22.34% |
| ❌ | Simulation | filter_list_primitive_short[Sparse] |
132.2 µs | 168.9 µs | -21.74% |
| ❌ | Simulation | filter_list_primitive_wide[Sparse] |
133.3 µs | 169.2 µs | -21.23% |
| ❌ | Simulation | take_filter_primitive_nullable_slice_mask_random_indices[16384, 1000] |
119.5 µs | 150.9 µs | -20.85% |
| ❌ | Simulation | filter_list_primitive_short[Clustered] |
134.2 µs | 168.6 µs | -20.4% |
| ❌ | Simulation | case_when_nary_3_conditions[1000] |
174.9 µs | 219.2 µs | -20.23% |
| ❌ | Simulation | case_when_nary_early_dominant[1000] |
174.8 µs | 216.1 µs | -19.1% |
| ❌ | Simulation | filter_list_primitive_short[Prefix] |
137.8 µs | 169.8 µs | -18.87% |
| ❌ | Simulation | filter_list_primitive_wide[Clustered] |
137.6 µs | 169.4 µs | -18.79% |
| ❌ | Simulation | case_when_nary_equality_lookup[10000] |
390.2 µs | 480 µs | -18.71% |
| ❌ | Simulation | filter_list_primitive_wide[Prefix] |
138 µs | 169.3 µs | -18.48% |
| ❌ | Simulation | filter_powerlaw_by_random[10000] |
32.3 µs | 37.9 µs | -14.66% |
| ❌ | Simulation | take_filter_primitive_slice_mask_sequential_indices[16384, 1000] |
53.7 µs | 61.5 µs | -12.69% |
| ❌ | Simulation | filter_ultra_sparse[250000] |
74.3 µs | 84.7 µs | -12.3% |
| ❌ | WallTime | mul_u32_nonnull_avx512 |
5.6 µs | 6.3 µs | -12.13% |
| ❌ | Simulation | case_when_without_else[10000] |
211.7 µs | 237.8 µs | -10.99% |
| ⚡ | Simulation | density_sweep_single_slice[0.9999] |
84.2 µs | 27.3 µs | ×3.1 |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ngates/batched-case-execution (2ae4de6) with develop (9d1b103)
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Polar Signals Profiling ResultsLatest Run
Powered by Polar Signals Cloud |
Benchmarks: FineWeb NVMe 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.024x ➖, 0↑ 1↓)
datafusion / parquet / ns (0.999x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.047x ➖, 0↑ 1↓)
duckdb / parquet / ns (1.013x ➖, 0↑ 1↓)
No file size changes detected. |
Benchmarks: TPC-H SF=10 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.988x ➖, 1↑ 0↓)
datafusion / parquet / ns (1.004x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.020x ➖, 0↑ 1↓)
duckdb / parquet / ns (1.005x ➖, 0↑ 0↓)
No file size changes detected. |
Benchmarks: Clickbench Sorted on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.003x ➖, 0↑ 0↓)
datafusion / parquet / ns (0.994x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.007x ➖, 0↑ 0↓)
duckdb / parquet / ns (0.999x ➖, 0↑ 0↓)
File Size Changes (100 files changed, -0.0% overall, 44↑ 56↓)
Totals:
|
Benchmarks: Statistical and Population Genetics 📖Commits: PR How to read Verdict and Engines
duckdb / vortex-file-compressed / ns (1.212x ❌, 0↑ 2↓)
duckdb / parquet / ns (1.009x ➖, 0↑ 0↓)
No file size changes detected. |
Benchmarks: Clickbench on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.999x ➖, 0↑ 0↓)
datafusion / parquet / ns (0.996x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.975x ➖, 2↑ 1↓)
duckdb / parquet / ns (0.983x ➖, 0↑ 1↓)
No file size changes detected. |
Benchmarks: FineWeb S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.942x ➖, 1↑ 0↓)
datafusion / parquet / ns (1.009x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.200x ➖, 0↑ 2↓)
duckdb / parquet / ns (1.001x ➖, 0↑ 0↓)
|
Benchmarks: TPC-H SF=1 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.003x ➖, 0↑ 0↓)
datafusion / parquet / ns (1.004x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.978x ➖, 2↑ 0↓)
duckdb / parquet / ns (1.002x ➖, 0↑ 0↓)
No file size changes detected. |
Benchmarks: PolarSignals Profiling 📖Commits: PR datafusion / vortex-file-compressed / ns (1.016x ➖, 0↑ 1↓)
No file size changes detected. |
Benchmarks: TPC-H SF=1 on S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.046x ➖, 1↑ 3↓)
datafusion / parquet / ns (1.057x ➖, 0↑ 2↓)
duckdb / vortex-file-compressed / ns (1.024x ➖, 0↑ 0↓)
duckdb / parquet / ns (1.041x ➖, 0↑ 1↓)
|
Benchmarks: TPC-DS SF=1 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.000x ➖, 0↑ 1↓)
datafusion / parquet / ns (1.003x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.001x ➖, 1↑ 3↓)
duckdb / parquet / ns (1.008x ➖, 1↑ 6↓)
No file size changes detected. |
| Mask::AllTrue(_) => Ok(Some(array.child().clone())), | ||
| Mask::AllFalse(_) => Ok(Some(Canonical::empty(array.dtype()).into_array())), | ||
| Mask::Values(_) => Ok(None), | ||
| mask @ Mask::Values(_) => { |
There was a problem hiding this comment.
can you make this a separate pr, also while you're at it fix MaskValues::last to use BitBuffer::last_set_index
There was a problem hiding this comment.
actually you can leverage contiguous_values_range here
|
|
||
| #[derive(Debug)] | ||
| struct ScalarFnUnaryFilterPushDownRule; | ||
| struct ScalarFnFilterPushDownRule; |
There was a problem hiding this comment.
can this be a separate pr?
|
|
||
| return Ok(Some(new_array)); | ||
| } | ||
| .map(|c| match c.as_opt::<Constant>() { |
There was a problem hiding this comment.
change this to call prepare_mask_for_reuse if nchildren > 1
Why
Fragmented CASE selections can repeatedly execute lazy branch expressions for tiny runs or individual rows. Batch those selections so branch evaluation stays columnar, while keeping the existing low-overhead approach for coarse runs.
Changes
Validation
cargo nextest run -p vortex-array: 3,453 passed, one skipped (including 55 CASE tests).cargo clippy -p vortex-array --all-targets --all-features -- -D warnings: passed.cargo +nightly fmt --allandgit diff --check: passed.cargo bench -p vortex-array --bench expr_case_when -- --sample-count 30: completed, including simple, n-ary, early-exit, fragmented, and coarse-run controls.On this revision, 65,536-row alternating decimal-product and string cases have medians of 165.5 µs and 157.1 µs. These are candidate-only measurements, not a fresh comparison against develop. Earlier paired experiments motivated the change, but also found small n-ary regressions; this is not a claim of uniform improvement.
Review considerations
The 128-rows-per-run heuristic is empirical. Compact assembly adds temporary output/permutation storage; no memory-saving claim is made. The shared scalar-function filter rule broadens selection pushdown, so its contract and deeply nested expression behavior deserve particular review.
Ported and validated with Codex. No dependency changes or unrelated Vortex experiments are included.