perf: specialize primitive sums for run-end arrays - #9823
Performance Regression: -11.87%
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 3 improved benchmarks
❌ 2 regressed benchmarks
✅ 2188 untouched benchmarks
🆕 35 new benchmarks
⏩ 218 skipped benchmarks1
Warning
Please fix the performance issues or acknowledge them on CodSpeed.
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | dict_canonicalize_gt_u8_avx512[16000000] |
6.8 ms | 11.3 ms | -39.91% |
| ❌ | WallTime | arrow_checked_add_u32_neon[16384] |
12.2 µs | 20.3 µs | -39.87% |
| ⚡ | WallTime | dict_canonicalize_gt_u8_neon[1000000] |
559.8 µs | 487.3 µs | +14.89% |
| ⚡ | Simulation | allocate_drop_arrow[0] |
456.9 ns | 402.7 ns | +13.45% |
| ⚡ | WallTime | dict_canonicalize_gt_u8_neon[16000000] |
9.4 ms | 8.3 ms | +12.92% |
| 🆕 | Simulation | grouped_runend_fallback[1, 1024] |
N/A | 738.5 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[1, 4] |
N/A | 830.8 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[1, 64] |
N/A | 763.8 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[128, 1024] |
N/A | 149.8 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[128, 4] |
N/A | 217.6 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[128, 64] |
N/A | 147.3 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[2, 1024] |
N/A | 443.9 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[2, 4] |
N/A | 518.4 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[2, 64] |
N/A | 454.6 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[8, 1024] |
N/A | 220.6 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[8, 4] |
N/A | 287.9 µs | N/A |
| 🆕 | Simulation | grouped_runend_fallback[8, 64] |
N/A | 222.2 µs | N/A |
| 🆕 | Simulation | grouped_runend[1, 1024] |
N/A | 455.4 µs | N/A |
| 🆕 | Simulation | grouped_runend[1, 4] |
N/A | 482.3 µs | N/A |
| 🆕 | Simulation | grouped_runend[1, 64] |
N/A | 461.4 µs | N/A |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ct/sum-runend (31e3f62) with develop (e3b8eb2)
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩