[GEODE-10646] Reduce up-front buffer allocation in the Lucene file output stream - #8075
Conversation
sboorlagadda
left a comment
There was a problem hiding this comment.
The chunking behavior looks preserved, and the existing large-write test plus
the new round-trip test provide useful coverage. Before approving the memory
optimization, could you add a test that detects reintroducing the eager 1 MiB
allocation and share a base/head allocation or GC comparison for small and
larger outputs? A fresh stream reaching 1 MiB now allocates 2040 KiB of buffer
arrays and recopies 1016 KiB during growth, although it saves substantially for
small files. If this is intended to address the integration-test OOM, please
also link evidence connecting that failure to these buffers; the existing
ticket discusses ClassGraph and ByteBuffersDirectory, a separate storage path.
|
Thanks for the review, @sboorlagadda. Added FileOutputStreamJUnitTest.testSmallFileAllocatesLessThanChunkSize, which measures thread allocation while writing a 100-byte file. It fails against develop's FileOutputStream (1,055,576 bytes allocated) and passes on this branch. Base/head allocation per stream in bytes (JDK 17, ThreadMXBean, best of 7, 1000-byte writes, open through close): This is consistent with your numbers: outputs reaching 1 MiB allocate about 1 MiB more per stream, from buffer growth. This PR isn't intended to address the integration-test OOM. To verify the effect with many streams open at once, I opened 452 streams on a map-backed FileSystem and wrote 2 KiB to each, using a 768 MiB heap. On develop each stream reserves a full 1 MiB buffer at open, and the 452 streams did not fit in the heap. On this branch the same streams retained about 6 MB. |
|
The stream now buffers each chunk in 8 KiB segments that are allocated only when needed and reused for later chunks. Each chunk is copied once into an exactly sized array, so the buffer never grows by copying. Allocation per stream in bytes (JDK 17, ThreadMXBean, best of 21 runs across 3 JVMs, 1000-byte writes, open through close):
From 1 MiB up, allocation is now within 2.5 KB of develop. Small and medium files keep the savings. I added |
Reduce up-front buffer allocation in the Lucene file output stream
For all changes, please confirm:
develop)?gradlew buildrun cleanly?