GH-3696: Cache ParsedVersion in FileMetaData and use it in fromParquetMetadata - #3700
GH-3696: Cache ParsedVersion in FileMetaData and use it in fromParquetMetadata#3700asifsmohammed wants to merge 5 commits into
Conversation
wgtmac
left a comment
There was a problem hiding this comment.
Thanks for fixing this! If this gets checked in, the PR that fixes malformed stats checking can be closed, right?
wgtmac
left a comment
There was a problem hiding this comment.
Thanks for pushing this cleaner direction. I think caching the parsed writer version in FileMetaData is the right foundation, but this PR is not a complete fix yet. It currently adds the cache, but does not migrate the production call sites that still parse created_by repeatedly. Please update the hot/footer/page/reader/rewrite paths to consume the cached ParsedVersion, and cover that behavior in tests.
| FileMetaData meta = | ||
| new FileMetaData(SCHEMA, Collections.emptyMap(), "parquet-mr version 1.12.0 (build abc123)"); | ||
|
|
||
| assertThat(meta.getWriterVersion()).isNotNull(); |
There was a problem hiding this comment.
These tests only cover the getter. They do not prove the repeated parsing problem is fixed. Please add coverage for at least one migrated production path using the cached ParsedVersion instead of reparsing createdBy.
d7c4f6d to
2f01073
Compare
…dant parsing Parse the createdBy version string once during FileMetaData construction and cache the result as a transient field. This avoids redundant VersionParser.parse() calls at every downstream call site (R×C times during footer decode alone).
2f01073 to
bc23cf5
Compare
- Change writerVersion to lazy computation on first getWriterVersion() call - Fixes deserialization correctness (transient fields recompute from createdBy) - Add writerVersionParsed flag to avoid retrying on parse failure - Document contract for distinguishing missing vs. unparseable in javadoc
@wgtmac I'd prefer to keep this PR focused on the caching foundation and migrate callers in a follow-up. The migration touches multiple files which is a larger change that's easier to review separately. The follow-up will include tests proving the production path uses the cached version. |
wgtmac
left a comment
There was a problem hiding this comment.
Thanks for the quick response!
I agree with keeping this PR small and moving the full migration to follow-ups. However, I do think it is worth migrating at least one production caller so this PR provides a concrete fix, not only an unused cache.
Please also narrow the title and description and remove “Closes #3696”, since the remaining paths still parse created_by repeatedly.
…zed init - Replace writerVersion + writerVersionParsed with single WriterVersionResult - Null field means not-yet-initialized, MISSING for null/empty createdBy - Double-checked locking with synchronized for thread-safe one-time init - Rethrow cached VersionParseException so callers preserve existing fallback - Use Strings.isNullOrEmpty for consistency with CorruptStatistics
Add shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName) overload to CorruptStatistics that uses the pre-parsed and cached SemanticVersion from ParsedVersion, eliminating redundant VersionParser.parse and SemanticVersion.parse calls in the R×C hot path. Refactor ParquetMetadataConverter.fromParquetMetadata to construct the hadoop FileMetaData before the row-group loop and extract the cached ParsedVersion once via getWriterVersion(). The loop now uses the ParsedVersion-based buildColumnChunkMetaData overload, avoiding per-column re-parsing. Falls back to the String-based path when getWriterVersion() throws VersionParseException to preserve exact logging parity.
| if (!writerVersion.hasSemanticVersion()) { | ||
| warnOnce("Ignoring statistics because created_by could not be parsed (see PARQUET-251): " + writerVersion); | ||
| return true; | ||
| } | ||
|
|
||
| SemanticVersion semver = writerVersion.getSemanticVersion(); |
There was a problem hiding this comment.
ParsedVersion eagerly parses and caches the SemanticVersion in its constructor, so getSemanticVersion() avoids the redundant SemanticVersion.parse(version.version) that the String-based overload previously performed on every call. The left and right spikes in flame graph are for parsing SemanticVersion twice.
|
Hi @wgtmac, I've addressed your feedback. |
|
Hi @wgtmac is it possible to speed up process to get this and subsequent PRs merged? It will bring significantly cost savings to our company |
wgtmac
left a comment
There was a problem hiding this comment.
Sorry for the delay! Thanks @asifsmohammed for improving this! I still have some comments. Will merge it after all those have been addressed.
| ParsedVersion version = VersionParser.parse(createdBy); | ||
| return shouldIgnoreStatistics(version, columnType); | ||
| } catch (RuntimeException | VersionParseException e) { | ||
| warnParseErrorOnce(createdBy, e); |
There was a problem hiding this comment.
| warnParseErrorOnce(createdBy, e); | |
| // couldn't parse the created_by field, log what went wrong, don't trust the | |
| // stats, but don't make this fatal. | |
| warnParseErrorOnce(createdBy, e); |
Let's keep the original comment.
| boolean isSet = formatStats.isSetMax() && formatStats.isSetMin(); | ||
| boolean maxEqualsMin = isSet ? Arrays.equals(formatStats.getMin(), formatStats.getMax()) : false; | ||
| boolean sortOrdersMatch = SortOrder.SIGNED == typeSortOrder; | ||
| // NOTE: See docs in CorruptStatistics for explanation of why this check is needed |
There was a problem hiding this comment.
Could we preserve these comments?
| * @param columnType the type of the column that this is checking | ||
| * @return true if the statistics may be invalid and should be ignored, false otherwise | ||
| */ | ||
| public static boolean shouldIgnoreStatistics(ParsedVersion writerVersion, PrimitiveTypeName columnType) { |
There was a problem hiding this comment.
This overload makes calls like shouldIgnoreStatistics(null, type) ambiguous; the cast in the new test demonstrates the source incompatibility. Could we use a distinct method name for the ParsedVersion path, including the new converter overloads?
There was a problem hiding this comment.
Added String createdBy as a parameter to the ParsedVersion overload, so this signature (ParsedVersion, String, PrimitiveTypeName) is no longer ambiguous with (String, PrimitiveTypeName). The createdBy parameter also solves the logging parity issue mentioned below comment. Converter overloads follow the same pattern.
| return true; | ||
| } | ||
|
|
||
| if (!writerVersion.hasSemanticVersion()) { |
There was a problem hiding this comment.
ParsedVersion has already swallowed SemanticVersionParseException here, so this no longer preserves the old warnParseErrorOnce(createdBy, e) behavior. Could we keep the original string and parse exception for this path?
There was a problem hiding this comment.
createdBy string is now passed as a parameter, and the !hasSemanticVersion() branch re-parses writerVersion.version to recreate the SemanticVersionParseException for warnParseErrorOnce(createdBy, e). This gives exact log parity (original string + stack trace). The re-parse only fires when the ParsedVersion fails to parse it, so zero performance impact on the hot path.
| String createdBy, Statistics formatStats, PrimitiveType type, SortOrder typeSortOrder) { | ||
| // create stats object based on the column type | ||
| return fromParquetStatisticsInternal( | ||
| CorruptStatistics.shouldIgnoreStatistics(createdBy, type.getPrimitiveTypeName()), |
There was a problem hiding this comment.
This evaluates shouldIgnoreStatistics even when stats are null or V2 min/max is used, so it may log “Ignoring statistics” and consume the one-shot warning when nothing is ignored. Could we keep this check inside the legacy min/max branch and add a regression test?
There was a problem hiding this comment.
Moved shouldIgnoreStatistics evaluation inside the V1 legacy min/max branch, V2 stats now bypass it entirely. Added testV2StatsDoNotTriggerCorruptStatisticsCheck regression test that verifies a corrupt writer version with V2 stats still gets valid min/max without consuming the one-shot warning.
- Add createdBy parameter to shouldIgnoreStatistics(ParsedVersion, ...) to resolve null ambiguity (distinct 3-param signature) and restore exact log parity by using the raw string in warnings. - Re-parse SemanticVersion in the !hasSemanticVersion() branch to recreate the exception for warnParseErrorOnce with stack trace. - Move shouldIgnoreStatistics evaluation inside the V1 legacy min/max branch so V2 stats never trigger the one-shot warning. - Restore original comments in both CorruptStatistics and ParquetMetadataConverter. - Add regression test for V2 stats with corrupt writer version.
Rationale for this change
VersionParser.parse(createdBy)is called from 7 production sites, all parsing the same constant string fromFileMetaData.getCreatedBy(). InfromParquetMetadata, this happens R×C times (once per column per row group) during footer metadata conversion. SinceFileMetaDataisconstructed once per file and already stores the
createdBystring, it is the natural place to parse once and cache the result.This PR caches the parsed version and migrates the first (and hottest) call site —
ParquetMetadataConverter.fromParquetMetadata— to use the cache, eliminating redundantVersionParser.parseandSemanticVersion.parsecalls from the R×C inner loop.Fixes #1 in this issue #3696
What changes are included in this PR?
getWriterVersion()toFileMetaDatawith lazy-init via an immutableWriterVersionResultholder (thread-safe, double-checked locking)shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName)overload toCorruptStatisticsthat uses the cachedSemanticVersionfromParsedVersiondirectlyParquetMetadataConverter.fromParquetMetadatato constructFileMetaDatabefore the row-group loop and use the cachedParsedVersioninbuildColumnChunkMetaDatagetWriterVersion()throwsVersionParseExceptionto preserve exact logging behaviorAre these changes tested?
Yes.
FileMetaDataTest— 6 tests covering valid, null, empty, unparseable version strings, and cachingCorruptStatisticsTest.testParsedVersionOverload— covers all branches of the newParsedVersionoverload including null, non-parquet-mr, empty version, invalid semver, corrupt, and fixed versionsTestParquetMetadataConverter— 70 existing tests pass (validates the refactoredfromParquetMetadatapath)Are there any user-facing changes?
No breaking changes. Adds new public methods:
FileMetaData.getWriterVersion()— returns cachedParsedVersion, throwsVersionParseExceptionfor unparseable stringsCorruptStatistics.shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName)— for callers that already have a parsed versionCloses #3601