Skip to content

ci: publish nightly SNAPSHOT jars to repository.apache.org - #5902

Open
andygrove wants to merge 6 commits into
apache:mainfrom
andygrove:publish-snapshot
Open

ci: publish nightly SNAPSHOT jars to repository.apache.org#5902
andygrove wants to merge 6 commits into
apache:mainfrom
andygrove:publish-snapshot

Conversation

@andygrove

@andygrove andygrove commented Sep 13, 2026

Copy link
Copy Markdown
Member

Which issue does this PR close?

Closes #5899.

Rationale for this change

Anyone who wants to try a fix or feature before the next release currently has to build Comet from source, including the Rust native library. Publishing nightly SNAPSHOT jars to the ASF snapshot repository gives users and downstream projects a ready-made artifact for the current development version, the same way Spark and Iceberg publish their snapshots.

The repository previously had a tag-triggered GHCR Docker publish workflow, removed in #4241. Its runs had failed for months, at first within a minute of starting and later as startup_failure from unpinned third-party actions, and it built arm64 under QEMU with caching disabled. This workflow avoids all three problems: it only uses actions/* and the repo's own composite actions, builds each architecture on a native runner, and caches the Cargo registry and Maven repository.

What changes are included in this PR?

A new scheduled workflow, .github/workflows/publish_snapshot.yml, with three jobs:

  • changes skips the nightly run when main has had no commits since the previous night, so Nexus does not accumulate identical snapshots. A manual dispatch always runs.
  • native builds libcomet.so for linux/amd64 on ubuntu-24.04 and linux/aarch64 on ubuntu-24.04-arm. Both build inside an ubuntu:20.04 container so the library links against glibc 2.31, the same baseline as the release builder in dev/release/comet-rm/Dockerfile, and a step fails the job if the library ever requires a newer glibc. The same make core-*-libs targets as the release are used, so the baseline CPU targets match released jars.
  • deploy places both libraries under spark/target/classes/org/apache/comet/linux/ the way build-release-comet.sh does, then runs ./mvnw deploy for the four default variants: Spark 3.4 and 3.5 with Scala 2.12, Spark 4.0 and 4.1 with Scala 2.13. Every variant is built with JDK 17, matching pr_build_linux.yml. -Dmaven.deploy.skip=false overrides the root pom so the parent pom is deployed too, since consumers need it to resolve the child poms and the release publishes it. Each variant is packaged and checked for both bundled native libraries before the publishing goal runs for it, so an incomplete jar fails the job instead of reaching consumers.

Credentials come from the NEXUS_USER and NEXUS_PW repository secrets that ASF Infra provisions for snapshot publishing, read into a settings.xml through ${env.*} rather than written to disk. The org.apache:apache parent pom already maps SNAPSHOT deploys to apache.snapshots.https, so no pom changes are needed. If the secrets are not yet configured on this repository, the deploy step fails at upload and an INFRA ticket is needed.

A dry_run input on workflow_dispatch runs the whole pipeline but ends with install instead of deploy and uploads the jars as workflow artifacts. It is also allowed on forks so the workflow can be exercised before a change lands.

The installation guide's snapshot-only section now explains where the snapshots live, which artifacts exist, that they are unreleased builds for testing only, and how to use one with spark-shell either by downloading the jar or via --packages with the snapshot repository. The workflows README lists the new standalone workflow.

How are these changes tested?

  • actionlint, prettier --check and dev/ci/check-ci-config.py pass locally.
  • A local ./mvnw deploy -Dmaven.deploy.skip=false -Pspark-4.1 -DaltDeploymentRepository=... into a file repository confirmed that the parent pom, comet-common, comet-spark and comet-spark-integration are all deployed with timestamped snapshot names.
  • A dry_run of the workflow on my fork built both native libraries in the Ubuntu 20.04 containers, passed the glibc check, built all four variants and verified that each jar bundles both libraries before publishing it: https://github.com/andygrove/datafusion-comet/actions/runs/34873305385

Cost per night from that run: 20 minutes for the amd64 native build, 17 for aarch64, and 18 minutes for the deploy job's four Maven builds, so roughly 55 runner-minutes and 39 minutes of wall clock. Packaging each variant before publishing it accounts for about 2 minutes of the deploy job; the second Maven invocation reuses the target directory the check ran against.

The first real publish will be a manual workflow_dispatch after this merges, followed by checking that the coordinates resolve from the snapshot repository.

@github-actions github-actions Bot added build Build environment enhancement New feature or request labels Sep 13, 2026
@andygrove

Copy link
Copy Markdown
Member Author

cc @kazuyukitanimura

Comment thread .github/workflows/publish_snapshot.yml Outdated
fi
for jar in jars/*.jar; do
for lib in linux/amd64 linux/aarch64; do
if ! unzip -l "$jar" | grep -q "org/apache/comet/$lib/libcomet.so"; then

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This can report a missing library when the entry is present. I reran this step against the four jars from your dry run in an Ubuntu 24.04 amd64 container: all eight checks failed with PIPESTATUS=141 0. grep -q exits after its match, leaving unzip to receive SIGPIPE, which pipefail treats as failure. The hosted run passed; this depends on pipe/consumer scheduling.

Letting grep consume the full listing passes all four jars and still rejects a copy with the aarch64 library removed:

Suggested change
if ! unzip -l "$jar" | grep -q "org/apache/comet/$lib/libcomet.so"; then
if ! unzip -l "$jar" | grep -F "org/apache/comet/$lib/libcomet.so" > /dev/null; then

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, and thanks for running it down. Reproduced locally against one of those dry-run jars: the old form failed 100 out of 100 runs with PIPESTATUS=141 0, so the hosted run passing was luck. Took your suggestion — grep now consumes the whole listing, 0 failures out of 100, and it still returns 1 for a library that is genuinely absent.

Comment thread .github/workflows/publish_snapshot.yml Outdated
# property overrides every module's setting. $profiles is a
# space-separated list of -P flags, so it must word-split.
# shellcheck disable=SC2086
JAVA_HOME="$jdk" ./mvnw -B "$GOAL" -DskipTests -Dmaven.deploy.skip=false $profiles < /dev/null

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we validate the packaged jars before deploying them? With GOAL=deploy, this call uploads each snapshot before the native-library check below runs. I confirmed the upload order by deploying the Spark 4.1 reactor to a local file repository.

If that check catches an incomplete jar, the job fails after the jar is already available to consumers. Please move validation ahead of publication so the deployed artifacts have passed the check.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You are right, the check was useless where it was. The loop now packages a variant, checks that jar, and only then runs the publishing goal for it, so nothing reaches Nexus before it has passed. The extra Maven invocation reuses the target directory the check ran against, and the parent pom pins project.build.outputTimestamp, so it re-creates the same jar — identical sha256 locally, and it cost 20s on top of a 75s package. I also dropped the separate count assertion in favour of failing when the jar glob does not match exactly one file.

One thing I decided not to do: variants are still published one at a time, so if the fourth fails its check the first three are already up. Making it all-or-nothing means packaging all four, then rebuilding each from clean to deploy it, and I did not think ~15 extra minutes a night was worth it for a snapshot that gets overwritten tomorrow. Happy to change it if you disagree.

Comment thread .github/workflows/publish_snapshot.yml Outdated
echo "::endgroup::"
done <<EOF
$JAVA_HOME_11_X64 -Pspark-3.4 -Pscala-2.12
$JAVA_HOME_17_X64 -Pspark-3.5 -Pscala-2.12

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just making sure, are AWS Labs folks are using Scala 2.12 for benchmarking?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just making sure, are AWS Labs folks are using Scala 2.12 for benchmarking?

Yes, according to https://awslabs.github.io/data-on-eks/docs/benchmarks/spark-gluten-velox-comet-benchmark

Add a scheduled workflow that builds the Comet native library for
linux/amd64 and linux/aarch64 inside an Ubuntu 20.04 container, matching
the glibc baseline of the release builder, and deploys SNAPSHOT jars for
Spark 3.4, 3.5, 4.0 and 4.1 to the ASF snapshot repository. The run
skips when main has not changed since the previous night, and a dry_run
dispatch builds and verifies the jars without touching Nexus so the
workflow can be exercised on a fork.

Document how to consume the snapshots in the installation guide.
apache#5897 dropped JDK 11 support: the Spark 3.4 profile now targets Java 17 and a
Maven enforcer rule rejects anything older, so the Spark 3.4 build in this
workflow would fail on the JDK 11 it was pinned to.
The native-library check ran after the whole build loop, so with GOAL=deploy
every jar was already uploaded by the time it ran. Package each variant, check
it, and only then run the publishing goal for that variant. The second Maven
invocation reuses the same target directory, and the parent pom pins
project.build.outputTimestamp, so it re-creates a byte-identical jar; locally
it took 20 seconds against a 75-second packaging run.

The check itself could also report a missing library for a complete jar:
`unzip -l | grep -q` leaves unzip killed by SIGPIPE once grep exits on the
match, and `pipefail` turns that into a failure. Letting grep read the whole
listing fixes it. Against a real 13 MB jar the old form failed 100 times out of
100 locally with PIPESTATUS=141 0, the new form none, and the new form still
reports a library that is genuinely absent.

Also guard the jar glob so a build that produces no plugin jar, or more than
one, fails with a clear message instead of publishing something unchecked. That
subsumes the separate count assertion.
@andygrove

Copy link
Copy Markdown
Member Author

Rebased on main and pushed the review fixes. Since @kazuyukitanimura approved, one thing changed beyond the review comments: #5897 dropped JDK 11, so the Spark 3.4 build in this workflow would now hit the enforcer rule. Every variant is built with JDK 17 and the second JDK install is gone.

@andygrove

Copy link
Copy Markdown
Member Author

@rich7420 I addressed feedback - could you take another look?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

build Build environment enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Publish nightly SNAPSHOT jars to repository.apache.org

3 participants