fix(guest): measure gpu-attestation after TPM key provisioning - #1410
Merged
Merged
Conversation
The gpu-attestation event commits to nvattest output made with a fresh random nonce, so its digest changes on every boot. It was extended into the runtime register before the app keys were requested, and the TPM key provider seals its seed to that register (SHA-256 PCR14 on GCP, SHA-384 PCR14 on AWS). The PCR14 value at unseal could therefore never match the value at seal, and a GPU CVM with the TPM key provider failed every boot after the first with a TPM policy error (0x99d). Keep the GPU gate (nvattest, policy, NVML checks, ready state) before key provisioning, so an unattested GPU still stops the boot before any key is released, and extend only the gpu-attestation event after the keys are provisioned, between boot-mr-done and key-provider. It still precedes system-ready, so a verifier replaying the event log sees the same payload bound to the quote. Regression from 1fbae7b (#789). Signed-off-by: Kevin Wang <wy721@qq.com>
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Root cause
A GPU CVM using the TPM key provider cannot reboot. Every boot after the first fails with
failed to unseal from TPM … TPM error: 0x0000099d(policy failure).gpu-attestationbefore the keys are requested. Itsevidence_sha256hashes nvattest output made with a fresh random nonce, so PCR14 at unseal can never equal PCR14 at seal.Fix
The GPU gate (nvattest, policy, NVML checks, ready state) still runs and fails closed before key provisioning. Only the
gpu-attestationevent is now extended after the keys are provisioned, betweenboot-mr-doneandkey-provider. PCR14 at the seal/unseal point then contains only deterministic events.Tradeoff:
system-ready, so a verifier that follows the documented rule ("exactly one pre-system-readygpu-attestation,evidence_sha256matches the boot-time bundle") keeps working. The payload schema does not change.mrAggregated(theboot-mr-donecutoff) no longer includes per-boot GPU evidence, so for GPU CVMs it is stable across boots.Rejected alternative: unsealing before the gate. On AWS, PCR8 is only measured after
boot-mr-done, and that option would also split key provisioning by provider.Verification
GCP
us-central1-a,a3-highgpu-1g(1x H100, CC on), TDX, SPOT, key providertpm. Dev mkosi image built rootless from this branch plus #1409 (needed for rootless builds) and #1411 (masks systemd's SRK units; it touches neither PCR 0/2/14 nor dstack's TPM handles).no sealed seed found, generating new seed, and the LUKS data disk was formatted and mounted. A marker file was written to/dstack/persistent.gcloud compute instances reset.unsealed root key seed from TPM (PCR policy: sha256:0,2,14).disk_crypt_key,k256_keyandenv_crypt_keyare identical to boot 1. The same data disk opened and the boot-1 marker is present.app_id,instance_idandcompose_hashare unchanged.boot-mr-doneis the same value (d04276ff…). The full event log replays to Info RTMR3, to RTMR3 in a fresh/v1/Attestquote and to TPM PCR14. There is exactly onegpu-attestationevent, afterboot-mr-done. SHA-256 of the/v1/Attestboot-time GPU evidence bundle equals itsevidence_sha256.0x99dafter reset.cargo clippy -p dstack-util -- -D warningsandcargo test -p dstack-util(114 passed).