From 110460d0ea5e990c32c99d1564cd312610f352b9 Mon Sep 17 00:00:00 2001 From: Kevin Wang Date: Thu, 24 Sep 2026 21:57:19 -0700 Subject: [PATCH] fix(guest): measure gpu-attestation after TPM key provisioning The gpu-attestation event commits to nvattest output made with a fresh random nonce, so its digest changes on every boot. It was extended into the runtime register before the app keys were requested, and the TPM key provider seals its seed to that register (SHA-256 PCR14 on GCP, SHA-384 PCR14 on AWS). The PCR14 value at unseal could therefore never match the value at seal, and a GPU CVM with the TPM key provider failed every boot after the first with a TPM policy error (0x99d). Keep the GPU gate (nvattest, policy, NVML checks, ready state) before key provisioning, so an unattested GPU still stops the boot before any key is released, and extend only the gpu-attestation event after the keys are provisioned, between boot-mr-done and key-provider. It still precedes system-ready, so a verifier replaying the event log sees the same payload bound to the quote. Regression from 1fbae7b25c (#789). Signed-off-by: Kevin Wang --- CHANGELOG.md | 1 + docs/attestation-tdx.md | 4 +-- ...s-attested-instance-security-evaluation.md | 2 +- docs/security/security-model.md | 6 +++- docs/tutorials/attestation-verification.md | 7 +++-- dstack/dstack-util/src/system_setup.rs | 30 ++++++++++++++----- 6 files changed, 35 insertions(+), 15 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c312372d4..6dd7b55fc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -39,6 +39,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - http-client: a caller can bound the response body (`http_request_bounded`, `PrpcClient::with_max_response_bytes`). Nothing is bounded by default — `dstack vmm logs --lines 100000` is a legitimate multi-megabyte fetch — but every client that talks to a guest agent opts in, in the gateway and in the VMM, because a CVM is untrusted and one of them polls on a timer against the whole fleet ### Fixed +- dstack-util: a GPU CVM using the TPM key provider can reboot again. The `gpu-attestation` event hashes `nvattest` output made with a fresh nonce on every boot, and it was extended into PCR14 before the TPM unsealed the seed sealed to PCR14, so every boot after the first failed with a TPM policy error. The GPU gate still runs before the keys are requested; only the event is now measured after key provisioning, after `boot-mr-done` and before `key-provider`. - verifier/kms/gateway: issuer certificates and CRLs named by a GCP TPM AK certificate are fetched only from an allowlist of hosts, including across redirects. Those URLs come from a certificate checked only after the fetch, so an unauthenticated `/verify` caller could make the verifier issue requests to any host, including internal ones. The default allows `privateca-content-*.storage.googleapis.com`, where Google's Private CA publishes them; `attestation.allowed_collateral_hosts` replaces it, and `*` matches within one DNS label - sdk: Python and JavaScript blockchain adapters now reject TLS-key responses. The deprecated conversion path read fixed PKCS#8 framing as private-key bytes, causing different TLS keys to derive the same Ethereum and Solana wallets. Use `get_key()` / `getKey()` for wallet keys. - data disks: discard now propagates through ZFS or ext4, dm-crypt, virtio-blk, and QEMU so encrypted qcow2 images release deleted blocks instead of growing with lifetime writes. Discard defaults on and can be disabled with `storage_discard: false` when allocation-pattern leakage is unacceptable; upgrading an existing ZFS pool also starts a one-time trim for historical free space diff --git a/docs/attestation-tdx.md b/docs/attestation-tdx.md index f9744e4ba..de1be2578 100644 --- a/docs/attestation-tdx.md +++ b/docs/attestation-tdx.md @@ -39,8 +39,8 @@ multiple parties and let each party verify its code independently without reconstructing the complete compose document. For a GPU launch, any `init-script-hash` events are followed by -`gpu-policy-hash` and, after successful NVIDIA attestation and policy -evaluation, `gpu-attestation`. The `gpu-policy-hash` payload is +`gpu-policy-hash`. After successful NVIDIA attestation and policy evaluation, +`gpu-attestation` follows `boot-mr-done` and precedes `key-provider`. The `gpu-policy-hash` payload is `SHA-256(JCS(requirements.gpu_policy))`, using `{}` when the field is omitted. The `gpu-attestation` payload is JSON containing the verified device count, CC/DevTools state, aggregate signed-claim `dbgstat` and `secboot`, and diff --git a/docs/aws-attested-instance-security-evaluation.md b/docs/aws-attested-instance-security-evaluation.md index 3575da342..7dbbb3841 100644 --- a/docs/aws-attested-instance-security-evaluation.md +++ b/docs/aws-attested-instance-security-evaluation.md @@ -35,7 +35,7 @@ NitroTPM path meets them today. | P1 | Verifiable platform root of trust | TDX/SNP/Nitro quote verification against vendor root; debug rejected; TCB surfaced | NitroTPM Attestation Documents are verified against the AWS Nitro Attestation PKI, including document timestamp sanity. AWS exposes no TDX/SNP-style TCB advisory field, so a verified attestation is normalized to `tcbStatus = "UpToDate"` and passes the standard authorization gate unchanged. | | P2 | Reproducible or independently computable base image measurement | meta-dstack rebuild plus `dstack-mr` computes `MRTD`/`RTMR0-2` | The unified `os/build.sh` flow emits the AWS image archive with `sha256sum.txt`, `digest.txt`, and `measurement.aws.cbor`; its output directory also contains the reference-PCR side-car `aws-pcrs.json`. `os_image_hash = sha256(sha256sum.txt)` is the same identity used on all platforms; the verifier recomputes it from the downloaded image directory (`dstack/verifier/src/verification.rs`). The hardening audit script `os/yocto/tools/aws/audit-aws-ec2-image-hardening.sh` checks the image for operator mutation channels. | | P3 | Boot command line and root filesystem integrity are measured | `RTMR1/2`, rootfs hash, dm-verity, measured initrd/cmdline | The UKI commits kernel, initrd, and embedded cmdline into `PCR4`; the rootfs is dm-verity-protected. `VmConfig.aws_measurement` is required and must bind `boot_pcr_digest = sha256(PCR4||PCR7||PCR12)` to the attested PCRs, so `PCR12` (external cmdline) is always part of the bound digest — a missing-PCR12 bypass is not expressible. Enforced in guest quote generation (`dstack/dstack-attest/src/attestation.rs`), `verify_os_image_hash_for_aws_nitro_tpm` (`dstack/verifier/src/verification.rs`), and the KMS pipeline via the same verifier check. | -| P4 | Runtime application identity is cryptographically bound | RTMR3 `compose-hash`, `app-id`, `instance-id`, `key-provider`; event log replay | SHA384 `PCR14` event-log replay is the authoritative binding (RTMR3 analogue; non-resettable). Launch events: `system-preparing`, `app-id`, `compose-hash`, zero or more ordered `init-script-hash` events, `instance-id`, `boot-mr-done`, `key-provider`, `storage-fs`, `storage-encrypted`, `system-ready` (`dstack/dstack-util/src/system_setup.rs`). GPU launches also include `gpu-policy-hash` and `gpu-attestation` before `instance-id`. `dstack-attest`, `dstack-verifier`, and KMS reject missing/mismatched PCR14 and bad replay. Optionally, the guest extends the raw `MrConfig` V2 `config_id` into `PCR8` once (`PCR8 = sha384(0^48 || config_id)`) so a lightweight third-party verifier can check compose hash + key provider without event-log replay; dstack's own verifier and KMS do not check PCR8. | +| P4 | Runtime application identity is cryptographically bound | RTMR3 `compose-hash`, `app-id`, `instance-id`, `key-provider`; event log replay | SHA384 `PCR14` event-log replay is the authoritative binding (RTMR3 analogue; non-resettable). Launch events: `system-preparing`, `app-id`, `compose-hash`, zero or more ordered `init-script-hash` events, `instance-id`, `boot-mr-done`, `key-provider`, `storage-fs`, `storage-encrypted`, `system-ready` (`dstack/dstack-util/src/system_setup.rs`). GPU launches also include `gpu-policy-hash` before `instance-id`, and `gpu-attestation` after `boot-mr-done` and before `key-provider`. `dstack-attest`, `dstack-verifier`, and KMS reject missing/mismatched PCR14 and bad replay. Optionally, the guest extends the raw `MrConfig` V2 `config_id` into `PCR8` once (`PCR8 = sha384(0^48 || config_id)`) so a lightweight third-party verifier can check compose hash + key provider without event-log replay; dstack's own verifier and KMS do not check PCR8. | | P5 | Challenge/liveness and caller key binding | `report_data` challenge or RA-TLS public key hash in quote | RA-TLS binds `report_data` to the TLS certificate public key. KMS key release is bound to the live RA-TLS handshake; external `/verify` callers supply and check their own `report_data` challenge. The low-level NitroTPM document verifier also rejects stale or far-future document timestamps. | | P6 | Secret release only to attested code | dstack KMS verifies attestation, checks auth policy, derives per-app keys | dstack KMS verifies the NitroTPM attestation, runs the same `verify_os_image_hash_for_aws_nitro_tpm` binding check as the verifier, builds `BootInfo` from verified boot PCRs plus PCR14 launch events, and checks auth policy before deriving app keys (`dstack/kms/src/main_service.rs`). AWS NitroTPM key release is gated behind the opt-in `aws_nitro_tpm_key_release` flag (default false in `kms.toml`). | | P7 | Key-release policy is not controlled by the untrusted account admin | KMS runs inside TEE; policy from auth API/contracts; KMS identity measured | Satisfied with dstack KMS or another verifiable secret authority outside the untrusted AWS account admin's control. A NitroTPM-backed dstack KMS keeps root material out of account-admin snapshots and clones. The policy backend (auth-simple in a trusted control plane, or on-chain `DstackKms`/`DstackApp`) must be outside the workload account admin's control. Same-account AWS KMS fails this property if the admin can change key policy, create grants, or call secret-bearing operations through a policy they control. | diff --git a/docs/security/security-model.md b/docs/security/security-model.md index 04e92b18a..67af2f070 100644 --- a/docs/security/security-model.md +++ b/docs/security/security-model.md @@ -126,11 +126,15 @@ For a successful TDX GPU launch, the GPU-relevant RTMR3 event order is: compose-hash init-script-hash (zero or more, in configured order) gpu-policy-hash -gpu-attestation instance-id boot-mr-done +os-image-hash (KMS key provider only) +gpu-attestation +key-provider ``` +The GPU gate itself runs before the app keys are requested, and a failed gate stops the boot before any key is released. Only the `gpu-attestation` event is extended after key provisioning. Its `evidence_sha256` covers `nvattest` output made with a fresh nonce on every boot, and the TPM key provider seals its seed to the runtime register (PCR14 on GCP and AWS). Extending a per-boot value before the unseal would make the sealed seed unrecoverable after a reboot. + `gpu-policy-hash` is emitted even for a GPU-less launch. `gpu-attestation` is emitted only after an attached GPU passes `nvattest`, the built-in checks, the optional Rego policy, and the NVML state checks. Its UTF-8 JSON payload has this shape: ```json diff --git a/docs/tutorials/attestation-verification.md b/docs/tutorials/attestation-verification.md index 1cc77c745..9f7c4cfd9 100644 --- a/docs/tutorials/attestation-verification.md +++ b/docs/tutorials/attestation-verification.md @@ -486,18 +486,19 @@ These are the standard events you'll see in the log: | `compose-hash` | SHA-256 of docker compose config | Should match `tcb_info.compose_hash` | | `init-script-hash` | SHA-256 of one init script; repeated in configured order (maximum 5) | Should match the independently approved script bytes | | `gpu-policy-hash` | SHA-256 of the JCS-canonicalized GPU policy (default `{}`) | Should match the expected `requirements.gpu_policy` digest | -| `gpu-attestation` | Verified GPU state, aggregate `dbgstat`/`secboot`, and digest of the boot-time `nvattest` JSON | Required for an attested GPU launch; verify as described below | | `instance-id` | Unique instance identifier | Should match `instance_id` from response | | `boot-mr-done` | Boot measurements complete | Marker event | | `os-image-hash` | Guest OS image hash | Should match `tcb_info.os_image_hash` | +| `gpu-attestation` | Verified GPU state, aggregate `dbgstat`/`secboot`, and digest of the boot-time `nvattest` JSON | Required for an attested GPU launch; verify as described below | | `key-provider` | Key provider type | e.g., `kms` | | `storage-fs` | Storage filesystem type | Storage configuration | | `storage-encrypted` | `1` when the data disk is LUKS-encrypted, `0` when it is not | `0` means the host can read the application's data at rest | | `system-ready` | System ready marker | Always present at end | For a successful GPU launch, the relevant order is `compose-hash`, any -`init-script-hash` events, `gpu-policy-hash`, `gpu-attestation`, `instance-id`, -and `boot-mr-done`. +`init-script-hash` events, `gpu-policy-hash`, `instance-id`, `boot-mr-done`, +then `gpu-attestation` after the app keys are provisioned and before +`key-provider`. After replaying the log to the quote's RTMR3, decode the JSON payload of `gpu-attestation` and compare its `evidence_sha256` with the SHA-256 digest of the boot-time GPU evidence bytes. Fetch those bytes by calling `/v1/Attest` diff --git a/dstack/dstack-util/src/system_setup.rs b/dstack/dstack-util/src/system_setup.rs index b1f277652..2aa9df286 100644 --- a/dstack/dstack-util/src/system_setup.rs +++ b/dstack/dstack-util/src/system_setup.rs @@ -2098,7 +2098,11 @@ impl Stage0<'_> { /// explicitly disabled — set the GPU ready state without verification. The /// optional Rego policy is always evaluated; when no attestation is /// performed, its claims-array input is empty. - async fn measure_gpu(&self) -> Result<[u8; 32]> { + /// + /// Returns the policy digest and, when a GPU was attested, the + /// `gpu-attestation` event payload. The caller measures that payload only + /// after key provisioning: see [`Stage0::setup_fs`]. + async fn measure_gpu(&self) -> Result<([u8; 32], Option>)> { let gpu_policy_hash = gpu::measure_gpu_policy(&self.shared.dir.app_compose_file())?; let gpu_policy = self @@ -2118,7 +2122,7 @@ impl Stage0<'_> { info!("application GPU Rego policy accepted an empty claims array"); } if inventory.nvidia == 0 { - return Ok(gpu_policy_hash); + return Ok((gpu_policy_hash, None)); } warn!( "requirements.gpu_policy.attest_gpu is false; setting GPU ready state without attestation" @@ -2127,7 +2131,7 @@ impl Stage0<'_> { if let Err(err) = gpu::set_gpu_ready_state(inventory.nvidia) { warn!("failed to set GPU ready state: {err:?}"); } - return Ok(gpu_policy_hash); + return Ok((gpu_policy_hash, None)); } let expected_devices = gpu::nvidia_gpu_count(inventory)?; if expected_devices == 0 { @@ -2135,7 +2139,7 @@ impl Stage0<'_> { if gpu_policy.rego.is_some() { info!("application GPU Rego policy accepted an empty claims array"); } - return Ok(gpu_policy_hash); + return Ok((gpu_policy_hash, None)); } self.vmm.notify_q("boot.progress", "attesting GPU").await; info!("verifying GPU TEE attestation"); @@ -2158,10 +2162,8 @@ impl Stage0<'_> { gpu_state.set_ready()?; let devtools = gpu_state.any_devtools(); let event = attestation.event(devtools)?; - emit_runtime_event("gpu-attestation", &event) - .context("failed to emit GPU attestation event")?; info!("GPU TEE attestation succeeded"); - Ok(gpu_policy_hash) + Ok((gpu_policy_hash, Some(event))) } } @@ -2261,6 +2263,8 @@ struct AppInfo { compose_hash: [u8; 32], gpu_policy_hash: [u8; 32], init_script_hashes: Vec>, + /// `gpu-attestation` payload, measured after key provisioning. + gpu_attestation_event: Option>, } struct Stage0<'a> { @@ -2978,7 +2982,7 @@ impl<'a> Stage0<'a> { for script_hash in &init_script_hashes { emit_runtime_event("init-script-hash", script_hash)?; } - let gpu_policy_hash = self + let (gpu_policy_hash, gpu_attestation_event) = self .measure_gpu() .await .context("failed to verify GPU TEE attestation")?; @@ -3012,6 +3016,7 @@ impl<'a> Stage0<'a> { compose_hash, gpu_policy_hash, init_script_hashes, + gpu_attestation_event, }) } @@ -3066,6 +3071,15 @@ impl<'a> Stage0<'a> { if app_keys.disk_crypt_key.is_empty() { bail!("Failed to get valid key phrase from KMS"); } + // The GPU gate already passed before the keys were requested; only its + // event is measured here. The payload hashes nvattest output made with + // a fresh nonce, so it differs on every boot. Extending it before the + // TPM key provider unseals would change PCR14, which the seed is sealed + // to, and the CVM could never unseal its keys again after a reboot. + if let Some(event) = &app_info.gpu_attestation_event { + emit_runtime_event("gpu-attestation", event) + .context("failed to emit GPU attestation event")?; + } self.verify_app(&app_info, &app_keys) .context("Failed to verify app")?;