diff --git a/CHANGELOG.md b/CHANGELOG.md index c312372d4..6dd7b55fc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -39,6 +39,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - http-client: a caller can bound the response body (`http_request_bounded`, `PrpcClient::with_max_response_bytes`). Nothing is bounded by default — `dstack vmm logs --lines 100000` is a legitimate multi-megabyte fetch — but every client that talks to a guest agent opts in, in the gateway and in the VMM, because a CVM is untrusted and one of them polls on a timer against the whole fleet ### Fixed +- dstack-util: a GPU CVM using the TPM key provider can reboot again. The `gpu-attestation` event hashes `nvattest` output made with a fresh nonce on every boot, and it was extended into PCR14 before the TPM unsealed the seed sealed to PCR14, so every boot after the first failed with a TPM policy error. The GPU gate still runs before the keys are requested; only the event is now measured after key provisioning, after `boot-mr-done` and before `key-provider`. - verifier/kms/gateway: issuer certificates and CRLs named by a GCP TPM AK certificate are fetched only from an allowlist of hosts, including across redirects. Those URLs come from a certificate checked only after the fetch, so an unauthenticated `/verify` caller could make the verifier issue requests to any host, including internal ones. The default allows `privateca-content-*.storage.googleapis.com`, where Google's Private CA publishes them; `attestation.allowed_collateral_hosts` replaces it, and `*` matches within one DNS label - sdk: Python and JavaScript blockchain adapters now reject TLS-key responses. The deprecated conversion path read fixed PKCS#8 framing as private-key bytes, causing different TLS keys to derive the same Ethereum and Solana wallets. Use `get_key()` / `getKey()` for wallet keys. - data disks: discard now propagates through ZFS or ext4, dm-crypt, virtio-blk, and QEMU so encrypted qcow2 images release deleted blocks instead of growing with lifetime writes. Discard defaults on and can be disabled with `storage_discard: false` when allocation-pattern leakage is unacceptable; upgrading an existing ZFS pool also starts a one-time trim for historical free space diff --git a/docs/attestation-tdx.md b/docs/attestation-tdx.md index f9744e4ba..de1be2578 100644 --- a/docs/attestation-tdx.md +++ b/docs/attestation-tdx.md @@ -39,8 +39,8 @@ multiple parties and let each party verify its code independently without reconstructing the complete compose document. For a GPU launch, any `init-script-hash` events are followed by -`gpu-policy-hash` and, after successful NVIDIA attestation and policy -evaluation, `gpu-attestation`. The `gpu-policy-hash` payload is +`gpu-policy-hash`. After successful NVIDIA attestation and policy evaluation, +`gpu-attestation` follows `boot-mr-done` and precedes `key-provider`. The `gpu-policy-hash` payload is `SHA-256(JCS(requirements.gpu_policy))`, using `{}` when the field is omitted. The `gpu-attestation` payload is JSON containing the verified device count, CC/DevTools state, aggregate signed-claim `dbgstat` and `secboot`, and diff --git a/docs/aws-attested-instance-security-evaluation.md b/docs/aws-attested-instance-security-evaluation.md index 3575da342..7dbbb3841 100644 --- a/docs/aws-attested-instance-security-evaluation.md +++ b/docs/aws-attested-instance-security-evaluation.md @@ -35,7 +35,7 @@ NitroTPM path meets them today. | P1 | Verifiable platform root of trust | TDX/SNP/Nitro quote verification against vendor root; debug rejected; TCB surfaced | NitroTPM Attestation Documents are verified against the AWS Nitro Attestation PKI, including document timestamp sanity. AWS exposes no TDX/SNP-style TCB advisory field, so a verified attestation is normalized to `tcbStatus = "UpToDate"` and passes the standard authorization gate unchanged. | | P2 | Reproducible or independently computable base image measurement | meta-dstack rebuild plus `dstack-mr` computes `MRTD`/`RTMR0-2` | The unified `os/build.sh` flow emits the AWS image archive with `sha256sum.txt`, `digest.txt`, and `measurement.aws.cbor`; its output directory also contains the reference-PCR side-car `aws-pcrs.json`. `os_image_hash = sha256(sha256sum.txt)` is the same identity used on all platforms; the verifier recomputes it from the downloaded image directory (`dstack/verifier/src/verification.rs`). The hardening audit script `os/yocto/tools/aws/audit-aws-ec2-image-hardening.sh` checks the image for operator mutation channels. | | P3 | Boot command line and root filesystem integrity are measured | `RTMR1/2`, rootfs hash, dm-verity, measured initrd/cmdline | The UKI commits kernel, initrd, and embedded cmdline into `PCR4`; the rootfs is dm-verity-protected. `VmConfig.aws_measurement` is required and must bind `boot_pcr_digest = sha256(PCR4||PCR7||PCR12)` to the attested PCRs, so `PCR12` (external cmdline) is always part of the bound digest — a missing-PCR12 bypass is not expressible. Enforced in guest quote generation (`dstack/dstack-attest/src/attestation.rs`), `verify_os_image_hash_for_aws_nitro_tpm` (`dstack/verifier/src/verification.rs`), and the KMS pipeline via the same verifier check. | -| P4 | Runtime application identity is cryptographically bound | RTMR3 `compose-hash`, `app-id`, `instance-id`, `key-provider`; event log replay | SHA384 `PCR14` event-log replay is the authoritative binding (RTMR3 analogue; non-resettable). Launch events: `system-preparing`, `app-id`, `compose-hash`, zero or more ordered `init-script-hash` events, `instance-id`, `boot-mr-done`, `key-provider`, `storage-fs`, `storage-encrypted`, `system-ready` (`dstack/dstack-util/src/system_setup.rs`). GPU launches also include `gpu-policy-hash` and `gpu-attestation` before `instance-id`. `dstack-attest`, `dstack-verifier`, and KMS reject missing/mismatched PCR14 and bad replay. Optionally, the guest extends the raw `MrConfig` V2 `config_id` into `PCR8` once (`PCR8 = sha384(0^48 || config_id)`) so a lightweight third-party verifier can check compose hash + key provider without event-log replay; dstack's own verifier and KMS do not check PCR8. | +| P4 | Runtime application identity is cryptographically bound | RTMR3 `compose-hash`, `app-id`, `instance-id`, `key-provider`; event log replay | SHA384 `PCR14` event-log replay is the authoritative binding (RTMR3 analogue; non-resettable). Launch events: `system-preparing`, `app-id`, `compose-hash`, zero or more ordered `init-script-hash` events, `instance-id`, `boot-mr-done`, `key-provider`, `storage-fs`, `storage-encrypted`, `system-ready` (`dstack/dstack-util/src/system_setup.rs`). GPU launches also include `gpu-policy-hash` before `instance-id`, and `gpu-attestation` after `boot-mr-done` and before `key-provider`. `dstack-attest`, `dstack-verifier`, and KMS reject missing/mismatched PCR14 and bad replay. Optionally, the guest extends the raw `MrConfig` V2 `config_id` into `PCR8` once (`PCR8 = sha384(0^48 || config_id)`) so a lightweight third-party verifier can check compose hash + key provider without event-log replay; dstack's own verifier and KMS do not check PCR8. | | P5 | Challenge/liveness and caller key binding | `report_data` challenge or RA-TLS public key hash in quote | RA-TLS binds `report_data` to the TLS certificate public key. KMS key release is bound to the live RA-TLS handshake; external `/verify` callers supply and check their own `report_data` challenge. The low-level NitroTPM document verifier also rejects stale or far-future document timestamps. | | P6 | Secret release only to attested code | dstack KMS verifies attestation, checks auth policy, derives per-app keys | dstack KMS verifies the NitroTPM attestation, runs the same `verify_os_image_hash_for_aws_nitro_tpm` binding check as the verifier, builds `BootInfo` from verified boot PCRs plus PCR14 launch events, and checks auth policy before deriving app keys (`dstack/kms/src/main_service.rs`). AWS NitroTPM key release is gated behind the opt-in `aws_nitro_tpm_key_release` flag (default false in `kms.toml`). | | P7 | Key-release policy is not controlled by the untrusted account admin | KMS runs inside TEE; policy from auth API/contracts; KMS identity measured | Satisfied with dstack KMS or another verifiable secret authority outside the untrusted AWS account admin's control. A NitroTPM-backed dstack KMS keeps root material out of account-admin snapshots and clones. The policy backend (auth-simple in a trusted control plane, or on-chain `DstackKms`/`DstackApp`) must be outside the workload account admin's control. Same-account AWS KMS fails this property if the admin can change key policy, create grants, or call secret-bearing operations through a policy they control. | diff --git a/docs/security/security-model.md b/docs/security/security-model.md index 04e92b18a..67af2f070 100644 --- a/docs/security/security-model.md +++ b/docs/security/security-model.md @@ -126,11 +126,15 @@ For a successful TDX GPU launch, the GPU-relevant RTMR3 event order is: compose-hash init-script-hash (zero or more, in configured order) gpu-policy-hash -gpu-attestation instance-id boot-mr-done +os-image-hash (KMS key provider only) +gpu-attestation +key-provider ``` +The GPU gate itself runs before the app keys are requested, and a failed gate stops the boot before any key is released. Only the `gpu-attestation` event is extended after key provisioning. Its `evidence_sha256` covers `nvattest` output made with a fresh nonce on every boot, and the TPM key provider seals its seed to the runtime register (PCR14 on GCP and AWS). Extending a per-boot value before the unseal would make the sealed seed unrecoverable after a reboot. + `gpu-policy-hash` is emitted even for a GPU-less launch. `gpu-attestation` is emitted only after an attached GPU passes `nvattest`, the built-in checks, the optional Rego policy, and the NVML state checks. Its UTF-8 JSON payload has this shape: ```json diff --git a/docs/tutorials/attestation-verification.md b/docs/tutorials/attestation-verification.md index 1cc77c745..9f7c4cfd9 100644 --- a/docs/tutorials/attestation-verification.md +++ b/docs/tutorials/attestation-verification.md @@ -486,18 +486,19 @@ These are the standard events you'll see in the log: | `compose-hash` | SHA-256 of docker compose config | Should match `tcb_info.compose_hash` | | `init-script-hash` | SHA-256 of one init script; repeated in configured order (maximum 5) | Should match the independently approved script bytes | | `gpu-policy-hash` | SHA-256 of the JCS-canonicalized GPU policy (default `{}`) | Should match the expected `requirements.gpu_policy` digest | -| `gpu-attestation` | Verified GPU state, aggregate `dbgstat`/`secboot`, and digest of the boot-time `nvattest` JSON | Required for an attested GPU launch; verify as described below | | `instance-id` | Unique instance identifier | Should match `instance_id` from response | | `boot-mr-done` | Boot measurements complete | Marker event | | `os-image-hash` | Guest OS image hash | Should match `tcb_info.os_image_hash` | +| `gpu-attestation` | Verified GPU state, aggregate `dbgstat`/`secboot`, and digest of the boot-time `nvattest` JSON | Required for an attested GPU launch; verify as described below | | `key-provider` | Key provider type | e.g., `kms` | | `storage-fs` | Storage filesystem type | Storage configuration | | `storage-encrypted` | `1` when the data disk is LUKS-encrypted, `0` when it is not | `0` means the host can read the application's data at rest | | `system-ready` | System ready marker | Always present at end | For a successful GPU launch, the relevant order is `compose-hash`, any -`init-script-hash` events, `gpu-policy-hash`, `gpu-attestation`, `instance-id`, -and `boot-mr-done`. +`init-script-hash` events, `gpu-policy-hash`, `instance-id`, `boot-mr-done`, +then `gpu-attestation` after the app keys are provisioned and before +`key-provider`. After replaying the log to the quote's RTMR3, decode the JSON payload of `gpu-attestation` and compare its `evidence_sha256` with the SHA-256 digest of the boot-time GPU evidence bytes. Fetch those bytes by calling `/v1/Attest` diff --git a/dstack/dstack-util/src/system_setup.rs b/dstack/dstack-util/src/system_setup.rs index b1f277652..2aa9df286 100644 --- a/dstack/dstack-util/src/system_setup.rs +++ b/dstack/dstack-util/src/system_setup.rs @@ -2098,7 +2098,11 @@ impl Stage0<'_> { /// explicitly disabled — set the GPU ready state without verification. The /// optional Rego policy is always evaluated; when no attestation is /// performed, its claims-array input is empty. - async fn measure_gpu(&self) -> Result<[u8; 32]> { + /// + /// Returns the policy digest and, when a GPU was attested, the + /// `gpu-attestation` event payload. The caller measures that payload only + /// after key provisioning: see [`Stage0::setup_fs`]. + async fn measure_gpu(&self) -> Result<([u8; 32], Option>)> { let gpu_policy_hash = gpu::measure_gpu_policy(&self.shared.dir.app_compose_file())?; let gpu_policy = self @@ -2118,7 +2122,7 @@ impl Stage0<'_> { info!("application GPU Rego policy accepted an empty claims array"); } if inventory.nvidia == 0 { - return Ok(gpu_policy_hash); + return Ok((gpu_policy_hash, None)); } warn!( "requirements.gpu_policy.attest_gpu is false; setting GPU ready state without attestation" @@ -2127,7 +2131,7 @@ impl Stage0<'_> { if let Err(err) = gpu::set_gpu_ready_state(inventory.nvidia) { warn!("failed to set GPU ready state: {err:?}"); } - return Ok(gpu_policy_hash); + return Ok((gpu_policy_hash, None)); } let expected_devices = gpu::nvidia_gpu_count(inventory)?; if expected_devices == 0 { @@ -2135,7 +2139,7 @@ impl Stage0<'_> { if gpu_policy.rego.is_some() { info!("application GPU Rego policy accepted an empty claims array"); } - return Ok(gpu_policy_hash); + return Ok((gpu_policy_hash, None)); } self.vmm.notify_q("boot.progress", "attesting GPU").await; info!("verifying GPU TEE attestation"); @@ -2158,10 +2162,8 @@ impl Stage0<'_> { gpu_state.set_ready()?; let devtools = gpu_state.any_devtools(); let event = attestation.event(devtools)?; - emit_runtime_event("gpu-attestation", &event) - .context("failed to emit GPU attestation event")?; info!("GPU TEE attestation succeeded"); - Ok(gpu_policy_hash) + Ok((gpu_policy_hash, Some(event))) } } @@ -2261,6 +2263,8 @@ struct AppInfo { compose_hash: [u8; 32], gpu_policy_hash: [u8; 32], init_script_hashes: Vec>, + /// `gpu-attestation` payload, measured after key provisioning. + gpu_attestation_event: Option>, } struct Stage0<'a> { @@ -2978,7 +2982,7 @@ impl<'a> Stage0<'a> { for script_hash in &init_script_hashes { emit_runtime_event("init-script-hash", script_hash)?; } - let gpu_policy_hash = self + let (gpu_policy_hash, gpu_attestation_event) = self .measure_gpu() .await .context("failed to verify GPU TEE attestation")?; @@ -3012,6 +3016,7 @@ impl<'a> Stage0<'a> { compose_hash, gpu_policy_hash, init_script_hashes, + gpu_attestation_event, }) } @@ -3066,6 +3071,15 @@ impl<'a> Stage0<'a> { if app_keys.disk_crypt_key.is_empty() { bail!("Failed to get valid key phrase from KMS"); } + // The GPU gate already passed before the keys were requested; only its + // event is measured here. The payload hashes nvattest output made with + // a fresh nonce, so it differs on every boot. Extending it before the + // TPM key provider unseals would change PCR14, which the seed is sealed + // to, and the CVM could never unseal its keys again after a reboot. + if let Some(event) = &app_info.gpu_attestation_event { + emit_runtime_event("gpu-attestation", event) + .context("failed to emit GPU attestation event")?; + } self.verify_app(&app_info, &app_keys) .context("Failed to verify app")?;