The platform's VM provisioner (giantswarm/vm-manager)
in the lab: KVM virtual machines with an instance metadata service, a vTPM
with measured boot, an immutable image and attestation, exposed to muster and
the agents as x_vm-manager_<tool> and to the portal as one more server of
the Agent Platform group. vm-manager is the sibling of agent-manager (agents)
and model-manager (models) — the write surface for VMs — and runs the way they
do: as a pod, the agent-platform chart's components.vm-manager. The kind
node is a privileged docker container, so the host's /dev/kvm and
/dev/vhost-vsock are in it, and the runtime hands them to the privileged pod
(nothing is mounted from the node; the chart has no hostPath); QEMU and swtpm
are the pod's children. The guest image the pod boots is an OCI artifact its
init container fetches into the pod's state volume: the one every vm-manager
release publishes (gsoci.azurecr.io/giantswarm/vm-manager-guest-image:<version>,
the chart default), or a local build the lab pushes into its registry.
platform.vmManager in agentlab.yaml:
platform:
vmManager:
enabled: true
imageDir: /home/you/projects/giantswarm/vm-manager/images/build # optional: a local guest image build (`make -C images`); empty boots the release's
devImages:
vm-manager: vm-manager:dev # optional: a build of the checkout's binary (`make docker-build`)agentlab configurereports whether this machine has the KVM devices (KVM /dev/kvm and /dev/vhost-vsock present — platform.vmManager on, the release's guest imageor… a local guest image build from …). A machine without them turns the key off with the reason; a machine with them keeps what the file says — the pod is a heavier piece of the lab than a model server, so it is never turned on by itself.--vm-manager[=false]pins it,--vm-manager-image-dir <dir>names a local build.agentlab up/agentlab platformrefuse the key on a node without the devices, with the fix. WithimageDirset,platformpushes the build into the lab registry (<cluster>-registry:5000, the one the Harness dev image uses; created when missing) asvm-manager-guest-image:devwithvm-manager image push— the same code the release pipeline runs — and records the manifest digest instate/vm-manager-guest-image.json. A rebuilt image is a new digest: the nextplatformrolls the pod, whose init container replaces the directory. Nothing is fixed atkind createany more.agentlab platformturnscomponents.vm-manageron in the chart's values: the meta chart's ownvm-manager:block brings the pinned Service name, OAuth againstglobal.identity(the lab Dex, private URLs and IPs) and the musterMCPServerwith theagent-platformtool group and forward-token auth; the lab addsguestImage(the lab registry, by digest, plain HTTP) when a local build is configured and a state claim (persistence.create: true, the node's local-path storage) so the fetched image, its golden PCR values and the lab's VM records survive a pod restart. Aplatform.devImages.vm-managerbuild is side-loaded into the node and swapped in like model-manager's (both containers of the pod run it). The run waits until muster reports the registration reachable.agentlab vm-manager-testis the proof (below).
An agentlab before 0.44 registered a vm-manager running on the host
(state/vm-manager.env, a lab-created MCPServer); agentlab platform
removes that registration before the install — the chart renders one of the
same name — and the environment file with it.
A freshly built image's policy.json carries PCR 11 (the UKI) and PCR 13
(the Kubernetes sysext) but no golden values for the firmware PCRs, and
vm-manager's verifier rejects every quote until they exist — so the proof
boots its VM with require_attestation: false and says so. The firmware is
the pod's OVMF (Ubuntu 26.04's, inside the vm-manager image), not the
host's: values recorded for another firmware fail every quote (golden mismatch on PCRs 0 and 7) and the VM never releases its user-data. Record
them once per guest image and vm-manager image, through the lab's identity,
inside the pod — the image directory is the pod's state volume, and the
pod's environment makes vm-manager image golden default to the pod's server
and directory:
# 1. one boot in learn mode: the pod accepts the golden PCRs it has no value for
# (an overlay in platform.valuesFiles, then `agentlab platform`)
cat > learn.yaml <<'YAML'
vm-manager:
vm:
learnGolden: true
YAML
agentlab login admin@lab.local # writes .token
kubectl --kubeconfig state/kubeconfig -n agent-platform port-forward svc/vm-manager 18080:8080 &
curl -s -H "Authorization: Bearer $(cat .token)" -H 'Content-Type: application/json' \
-d '{"name":"golden-bringup","cpus":1,"memory_mib":1024,"disk_gib":8,"require_attestation":true,"wait_for":"ready"}' \
http://127.0.0.1:18080/api/v1/vms | jq -r .id
# 2. write the VM's verified ready-stage quote into the image's policy.json, inside the pod
kubectl --kubeconfig state/kubeconfig -n agent-platform exec deploy/vm-manager -- \
vm-manager image golden giantswarm-vm-base_0.1.0 --from-vm <id> --token "$(cat .token)"
# 3. delete the VM, drop the overlay, `agentlab platform` (the pod restarts and reads the policy): every later boot is compared
curl -s -X DELETE -H "Authorization: Bearer $(cat .token)" http://127.0.0.1:18080/api/v1/vms/<id>From then on list_images reports the golden values, agentlab vm-manager-test boots its VM with require_attestation: true and reads both
quotes verified. The values live in the pod's state claim: they survive pod
restarts, and a new guest image (another digest — a rebuilt local build, or
a vm-manager release with another image) replaces the directory, so image golden again; a new vm-manager image with another OVMF build changes PCRs 0
and 2-4, the same. To keep them for a local build, copy the pod's
policy.json back into the checkout's images/build (kubectl cp from
/var/lib/vm-manager/images/policy.json) before the next push: it travels
inside the artifact.
agentlab vm-manager-test [email] [--skip-vm] [--vm-timeout 6m]
- The identity boundary, against the pod's Service from inside the
cluster (a probe pod, the way muster dials it).
GET /api/v1/hostwithout a token answers 401; with the person's Dex id_token — the very token muster forwards — it answers the capability report (ready,missing, the kernel, CPUs, memory). A vm-manager that answers anonymously (OAuth off in the release's values), or refuses the lab's token (another issuer, CA or audience), fails here with the fix. - Through muster as the person. The
x_vm-manager_*tools are aggregated — the core set of nineteen — with vm-manager's annotations intact (get_hostread-only,delete_vmdestructive: what a model reads before it calls);get_hostnames the same host as the direct call;list_imagesandlist_networksanswer (an image without golden PCR values is reported as such). - The registration. The chart's
MCPServercarries the agent-platform tool group and reads Connected. - The lifecycle (unless
--skip-vm, and only on a host that reportsreadywith an image):create_vmon the newest image (1 CPU, 1 GiB, an 8 GiB disk,wait_for: none), the states followed withget_vm— the installer boot, the installed boot,READY=1over vsock — untilready, with the install and boot durations;get_vm_attestationwith both stages verified when the image's policy carries golden values (untilvm-manager image goldenrecorded them the VM boots withrequire_attestation: false, and the proof says so); the last console lines;delete_vm, andlist_vmswithout it. A failed boot carries vm-manager'slastErrorand the console tail. A leftover of an interrupted run (agentlab-vm-test-*) is removed first.
The lab's platform-test keeps its tool-group check: the chart's servers
(agent-manager, model-manager, vm-manager) carry the platform group from their
templates; nothing lab-created may claim it (the fake fleet carries
infrastructure, the OAuth fixture no label).
Through the lab muster as an MCP server (https://muster.127.0.0.1.nip.io/mcp):
x_vm-manager_get_host first, then x_vm-manager_list_images and
x_vm-manager_list_networks, then x_vm-manager_create_vm — the server's
own instructions say the order. In the portal, the MCP servers page lists
vm-manager under Agent Platform; an agent with a toolset that admits it can
provision VMs as the person who asked. agentlab logs vm-manager follows the
pod (--verbose is on in the lab's values for a dev image).
The pod's VMs end with the pod: a agentlab platform that rolls the
Deployment (a new dev image, a new guest image digest, changed values) stops
every VM; their records and disks stay on the state claim for the next
start_vm. What is not in the lab: a network policy around the pod (the
lab runs without policies).