Skip to content

kerf: Run on arm64 hosts and publish an arm64 wheel - #20

Open
congwang-mk wants to merge 5 commits into
mainfrom
arm64
Open

congwang-mk wants to merge 5 commits into
mainfrom
arm64

Conversation

@congwang-mk

Copy link
Copy Markdown
Contributor

kerf on main stops at "Could not read APIC IDs" on an arm64 host, since it reads physical CPU IDs from /proc/cpuinfo. This makes kerf work on arm64 and starts publishing an arm64 wheel along with it.

Changes

  • MPIDR topology: on arm64 a CPU's physical ID is its MPIDR affinity value. kerf reads it from the CPU's device tree node on a DT host, and from the MADT GICC entries (matched through the CPU's processor UID) on a UEFI/ACPI host. logical_to_physical() picks whichever the machine has, and both init's CPU validation and NUMA placement use it. These sysfs files stay around for CPUs the pool has taken, unlike the /proc/cpuinfo entries on x86.
  • Pseudo-NMIs: kerf load appends irqchip.gicv3_pseudo_nmi=1 to every arm64 spawn's command line. The host's stop interrupt only reaches a CPU spinning with interrupts off when pseudo-NMIs are on, so without it a hung spawn CPU can never be taken back. The parameter goes last so it overrides a command line that turns it off, and it is not added twice.
  • Boot CPU: kerf init assumed the host's boot CPU had physical ID 0. An arm64 server can boot on MPIDR 0x10000 and have no CPU 0 at all, and x86 does not promise the BSP APIC ID 0 either. init now reads the boot CPU's ID from logical CPU 0 and always keeps it with the host, warning if the request named it.
  • Messages: CPU IDs are called "MPIDR" on arm64 and "APIC ID" elsewhere in errors, warnings and verbose output.
  • arm64 wheel: the release build job is a matrix over ubuntu-latest and ubuntu-24.04-arm. Each builds its own static kerf-init and retags its wheel py3-none-manylinux_2_17_<arch>.manylinux2014_<arch>.musllinux_1_1_<arch>. The sdist is built once, on x86_64, and publishing collects every job's dist-* artifact.

Testing

  • Topology reading was checked against QEMU virt guests booted from a device tree and from UEFI/ACPI, before and after CPUs were moved into the pool.
  • python3 -m pytest -q: 416 passed, 1 skipped. pylint is clean at 10.00, and release.yml passes actionlint.
  • The arm64 release leg has not run yet; the first tag after this merges will be its first run.

🤖 Generated with Claude Code

Multikernel names a CPU by its physical ID, the APIC id on x86, and
kerf reads those from /proc/cpuinfo. On arm64 the ID is the MPIDR
affinity value and /proc/cpuinfo does not have it, so "kerf init"
stopped at "Could not read APIC IDs" before doing anything.

Read the MPIDR from the CPU's device tree node where the host booted
from one, and from the MADT's GICC entries, matched through the
processor UID of the CPU's firmware node, on a UEFI/ACPI host.
logical_to_physical() picks whichever the machine has, and both the
CPU validation of init and the NUMA placement use it. Unlike the
/proc/cpuinfo entry on x86, the sysfs files stay around for a CPU that
the pool has taken.

Loading and starting an instance already had the arm64 syscall
numbers, and an Image file goes to kexec_file_load() as it is. The
rdtsc module only exists on x86, where it remains a dependency; exec
merely prints a timestamp with it.

Checked against QEMU virt guests booted from a device tree and from
UEFI/ACPI, before and after CPUs were moved into the pool. kerf itself
has not been run on arm64 yet.

Signed-off-by: Cong Wang <cwang@multikernel.io>
On arm64 the host has one way to force down a spawn that no longer
answers: a stop interrupt. A CPU that spins with interrupts off only
takes it when its kernel runs with pseudo-NMIs, which is a boot
parameter and off by default. Without it such a CPU can never be
stopped, and the host refuses to reuse it until the machine reboots.

That makes irqchip.gicv3_pseudo_nmi=1 a requirement of multikernel
rather than a choice, so kerf load appends it to every spawn's command
line on arm64. It goes last, where it wins over a command line that
turns it off, and is not repeated when it is already in effect. The
spawn kernel has to be built with CONFIG_ARM64_PSEUDO_NMI for it to
mean anything.

Signed-off-by: Cong Wang <cwang@multikernel.io>
Every message about a CPU ID called it an APIC ID, which is the x86
name. On arm64 the ID is the MPIDR, so an arm64 user was told that
their CPUs had "APIC IDs" such as 65536.

Add cpu_id_name(), "MPIDR" on arm64 and "APIC ID" elsewhere, and use it
in the errors, warnings and verbose output of init and create and in
resource validation. The --cpus help texts are fixed strings, so they
name both.

Signed-off-by: Cong Wang <cwang@multikernel.io>
init assumed that the host's boot CPU had physical ID 0. Where the pool
request covered every CPU, it moved ID 0 back to the host. In every
other case it passed the boot CPU to the kernel, which cannot take it
into the pool. Physical ID 0 need not be the boot CPU, either: an arm64
server can boot Linux on MPIDR 0x10000 and have no CPU 0 at all, and
x86 does not promise the BSP APIC ID 0.

Read the boot CPU's physical ID from logical CPU 0 and always keep it
with the host, with a warning whenever the request named it. The
kernel now rejects a baseline that includes the boot CPU, so the
request should not reach it.

Signed-off-by: Cong Wang <cwang@multikernel.io>
The wheel carries a native kerf-init, so an arm64 host could only
install from the sdist, which needs musl-gcc. Build the wheel on native
x86_64 and arm64 runners and retag each for its own architecture. The
sdist is built once, on x86_64, and publishing collects every job's
distributions.

Signed-off-by: Cong Wang <cwang@multikernel.io>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant