Host preparation and recovery¶
New and existing appliances use one post-install workflow. The installer
provides Ubuntu, K3s, the dashboard, diagnostic helpers and a root-owned local
worker. It does not contain a second GPU package-selection or reboot path.
The base kernel must already boot the computer and reach storage/network;
post-install preparation cannot repair an installer that cannot boot.
For an installer-created NVIDIA display host, the base host playbook separately
installs a pinned display-owning driver and schedules one reboot after successful
first convergence, so the physical setup/TUI console remains available. The
same host role keeps nvidia-persistenced active once the driver is usable. The
GPU Operator's CDI device specifications refer to its Unix socket, so an
inactive service can prevent newly created NVIDIA model containers from starting
even when the operator and nvidia-smi appear healthy. Check
systemctl is-active nvidia-persistenced and
test -S /run/nvidia-persistenced/socket before changing the GPU Operator.
The media does not select the driver or schedule that reboot. CPU-only and AMD-only
hosts skip it; optional AMD kernel/profile changes remain administrator-confirmed.
New USB media use Ubuntu 26.04 LTS and its native generic kernel for both the
installer and installed system. This base-install choice is separate from the
reviewed GPU package profiles below; it does not upgrade an existing appliance
or change its distribution repositories. Ubuntu 24.04 host preparation remains
available for existing installations. See
installer kernel selection.
Fresh K3s installations are pinned through the overridable Ansible default
k3s_version: v1.36.4+k3s1, rather than resolving a moving install channel.
This release includes the native containerd v3 configuration drop-in imports
used by the NVIDIA toolkit. The installation task retains its
creates: /usr/local/bin/k3s guard: ordinary host convergence does not upgrade
an existing cluster. Existing clusters require a separately planned K3s
upgrade and confirmation that the generated containerd configuration imports
config-v3.toml.d/*.toml before adopting the new NVIDIA configuration. An override
must preserve that prerequisite; choosing another K3s version is not a claim of
GPU support. See the pinned K3s release.
User workflow¶
System → Model cache uses this same host-operation boundary for explicit cleanup of downloaded models. It never removes credentials or container images, and refuses cleanup while local model deployments can use the cache. See model cache management for scope and upgrade requirements.
Ethernet and Wi-Fi use the same bounded host-operation mechanism through System → Settings → Network, adding independent local timeout and boot recovery. See network management. Network, power and GPU changes never execute concurrently on the same host.
Open System → Hardware → GPU nodes → Host preparation. GPU operators are shown first. Each node contains its kernel plan, preparation acknowledgement, advanced profile override and collapsed GPU memory configuration. Help and diagnostic details are available through info icons, while action warnings and exact-host confirmations remain explicit. The local worker periodically inspects the OS, running kernel and all PCI display GPUs. It publishes a versioned, host-bound plan, including exact package changes and whether a reboot is necessary. Inspection itself never installs packages or restarts the host.
For a matching reviewed plan, acknowledge the experimental hardware profile,
select Review hardware preparation, and type the computer's exact name in
the confirmation dialog. The confirmation authorizes the displayed packages,
one orderly restart if required, and AMD GPU profile activation. The
local root worker runs the existing Ansible preparation role, persists its
progress, and resumes verification after reboot. Closing the dashboard does
not cancel an accepted operation. Existing working kernels are not removed.
After host verification, Registering waits for fresh eligible hardware and an
allocatable Kubernetes GPU. This completes preparation without an engine smoke
test. Engine validation is an optional manual action in System → Hardware;
no test-model or inference-image download is started by host preparation.
The collapsed Advanced · AMD runtime profile section is a manual override,
not another required setup step: preparation already activates the matching
runtime profile. The override changes no host packages or kernel.
Package and power operations target only the selected computer. GPU runtime
profile selection uses the existing AMD ModuleActivation, which
applies to matching cluster Nodes; the confirmation explicitly includes that
scope. Concurrent preparation requests must not overwrite a changed module
decision: the worker detects configuration changes and stops its profile
handoff. Runtime profile selection is still cluster-wide. The node's Verify
Ollama and Verify vLLM buttons request only that engine on that node;
per-model runtime placement remains a separate extension.
The shipped x86-64 Strix Halo (1002:1586) host package profiles are:
| Host OS | Profile / version | Exact package | Target kernel |
|---|---|---|---|
| Ubuntu 26.04 | strix-halo-ubuntu-26.04 / 1 |
linux-generic=7.0.0-31.31 |
7.0.0-31-generic |
| Ubuntu 24.04 | strix-halo-ubuntu-24.04 / 1 |
linux-generic-hwe-24.04=7.0.0-31.31~24.04.1 |
7.0.0-31-generic |
The 26.04 package pin comes from the official
resolute-updates package listing.
It is a bounded additional preparation plan, not an OS upgrade or a hardware
certification. The new Ubuntu 26.04 installation path still requires local
hardware and inference acceptance; both profiles remain experimental.
An already working native 7.0 kernel needs no package change or restart: explicit
preparation only activates the matching GPU profile. A suitable kernel with a
failing driver instead remains blocked for diagnosis. The worker does not
reinstall the pinned package or downgrade a newer working kernel.
NVIDIA, Intel and other AMD devices without this preparation requirement
retain their existing operator path; no matching profile is not proof of GPU
support. Profiles must be maintained and revalidated for future security updates;
the worker never selects arbitrary latest releases. The Ansible package gate
allows the linux-generic meta-package only on 26.04, and the 24.04 HWE
meta-package only on 24.04. It never adds third-party package repositories.
Mixed systems and experiment mode¶
All GPUs inside one computer share its kernel. Different cluster Nodes have independent kernels. The normal path blocks a Strix Halo kernel change on an unreviewed mixed/multi-GPU computer, without disabling working vendor operators. A working Strix Halo host driver alongside another vendor can still undergo AMD engine validation without changing its kernel.
When a shipped host profile matches the OS/architecture and required Strix Halo device but the additional GPUs are unreviewed, administrators can enable Experiment mode. The UI shows the full detected GPU list and the proposed profile. A separate acknowledgement and exact-host confirmation are required. Other GPU drivers, graphics output or connectivity may break; have physical console access and the previous kernel available in the bootloader. There is no claim of automatic bootloader rollback or remote power recovery.
Experiment mode permits only the displayed, shipped package plan. It cannot
accept an arbitrary kernel URL, shell command, package name, repository, ROCm
installer or architecture override. Unknown OS/architecture and incomplete PCI
inventory remain blocked; a new package profile needs code review first. It does
not bypass inference eligibility. Multi-AMD layouts not supported by the current
node-scoped plugin profile may finish as PreparedUnverified: host preparation
is distinct from GPU availability. Successful AMD smoke tests do not certify
other vendors or the entire mixed system. Tests remain local evidence, never
automatic additions to the public compatibility catalog.
Fixed and dynamic GPU memory¶
System → Hardware → GPU nodes → GPUs → GPU Configuration AMD → Shared GPU memory provides two administrator controls
on one supported Strix Halo GPU with compatible kernel/firmware evidence,
including hosts with additional NVIDIA GPUs bound to the nvidia driver:
- Fixed GPU reservation (firmware): a discrete slider containing only the
options reported by that computer's
uma/carveout_options. This RAM is carved out at boot and is unavailable to Linux. - Dynamic GPU memory limit: a slider in 1 GiB steps for the TTM ceiling on GPU use of shared Linux RAM. It is not a reservation, an additional memory pool, or guaranteed free memory. CPU workloads compete for that RAM.
The panel separates current values from the unsubmitted draft. Changing a
slider sends no host operation. A non-step-aligned kernel default can have a
rounded draft, but that alone cannot enable submission. If the active limit
exceeds the allowed maximum for the current firmware reservation, the panel
instead offers a limited corrective draft for review without another slider
movement. It still requires exact-host confirmation and never applies itself.
The preview estimates Linux RAM as current MemTotal plus the old fixed reservation minus the new
fixed reservation. The dynamic ceiling must leave at least 16 GiB outside
GPU dynamic allocations; this allowance is not a Kubernetes memory reservation.
Actual post-boot MemTotal, not this projection, governs the second stage.
The worker selects the Strix Halo device by its verified PCI address; NVIDIA
VRAM is neither reconfigured nor added to the shared-memory budget. The TTM
limit is system-wide, so additional AMD GPUs, other GPU vendors, missing driver
bindings and NVIDIA GPUs using nouveau or vfio-pci remain blocked. Complete
PCI inventory and driver bindings are rechecked at approval and after each
restart. After an approved memory reboot, an unchanged but not-yet-bound NVIDIA
GPU gets up to ten minutes for its operator-managed driver to start; no further
memory write or restart runs during that wait. The PCI hardware identity must
remain unchanged; full driver identity is compared once the driver has bound.
A changed companion GPU, wrong driver binding or timeout stops the operation. This memory
workflow never changes the kernel, NVIDIA driver or operator configuration;
ordinary preparation/experiment-mode safety rules remain separate.
Select Review memory configuration and enter the exact computer name in the final disruption confirmation to apply. A fixed-reservation change may require up to two restarts: first apply and verify the firmware choice, then apply the dynamic limit through the existing Ansible role and verify it after another boot. A dynamic-only change requires one restart. All workloads on that computer are interrupted; no migration or automatic firmware rollback is promised. Keep physical console/recovery access available. Closing the browser does not cancel an accepted operation.
The worker independently checks the reviewed hardware/configuration identity, actual firmware options, active memory values and safety allowance. Missing firmware controls, mixed/unknown GPU layouts, incomplete/stale evidence and competing boot/modprobe overrides disable this flow. Hardware experiment mode does not override those memory guards. A memory operation installs no kernel or packages, changes no inference profile and starts no GPU validation workload. Its success confirms the requested memory settings, not inference compatibility, performance, cgroup accounting or successful model execution.
The hardware fingerprint excludes only the configurable UMA VRAM capacity: changing that capacity is the intended operation, not evidence of a replaced GPU. PCI identities, architecture, driver, kernel, OS and firmware versions remain checked. The desired firmware reservation and dynamic limit are verified separately against their exact approved values after reboot. If a previous operation stopped after applying the firmware choice, its old dynamic limit may still be active. Diagnose the failure, refresh the current evidence and confirm a new plan; terminal failed requests are never replayed automatically.
The next-boot dynamic setting is owned in
/etc/modprobe.d/90-magicstick-ttm.conf; the role refreshes initramfs. Neither a
live TTM-only write nor the deprecated amdgpu.gttsize override is used. Firmware
option indices are resolved locally; browser requests contain no device path or
shell command. Linux UMA controls
and AMD shared-memory guidance.
Restart and shut down¶
Administrators see a separate Computer power tab inside System, immediately
after System Status (#/system/power), with Restart computer and
Shut down computer. These controls are only mounted on this tab. Viewers and
operators cannot access it, including through a direct link. If several managed
computers are reported, select
the target explicitly. Both buttons require typing its exact name and accepting
interruption of all services/workloads on that computer. They are unavailable
for stale/offline workers or while another operation is active.
The worker uses normal systemd shutdown -r +1 or shutdown -P +1: approximately
one minute after local acceptance, not an immediate forced reset. Save work
before confirming; there is no automatic Pod migration or zero-downtime promise.
The dashboard distinguishes accepted, scheduled, and new boot observed.
It cannot prove physical power-off while the machine is unreachable, and does
not repeatedly send a power request. Power-on requires local action or separately
configured remote power management. systemd shutdown manual
Execution and security boundary¶
magicstick-host-management.timerruns after boot and every 15 seconds after the previous tick completes. It waits for cloud-init's base installation to finish. No installation-time consent is inferred.- Root-owned implementation and profiles live in
/usr/local/lib/magicstick/host-management/. Local progress is atomically persisted in/var/lib/magicstick/host-management/state.json(root only). - The worker publishes
appliance.magicstick.dev/host-managementon its local Node. API availability requires matching Node UID, boot ID and kernel plus evidence no more than three minutes old. GET /api/host-managementis viewer-readable. Administrator-onlyPOST /api/host-management/operationsrequires existing authentication, same-origin/CSRF checks, exact-host confirmation, current Node/boot/plan IDs, explicit disruption consent, and a unique 32-hex request ID. No secrets or executable content are accepted. The actor subject is hashed for the request.- One immutable namespaced
HostOperationper Node serializes browser requests. New requests expire after five minutes if not accepted locally. Active requests cannot be replaced; terminal replacement uses a resource UID precondition. The worker independently rechecks all inputs, and remembers processed IDs. - The dashboard may create/read/delete these requests in
ai-system; it cannot write their status, patch Nodes or execute host commands. No privileged dashboard Pod, hostPath or network-facing root agent is introduced. The worker uses the existing root-only local K3s kubeconfig; cluster/root administrators remain trusted. Arbitrary external Kubernetes installs and agent-only Nodes without a local management credential do not silently gain host control. - Host convergence and host operations share
/var/lib/magicstick/host-management/maintenance.lockinside a root-only directory. Convergence skips an active operation and a scheduled shutdown; it does not compete with APT preparation. - APT uses exact approved versions, signed configured repositories, no downgrades and no package removals. The prepare playbook is the same bounded Ansible implementation used for explicit local maintenance. Ansible APT contract
Progress and recovery¶
Typical preparation: Preparing → RebootScheduled → Verifying → Registering
→ Succeeded. A host that needs no restart enters registration directly.
Succeeded means fresh host/driver evidence and an eligible registered GPU,
not an engine smoke test, production model acceptance or KubeAI runtime-image
adoption. Legacy in-progress Validating operations now complete from the same
registration checks. Optional engine diagnostics remain visible under
GPU compatibility.
Memory configuration uses Preparing → RebootScheduled → Verifying, with a
second bounded cycle when both firmware reservation and TTM need changing.
Root-owned state records the stage and reboot count before side effects. An
interrupted/uncertain write is never automatically replayed. Changed hardware,
unexpected actual memory or a conflicting local setting stops the operation;
inspect the journal before issuing a new confirmed request.
Rejected, Interrupted, Failed and PreparedUnverified are terminal. An
uncertain package execution, changed boot/profile, wrong kernel after reboot,
expired request or failed engine test does not trigger an automatic retry,
another kernel change, or a reboot loop. Inspect and submit a new confirmed
request only when appropriate. GPU validation waits at most one hour, allowing
for large image downloads; host driver verification waits five minutes.
sudo systemctl status magicstick-host-management.timer
sudo journalctl -u magicstick-host-management --since '-30 minutes'
sudo magicstick-gpu-preflight --json
kubectl -n ai-system get hostoperations
Do not delete an active request or its root-owned state to force a retry. If an administrator manually removes Kubernetes execution state, the local worker stops rather than guessing whether a power/package operation already happened. Retain the journal for audit and review local recovery before clearing state.
Local verification¶
PYTHONDONTWRITEBYTECODE=1 python3 -m unittest discover \
-s magic-host/roles/host-management/tests
PYTHONDONTWRITEBYTECODE=1 python3 \
magic-host/roles/host-management/tests/rancher_contract_test.py \
--context rancher-desktop
The second test explicitly targets local Rancher, creates an isolated namespace and the CRD only if absent, and removes its own resources afterwards. It verifies real Kubernetes validation, immutability, RBAC and status persistence with a fake power executor. It neither installs a host worker nor patches/restarts any Node. Unit tests cover package/reboot decisions and post-boot continuation; physical power loss, bootloader fallback and actual mixed-GPU compatibility still require a separately approved hardware test.