Boot Interfaces and DPU Policies
This guide explains how NICo decides which interface a host boots from, how a host’s DPUs are managed, and how operators configure both through the Expected Machines table. It is the deep companion to Ingesting Hosts: that page covers the end-to-end ingest flow and the basic expected_machines.json; this page covers the per-host and per-NIC knobs (dpu_policy, host_nics), what the defaults do when you set nothing, and how a boot device is chosen and applied behind the scenes.
For the DHCP and network-segment substrate these knobs sit on (how a relay’s giaddr maps to a segment), see IP and Network Configuration.
Who should read this. Operators configuring hosts for ingestion, and anyone debugging “why did this host boot from that interface?” Most hosts need no configuration here — the defaults handle the common managed-DPU case. Reach for the knobs in Section 2 and Section 3 only for
nic,ignore, or integrated-NIC hosts. Sections 6–7 explain the machinery when you need to trace a problem.
1. The mental model: two independent axes
Historically, two separate decisions were conflated into “what kind of host is this.” Keep them apart:
A “normal” host couples them (managed DPU + boot through that DPU). But they are independent: you can keep a host’s DPUs managed and still boot it from a plain integrated NIC. This guide treats the two axes separately, because the configuration knobs are separate.
The policy is also separate from the card’s current hardware state. NICo models
operator intent as HostDpuPolicy::{Manage, Nic, Ignore} and the mode
reported by a BlueField as BlueFieldOperatingMode::{Dpu, Nic}. Existing
protobuf and Redfish boundaries retain the legacy NicMode name for
compatibility, but model code does not use that name for observed state.
Network segment types
A host’s management network lives on one of a few segment types. Which one depends on how the host boots:
Rule of thumb: the host-OS segment follows the boot interface’s mode — boot via DPU ⇒ Admin; boot via a plain NIC ⇒ HostInband. A non-DPU NIC is never on the Admin segment, because Admin is the DPU overlay.
The boot (primary) interface
In NICo the boot interface and the primary interface are the same thing by construction: a host’s primary interface is the one its boot order targets. A boot interface is a (MAC, Redfish interface id) pair — the MAC identifies the NIC on the wire; the Redfish id lets NICo set the boot order on the BMC. NICo refers to this pair as the host’s MachineBootInterface.
2. Configuring via Expected Machines and the defaults
Boot and DPU configuration is declarative: you describe the host in the Expected Machines table, and site-explorer + the machine-controller make it so during ingestion. For the table’s basics (schema, replace-all upload, credentials), see Ingesting Hosts → Add Expected Machines. This section covers only the boot/DPU fields.
If you set nothing (the default)
Most hosts with DPU hardware need zero boot/DPU configuration. Outside rack-manager deployments, when neither the host nor the site sets dpu_policy, and the host has no host_nics declaration:
- The effective
dpu_policyresolves tomanage— NICo ingests and manages the host’s DPUs, and the host boots through its primary DPU on the Admin network. - Site-explorer auto-selects the boot interface: the lowest-PCI DPU host-PF (the NIC a DPU presents to the host).
- The host’s IP comes from whichever segment its DHCP relay lands in (see IP and Network Configuration).
So a standard DPU host is handled entirely by defaults. A host without DPU hardware must instead use dpu_policy: ignore and declare a primary HostInband NIC as described in 3.2. The knobs below exist for the hosts that don’t fit the standard-DPU mold.
rack_management_enabled is a separate deployment-mode setting, not another
DPU policy, and it does not override this resolution. Rack-manager deployments
that operate DPUs as NICs must also set the site-wide dpu_policy (or each
host’s policy) to nic. If rack management is enabled while both policy
levels remain unset, the effective policy is still manage and the Admin-network
default above still applies.
dpu_policy
Resolution order: per-host nic and ignore policies override the site. For backward compatibility, per-host manage (like an omitted per-host policy) defers to the site-wide [site_explorer] dpu_policy; if the site policy is also unset, the result is manage.
Compatibility declarations remain accepted when deserializing configuration and admin JSON input: the previous dpu_policy value use_as_nic remains accepted, and the legacy dpu_mode key with values dpu_mode, nic_mode, and no_dpu maps to manage, nic, and ignore, respectively. The admin CLI likewise accepts the previous use-as-nic value and legacy --dpu-mode dpu-mode|nic-mode|no-dpu forms, but new automation should use --dpu-policy manage|nic|ignore.
The Forge RPC intentionally retains ExpectedMachine.dpu_mode (field 16,
DpuMode) as its stable compatibility surface. Direct gRPC clients continue to
send DPU_MODE, NIC_MODE, or NO_DPU; NICo translates those values at the
RPC boundary to the internal manage, nic, and ignore policies.
HostDpuPolicy and dpu_policy are model, configuration, admin-CLI, and admin
JSON vocabulary rather than Forge protobuf symbols. Responses translate
non-default policies back through dpu_mode; the default manage policy might
leave that field unset.
host_nics (per-NIC declaration)
The optional host_nics array declares specifics for individual host NICs. Each entry (ExpectedHostNic):
What
network_segment_typeactually does. A NIC’s segment is normally determined by its DHCP relay: NICo picks the segment whose prefix contains the relay address. Where segment prefixes nest or overlap — for example a/27HostInband segment inside a/24underlay — one relay matches several segments.network_segment_typenarrows that to the segment of the named type. If a relay maps unambiguously to one segment (the common case), this field is unnecessary.
Admin JSON (an Expected Machine entry):
CLI (single host):
(em is the alias for expected-machine.)
3. Scenarios
Concrete recipes for the cases beyond the default. All assume the rest of the Expected Machine entry (BMC credentials, serial) is filled in as usual.
3.1 Standard DPU host
With the site policy unset or set to manage, there is nothing to configure. DPUs are managed and the host boots through the primary DPU on Admin.
3.2 Zero-DPU host (no DPU hardware)
A plain server with one or more host NICs and no DPU. Declare ignore and mark the boot NIC primary:
The host boots from that NIC on HostInband and gets its IP from central NICo DHCP.
3.3 DPU in NIC mode
The host has DPU hardware, but you want it treated as a plain NIC (not managed). Declare nic. Site-explorer still explores the DPU (and will issue the physical mode flip — see 3.5) but does not link it as a managed machine; the host boots HostInband:
3.4 Boot an integrated NIC while keeping the DPUs managed
This is the case where the two axes genuinely diverge: the host has cabled, explorable DPUs you want managed (for the data plane), but you want the host OS to boot from a plain integrated NIC rather than through a DPU. First ensure the site-wide policy is unset or manage; a per-host manage declaration inherits the site and cannot override nic or ignore. Then leave the host policy unset (or explicitly set manage) and mark the integrated NIC primary on HostInband:
NICo keeps the DPUs explored, linked, and underlay-addressed (running agents for the data plane), but the host boots from the integrated NIC. The DPU-backed admin links are kept but go dormant — the host’s admin/boot path is the HostInband NIC.
Previously this required selecting an unmanaged-DPU policy, which threw away DPU management to get integrated boot. The two are now decoupled.
3.5 Flipping a DPU to NIC mode
To change a host that’s already ingested (e.g. from managed-DPU to NIC mode), update its Expected Machine policy with --dpu-policy, then force-delete and let it re-ingest so site-explorer re-explores and applies the new physical mode:
See the Force Delete playbook for the full re-ingest procedure. NICo preserves the host’s boot-interface Redfish id across the deletion gap via the retained-boot-interface mechanism (Section 7.3), so the host can be re-targeted for boot before a fresh exploration completes. After a flip, you can re-apply the resolved boot interface with one click via Restore Boot Interface in the web UI (Section 5).
The mode change itself takes effect only across a power cycle: The queued Mode.Set stays staged on the DPU until the host power-cycles. Site-explorer issues that power cycle automatically on every vendor, trying the standard Redfish PowerCycle reset first and escalating to a cold ACPowercycle on platforms that refuse it (HPE, Lenovo, Supermicro, and GBx00 systems). Repeat cycles are rate-limited, so a host mid-flip can take a pass or two of exploration to converge. If both reset types are refused, or the DPUs still report the old mode after cycling, the host surfaces the manual_power_cycle_required pairing blocker and waits for an operator (refer to Section 8).
4. admin-cli and gRPC reference
All of these are admin-only; the Forge gRPC service enforces admin authorization. The nico-admin-cli commands are thin wrappers over the listed Forge RPCs.
Expected machines
Boot interface / primary interface
Setting the DPU-first boot order directly (by MAC) is also exposed as a one-off action through the web UI (Section 5). Under normal operation the machine-controller sets the boot order automatically during ingestion (Section 6).
Ingestion control
5. Web UI
The NICo admin web UI (/admin/…) is primarily for visibility, with a focused set of boot/ingestion actions. There is no DPU-policy control in the UI — change dpu_policy via Expected Machines (CLI/JSON) as in Section 2.
View:
/admin/machineand/admin/machine/{id}— machine inventory and detail, including each interface’s primary indicator, MAC, segment, and attached-DPU id./admin/dpuand/admin/dpu/versions— DPU inventory, associated host, and version info (read-only)./admin/expected-machine— a status board of expected vs. unexpected machines, with tabs for Completed / Unseen / Unexplored / Unlinked / Unexpected. (Read-only; entries are defined via the CLI/JSON.)/admin/explored-endpoint— discovered BMC endpoints with their preingestion state, last-exploration latency, and errors.
Act:
- Set DPU First Boot Order (by MAC) — on a machine or explored endpoint.
- Restore Boot Interface — one-click re-apply of the host’s resolved boot interface. It uses the machine’s designated primary once managed, or site-explorer’s automatic default before that. Handy right after a DPU↔NIC-mode flip.
- Machine Setup — prepare an endpoint for ingestion (optionally with a boot-interface MAC; Dell endpoints require it).
- Endpoint controls — Re-Explore, Refresh, Clear Last Error, Pause/Resume Remediation, plus power, Secure Boot, lockdown, and BMC-reset actions.
6. Behind the scenes: how a boot device is chosen and set
Boot configuration spans two components: site-explorer discovers the host and records what it should boot from; the machine-controller state machine applies it.
The host lifecycle
A host moves through these states (see the Managed Host State Diagrams for the full picture):
Site-explorer creates the host in Created (DPUs, if any, in DpuDiscoveringState) and records the boot predictions (below). The machine-controller picks it up and drives it through HostInit, where it configures BIOS and sets the boot order.
Resolving the boot interface
At each boot-config step the controller resolves the target via load_boot_predictions → boot_interface_target, with this precedence:
- The host’s own owned interface row (
machine_interfaces), if it has a boot interface — used once the host has taken its first DHCP lease. - Otherwise, the boot prediction (
pick_boot_prediction) — used before the first lease. - Otherwise, a classification:
- AwaitingNic — a zero-DPU host using
ignoreornicwhose boot NIC hasn’t appeared yet; wait. - Missing — a host with managed DPUs has neither an owned primary interface nor a usable prediction; a fault to investigate. This includes a managed-DPU host declared to boot from an integrated HostInband NIC if that declaration did not produce a prediction.
- AwaitingNic — a zero-DPU host using
Key timing. A host has no
machine_interfacesrow until its first DHCP lease. Predictions are what let the controller configure boot before that lease. Once the host leases and the prediction is promoted to an owned row, the owned row supersedes the prediction.
Applying the boot order
configure_host_bios(atWaitingForPlatformConfiguration) calls Redfishmachine_setupwith the resolved boot interface; on Dell this schedules a BIOS job (WaitingForBiosJob).PollingBiosSetupverifies the BIOS settings took.SetBootOrdertargets the resolved primary interface via Redfish. In the default managed-DPU topology that is the primary DPU host-PF; for a declared integrated primary or a host usingignoreornic, it is the resolved HostInband interface. A “no DPU” response from the BMC is expected and treated as success only for a host with no managed DPUs.- On a reprovision repair,
check_host_boot_configre-checks BIOS + boot order and only remediates if they drifted.
7. The boot-interface data model
The boot interface flows through three tables: predicted → managed → retained.
7.1 Predicted (predicted_machine_interfaces)
Site-explorer mints a prediction per declared host NIC before the host’s first DHCP lease. A prediction carries machine_id, mac_address, network_segment_type, the operator’s declared primary intent, and the boot_interface_id (the Redfish EthernetInterface.Id, captured from the exploration report once available). Predictions are what the controller uses to configure boot pre-lease.
7.2 Managed (machine_interfaces) — promotion
When the host first DHCPs on a predicted NIC, move_predicted_machine_interface_to_machine promotes the prediction into an owned machine_interfaces row:
- The row is created (or an existing static-preallocation row is reused) and associated with the machine.
primary_interfaceis set from the prediction’s declared intent; if it’s primary, any prior primary on the machine is demoted first (so exactly one primary survives).- The
boot_interface_idis resolved by precedence: prediction > existing row value > retained (see below). - The prediction is deleted — the owned row is now authoritative.
The owned table is Store B: the authoritative source of truth once a host is owned, kept current by per-exploration updates.
7.3 Retained boot interfaces
The Redfish boot-interface id is the one fact a MAC cannot always rediscover after deletion (a DPU/NIC-mode flip can drop the MAC from BMC reports while the id stays stable; a re-ingested host needs to be targeted for boot before a fresh exploration). So:
- On deletion of a
machine_interfacesrow, itsboot_interface_idis upserted intoretained_boot_interfaces(keyed by MAC; newest wins). - On creation of a new row, any retained id for that MAC is consumed and applied — provided it’s within the configured
retained_boot_interface_window(default: no expiry, i.e. retained forever; set a window to bound recycled-MAC reuse).
This is what carries a host’s boot target across a force-delete / re-ingest gap.
7.4 Selection precedence
The same precedence applies wherever NICo picks a boot interface, over owned rows, predictions, or the explored report respectively:
Store A vs. Store B: before a host is owned, site-explorer records a pre-ownership boot default on the explored endpoint (explored_endpoints.boot_interface_mac/_id, via fetch_host_primary_interface_mac). Once owned, machine_interfaces (Store B) is authoritative. A declared primary wins in both stores, so the boot interface is consistent across the ownership handoff.
8. Verifying and troubleshooting
Check a host’s boot interface:
The interfaces section shows each NIC’s MAC, segment, and which one is primary (the boot interface). The web UI machine-detail page shows the same with a primary indicator.
To inspect the boot interface itself — every store (managed, predicted, explored, retained) side by side, the effective interface the system would select, and a flag when the stores disagree — use boot-interface show <machine-id>; boot-interface candidates <machine-id> narrows it to the candidate NICs and the picks among them. Both are read-only (Section 4).
Common situations:
For ingestion/pairing diagnostics generally, see Ingesting Hosts → Troubleshooting.
Related pages
- Ingesting Hosts — the end-to-end ingest flow and the base
expected_machines.json. - IP and Network Configuration — network segments, DHCP relay/
giaddr→ segment matching. - Force Delete — the re-ingest procedure used when flipping DPU modes.
- Managed Host State Diagrams — the full host state machine.
- DPU Lifecycle Management — DPU OS install, firmware, health, reprovision.