October 2, 2026

The One-Flag Fix: A Bare-Metal Restore That Took a Day and Ended in a Single QEMU Setting

A field story from a BDR acceptance test — NAKIVO Backup & Replication v11.2.2, Proxmox VE, a Dell Windows 11 laptop, and a boot failure that lied about its identity for most of a day.

TL;DR

We backed up a real Windows 11 laptop with NAKIVO, restored the image to a Proxmox VE virtual machine as a bare-metal recovery (BMR) test, and the restored copy would not boot — looping in "Preparing Automatic Repair" with no stop code and no repair log. After a day of boot forensics (UEFI, BCD, registry hives, storage drivers, boot journals), the root cause was neither the backup, the restore, the disk geometry, nor any Windows component:

Proxmox's default QEMU CPU model (kvm64) sits below Windows 11's x86-64-v2 baseline (no SSE4.2/POPCNT). The Windows kernel failed instantly and silently, before painting a single pixel — and every symptom pointed somewhere else.

The fix was one line:

qm set 900 --cpu host

The restored image then booted clean to the Windows 11 lock screen.

The backup was never in doubt — file-level restore with verified content had already passed. But the boot test is the acceptance gate for bare-metal recovery, and it became the longest single-problem day of the project. This is the full account: every wrong turn, the evidence that killed each wrong theory, and the checklist that would have saved the day.

The setup

The players:

RoleWhatNotes
Protected machineDell Windows 11 laptop, 500 GB NVMe~356 GB used on C: — size the target from used space, never disk size
Backup productNAKIVO Backup & Replication v11.2.2 (trial)Director on an Ubuntu VM, agent-based physical backup
BDR hostHP Z6 G4 tower, Proxmox VE 9ZFS root mirror (2× 512 GB NVMe), 8 TB Exos scratch drive for restore tests
SandboxVM 900 "bmr-sandbox"Isolated: no NIC (never bridge a restored copy onto the LAN while the original exists), image on the scratch drive
The backupFirst full ≈ 356 GB source → ~293 GB in the repoDedupe/compression as expected; nightly incrementals after

The acceptance test: prove the whole chain — laptop → agent → tower repository → bare-metal restore into a VM → that VM boots. A backup that restores to files but never proves it boots is only half a BDR product.

The restore itself was almost anticlimactic: 28 minutes from the Director to a 500 GB sparse image on the scratch drive. The problems all came after.

The issues, in order

1. The scratch drive didn't exist (BIOS-disabled SATA controller)

The 8 TB Seagate was installed, seated, cabled — and completely invisible to Proxmox: no lsblk entry, no SATA controller on the PCI bus at all (/sys/class/scsi_host/ empty, no 0106 class device in lspci).

Lesson: an invisible SATA drive on an HP workstation is a controller diagnosis before it's a cable diagnosis. The Z6 G4 ships the sSATA controller disabled (and the adjacent sSATA RAID mode checked). Fix was in F10 BIOS: Advanced → System Options → enable the sSATA controller in AHCI mode.

2. The tower stopped booting (polluted UEFI boot order + a corrupt firmware entry)

After the BIOS session, the tower hung at boot. Two stacked causes:

  • Adding hardware let new entries (the fresh Seagate with stale boot blocks, plus install-day USB) insert themselves into the UEFI boot order ahead of the Proxmox NVMe mirror.
  • A genuinely corrupt firmware entry (Boot000A) claimed to be a 1024 GB NVMe with the wrong serial number and an all-zero device path — a firmware lie that could never boot.

Fix: move both NVMe mirror members to the top of the boot order (each member's ESP is a complete identical boot copy — that's the point of the mirror), delete the garbage entry, refresh both ESPs with proxmox-boot-tool. Unattended reboot verified before the long restore — never point 373 GB of restore work at a host whose boot you haven't proven.

3. Getting the recovery ISO was a reverse-engineering project

The bare-metal recovery environment is a bootable ISO that the Director generates for itself — it embeds the Director's address so the recovery wizard's browser can find home. The obvious API route (plain GET) returned HTML. The working path turned out to be a form-encoded RPC POST (BmrManagement.getIso) — a 2.57 GB download that completed in 45 seconds once the shape was right.

4. The restore: 28 minutes, textbook

The recovery environment boots a Linux desktop with a browser, the wizard is five clicks (cert override → Director login → machine → savepoint → target disk), and the image materialized to the scratch drive at wire speed. Verification immediately after, mounting the image read-only:

  • GPT present, 512/512 logical/physical sector geometry (matches the source NVMe)
  • Partition layout intact: ESP 100 MB, MSR 16 MB, C: ≈ 476 GiB, WinRE ≈ 982 MB
  • System32: 4,825 entries. bootmgfw.efi present on the ESP, correct size.
  • No hibernation file poisoning the image

Everything that can be verified without booting said: this image is complete and structurally perfect.

5. Then Windows looped. Silently.

Boot test: OVMF → ESP → bootmgr → winload — the firmware chain worked — then "Preparing Automatic Repair", reboot, repeat. No stop code. No bluescreen. Automatic Repair never produced a verdict and (it turned out) never wrote SrtTrail.txt. The loop painted exactly two screens: firmware splash, repair text.

This is where the day began.

6. Wrong theory #1 — the storage driver (plausible, and wrong)

The leading theory: the Dell boots its storage through Intel VMD (iaStorAC.sys), the QEMU sandbox presents an AHCI controller, so Windows hits INACCESSIBLE_BOOT_DEVICE and loops into repair. This theory had everything — vendor-specific driver binding, a sandbox that presents different hardware, an industry-standard failure shape.

The fix was staged and technically elegant: present the restored image to the VM as an emulated NVMe device (QEMU args: building an NVMe controller + namespace backed by the image file), so Windows' in-box stornvme driver takes it with no injected drivers.

The UEFI shell even proved the emulation worked — the firmware mapping table showed NVMe(0x1,...) with the full Windows layout. Manually launching bootmgfw.efi from the shell (worked around OVMF's stored boot entries still pointing at the old AHCI device path — issue #6.5) led to... the same Automatic Repair loop.

The storage theory survived every test thrown at it. It was also not the disease.

7. Self-inflicted trap — interrupting Automatic Repair

An honest confession that cost hours: earlier diagnostic cycles ended with hard VM stops mid-repair-loop. That hard-dirties the restored NTFS and prevents WinRE's Automatic Repair from ever completing a pass — so the environment was never allowed to either self-heal or name the failure. The discipline that fixed this layer: one clean, uninterrupted boot, screen captures every ~3 seconds for the first minutes, then let it run.

8. The decisive instrument — BOOTSTAT.DAT

Windows' boot manager keeps its own journal: C:\Windows\BOOTSTAT.DAT. Decoding it offline (the image mounted read-only via qemu-nbd) produced the single most valuable fact of the day:

The OS loader launched 41 times. Every attempt returned 0xC0000001 — instantly, before any screen painted.

That status and that speed pattern change the entire diagnosis:

  • It's not INACCESSIBLE_BOOT_DEVICE (that fails later, with a visible bluescreen)
  • It's not a dirty filesystem, not a BCD problem, not a missing driver file
  • The kernel died at handoff — every time, instantly, silently

9. The layer audit that cleared everything else

With the failure localized to "instant death at kernel handoff," a full offline audit exonerated, one by one: all boot-critical files present (winload, drivers, registry hives), NTFS clean, ESP intact, BCD entries valid with correct device blobs, all storage drivers at boot-start (the VMD theory's own drivers were fine), firmware variables rebuilt fresh (NVRAM test — its own gotcha: a fresh restore writes new partition GUIDs, so stale efidisk NVRAM entries are all dead and OVMF drops to the UEFI Shell).

Everything clean. Loader still dying instantly. That combination — perfect files, instant silent death — is the fingerprint of the CPU instruction-set floor.

10. The answer: the CPU model nobody set

Proxmox VE's default CPU model, when nothing is configured, is kvm64 — a deliberately generic, lowest-common-denominator model. It is invisible in qm config (unset = not printed), which is exactly why it hid through a day of config reviews. It only shows in the effective command line:

qm showcmd 900 | grep -oE "\-cpu [^ ]+"

kvm64 predates and lacks SSE4.2 and POPCNT — both mandatory in the x86-64-v2 baseline that Windows 11 is built against. The moment the Windows kernel starts, it touches an instruction the vCPU doesn't have, and dies. Instantly. Silently. No bluescreen, no log, no repair verdict — the machine never gets far enough to produce any.

The fix:

qm set 900 --cpu host

Pass the host's real CPU features through. The next boot: firmware → bootmgr → winload → Windows 11 lock screen. First try. The image that had "failed" all day was perfect from the first restore.

The side quests worth their own notes

OneDrive folders that weren't empty — but looked empty. Verifying restored file content, OneDrive-synced trees read as empty directories when the image is mounted with the FUSE ntfs-3g driver — it silently can't traverse OneDrive's cloud-only reparse points. The kernel ntfs3 driver traverses them fine and the full ~200 GB tree appears. A disk image cannot contain cloud-only files (they never lived on disk) — but hydrated trees that ARE on disk must be read with the right driver, or a good backup looks like a data-loss event.

The shrinking image that wasn't losing data. After restore, the sparse image's allocated size shrank — TRIM punches holes where the guest freed space. Sparse allocation is not data; verify with content, not file size.

Trust the hash, verify the corpus. The file-level restore proof zip passed every transport check (SHA256 exact match, clean archive test) and still warranted opening a file and reading it — because a verified transport of the wrong folder is still the wrong folder. Confirm exact source paths before extracting.

Firmware-level proofs beat boot-cycle evidence. The UEFI Shell's device mapping table proved the NVMe emulation worked with zero boot cycles spent. When debugging boot, look for the observation that removes a whole theory in one look.

The checklist that would have saved the day

For anyone restoring a Windows image into a Proxmox/QEMU sandbox, in the order they should be checked:

  1. CPU model — FIRST. qm showcmd <vmid> | grep -oE "\-cpu [^ ]+". Empty or kvm64 with a Win11 (or any x86-64-v2-era) guest = stop, set --cpu host, retest. Ten seconds. We spent a day proving this the hard way.
  2. Uninterrupted boot discipline. Never hard-stop mid-Automatic-Repair; hard stops dirty NTFS and erase WinRE's chance to tell you anything.
  3. BOOTSTAT.DAT before theories. The boot manager's own journal names the failing stage and the NTSTATUS — one offline read replaces hours of guesswork.
  4. UEFI: fresh restore = fresh NVRAM. New partition GUIDs strand stored boot entries; expect the UEFI Shell and know the fs0:\efi\microsoft\boot\bootmgfw.efi launch.
  5. Storage-driver theories last, not first. VMD-vs-AHCI is a real failure mode, but it produces visible INACCESSIBLE_BOOT_DEVICE symptoms. A silent instant fail-fast isn't it.
  6. Read OneDrive trees with ntfs3, not ntfs-3g — or your file-level proof will under-report.
  7. Never bridge a restored copy onto the production LAN while the original lives — duplicate-identity damage (USN/SID) is worse than any test value.

What it means for BDR practice

  • The acceptance test is the product. "Restore completed" is a transport metric; "boots to login" is the deliverable. Budget acceptance-test day as its own project day, not the tail of install day.
  • Restore RTO was 28 minutes for ~356 GB used — the restore side of the chain is fast and boring, exactly what you want. All the drama was in the sandbox presentation of the restored image, which real bare-metal restores to original hardware never touch.
  • Every failure was environmental, none was the backup. The final image booted first-try once the sandbox stopped sabotaging it. That's the strongest possible statement about the backup itself — but you only earn that sentence by finishing the boot test.
  • Silent fail-fasts are an instrument-availability problem. With BOOTSTAT.DAT decode in the kit from hour one, this is a 30-minute diagnosis. Build the instrument before the fire.

Test environment: NAKIVO Backup & Replication v11.2.2 trial, Proxmox VE 9 on HP Z6 G4 (ZFS root mirror), Dell Windows 11 laptop, September 2026. Restored image: 500 GB sparse raw on 8 TB Exos scratch, VM 900, OVMF/q35, no NIC.

The One-Flag Fix — A Bare-Metal Restore That Took a Day and Ended in a Single QEMU Setting | Tech Connect Arizona