October 8, 2026

NAKIVO Never Documented Their API: How I Automated a Whole BDR Stack Anyway

A field story from a Proxmox backup tower, a Synology NAS mid-RAID-rebuild, and an undocumented JSON-RPC API with 117 action groups: five failure rungs to get one backup running, a repository that refused to exist, three silent failures at 0%, and a VM that booted straight off the backup in about 40 seconds.

The setup

I build backup and recovery (BDR) stacks for clients. The current project: a Proxmox tower, a Synology NAS as the backup repository over a private link, and NAKIVO Backup & Replication as the software that ties it together. NAKIVO has a web UI called the Director. In hypervisor terms, the Director is the vCenter: it holds the job configs, the schedule, the catalog of backups. It moves no backup data itself. The transporter (their worker process) does the actual disk I/O wherever data flows.

Here's my constraint, and it's the reason for everything in this article: I run a headless AI assistant that operates my infrastructure. It can SSH into anything, it can run scripts on a schedule, it can watch counters and alert me. What it cannot do is click through a web wizard at 2 AM. So for me, "we'll just do it in the GUI" is not a plan. A GUI action is a step that can't be scheduled, can't be monitored, can't heal itself, and can't be reused on the next client site. The API was the whole project. The GUI was just where the errors went to hide.

That line is the thesis of this story: the GUI teaches you what the product does. The API is where you find out what it actually accepts.

NAKIVO does document an API, in the sense that there's a page somewhere describing a REST-ish interface for a few operations. What they do not document is the interface their own web UI uses, which is a different thing entirely and covers everything. This article is about that one.

First: finding the API nobody documents

The first hint came from watching the network tab while clicking around the Director UI. Every click fired a POST to the same endpoint: https://<director>:4443/c/router. Not REST paths. Not /api/v1/jobs. One endpoint, and the payload tells the server what you want:

{"action": "JobManagement", "method": "run",
 "data": [{"jobIds": [1], "runType": "ALL"}],
 "type": "rpc", "tid": 42}

That's Ext.Direct, a JSON-RPC dialect that Sencha (Ext JS) applications speak. The web UI is a Sencha app. Which means everything the UI can do, that endpoint can do, if you can figure out the shapes. And there's a cheat sheet sitting in the open:

GET https://<director>:4443/c/api

That returns the full descriptor: every action group, every method, and how many positional arguments each one takes. On this build: 117 action groups. Two things to know about the descriptor. One, it's JSONP-style, so the response starts with Ext.app.REMOTING_API={...} and you strip that wrapper before parsing. Two, it shifts between releases, so re-pull it on every version bump. The method names and argument counts are the authoritative list for the version you're running.

Login is the obvious first call, and it's where you learn the API's personality: AuthenticationManagement.login([user, password, remember]), positional arguments, sets a session cookie you carry on every call. One account for automation, credentials staged in a file with 0600 permissions. Repeated bad passwords increment a failedAttempts counter toward lockout, so never guess: if you don't know a shape, don't brute force it, go read something else.

The login response includes a field called productConfigured. First time in, it was false. That does not mean your account is broken. It means the first-run setup wizard hasn't finished: no inventory added, no repository created, no backup configured yet. The UI locks most of its menus even for a full administrator until that's true. Fix it through the API (add inventory, create a repo, run one backup) and the flag flips. Nothing to reinstall, no license to blame. That one cost me an evening of staring at a UI that had locked me out of itself.

Talking to it: the basics that still cost three rounds

A few traps are structural and worth naming before the war stories, because they explain half the pain.

Responses live under data, never result. Every successful call returns {"result": null, "data": {...}}. If you check resp["result"] like a normal JSON-RPC client, every healthy call looks like a failure. I burned three rounds re-probing argument shapes that were already correct because I was reading the wrong key. On first contact with any method: dump the whole response raw and look at the keys before you write a single assertion.

Arguments are positional arrays, and the count matters. The descriptor tells you each method's argument count. run takes one argument: an object with jobIds inside it. A bare [1] gets rejected. getJobShortInfo takes one argument that must itself be a list, so a bare id returns null. listVirtualMachines takes one argument that is an array of vid strings, so the data field double-wraps: [["PROXMOX_HOST-1"]]. There's no rule you can memorize. There's only: read the descriptor, count the positions, try the shape, and when it fails, go read the real error (more on where errors actually live in a minute).

The error strings are a decoder ring. Three patterns cover most rejections:

  • empty value in X means field X is missing or unmapped. Add exactly that key.
  • The parameter value is invalid means the top-level shape is wrong (arity, or a wrapper where an array belongs, or vice versa).
  • A NullPointerException naming a getter, like JobOptions.getGenerateMac() is null, is the server politely telling you the exact DTO field you forgot to fill.

That last one becomes the hero of the final chapter.

The first backup job: five failure rungs, in order

Goal one was a real backup job through the API: a Windows 11 laptop, physical machine, agent-based backup into a repository on the tower. NAKIVO's own UI builds this in a five-screen wizard. Through the API it took five failure rungs, each one costing a cycle, in this order:

Rung 1: the destination that doesn't exist. The repository list returns vid: null for every repository (a quirk of the listings call), so there's no ID to point the job at. The identifier the runtime actually wants is the string BACKUP_REPOSITORY-<id>, concatenated by hand. I found the format by mining their own code: a SQL fragment that reads targetStorageVid = concat('BACKUP_REPOSITORY-', br.id). Omit the destination and the run dies in seconds with SelectTargetBackupRepository: Cannot find the "null" item. Which is a genuinely great error message, by NAKIVO standards. Enjoy it, it's the last friendly one.

Rung 2: the source that isn't a source. The discovery process (you feed the Director your Proxmox or Windows host credentials, it scans and lists machines) produces items with their own IDs. The job editor happily accepts a discovery-item ID as the job's source. At runtime the run dies instantly: PhysicalDiscoveryItem cannot be cast to SourceItem. And here's the part that turns an error into an ordeal: after that cast failure, the Director silently retries every 15 minutes while the job status stays RUNNING. It looks alive. Zero bytes move. The correct source is the machine's inventory ID (PM-<id>), which you get off the discovery item's targetVid field.

Rung 3: mappings. A backup job needs a per-disk mapping list: which physical disks of the machine go into the backup. The API lets you save a job with an empty mappings array. The run then dies in pre-flight with no disk mappings to process. The disk list comes from an inventory call that returns extendedInfo.disks[], one entry per disk, and you build one mapping per disk.

Rung 4: the enabled flag. The save envelope has an outer enabled field. Save without it and the job reads back enabled: null, which is to say disabled. Every subsequent run call then sits in WAITING_DEMAND doing absolutely nothing, with no error, because a disabled job isn't an error, it's a choice you didn't know you made. After every save: read back the flag, fix it, call enable.

Rung 5: the edit lock. Try to save a job while it's running and you get This job is locked by the run action. Order matters: stop, wait a few seconds, save, run again.

When the laptop's first full finally ran clean, it wrote 292.9 GB into the repository (356 GB used on the source, FAST compression doing its job). And then the job reported FAILED. With one object OK and one object failed. The failed object was a zombie: a stale leftover object from an earlier edit, still attached to the job, that can never succeed. It burned its retry ladder (15, 30, 45 minutes) and stamped the run FAILED at the same timestamp the real machine's "processing has finished" event landed. Two events, same moment: one says finished, one says failed. Both messages at once. Very helpful.

The heal is API-clean: mark the stale object with nextRunAction: "DELETE_KEEP_TARGET" (an enum I mined out of their JAR, one of four variants, this one keeps the backup data and only unlinks the job object), save, run. Object count went 2 to 1, verified by read-back, and the re-run finished clean. The standing lesson, learned the hard way: never report DONE or FAILED from the job-level flag alone. Check the per-object results, then count the real objects, then look at the disk. The disk doesn't do politics.

That's the general truth-meter problem with this whole API, and it deserves its own paragraph. The API's progress field is per-thousand of the current step, not the job: I watched it climb past 1661 (that's 166%) mid-run and later saw it go backwards, 6771 to 5036, with no run change in between. There are no byte-level run stats anywhere in the API. The only honest number is file growth in the repository directory itself: a du delta on the repo host. Wire meter says one thing, the Director says another, and the disk says the truth. The disk always says the truth.

One more field anecdote because it's the kind of thing that only shows up in the field: stop a first full mid-flight and the repository shrinks about 11 GB, because the un-committed tail rolls back on stop. "Dedup keeps what's sent" only applies to data that made it into a completed savepoint. And switch the source machine's network mid-run and the run dies (the agent is a listener, it never calls home with a new address), then the restarted run re-indexes the rolled-back tail for several minutes while the wire already shows line rate. Plan network switches before the first full, not during it. I did not plan. You're welcome.

The NAS chapter: a repository that refused to exist

Next chapter: the Synology NAS as the repository target, over a private 10-gig link between tower and NAS. NAKIVO supports NFS-share repositories. Creating one through the API is a lesson in how much meaning hides in a field named "optional."

The DTO, assembled one rejection at a time: type NFS_SHARE, repositoryType: "NORMAL" (omit it and the error names exactly this field: a not-null violation, at least it's honest), path: "<host>:<export>" in plain host:path form (not an nfs:// URL), compression: "FAST" as a string (pass a boolean like a sane person and you get "parameter value is invalid" with no further commentary), dataStorageType: "FOREVER_INCREMENTAL", and a transporter ID.

Then the fun starts. The repository stuck at CREATE_PENDING forever. The Director's event log stayed generic. The real error turned out to live in the transporter host's log, not the Director's: Mount script failed (32), then Failed to initialize storage (71). First trap: the export was locked to the tower's private-link IP (correct security), but the repository was assigned to the onboard transporter, whose host is the Director VM, which has no path to that export. Nothing ever left the box. Repoint the repository to the transporter whose host can actually reach the export and re-read to verify it stuck.

Second trap was the good one. Root could mount the NFS export and write to it fine (Synology's export allows root writes). But every NFS repository NAKIVO created still looked broken. The reason: DSM's share ACL was d---------+, administrators only. The transporter's read/write probe doesn't run as root, it runs as uid 997, and uid 997 could not even traverse into the share. Root writes prove nothing about the product's actual probe. The fix was chmod 777 on the share root (fine here: the export is locked to one IP on an isolated link), plus deleting a stale NakivoBackup directory that the mount script treats as fatal (it exits 101 if the directory already exists), plus removing and re-creating the repository, because the repair function does not unstick CREATE_PENDING. It just doesn't.

While that fight was happening, Synology DSM 7.4 turned out to serve NFSv2/v3 only, so mount options got pinned to vers=3. That combination (NAKIVO's NFS handler, DSM 7.4's v3-only, share ACLs) was the single meanest knot in the whole project, and the honest summary is: the error that names the problem is always one host away from the error that gets shown.

Then, because the story needed it, the NAS pool went into a full RAID-6 rebuild (five 8 TB drives, resync measured at about 55 to 61 MB/s, hundreds of minutes on the clock). Standing rule I've adopted for exactly this situation: never run backup jobs against a NAS target whose pool is mid-resync. The rebuild window is single-failure-tolerance; you don't park your only safety net on top of it. Reads are cheap though, which fact becomes important in the final chapter.

Three failures at 0%, no error anywhere

With the repository finally healthy and the resync grinding along, I built the Proxmox backup job: three VMs, NAS repository, per-disk mappings wired. Fired it. FAILED at 0%. No event. No transporter dispatch. The transporter's own log had nothing but "Idle timeout." Ran it again. Same. A third time. Same.

This is the failure mode that makes undocumented APIs expensive: not the errors, the absence of errors. The RPC layer happily accepts a job it will later refuse to run, tells nobody, and leaves the transporter sipping coffee. If your only windows into the system are the API and the event log, you're stuck.

The unlock came from a change of altitude: I had SSH access to the Director's own VM (it's a small Ubuntu box I built myself, cloud-init, my key in it from day one). Its log directory carries the whole story. The run plan is logged action by action, and the failure reason is written where the UI never looks:

[JOB-2/TEST dummy-vm01 to NAS-BDR01]

Root cause, confirmed in the logs: the per-disk mappings I thought were wired were empty on the object the pre-flight actually inspected. Same rung as the physical job's rung 3, wearing a different hat. The mappings for Proxmox VM jobs use the VM's live disk inventory (PROXMOX_DISK-<n> with a sourceIdentifier like local-zfs:vm-101-disk-0), and empty mappings equal silent pre-flight rejection. Fix, re-run: activity completed clean. 463 files landed in the repository, 775 MB on disk for a 3.5 GB virtual machine after dedup and compression. First backup on the rebuilt NAS.

A silent null is not "no error." It is an error the API declines to report. When an API fails quietly, the truth is always in the logs on the box that runs the code. Get shell access to that box on day one, not as a last resort.

One more discovery while I was in there, because it made the wait productive: the run call accepts a sourceVids list with runType: "SPECIFIC", so you can run one machine out of a multi-VM job. That's how a 3.5 GB test VM got its first full while a 128 GB first full for the other two VMs waited behind the RAID rebuild. Undocumented, and it's exactly the knob you want when your repository host is mid-rebuild and you want to keep testing without stacking a real backup on a rebuilding pool.

The finale: booting a VM straight off the backup

The whole point of a BDR product is the recovery side. NAKIVO's instant recovery for hypervisor VMs is called Flash VM Boot, and the mental model is a linked clone off the backup datastore: the hypervisor registers a temp VM whose disk is served read-only out of the backup repository, writes go to a throwaway redo on the hypervisor's own storage, and the backup itself is never modified. Boot in seconds-to-minutes. Then you choose: keep it running, permanently commit it (live-migrate the disk to real storage), or discard.

For Proxmox, this feature is not a documented API call either. But the descriptor had a job type called FLASH_BOOT, and the recovery actions in the catalog had names like checkAndCreateRecoverySession. So: bisect ladder time. Every attempt that failed silently went to the Director's logs, every log line named the next missing piece, and four rungs got climbed in two evenings:

Rung 1: the four-slot target. The flash-boot job's container target is an array with exactly four slots: boot host, storage for the redo, network, resource pool. Send three slots and the rejection (server-side only, of course) says FeatureNotSupportedException: bad LU target count, got 3, 4 expected.

Rung 2: the poison field. The backup job uses targetVid on the object for its destination. Flash boot wants the boot host in a different field, targetLuVid, and if targetVid is present at all the save rejects silently. Two job types, same API, one field that flips meaning and one that turns fatal.

Rung 3: options can't be null. Omit the options block entirely and the run dies with JobOptionsDto ... getPostScriptExecutionMode() because "jobOptionsDto" is null. So the options object is mandatory, not decorative.

Rung 4: the one NPE that took the longest. With everything else filled, the run died one second in: JobOptions.getGenerateMac() is null. A boolean. The field that says "generate a new MAC address for the temp VM." The GUI wizard fills it in for you so quietly you'd never know it exists. The API leaves it null and the run path steps on it.

I stopped guessing at rung 4 and went to the source, literally: extracted the flash-boot code classes out of NAKIVO's own JAR on the Director VM and walked the constant pools for every JobOptions getter the Proxmox flash-boot path actually reads. Eight fields. I pre-filled all eight with valid values, enum strings mined the same way, and the next run went green on the first try.

Then this happened, and it's the part I keep re-reading: about 35 to 45 seconds after I fired the run, a temp VM appeared on the Proxmox host, powered on, booting from a backup that lives on a NAS whose RAID array was still rebuilding underneath it. VMID 103 on the tower. Its disk config pointed at a transporter-served block device: the backup repository serving blocks over the private link. Its writes landed on local NVMe. Its NIC was deliberately NOT_CONNECTED: a recovered copy never touches the production LAN while the original exists, that's the iron rule and the API let me build it in from the start.

Boot evidence without a console (the temp VM registers with a serial-only video config and no serial socket, and this particular image has no guest agent, so screenshot paths were all dead): the QEMU process pulled about 47 MB through the transporter disk at boot, kernel and initramfs and rootfs, then sat idle. The NAS private-link traffic went flat in steady state. The RAID resync didn't blink. The Director's own log declared the job state "RUNNING_VMS," which is its way of saying: VM's up, I'm watching it, your move.

Teardown is one call: stop the job and the temp VM is destroyed, changes discarded, backup untouched. That took seconds, verified the temp VM vanished, and closed the loop. A backup that restores to files is a file server. A backup you can boot in under a minute is a parachute.

The options, field by field

This is the section I wish existed when I started. The options block on a job is where the API's two personalities live, so here's the full walkthrough, from save-time to run-time.

Layer 1, save-time validation. The server validates a handful of fields at save and rejects nulls on exactly those. profilingEnabled is the famous one on backup jobs (null it and you get a property exception; false is the working value). The save layer is tolerant: most fields can be omitted and the server fills defaults. That tolerance is a trap in itself, because it trains you to send minimal payloads.

Layer 2, run-time reads. The run path reads a different set of fields, and for flash boot jobs the server defaults nothing. This is the asymmetry that cost the final day: a job that saves fine can die at run time on a field the save layer never asked about. The eight getters the Proxmox flash-boot path actually reads, mined from their bytecode:

FieldValues that workWhat it does
generateMactrue (Boolean, non-null)New random MAC for the temp VM. Null = the NPE that killed a whole evening. The wizard fills this silently.
powerVmsOntruePower the recovered copy on after registration.
namingTypeASIS / APPEND_AFTER / USE_FROM_REPLICA_INFOTemp VM naming. ASIS keeps the given name. Observed quirk: the temp VM registered under the source VM's name anyway.
vmVerificationTypeNONE / SCREENSHOT_VERIFICATION / BOOT_VERIFICATIONPost-boot verification stage. NONE for a controlled test.
encryptionModeNONE / NORMALDrives the internal use-backup-encryption flag; NONE keeps the encryption-password lookup out of the path.
transporterModeAUTO / SAN / LAN / HOT_ADDHow the data path picks its transporter. AUTO everywhere.
profilingEnabledfalseThe save-time load-bearing null trap.
preScriptExecutionMode / postScriptExecutionModeNEVER / ALWAYSPre/post job hooks. NEVER. The missing-options NPE names the post-script getter, which is how you learn these exist.

Two fields turned out to be harmless insurance: recoveryType (SYNTHETIC / PRODUCTION) and the script modes are not read at the object level on this build. I set them anyway and I'd do it again: pre-filling a field with a valid enum string costs nothing; discovering the next NPE costs a log round-trip.

Layer 3, the rule underneath it all. The enum strings must match the product's own enums exactly, and those live in the product's JAR files, not in any documentation. Every enum is an inner class with ALL_CAPS values, extractable with a 30-line Python zipfile pass over the constant pool. NAMING_TYPE {ASIS, APPEND_AFTER, USE_FROM_REPLICA_INFO}, VM_VERIFICATION_TYPE {NONE, SCREENSHOT_VERIFICATION, BOOT_VERIFICATION}, and so on. When you send a string the server doesn't recognize, you get the generic "parameter value is invalid," which is the API's way of saying "close, but I'm not going to tell you the list."

What I'd tell someone else

  • The API is the UI's own chatter. If there's no docs page for what you need, open the browser dev tools, click the thing once, and read the request. Then find the descriptor at /c/api and you have the full surface with argument counts.
  • Every error string names the next thing to mine. empty value in X = add field X. The parameter value is invalid = your top-level shape is wrong. An NPE naming SomeDto.getFoo() = fill exactly that field. The errors are terse, not useless.
  • When it fails with no error at all, stop poking the API. SSH into the box that runs the product and read its logs. Run plans, stack traces, and real rejection reasons live there. This one habit was the difference between three silent failures and a solved bug.
  • Mine the client, not the docs. The UI's JavaScript bundle contains the exact call shapes the product itself uses (including the positional-argument quirks the descriptor hides). The JARs contain every DTO field and enum value. Both are on the box you already have shell access to.
  • If the API won't give you bytes, count them yourself. Progress percentages lie by design here (per-step, phase-weighted, occasionally backwards). A du delta on the repository directory is the only number that has never lied to me.
  • Never trust a run flag alone. Per-object results, real object count, disk growth: in that order. Zombie objects stamp completed runs FAILED and disabled jobs sit in silence.
  • Version-pin everything you learn. The descriptor shifts between releases, and behavior shifted under me between weeks on the same version. Re-pull /c/api on every bump and re-verify one known call before trusting anything else.

Still on the bench

The saga isn't over. Still queued:

  • Flash boot across all VMs, not just the 3.5 GB test dummy (the Director itself is one of them: booting it means a second Director instance, so it stays unplugged like everything else).
  • The remaining first-fulls for the production VMs, gated behind the RAID-6 resync finishing.
  • Live migration test: the keep-it-running path after flash boot (permanently committing the recovered copy by migrating its disk to real storage). Queued for the end.
  • The bare-metal-restore (physical machine) side is its own saga with its own one-flag fix, already written up: a restore that took a day and ended in a single QEMU setting.
  • Unsolved curiosities, documented for the next person: the support-bundle API call rejects an empty "command" argument and I never found the right one; the H2 database on the Director is AES-encrypted with a key-derivation I haven't cracked (didn't need to, the logs gave up everything).

This one stays open. When we find more information, we'll update this article with it.

Test environment: HP Z6 G4 tower with Proxmox VE 9.2.2 (ZFS root mirror, boot from NVMe), NAKIVO Backup & Replication v11.2.2 trial running on an Ubuntu 24.04 VM (2 vCPU, 8 GB RAM) on that tower, Synology DS1823xs+ on DSM 7.4 with 5x8 TB in RAID 6 serving NFSv3 over a private link, test VMs from Ubuntu cloud images. September to October 2026. All backup runs during this story were test VMs on a blank repository target: zero production data exposure, and the production first-fulls stayed gated behind the RAID rebuild on purpose.

NAKIVO Never Documented Their API: How I Automated a Whole BDR Stack Anyway | Tech Connect Arizona