PVE: A Storage Deep Dive – LVMThin vs ZFS vs EXT4

Testing different Storage Architectures against a cheap consumer SSD

When I first spun up my virtualization host, I knew that storage configuration would make or break my performance. But nothing prepared me for just how wildly my benchmarks would swing based purely on software definitions. Using the exact same underlying physical SSD, I ran a series of tests to uncover how different virtual controllers, file systems, and storage architectures handle real-world disk operations.
Here is the raw data from my benchmarking journey, followed by a breakdown of what I learned, and a definitive pros and cons guide for every setup.

📊 The Cumulative Benchmarks: My Raw Results

To keep the comparisons fair, I tracked sequential and random operations across four distinct iterations of my storage stack.
1. The Baseline: LVMthin with a Virtual SATA Controller (64 MiB Test)
My journey started here, and the results were a painful wake-up call to the limits of legacy hardware emulation.
[Read]  SEQ Q8T1:  429.221 MB/s  |  SEQ Q1T1:  318.640 MB/s  |  RND Q32T1:  13.760 MB/s  |  RND Q1T1:   5.507 MB/s
[Write] SEQ Q8T1:   44.661 MB/s  |  SEQ Q1T1:   16.980 MB/s  |  RND Q32T1:   2.382 MB/s  |  RND Q1T1:   1.655 MB/s
2. The Protocol Flip: LVMthin upgraded to VirtIO SCSI Single (64 MiB Test)
By changing a single dropdown menu in my VM settings from SATA to VirtIO SCSI, I unbottlenecked the communication protocol.
[Read]  SEQ Q8T1:  493.329 MB/s  |  SEQ Q1T1:  383.567 MB/s  |  RND Q32T1:  88.362 MB/s  |  RND Q1T1:  12.480 MB/s
[Write] SEQ Q8T1:   90.589 MB/s  |  SEQ Q1T1:   21.391 MB/s  |  RND Q32T1:   2.403 MB/s  |  RND Q1T1:   1.879 MB/s
3. The RAM Illusion: ZFS with VirtIO SCSI Single (Cached 64 MiB Test)
Next, I swapped my backend volume manager to ZFS. With a tiny test file size, I accidentally benchmarked the sheer speed of my host’s RAM via the ZFS ARC cache.
[Read]  SEQ Q8T1: 4401.052 MB/s  |  SEQ Q1T1: 1921.489 MB/s  |  RND Q32T1:  84.354 MB/s  |  RND Q1T1:  19.714 MB/s
[Write] SEQ Q8T1:  205.138 MB/s  |  SEQ Q1T1: 1791.981 MB/s  |  RND Q32T1:  68.788 MB/s  |  RND Q1T1:  17.412 MB/s

4. The True Hardware Limit: ZFS with VirtIO SCSI Single (Raw 8 GiB Test)
By cranking the test file size up to 8 GiB, I forced the data straight through the RAM buffers and onto the raw NAND flash chips. This represents my real physical drive performance under ZFS.
[Read]  SEQ Q8T1:  739.249 MB/s  |  SEQ Q1T1:  715.466 MB/s  |  RND Q32T1:  77.281 MB/s  |  RND Q1T1:  17.737 MB/s
[Write] SEQ Q8T1:  382.726 MB/s  |  SEQ Q1T1:  219.378 MB/s  |  RND Q32T1:  49.969 MB/s  |  RND Q1T1:  15.344 MB/s
5. The Image File Contender: qcow2 on an EXT4 Host File System (64 MiB Test)
Finally, I tested the classic approach—placing a qcow2 virtual disk file on top of a standard Linux EXT4 partition.
[Read]  SEQ Q8T1:  189.151 MB/s  |  SEQ Q1T1:  154.374 MB/s  |  RND Q32T1:  41.495 MB/s  |  RND Q1T1:  10.420 MB/s
[Write] SEQ Q8T1:  124.732 MB/s  |  SEQ Q1T1:  115.363 MB/s  |  RND Q32T1:   2.419 MB/s  |  RND Q1T1:   1.726 MB/s

🔍 Structural Deep Dive: Pros, Cons, and My Key Architectural Takeaways

Every storage backend has a specific target audience. Based on my benchmarking data and the structural architecture of these systems, here is how they stack up.

Setup A: LVMthin (Block Storage)

LVMthin acts as a pure logical volume manager natively inside the Linux kernel. It gives the VM direct block-level access to the storage layer without an intervening file system.
    • The VirtIO SCSI Upgrade Difference: My data showed a 542% explosion in random reads (from 13.76 MB/s to 88.36 MB/s) just by moving away from the virtual SATA controller. SATA forced my VM into a shallow, serialized hardware queue. Upgrading to VirtIO SCSI unlocked deep parallel queue depths (Q=32), passing commands straight down to the kernel block layer efficiently.

Pros
    • Ultra-Low Overhead: It consumes almost zero host RAM and negligible CPU power.
    • Zero Write Amplification: Blocks are written directly to the SSD, maximizing the physical lifespan (TBW) of consumer drives.

Cons
    • No Cache Cushioning: Because it lacks an aggressive RAM cache layer, your performance is entirely capped by the physical drive hardware. My random writes hit a miserable 2.4 MB/s wall in both tests because the raw SSD couldn’t keep up with heavy random write thrashing.
    • Feature Light: It lacks advanced data integrity features like inline file self-healing.


Setup B: ZFS (Advanced Copy-on-Write File System)

ZFS merges the file system and volume manager into a single entity. It uses a Copy-on-Write (CoW) design, meaning existing data is never overwritten; modifications are written to entirely new blocks, and metadata addresses are updated afterward.
    • The RAM vs. Real Disk Difference: In my 64 KiB test, ZFS hid my operations completely inside the Adaptive Replacement Cache (ARC) and transaction memory, showing an insane 4.4 GB/s read rate. Once I used an 8 GiB file to smash past the RAM barrier, I saw my true physical drive speeds (~740 MB/s Read / ~380 MB/s Write).
    • The Random Write Fix: Remarkably, ZFS managed to boost my raw un-cached random writes up to 49.97 MB/s (compared to LVMthin’s 2.4 MB/s). This is because ZFS groups messy random writes together in memory and flushes them to the physical disk as a clean, orderly sequential stream.

Pros
    • Unmatched Burst Speed: Turns unused host RAM into an ultra-fast read/write cache buffer.
    • Enterprise Features: Offers built-in data checksums to prevent silent data corruption (“bit-rot”), easy live replication, and zero-latency snapshots.

Cons
    • Resource Hungry: ZFS expects a steep memory tax (often quoted as 1GB of RAM per 1TB of storage) to keep its ARC cache happy.
    • The SSD Lifespan Killer: The Copy-on-Write mechanism causes heavy write amplification. On consumer-grade SSDs without power loss protection, ZFS constantly flushes transaction metadata, which can burn through the drive’s endurance metrics in a fraction of its normal lifespan.


Setup C: qcow2 on EXT4 (File-on-File Storage)

This architecture represents a double-layer virtualization penalty. The hypervisor creates a structured image file (.qcow2), which handles its own internal allocations, snapshots, and clusters, and then places that file on top of a standard host Linux file system like EXT4.
    • The Double-Abstraction Penalty: My benchmarks reveal that this setup struggled heavily with throughput. Sequential reads completely tanked out to just 189 MB/s, and sequential writes crawled at 124 MB/s. Because every single block request from the VM has to be translated twice—first by the qcow2 file allocation table and then by the host’s EXT4 file tracker—latency skyrocketed to 44–66ms, severely choking sequential bandwidth.

Pros
    • Extreme Portability: The entire VM drive is just a single file on the host. You can easily copy, move, back up, or email a .qcow2 file across systems.
    • Thin-Provisioning Simplicity: The file grows dynamically on disk automatically as the VM consumes space, making storage management incredibly intuitive.

Cons
    • Severe Throughput Bottlenecks: Easily the slowest sequential performer of the group due to file system translation overhead.
    • Terrible Random Writes: Like LVMthin, it lacked the intelligent memory-grouping mechanics of ZFS, collapsing down to a sluggish 2.4 MB/s on parallel random 4KiB writes.


⚠️ Important Architectural Caveats & Exceptions

Storage performance does not exist in a vacuum. Before choosing a design based on my charts, you must consider your underlying host hardware infrastructure.
1. Hardware RAID Controllers vs. ZFS (The Golden Rule)
If your server utilizes a physical Hardware RAID controller (e.g., a Dell PERC or HPE Smart Array), do not use ZFS.
    • Why: ZFS requires total, direct control over the physical disk topology to manage cache flushing and data integrity.
    • The Conflict: If you place ZFS on top of a hardware RAID logical volume, the hardware controller will hide the physical disk health from ZFS. Even worse, the controller’s onboard cache will conflict with the ZFS ARC, severely increasing the risk of cataclysmic data corruption during an unexpected power loss.
    • The Winner: For Hardware RAID arrays, LVMthin or EXT4 are the proper structural choices.

2. Battery-Backed Write Cache (BBWC) Controllers
If your host server features an enterprise RAID card paired with a physical Battery Backup Unit (BBU) or supercapacitor flash cache, the playing field changes entirely.
    • Why: A BBWC allows the host to immediately report a write operation as “complete” the millisecond it hits the controller’s physical onboard RAM cache, safely knowing that if the power cuts out, the battery will preserve the data until reboot.
    • The Impact: If I had a BBWC under my LVMthin or qcow2 tests, that terrible 2.4 MB/s random write bottleneck would have completely vanished, skyrocketing up into hundreds of megabytes per second. It essentially provides ZFS-like memory caching speeds without the software complexity or RAM overhead.


🏁 Final Verdict: Which One Should You Build?

    • Choose ZFS if you have plenty of host RAM to spare, you want elite enterprise data safety features, you rely on rapid multi-node backup replication, and your storage drives are connected to an HBA (Host Bus Adapter) in IT/Passthrough mode.
    • Choose LVMthin if you are using consumer-grade hardware or mini-PCs, you want to maximize the physical lifespan of your SSDs, you need every single drop of host RAM allocated to running your actual VMs, or you are deploying on top of a Hardware RAID card.
    • Choose qcow2 on EXT4 if absolute peak performance takes a back seat to sheer simplicity, storage portability, and easy file-level drive movement between isolated hypervisors.

I’m still trying to figureout the root cause of Veeam restore being unreal poor on a LVMThin target. Stay tuned!

Migrating Veeam as a VMware VM to a Proxmox VM

Running Veeam as a Windows VM on Proxmox VE

With Broadcom completely upending the virtualization landscape, like many admins, I’ve been heavily evaluating Proxmox VE (PVE) as a viable enterprise alternative. While Veeam now natively supports PVE, I wanted to see what happens when you run the Veeam Backup & Replication (VBR) server as a Windows VM hosted directly on a PVE node, rather than on dedicated hardware.
Here is the technical reality of this architecture, the hypervisor mechanics, and the exact step-by-step process I used to execute a zero-duplication migration from ESXi into a tight local LVM-Thin storage layout.

1. The Core Architecture: How VBR Operates Inside PVE

When I deployed VBR into a Windows VM on Proxmox, the fundamental architecture shifted away from what I was used to in the VMware world:
  • The VBR VM is just the brain: The Windows VM hosting Veeam acts strictly as the management plane, controlling schedules, maintaining the PostgreSQL database, and exposing the console.
    – Works the same way as a VMware VM, it just can’t act as it’s own VMware hotadd node when on PVE VM
  • The PVE Worker handles the heavy lifting: Because PVE lacks a monolithic management layer like vCenter, Veeam requires me to add each individual node in the cluster to the console. Veeam then provisions a lightweight, Linux-based Worker VM onto each host to act as a stateless data mover.
  • The Control vs. Data Paths: When a job triggers, my Windows VBR VM talks directly to the Proxmox API to initiate a hypervisor-level snapshot. Once the snapshot is ready, VBR instructs the local Worker VM to mount that snapshot, compress the data blocks, and stream them directly to my backup repository.
                  ┌─── [ Veeam VBR Server (Windows) ] ───┐
                  │                                      │
                  ▼ (1. Orders Snapshot via API)         ▼ (3. Orders Data Processing)
       [ Proxmox VE Host ]                       [ Veeam PVE Worker VM ]
                  │                                      │
                  ▼ (2. Triggers Quiescing)              │
       [ Target VM + QEMU Agent ]                        │
                  │                                      │
                  └─────── (4. Reads Snapshot Data) ─────┘

2. Technical Friction Points & Lessons Learned

While the setup works, moving from ESXi to PVE exposed several architectural limitations and maturity gaps that I had to account for.
The Infuriating VLAN Blind Spot in the Worker Wizard
This is one of the clearest signs that Veeam’s Proxmox integration is still green. When deploying the Linux-based Worker VMs across my cluster nodes, the Veeam Deployment Wizard lacks any field to define a VLAN ID. It lets me select the network bridge (e.g., vmbr0), but that’s it.

  • The Hack: Because my management network is VLAN-tagged, the newly deployed Worker VMs immediately drop offline upon creation because they can’t get an IP address. I have to manually jump into the Proxmox GUI, open the hardware settings of the newly provisioned Worker VM, inject the VLAN tag into its virtual NIC, and reboot it.
  • The Catch: Every time Veeam pushes an update that redeploys or patches these Worker VMs, this manual configuration gets wiped out. It requires immediate administrative intervention in the PVE GUI to fix the network interfaces after every lifecycle update.

3. The Migration Strategy: Zero-Duplication Streaming to LVM-Thin

Moving my Veeam server presented a unique chicken-and-egg problem. To make matters worse, my old ESXi lab layout was messy: my Veeam backup repository wasn’t an independent iSCSI target—it was a raw, bloated 1.2 TiB VMDK sitting directly on local mechanical spindle datastores.
To avoid the pain of building a fresh Windows OS, configuring IPs, and rebuilding my environment from scratch, I chose to use the Self-Backup and Restore method. I backed up just the C: drive of the Veeam VM, restored it to Proxmox, and brought the compute layer online.
However, migrating that massive data disk into a tight, local LVM-Thin pool presented a storage Catch-22: LVM-Thin does not use files like VMDK or QCOW2. It is a block-level storage architecture. I couldn’t just copy the file over first because my PVE host had no local path large enough to store a temporary 1.2 TiB file.
To bypass this hurdle and ensure zero file duplication, I streamed the conversion directly over the network from the ESXi host using the command line. Which involves adding the ESXi host to PVE, then using PVEs CLI.

Phase 1: Pre-Migration Trim (Dropping the Bloat)

Inside the Windows Veeam VM, the OS reported only 600GB of actual data in use, meaning I had over 200GB of “ghost bloat” on the 808GB allocated ESXi VMDK. Before shutting down the VM, I ran a quick optimization check to ensure Windows explicitly issued zero-discard commands to the hypervisor:
powershell
Optimize-Volume -DriveLetter D -Defrag -Verbose
 Optimize-Volume -DriveLetter D -ReTrim -Verbose
(After trimming, my actual active file space dropped to exactly 519 GB).

Phase 2: Live Network Streaming via qm importdisk

When you add an ESXi storage provider to Proxmox, PVE transparently maps the remote files to a hidden local runtime directory on the PVE host. This allows you to target the file directly with native Proxmox CLI utilities without copying anything locally.
Add ESXi Storage:

  • Log in to the Proxmox Web UI.
  • Go to Datacenter > Storage > Add and select ESXi.
  • Enter a unique ID, your ESXi host Server IP/hostname, and your ESXi administrator Username and Password.
  • Check Skip Certificate Verification if your ESXi node uses a self-signed certificate.
  • Click Add.
I created a dummy placeholder virtual disk shell on my newly restored Proxmox Veeam VM (let’s assume VM ID 108), hopped into the Proxmox CLI, and initiated a fileless stream:
qm importdisk 108 /run/pve/import/esxi/desktopesxi/mnt/ha-datacenter/Desktop-2TB/Veeam/Veeam.vmdk msiNM620

Beating the Scary LVM-Thin Warnings
During the transfer, Proxmox threw a terrifying set of warnings:
WARNING: Sum of all thin volume sizes (<1.53 TiB) exceeds the size of thin pool nvme-nm620/nvme-nm620 and the size of whole volume group (<953.87 GiB).
Because the 1.2 TiB virtual boundary of the VMDK was larger than my physical 953 GiB NVMe drive, LVM-Thin immediately flagged that I was overprovisioning.
However, because we trimmed the drive beforehand, the qm importdisk zero-detection engine kicked in. As the empty space traveled over the LAN, the engine detected the logical zeros and completely dropped them instead of writing them to disk.
The progress bar crept through the full 1.2 TiB virtual structure, but it stopped consuming physical blocks when it hit empty space. It finished flawlessly at 100%, consuming an exact physical footprint of 518.88 GB on my NVMe pool—matching my source data perfectly without a single byte of temporary file bloat.

4. The Finish Line

To wrap up the migration, I executed the final hookups:
  1. Attach the Block Storage: Inside the PVE GUI for VM 108, I went to Hardware, double-clicked the newly created Unused Disk, set it to SCSI, and checked the Discard box (ensuring that future Veeam deletions will actively reclaim space on my LVM-Thin pool).
  2. Mount the ISO & Fix the NIC: Because the restored Windows VM lacked KVM network drivers, it booted up offline. I used scp to copy the virtio-win.iso from my other cluster node over to this host’s local directory (/var/lib/vz/template/iso/), mounted it to the VM’s virtual CD drive, opened Device Manager, and updated the Ethernet Controller driver.
  3. Rescan & Run: I brought the data disk Online in Windows Disk Management (which instantly recognized the original NTFS/ReFS structure perfectly), opened the Veeam console, and ran a Repository Rescan.
Within minutes, Veeam mapped my existing backup chains, and my entire infrastructure was fully functional on Proxmox VE with zero data loss and zero configuration rebuilding.
This was nice but I didn’t move the backup data to a dedicated block level storage, did I… No… will I… mhmmm maybe eventually, but I’m surprised how well the conversation went, another step to a VMware -> PVE migration. Also, the backup jobs moved from hotadd to NBD cause the Veeam server is not longer on a ESXi host. And you can’t map a backup set when re-adding a backup job sourcing a PVE host.

BONUS

When I tried running a Full VM Restore from my VMware backups directly onto an LVMthin pool on my high-speed SSD. The restore immediately cratered to a miserable 3 MB/s, pinning my host at a massive iowait percentage. Same problem in my previous blog post when I migrated my email server.
When I cancelled the job in Veeam, the console said “Failed,” but the helper VM refused to die, and the host remained completely paralyzed. Even running a basic Linux sync command via SSH completely hung and stole my terminal shell.
Here is the juicy reality of what happened under the hood:
  • The Block-Level Choke: Veeam streams raw, unallocated data blocks during standard restores. On an un-provisioned LVMthin target, this forces the Proxmox host to perform real-time metadata lookups and extend block structures on every single write loop. This created a massive write-amplification loop that totally choked the SSD controller queue.

  • The Unkillable “D-State”: Hard-killing the VM mid-stream locked the dm-thin kernel driver into an uninterruptible sleep state (Linux D-state). Because the storage layer was frozen waiting on hardware, the kernel completely ignored standard ACPI shutdown signals.
How I broken it loose without pulling the physical power cord:
With the storage layer totally unresponsive and sync frozen, standard reboot commands were completely useless. I had to open a secondary SSH shell and talk directly to the Linux kernel using aggressive, low-level emergency triggers.
First, I enabled the kernel’s magic SysRq interface:
echo 1 > /proc/sys/kernel/sysrq
Then, I sent the emergency hardware reboot signal (b):
bash
echo b > /proc/sysrq-trigger

This acted as a software-level reset button, forcing the motherboard to instantly reboot without trying to unmount or flush the corrupted, frozen storage queues. On startup, the LVMthin subsystem automatically repaired its metadata boundaries and the host came back up healthy.
I found restoring to the same SSD configured as a DIR Ext4 storage had to problem and restored the VM in 5 minutes. So, not sure how to properly handle a Veeam restore to a LVM-thin without having this problem.

Proxmox Cluster Findings

Proxmox Cluster Findings

When I set up my new MSI PVE host, I decided to use the “Golden Blueprint” for storage: installing the base Proxmox OS on a cheap SATA SSD to shield it from heavy log writes, leaving my lightning-fast NVMe drive completely clean and unpartitioned as a dedicated LVM-Thin block storage pool for my VMs and containers.
But things got complicated when I tried to link my new host to my old host. Here is what I learned while wrestling with Proxmox cluster logic, VMware design differences, and hardcoded network traps.

1. The Populated Host Conundrum

My initial plan was straightforward: create a fresh cluster on my pristine new host, and then have my old host (which currently runs all my active workloads) join it. I quickly hit a hard wall. Unlike VMware vCenter—which effortlessly imports populated ESXi hosts and reindexes VM identifiers on the fly—Proxmox operates on a decentralized, shared-filesystem architecture (pmxcfs).
Because Proxmox uses hardcoded VMIDs (like 100, 101) to define physical storage paths and configuration maps, a host cannot join an existing cluster if it has any virtual machines on it. Joining a cluster completely overwrites the local /etc/pve directory, which would instantly orphan any existing virtual disks.
The Fix: I flipped the workflow. I created the cluster directly on my old host first. Because it initialized the cluster, its existing VMs remained perfectly safe.

2. Leaving a Cluster Requires the CLI

Before I figured out the correct order of operations, I had already initialized a test cluster on my new node. When I went to undo it, I discovered there is no “Delete Cluster” button in the Proxmox Web GUI. Because tearing down a cluster can cause severe data corruption if done incorrectly, Proxmox forces you into the shell.
To completely reset my new node back to standalone mode without reinstalling the entire OS, I had to open the terminal and force-clear the configuration database using these steps:
# Stop cluster synchronization services
systemctl stop pve-cluster
systemctl stop corosync

# Force the configuration filesystem into a temporary local mode
pmxcfs -l

# Permanently erase the old cluster registries
rm -f /etc/pve/corosync.conf
rm -rf /etc/corosync/*

# Kill the background lock process and restart standalone services
killall pmxcfs
systemctl start pve-cluster

3. The Hardcoded IP/VLAN Trap

While running pmxcfs -l, I noticed the terminal spat out an old IP address (192.168.0.68) that I had previously removed when migrating the host to a new VLAN-tagged management network.
I realized that Proxmox doesn’t query active network interfaces to resolve its own name; it relies entirely on a static registry. The old IP address was still hardcoded inside my /etc/hosts file. I used nano /etc/hosts to update the line to my new VLAN IP (172.16.21.20) so the host could resolve itself properly. (Blog post covering IP change updated)

4. Breaking the Handshake Hang

When I finally went to join my clean new node (172.16.21.20) to the old node’s cluster (172.16.21.60), the Web GUI installer completely hung on the message: “Request addition of this node.”
However, in classic homelab fashion, a simple web browser refresh cleared the stalled API cache, the handshake successfully completed on its own, and both nodes cleanly populated into a single sidebar.

5. Local Disks Show Up on the Wrong Nodes

Right after my two nodes successfully clustered, I noticed something deeply alarming in the sidebar UI: my old host’s massive local storage drive (mass-storage) was suddenly showing up underneath my shiny new msi-pve node.
I knew for a fact it wasn’t shared storage—there were no NFS, SMB, or iSCSI links connecting them. The drive was physically plugged into the old hardware. So why was my new host pretending it owned it?
The Cause: Global Cluster Definitions
This is one of Proxmox’s quirky design traits. Proxmox stores every single storage path inside a single, cluster-wide configuration file (/etc/pve/storage.cfg). The moment my new node joined the cluster, it downloaded this file and blindly copied the layout.
By default, Proxmox assumes any storage listed in that file is accessible by all nodes unless you explicitly state otherwise. It draws the icon under every host in the sidebar, creating a dangerous phantom placeholder. If I had tried to spin up a VM on my new node and targeted that ghost storage, the deployment would have crashed instantly with activation errors.
Luckily the fix, correcting this is incredibly simple and doesn’t require the command line:
    1. I clicked on Datacenter at the very top of the sidebar.
    2. I went to Storage, highlighted the phantom mass-storage pool, and clicked Edit.
    3. I found the Nodes dropdown—which was completely blank (Proxmox-speak for “Allow All Nodes”).
    4. I changed it to explicitly select only my old host (g9-pve) and saved it.
The second I clicked OK, the phantom icon vanished from underneath my new node, keeping my environment clean and preventing any catastrophic accidental deployments.

6. Migration Failure due to VMs bound virtual NIC

Coming from VMware, hitting a hard wall during a migration because of a network name mismatch feels entirely unnecessary. In VMware—even on standard vSwitches without distributed virtual switching—the migration wizard natively handles network remapping. If a target host doesn’t have a matching network, the wizard stops and lets you choose a new path on the fly.
This highlights a fundamental architectural difference in how these two hypervisors handle networking:
  • VMware uses Port Groups (VMPGs): This creates a layer of abstraction. The VM connects to a named Port Group, and that group handles the VLAN tagging down at the vSwitch level.
  • Proxmox uses Direct Bridging: By default, Proxmox expects you to point the VM’s virtual NIC directly at a specific host bridge (like vmbr1) and explicitly type the VLAN tag directly into the VM’s network device settings.
Because Proxmox binds the VM directly to a specific host bridge string rather than an abstract port group, its migration wizard is completely rigid. If the destination host doesn’t have an identically named bridge, the migration fails with an error instead of letting you remap it during the transfer. To make migrations seamless, you have no choice but to ensure your Linux bridge names match exactly across every single node in your cluster.

7. Migrating an LXC

With a container healthy, I attempted to clone it over to the new host’s NVMe drive. The Proxmox UI threw a sudden error:

Full clone of a running container is only possible from a snapshot (500)
Because containers share live kernel space, they cannot be live-cloned without a frozen snapshot boundary. I attempted to take a snapshot, only to hit a secondary brick wall:

The current guest configuration does not support taking new snapshots
This error points directly to the underlying physical storage. While advanced file systems like ZFS and LVM-Thin support snapshots out of the box, traditional file-based directory storages and thick LVM volume groups do not. I couldn’t take a snapshot, which meant I couldn’t live-clone. A cold migration was my only path forward, or so I thought…

Overriding the Migration Wizard with Manual Backup Streams

I shut the container down and clicked Migrate, but Proxmox aborted yet again:

ERROR: migration aborted: storage 'Mass-Storage' is not available on node 'msi-pve'
When dealing with local storage migration, the standard Proxmox wizard stubbornly expects the exact same storage pool name to exist on the target host. Because my old host’s 4TB drive was named Mass-Storage, and my new host only possessed an unmapped LVM-Thin layout, the handshake failed.
To bypass the rigid migration UI entirely, I resorted to a manual Secure Copy (SCP) backup stream. First, I shut down the source container on g9-pve and took a standard uncompressed cold backup file to my local directory. Then, because LVM block devices don’t show up via standard df -h folders, I targeted Proxmox’s universal root backup directory (/var/lib/vz/dump/) and pushed the data over the network via SCP:
bash
scp /mnt/pve/mass-storage/dump/vzdump-lxc-105-2026_09_21-19_57_35.tar.zst root@172.16.21.20:/var/lib/vz/dump/

The Final Handshake

Once the transfer finished, the backup file appeared beautifully under the new host’s local storage tab in the Web UI. I hit Restore, pointed the destination target directly into the brand-new 1TB NVMe LVM-Thin pool, and let it extract.
To guarantee no IP collisions occurred on my network while validating the data, I went into the original source container on g9-pve and checked the “Disconnected” flag on its network interface. With the original container safely blinded, I fired up the newly restored container on msi-pve. I adjusted its hardware bridge settings from vmbr1 to the new host’s active vmbr0 bridge, and the pings immediately started flowing.

The service came up perfectly healthy, allowing me to finally go back and permanently shut down the original source container on the old hardware.

8. Proxmox “Cluster Join Aborted – No Quorum” Errors

When adding a new node to an existing Proxmox VE cluster, you might encounter this frustrating terminal error:
An error occurred on the cluster node: cluster not ready - no quorum? TASK ERROR: Cluster join aborted!

Here is exactly why this happens and how to bypass it safely.

🔍 The Root Cause
Proxmox requires a strict majority (>50%) of voting nodes to be online and communicating via Corosync to allow administrative changes.
    • If a node in your existing cluster is offline, powered down, or network-isolated, the remaining online nodes lose quorum.
    • To prevent accidental data corruption, Proxmox drops into a protective, read-only state and blocks any new nodes from joining.

🛠️ The Fix: Bypassing Quorum
If you cannot easily boot the offline node back up, you can manually force the surviving node to lower its expected vote count.
Step 1: Force Expected Votes
SSH into your active cluster node and run the following command to temporarily lower the expected vote threshold to 1:
pvecm expected 1
This instantly restores cluster write-privileges to the remaining online node.
Step 2: Join the New Node
Return to your new node (via CLI or the Web GUI) and re-run the cluster join process. The join operation will now succeed.
Step 3: Boot the Offline Node
Once the new node has safely integrated into the cluster, power your original offline node back on. It will automatically check in with the rest of the network, detect the new node, and sync the updated cluster configuration seamlessly.
⚠️ Important Safety Checklist
Before running pvecm expected 1, always verify your storage layout:
  • Shared Storage (Ceph, NFS, iSCSI): Ensure the offline node is completely powered off. If it boots up isolated, it could cause a “split-brain” scenario and corrupt shared virtual disks.
  • Local Storage Only: If your cluster strictly uses local disks, the risk is near zero. The offline node cannot physically touch or corrupt data on the other hosts.

If you do Bring back your old host and didn’t remove it from the cluster, it may not be happy with you and remain in its own island cluster and not play nice with your old hosts…

9. Proxmox Cluster Split-Brain and Corosync Version Mismatch

When a Proxmox VE node is taken offline and changes are made to the remaining cluster (such as adding a new host), bringing that offline node back online can trigger a silent cluster failure. The node will refuse to rejoin, and the Proxmox Web UI will show the hosts as completely disconnected from one another.

Here is a breakdown of why this happens and how to safely force the node back into the cluster.

The Problem: The “Chicken and Egg” Loop
Proxmox relies on two core components to maintain a cluster:
    1. Corosync: The underlying network messaging engine that handles cluster quorum and node handshakes.
    2. pmxcfs (Proxmox Cluster File System): The distributed filesystem that mirrors configuration files (like /etc/pve/corosync.conf) across all nodes.

This creates a dependency loop: pmxcfs needs Corosync to sync files across the network, but Corosync needs to read pmxcfs to get its startup configuration.
If a node is offline when the active cluster undergoes an update, the cluster increments its configuration version (e.g., from config_version: 2 to 3). When the offline node boots back up, a security mechanism triggers:
    1. Corosync reads its local file and sees Version 2.
    2. It reaches across the network and hears Version 3 from the active cluster.
    3. To prevent a “split-brain” scenario (where two isolated parts of a cluster might blindly overwrite each other’s storage configs), Corosync panics and immediately exits with the error:
      [CMAP] Received config version (3) is different than my config version (2)! Exiting

Because Corosync crashes, the network pipe never initializes. Because the pipe never initializes, the node can never automatically download the updated file from the healthy cluster.

The Fix: Trick Corosync to Open the Pipe
To resolve this, you must force the local node’s configuration filesystem into read/write mode, manually spoof the configuration version number to match the cluster, and restart the services.
Step 1: Kill Zombie Processes & Force Local Mode
Because the cluster is offline, Proxmox locks /etc/pve into a read-only state. Clear out any stuck systemd loops and force it to run locally:
# Kill any hung cluster processes holding filesystem locks
killall -9 pmxcfs corosync pve-cluster

# Stop the systemd tracking services
systemctl stop pve-cluster corosync

# Force the cluster filesystem to run in independent local mode
pmxcfs -l
Step 2: Manually Update the Config Version
With the filesystem unlocked, edit the configuration file to match the version expected by the active cluster:
nano /etc/pve/corosync.conf
    • locate the config_version: line in the file.
    • Change the number to match the cluster’s current version (e.g., change 2 to 3).
    • Save and exit (Ctrl+O, Enter, Ctrl+X).

Step 3: Restart Services
Now that the version numbers match on paper, kill your manual filesystem override and let systemd start the cluster services normally:
# Clear the manual local mode instance
killall -9 pmxcfs

# Start the cluster services sequentially
systemctl start pve-cluster
systemctl start corosync
The Result: The True Config Wins
By manually matching the version numbers, Corosync is tricked into thinking the nodes are synchronized, allowing it to safely open the network tunnel.
The moment the connection stabilizes, Proxmox’s internal generation rules take over. It recognizes that the active cluster holds the true, updated configuration file. The cluster immediately pushes the complete configuration down to the returning node, gracefully overwriting your temporary edit and cleanly reintegrating the host.
You can verify the fix directly from the CLI:
pvecm status

If the above didn’t work cause your host was offline when you brought this one back online, and that host already has the same version number, but get connection issues in the GUI check out the next section.

10. Wiping Stale State and Bypassing SSL Mismatches

When a Proxmox VE cluster shifts structurally while a specific node is offline (such as adding new nodes or changing quorum dynamics), that offline node can fall too far out of sync. Even if configuration version numbers match, the node may completely fail to authenticate with the live cluster due to outdated cryptographic tokens or security states.
Here is how to safely wipe a node’s stale cluster data, preserve existing virtual machines, and force-join it back into production using an SSH fallback.
The Problem: Authentication Locks & The SSL 500 Error
When a node falls severely behind the cluster’s state, trying to rejoin it often leads to a dead end. To fix this, you must first reset the node back into a standalone host by purging its old configuration database files:
# Stop the stuck cluster loops
systemctl stop pve-cluster corosync
killall -9 pmxcfs corosync

# Force the configuration filesystem to open up locally
pmxcfs -l

# Purge the outdated corosync database files entirely
rm /etc/pve/corosync.conf
rm -rf /etc/corosync/*
rm -rf /var/lib/corosync/*

# Restart the local filesystem daemon as a clean standalone service
killall -9 pmxcfs
systemctl start pve-cluster
Once the node is wiped and returned to a standalone state, running a standard pvecm add [CLUSTER_IP] often triggers two scary security blocks:
  1. The Guest Warning: * this host already contains virtual guests.
  2. The SSL Failure: 500 Can't connect to [IP]:8006 (hostname verification failed).
Understanding the SSL 500 Error
By default, Proxmox attempts to execute cluster joins using its secure HTTPS API on port 8006. The target node serves an SSL certificate that matches its actual hostname (e.g., g9-pve), but the joining node is hitting it via a raw IP address. Because the IP address does not match the name listed on the security certificate, the API strictly rejects the connection and throws the 500 error.
The Solution: Forcing and Bypassing via SSH
To safely resolve both issues, use the following advanced join command:
pvecm add [IP of neigbouring host] --force --use_ssh
Why This Works:
  • The --force Flag: This tells Proxmox to intentionally ignore the virtual guest warning. Your VM’s raw virtual disks sitting on local storage are 100% safe and are never modified or deleted. The flag simply allows Proxmox to safely merge the node’s local configuration directory into the cluster’s unified database.
  • The --use_ssh Flag: This is the magic bullet for the certificate error. It instructs Proxmox to completely bypass the strict HTTPS API verification loop. Instead, it creates a secure tunnel over standard SSH to pull the fresh cluster schema, synchronize cryptographic keys, and launch Corosync cleanly.

The Result
The joining host will pull the clean, master configuration down, automatically register itself into the cluster map, and instantly turn green across your unified Proxmox dashboard.
Your existing VMs will reappear right where they belong under the node’s localized tree in the Web UI, fully integrated and ready to manage.

Setting up a Proxmox host

Install Proxmox

Hardware

Step 1) Pick Hardware. Important is CPU support. Mostly ARM or x86_64.
– My host a HPE DL160 G9.

Software

Step 2) Install using appropriate installer image.
– In my case x86_64 version 9.2
– I installed to an internal 32GB sd card.

Storage (Physical)

Step 3) VM Storage.
I’ve discussed this in the past specially when it comes to shared storage options. There you can see a picture of all the options available to Proxmox, and if it supports snapshots or if its shared. For ease sake of this post we’re going to stick to local storage.  With the minted information from that chart alone ZFS would seem the winner, however….

Choosing the right storage architecture for a virtualization host requires a careful balance between resource allocation, hardware capabilities, and performance goals. For this build—featuring 50 GB of memory, an HPE B140i controller running in SATA AHCI pass-through mode, and a mix of SSDs and a mechanical drive—maximizing raw performance and preserving system RAM for virtual machines is our primary objective. By selecting LVM-Thin instead of ZFS, we bypass the heavy computational and memory overhead of a Copy-on-Write filesystem, ensuring that nearly all 50 GB of RAM remains dedicated strictly to our workloads. The design stripes multiple solid-state drives into a single, high-performance LVM-Thin volume group to multiply IOPS and throughput for VM boot disks. Meanwhile, the standalone 4TB spindle drive is formatted as a standard, zero-RAM-footprint Linux directory to act as an isolated target for possible Proxmox backups and static file shares. This hybrid, LVM-centric approach eliminates storage controller bottlenecks, maximizes the lifespan and speed of our SSDs, and relies on a robust backup strategy rather than restrictive hardware or software redundancy.

After messing around about an hour, I found out the reason I wasn’t seeing the drives was due to a controller configuration (it was already set to Sata AHCI support mode), which I double verified by seeing the disks and running dd commands against them to get sequential performance numbers. The reason, was cause apparently in this mode drives are not hot swapable.

Could attempt a manual rescan via the shell backend, I guess but as noted there. “If your SATA controller supports hot swap, it should “just work(tm).”

<rant> Stupid ass HP, always causing me to waste my life away cause of their stupid ass storage controller and firmware/driver choices.. ughhh </rant>

Turn off Swap

SIDE QUEST! Congratulations you just entered a side quest on your way to configuring your storage for your PVE hypervisor. SWAP!

An critical optimization step for any Proxmox host booting from flash media is managing the system’s swap space. By default, the Debian-based Proxmox installer creates a virtual memory swap partition directly on the boot drive. When running Proxmox from an internal SD card, leaving swap enabled is a hardware hazard; Linux will continuously shift idle processes onto the card, exhausting its low write-endurance and risking boot environment corruption. Because this host boasts a healthy 50 GB of physical RAM, we immediately disabled and removed the default swap volume to shield the SD card from unnecessary wear. However, completely lacking a swap space can lead to kernel instability under unexpected memory spikes. Our strategy resolves this by re-establishing a dedicated swap space directly on our new solid-state storage tier. Crucially, this swap will not be placed inside the dynamic LVM-Thin pool—which can cause file system deadlocks and severe latency—but will instead be carved out as a fixed, pre-allocated ‘Thick’ LVM volume. This hybrid approach ensures the SD card remains read-heavy and protected, while giving the host an ultra-fast, safe SSD safety net without sacrificing valuable system memory to ZFS.

1. Turn off active swap immediately

swapoff -v /dev/mapper/pve-swap

2. Stop it from turning back on when you reboot

Open your filesystem table:
nano /etc/fstab
Look for the line that mentions pve-swap. It will look similar to this:
/dev/pve/swap none swap sw 0 0
Add a # at the very beginning of that line to comment it out and disable it permanently:
# /dev/pve/swap none swap sw 0 0

3. Delete the volume entirely (Optional but recommended)

To ensure the OS never touches it again, remove the logical volume entirely:
bash
lvremove /dev/pve/swap
Now with swap off we can finally build our LVM groups and move the swap to the SSDs.

Back to Storage (Logical)

Step 1: Create the Physical Volumes (PV)

First, we tell LVM that these three specific SSDs are ready to be used as raw storage building blocks.
pvcreate /dev/sda /dev/sdb /dev/sdc
Expected output: Physical volume "/dev/sda" successfully created. x3

Step 2: Combine them into a Volume Group (VG)

Now, we pool those three independent drives into one large, unified storage pool. We will name this group pve-fast.
vgcreate pve-fast /dev/sda /dev/sdb /dev/sdc
Expected output: Volume group "pve-fast" successfully created.

Step 3: Verify the Master Pool

To confirm everything was combined properly and to check your exact total available space, run:
bash
vgs pve-fast
You should see pve-fast listed with 3 physical volumes (#PV) and a total size that roughly equals the combined capacity of your three SSDs.
root@g9-pve:~# pvcreate /dev/sda /dev/sdb /dev/sdc
Physical volume "/dev/sda" successfully created.
Physical volume "/dev/sdb" successfully created.
Physical volume "/dev/sdc" successfully created.
root@g9-pve:~# vgcreate pve-fast /dev/sda /dev/sdb /dev/sdc
Volume group "pve-fast" successfully created
root@g9-pve:~# vgs pve-fast
VG #PV #LV #SN Attr VSize VFree
pve-fast 3 0 0 wz--n- <670.70g <670.70g

Creating the Volume Group (VG) only defines the boundaries of your master pool. It tells LVM: “You are allowed to use the storage blocks inside sda, sdb, and sdc.” It does not decide how data is laid out yet.

The master pool itself is neutral. The choice between Linear or Striped happens entirely in the next step when we create the Logical Volumes (LVs) inside that pool. Why it’s like this, I dunno, I’m just here to figure out how it works.

Back to Swap

We will allocate 4 GB of space for this safety net. We will use the -i 3 flag to guarantee that any memory swapped to disk is interleaved across all three SSD controllers simultaneously for maximum throughput.
Run these four commands sequentially in your Proxmox CLI:

1. Create the Striped Logical Volume

We will carve out a new volume named fast-swap from your pve-fast volume group.
lvcreate -L 4G -i 3 -I 64k -n fast-swap pve-fast

    • -L 4G: Allocates exactly 4 Gigabytes of space.
    • -i 3: Forces the volume to stripe data across exactly 3 physical disks (RAID0 behavior).
    • -I 64k: Sets the optimal block stripe size for performance.

2. Format the Volume for Swap

Now we tell the operating system to format this new striped block device specifically as Linux swap space.
mkswap /dev/pve-fast/fast-swap

3. Activate the New Swap Space

Turn on the newly created SSD swap space right now so the system can begin utilizing it.
swapon /dev/pve-fast/fast-swap

4. Make it Permanent Across Reboots

We need to register this new location in your system’s filesystem table so it mounts automatically every time the server turns on. Run this command to append the new rule to your configuration file:
echo '/dev/pve-fast/fast-swap none swap sw 0 0' >> /etc/fstab

Verify the Configuration

To verify that your swap is active, running at top speed, and no longer touching your 32 GB SD card, run:
swapon --show
You should see /dev/mapper/pve--fast-fast--swap listed as the sole active device, confirming your SD card is safe.
root@g9-pve:~# lvcreate -L 4G -i 3 -I 64k -n fast-swap pve-fast
Rounding size 4.00 GiB (1024 extents) up to stripe boundary size <4.01 GiB (1026 extents).
Logical volume "fast-swap" created.
root@g9-pve:~# mkswap /dev/pve-fast/fast-swap
Setting up swapspace version 1, size = 4 GiB (4303351808 bytes)
no label, UUID=a6ec7580-dc9c-40de-8405-b9f8c66a41fd
root@g9-pve:~# swapon /dev/pve-fast/fast-swap
root@g9-pve:~# echo '/dev/pve-fast/fast-swap none swap sw 0 0' >> /etc/fstab
root@g9-pve:~# swapon --show
NAME TYPE SIZE USED PRIO
/dev/dm-0 partition 4G 0B -1
root@g9-pve:~# fdisk -l | grep mapper
Disk /dev/mapper/pve-root: 12.8 GiB, 13740539904 bytes, 26836992 sectors
Disk /dev/mapper/pve--fast-fast--swap: 4.01 GiB, 4303355904 bytes, 8404992 sectors
K we’re finally done with swaps. I asked AI why I wasn’t seeing any partitions on the SSDs after allocating them to the LVM pool. it stated:

“When we ran the pvcreate command earlier, LVM took over the raw, unpartitioned block devices (sda, sdb, sdc) directly, rather than writing a traditional partition table like MBR or GPT.

Why LVM Prefers No Partition Layout
Tools like fdisk and gparted look for standard partition tables to tell them where a drive starts and ends. LVM bypasses this entirely:
    1. Direct Disk Ownership: LVM wipes out the standard partition header area and writes its own metadata directly to the very first sectors of the raw disk.
    2. Simplified Layout: Because we are striping these drives together, creating individual partitions (like sda1, sdb1) is an unnecessary extra layer of management. By giving LVM the raw disks, it can seamlessly manage the entire drive boundary.

How to Properly View Your LVM Layout
Because fdisk doesn’t natively parse LVM metadata structures, it will report that the disks look empty. To see the true reality of your storage, you must use LVM-specific inspection commands instead:
    • To see the raw disks LVM owns: pvdisplay or pvs
    • To see your combined master pool: vgdisplay or vgs
    • To see the striped swap volume we just carved out: lvdisplay or lvs"

Take this with a grain of sand I do not know if its true but seems that way.

root@g9-pve:~# fdisk -l
Disk /dev/sda: 223.57 GiB, 240057409536 bytes, 468862128 sectors
Disk model: KINGSTON SA400S3
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 512 bytes
I/O size (minimum/optimal): 512 bytes / 512 bytes
Disk /dev/sdb: 223.57 GiB, 240057409536 bytes, 468862128 sectors
Disk model: KINGSTON SA400S3
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 512 bytes
I/O size (minimum/optimal): 512 bytes / 512 bytes
Disk /dev/sdc: 223.57 GiB, 240057409536 bytes, 468862128 sectors
Disk model: KINGSTON SA400S3
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 512 bytes
I/O size (minimum/optimal): 512 bytes / 512 bytes
root@g9-pve:~# pvs
PV VG Fmt Attr PSize PFree
/dev/sda pve-fast lvm2 a-- <223.57g 222.23g
/dev/sdb pve-fast lvm2 a-- <223.57g 222.23g
/dev/sdc pve-fast lvm2 a-- <223.57g 222.23g
/dev/sde3 pve lvm2 a-- <29.22g 3.63g
root@g9-pve:~# vgs
VG #PV #LV #SN Attr VSize VFree
pve 1 2 0 wz--n- <29.22g 3.63g
pve-fast 3 1 0 wz--n- <670.70g 666.69g
root@g9-pve:~# lvs
LV VG Attr LSize Pool Origin Data% Meta% Move Log Cpy%Sync Convert
data pve twi-a-tz-- <10.79g 0.00 1.58
root pve -wi-ao---- <12.80g
fast-swap pve-fast -wi-ao---- <4.01g

Back to Storage

So now I just need another logical volume for the VM high speed OS vHDDs.
When I tried to assign 100% of the remaining space to the VM data volume, I hit a common LVM roadblock: the thin-pool conversion failed due to insufficient free space (0 extents). This happens because an LVM-Thin pool needs a tiny bit of unallocated space left over to build its hidden metadata index for tracking snapshots. To fix this, I deleted the raw volume and recreated it using 99%FREE of the remaining pool instead. This small tweak left plenty of breathing room for the tracking database while keeping the 3-disk stripe perfectly aligned.

1: Recreate it with 99% of the pool space

By allocating 99%FREE instead of 100%FREE, we guarantee there is plenty of room left over for the metadata engines while still satisfying the stripe alignment requirements.
lvcreate -l 99%FREE -i 3 -I 64k -n fast-data pve-fast

2: Convert it to a Thin Pool

lvconvert --type thin-pool pve-fast/fast-data
Once it says successfully converted, run the final step to link it to your Proxmox dashboard:
pvesm add lvmthin Striped-SSDs --vgname pve-fast --thinpool fast-data

The conversion went through smoothly, and the high-speed storage tier is now online in the Proxmox GUI under the name Striped-SSDs

Quick Sequential I/O test:

root@g9-pve:~# swapoff /dev/pve-fast/fast-swap
root@g9-pve:~# dd if=/dev/zero of=/dev/pve-fast/fast-swap bs=1M count=2000 status=progress conv=fdatasync
2000+0 records in
2000+0 records out
2097152000 bytes (2.1 GB, 2.0 GiB) copied, 3.59334 s, 584 MB/s
root@g9-pve:~# swapon /dev/pve-fast/fast-swap
swapon: /dev/mapper/pve--fast-fast--swap: read swap header failed
root@g9-pve:~# mkswap /dev/pve-fast/fast-swap
Setting up swapspace version 1, size = 4 GiB (4303351808 bytes)
no label, UUID=aa6557d6-626b-4680-9741-2b63a1f55a13
root@g9-pve:~# swapon /dev/pve-fast/fast-swap

More Storage

Yes even more storage, while we used LVM to stripe across our 3 SSDs. We are going to use Ext4 on the 4TB Drive to host ISOs, or large disk virtual drives on the VMs.

1. Create a Standard Partition Table

We will write a clean, modern GPT partition table to the raw drive and create a single partition that takes up 100% of the 4TB space.
parted -s /dev/sdd mklabel gpt mkpart primary ext4 0% 100%

2. Format the Partition as Ext4

Now, we format that fresh partition (/dev/sdd1) with the standard Linux Ext4 filesystem. This handles sequential data streams beautifully on mechanical platters.
mkfs.ext4 /dev/sdd1

3. Create a Mount Point and Mount the Drive

We will create a permanent folder on your host OS and mount the physical drive into it.
mkdir -p /mnt/pve/mass-storage
mount /dev/sdd1 /mnt/pve/mass-storage

4. Make the Mount Permanent Across Reboots

To make sure Debian hooks this drive back up every time the server boots, we add its unique identification to your filesystem table (fstab). Run this command to fetch the drive’s unique ID and automatically write the mount rule:
echo "/dev/sdd1 /mnt/pve/mass-storage ext4 defaults,noatime,nofail 0 2" >> /etc/fstab
(Note: noatime eliminates unnecessary write cycles to track when files are read, and nofail ensures your Proxmox host still boots perfectly even if the 4TB drive is unplugged).

4. Register the storage for the Proxmox GUI

Finally, run this command to tell Proxmox that this folder is ready to accept backup files, ISOs, and VM disks:
pvesm add dir Mass-Storage --path /mnt/pve/mass-storage --content backup,iso,images

Verify Your Entire Server Storage Layout

Now that everything is fully configured, your storage is split perfectly into two distinct, high-efficiency worlds. If you run:
df -h /mnt/pve/mass-storage
root@g9-pve:~# parted -s /dev/sdd mklabel gpt mkpart primary ext4 0% 100%
root@g9-pve:~# mkfs.ext4 /dev/sdd1
mke2fs 1.47.2 (1-Jan-2025)
Creating filesystem with 976754176 4k blocks and 244195328 inodes
Filesystem UUID: 84d88803-a246-4067-8cec-3a93ba188169
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
4096000, 7962624, 11239424, 20480000, 23887872, 71663616, 78675968,
102400000, 214990848, 512000000, 550731776, 644972544

Allocating group tables: done
Writing inode tables: done
Creating journal (262144 blocks): done
Writing superblocks and filesystem accounting information: done

root@g9-pve:~# mkdir -p /mnt/pve/mass-storage
root@g9-pve:~# mount /dev/sdd1 /mnt/pve/mass-storage
mount: (hint) your fstab has been modified, but systemd still uses
the old version; use 'systemctl daemon-reload' to reload.
root@g9-pve:~# systemctl daemon-reload
root@g9-pve:~# echo "/dev/sdd1 /mnt/pve/mass-storage ext4 defaults,noatime,nofail 0 2" >> /etc/fstab
root@g9-pve:~# dd if=/dev/zero of=/mnt/pve/mass-storage/zerofile bs=1M status=progress conv=fdatasync
78979792896 bytes (79 GB, 74 GiB) copied, 293 s, 270 MB/s
root@g9-pve:~# dd if=/dev/zero of=/mnt/pve/mass-storage/zerofile bs=1M status=progress oflag=direct
2867855360 bytes (2.9 GB, 2.7 GiB) copied, 33 s, 86.9 MB/s

Summary: Proxmox VE Storage Architecture: Maximizing Performance and RAM with a Non-Cached Controller

When designing local storage for a hypervisor host, the default answer is often to throw everything into a ZFS pool. However, storage architecture should never be a one-size-fits-all solution. For this Proxmox VE build—featuring 50 GB of physical RAM, an HPE B140i controller running in SATA AHCI pass-through mode, a trio of SSDs, and a single 4TB mechanical spindle drive—we chose a hybrid, LVM-centric approach designed specifically to prioritize raw performance and maximize available system memory for virtual workloads.

Phase 1: Protecting the Boot Media and Reclaiming Swap

The server boots Proxmox from an internal 32 GB SD card. By default, the Debian-based installer creates an active swap volume directly on the boot drive. Because SD cards utilize low-endurance flash memory, leaving an active swap partition on this media is a hardware hazard that would quickly wear out the card. Furthermore, with 50 GB of physical RAM available, host-level swapping should be incredibly rare.
We immediately disabled and purged the default pve-swap volume from the SD card. To preserve a host safety net without introducing latency or deadlocks, we moved the swap partition to our incoming solid-state pool. Crucially, this swap space was created as a fixed, pre-allocated “Thick” LVM volume rather than being nested inside a thin data pool, ensuring kernel stability under unexpected memory spikes.

Phase 2: The Performance Tier (3-Disk LVM-Thin Stripe)

To give our VM operating systems maximum IOPS and unthrottled throughput, we grouped our three zeroed SSDs (sda, sdb, sdc) into a single LVM Volume Group named pve-fast.
Because data safety is managed via a strict backup strategy rather than local fault tolerance, we chose to strip data evenly across all three disks using LVM’s interleaved striping parameter (-i 3). This acts as a high-efficiency software RAID0 array directly inside the Linux kernel. During configuration, we encountered a classic LVM hurdle: allocating 100% of the remaining pool to a raw data container left 0 extents behind for the metadata tracker, causing the LVM-Thin conversion to fail. Re-provisioning the container at 99%FREE provided the required breathing room for the tracking database while maintaining perfect alignment across the three controllers.
This performance tier consumes virtually 0 MB of host RAM, leaving almost all 50 GB available for our VMs. A sequential write test using dd directly against the raw striped blocks clocked in at a blistering 584 MB/s, successfully compounding the bandwidth of our independent controllers.

Phase 3: The Mass Storage Tier (Zero-RAM Spindle Directory)

For our 4TB mechanical drive (sdd), we chose to completely bypass LVM and ZFS, formatting it directly as a standard Ext4 Directory partition mapped straight to the Proxmox dashboard.
Using ZFS here would have starved our host by demanding a massive chunk of RAM for its ARC cache, while its Copy-on-Write architecture would have severely choked write performance on a controller lacking a battery-backed hardware cache. LVM-Thin was also discarded for this drive; Proxmox backup files require a standard filesystem folder, and thin block-level provisioning creates massive physical fragmentation on spinning platters over time.
By using a standard Ext4 directory, we can provision large virtual hard disks (vHDDs) for backup servers like Veeam using the QCOW2 file format. QCOW2 handles thin-provisioning intelligently at the virtual file level, preventing the hypervisor from scattering blocks chaotically across the physical disk.

Phase 4: Benchmarking and the Reality of Caching

We ran two distinct write benchmarks against our newly formatted 4TB Ext4 storage directory to observe how the operating system handles a mechanical drive:
  1. The Buffered Test (conv=fdatasync): After an initial RAM-buffered burst, the sequential write stream stabilized at an impressive 270 MB/s. This is significantly faster than the drive’s raw hardware capability. The boost is entirely driven by Ext4 optimizations like Delayed Allocation (delalloc) and sequential extents, which neatly arrange incoming data blocks on the fast, outer edge of the empty platter.
  2. The Direct I/O Test (oflag=direct): To expose the raw physical limits of the drive, we bypassed the Linux kernel’s RAM page cache completely. Stripped of its file-system optimizations, the performance leveled off at 100 MB/s, exposing the exact mechanical floor of the spindle and proving how vital the filesystem’s caching layer is for normal operation.

The Power-Safety Tradeoff

The 270 MB/s buffered speed comes with an engineering tradeoff: write safety. Because our AHCI pass-through controller lacks a physical battery-backed write cache, any data floating in the host’s volatile RAM cache during a sudden power outage will be lost.
While this risk would be unacceptable for a live production database, it is perfectly suited for this specific architecture. The 4TB tier is dedicated strictly to static ISOs and compressed Veeam backup repositories; a power failure simply invalidates a running backup job, which can easily be restarted once the system boots back up. To completely mitigate this, the host will be plugged into an Uninterruptible Power Supply (UPS) integrated with automated shutdown software, ensuring all memory caches are safely flushed to the physical disks before the server powers down.
This finished architecture leaves us with a highly optimized, dual-tier environment: a blazing fast 584 MB/s SSD stripe for active VMs, a highly efficient 270 MB/s sequential mass storage folder for backups, and a completely unburdened 50 GB pool of RAM dedicated entirely to running workloads.
This is the bare basics of setting up a Proxmox server. Things we haven’t covered yet are networking, updating, clustering, shared storage, managing VMs, etc. These will be covered in the upcoming blog posts. This one is just the fundamental requirement to all those other topics. This is just the foundation. Hope this helps someone.
BONUS MATERIAL!!!!
If you read this far amazing, you may wonder how to figure out how to know how much actual disk space a vm’s disk is using when configured on a LVM thin. well the GUI won’t tell you. You can run “lvs” and do math… why a native command gives you this data in a percentage instead of actual size? Beats me.. but here paste this into the shell to create a better new command “lvu” which I call “logical Volume Usage”
cat << 'EOF' >> ~/.bashrc
alias lvu="lvs -o lv_name,lv_size,data_percent --noheadings --units g | awk '{
name = \$1;
alloc = \$2;
pct = \$3;
gsub(/[A-Za-z]/, \"\", alloc);
gsub(/%/, \"\", pct);
if (pct == \"\" || pct == 0) {
used = alloc;
} else {
used = (alloc * pct / 100);
}
printf \"%-16s | Allocated: %6.2f G | Used Space: %6.2f G\n\", name, alloc, used
}'"
EOF
source ~/.bashrc

‘cat << ‘EOF’ >> ~/.bashrc
alias lvu=”lvs -o lv_name,lv_size,data_percent –noheadings –units g | awk ‘{
name = \$1;
alloc = \$2;
pct = \$3;
gsub(/[A-Za-z]/, \”\”, alloc);
gsub(/%/, \”\”, pct);
if (pct == \”\” || pct == 0) {
used = alloc;
} else {
used = (alloc * pct / 100);
}
printf \”%-16s | Allocated: %6.2f G | Used Space: %6.2f G\n\”, name, alloc, used
}'”
EOF
source ~/.bashrc’

 

Now just type “lvu”

root@g9-pve:~# lvu
data                         | Allocated: 10.79 G       | Used Space: 10.79 G
root                          | Allocated: 12.80 G       | Used Space: 12.80 G
fast-data               | Allocated: 660.02 G   | Used Space: 8.12 G
fast-swap             | Allocated: 4.01 G          | Used Space: 4.01 G
vm-100-disk-0 | Allocated: 0.00 G          | Used Space: 0.00 G
vm-100-disk-1  | Allocated: 32.00 G       | Used Space: 8.12 G

Why this isn’t a native command, also beats me.