PVE: A Storage Deep Dive – LVMThin vs ZFS vs EXT4

Testing different Storage Architectures against a cheap consumer SSD

When I first spun up my virtualization host, I knew that storage configuration would make or break my performance. But nothing prepared me for just how wildly my benchmarks would swing based purely on software definitions. Using the exact same underlying physical SSD, I ran a series of tests to uncover how different virtual controllers, file systems, and storage architectures handle real-world disk operations.
Here is the raw data from my benchmarking journey, followed by a breakdown of what I learned, and a definitive pros and cons guide for every setup.

📊 The Cumulative Benchmarks: My Raw Results

To keep the comparisons fair, I tracked sequential and random operations across four distinct iterations of my storage stack.
1. The Baseline: LVMthin with a Virtual SATA Controller (64 MiB Test)
My journey started here, and the results were a painful wake-up call to the limits of legacy hardware emulation.
[Read]  SEQ Q8T1:  429.221 MB/s  |  SEQ Q1T1:  318.640 MB/s  |  RND Q32T1:  13.760 MB/s  |  RND Q1T1:   5.507 MB/s
[Write] SEQ Q8T1:   44.661 MB/s  |  SEQ Q1T1:   16.980 MB/s  |  RND Q32T1:   2.382 MB/s  |  RND Q1T1:   1.655 MB/s
2. The Protocol Flip: LVMthin upgraded to VirtIO SCSI Single (64 MiB Test)
By changing a single dropdown menu in my VM settings from SATA to VirtIO SCSI, I unbottlenecked the communication protocol.
[Read]  SEQ Q8T1:  493.329 MB/s  |  SEQ Q1T1:  383.567 MB/s  |  RND Q32T1:  88.362 MB/s  |  RND Q1T1:  12.480 MB/s
[Write] SEQ Q8T1:   90.589 MB/s  |  SEQ Q1T1:   21.391 MB/s  |  RND Q32T1:   2.403 MB/s  |  RND Q1T1:   1.879 MB/s
3. The RAM Illusion: ZFS with VirtIO SCSI Single (Cached 64 MiB Test)
Next, I swapped my backend volume manager to ZFS. With a tiny test file size, I accidentally benchmarked the sheer speed of my host’s RAM via the ZFS ARC cache.
[Read]  SEQ Q8T1: 4401.052 MB/s  |  SEQ Q1T1: 1921.489 MB/s  |  RND Q32T1:  84.354 MB/s  |  RND Q1T1:  19.714 MB/s
[Write] SEQ Q8T1:  205.138 MB/s  |  SEQ Q1T1: 1791.981 MB/s  |  RND Q32T1:  68.788 MB/s  |  RND Q1T1:  17.412 MB/s

4. The True Hardware Limit: ZFS with VirtIO SCSI Single (Raw 8 GiB Test)
By cranking the test file size up to 8 GiB, I forced the data straight through the RAM buffers and onto the raw NAND flash chips. This represents my real physical drive performance under ZFS.
[Read]  SEQ Q8T1:  739.249 MB/s  |  SEQ Q1T1:  715.466 MB/s  |  RND Q32T1:  77.281 MB/s  |  RND Q1T1:  17.737 MB/s
[Write] SEQ Q8T1:  382.726 MB/s  |  SEQ Q1T1:  219.378 MB/s  |  RND Q32T1:  49.969 MB/s  |  RND Q1T1:  15.344 MB/s
5. The Image File Contender: qcow2 on an EXT4 Host File System (64 MiB Test)
Finally, I tested the classic approach—placing a qcow2 virtual disk file on top of a standard Linux EXT4 partition.
[Read]  SEQ Q8T1:  189.151 MB/s  |  SEQ Q1T1:  154.374 MB/s  |  RND Q32T1:  41.495 MB/s  |  RND Q1T1:  10.420 MB/s
[Write] SEQ Q8T1:  124.732 MB/s  |  SEQ Q1T1:  115.363 MB/s  |  RND Q32T1:   2.419 MB/s  |  RND Q1T1:   1.726 MB/s

🔍 Structural Deep Dive: Pros, Cons, and My Key Architectural Takeaways

Every storage backend has a specific target audience. Based on my benchmarking data and the structural architecture of these systems, here is how they stack up.

Setup A: LVMthin (Block Storage)

LVMthin acts as a pure logical volume manager natively inside the Linux kernel. It gives the VM direct block-level access to the storage layer without an intervening file system.
  • The VirtIO SCSI Upgrade Difference: My data showed a 542% explosion in random reads (from 13.76 MB/s to 88.36 MB/s) just by moving away from the virtual SATA controller. SATA forced my VM into a shallow, serialized hardware queue. Upgrading to VirtIO SCSI unlocked deep parallel queue depths (Q=32), passing commands straight down to the kernel block layer efficiently.
Pros
  • Ultra-Low Overhead: It consumes almost zero host RAM and negligible CPU power.
  • Zero Write Amplification: Blocks are written directly to the SSD, maximizing the physical lifespan (TBW) of consumer drives.
Cons
  • No Cache Cushioning: Because it lacks an aggressive RAM cache layer, your performance is entirely capped by the physical drive hardware. My random writes hit a miserable 2.4 MB/s wall in both tests because the raw SSD couldn’t keep up with heavy random write thrashing.
  • Feature Light: It lacks advanced data integrity features like inline file self-healing.

Setup B: ZFS (Advanced Copy-on-Write File System)

ZFS merges the file system and volume manager into a single entity. It uses a Copy-on-Write (CoW) design, meaning existing data is never overwritten; modifications are written to entirely new blocks, and metadata addresses are updated afterward.
  • The RAM vs. Real Disk Difference: In my 64 KiB test, ZFS hid my operations completely inside the Adaptive Replacement Cache (ARC) and transaction memory, showing an insane 4.4 GB/s read rate. Once I used an 8 GiB file to smash past the RAM barrier, I saw my true physical drive speeds (~740 MB/s Read / ~380 MB/s Write).
  • The Random Write Fix: Remarkably, ZFS managed to boost my raw un-cached random writes up to 49.97 MB/s (compared to LVMthin’s 2.4 MB/s). This is because ZFS groups messy random writes together in memory and flushes them to the physical disk as a clean, orderly sequential stream.
Pros
  • Unmatched Burst Speed: Turns unused host RAM into an ultra-fast read/write cache buffer.
  • Enterprise Features: Offers built-in data checksums to prevent silent data corruption (“bit-rot”), easy live replication, and zero-latency snapshots.
Cons
  • Resource Hungry: ZFS expects a steep memory tax (often quoted as 1GB of RAM per 1TB of storage) to keep its ARC cache happy.
  • The SSD Lifespan Killer: The Copy-on-Write mechanism causes heavy write amplification. On consumer-grade SSDs without power loss protection, ZFS constantly flushes transaction metadata, which can burn through the drive’s endurance metrics in a fraction of its normal lifespan.

Setup C: qcow2 on EXT4 (File-on-File Storage)

This architecture represents a double-layer virtualization penalty. The hypervisor creates a structured image file (.qcow2), which handles its own internal allocations, snapshots, and clusters, and then places that file on top of a standard host Linux file system like EXT4.
  • The Double-Abstraction Penalty: My benchmarks reveal that this setup struggled heavily with throughput. Sequential reads completely tanked out to just 189 MB/s, and sequential writes crawled at 124 MB/s. Because every single block request from the VM has to be translated twice—first by the qcow2 file allocation table and then by the host’s EXT4 file tracker—latency skyrocketed to 44–66ms, severely choking sequential bandwidth.
Pros
  • Extreme Portability: The entire VM drive is just a single file on the host. You can easily copy, move, back up, or email a .qcow2 file across systems.
  • Thin-Provisioning Simplicity: The file grows dynamically on disk automatically as the VM consumes space, making storage management incredibly intuitive.
Cons
  • Severe Throughput Bottlenecks: Easily the slowest sequential performer of the group due to file system translation overhead.
  • Terrible Random Writes: Like LVMthin, it lacked the intelligent memory-grouping mechanics of ZFS, collapsing down to a sluggish 2.4 MB/s on parallel random 4KiB writes.

⚠️ Important Architectural Caveats & Exceptions

Storage performance does not exist in a vacuum. Before choosing a design based on my charts, you must consider your underlying host hardware infrastructure.
1. Hardware RAID Controllers vs. ZFS (The Golden Rule)
If your server utilizes a physical Hardware RAID controller (e.g., a Dell PERC or HPE Smart Array), do not use ZFS.
  • Why: ZFS requires total, direct control over the physical disk topology to manage cache flushing and data integrity.
  • The Conflict: If you place ZFS on top of a hardware RAID logical volume, the hardware controller will hide the physical disk health from ZFS. Even worse, the controller’s onboard cache will conflict with the ZFS ARC, severely increasing the risk of cataclysmic data corruption during an unexpected power loss.
  • The Winner: For Hardware RAID arrays, LVMthin or EXT4 are the proper structural choices.
2. Battery-Backed Write Cache (BBWC) Controllers
If your host server features an enterprise RAID card paired with a physical Battery Backup Unit (BBU) or supercapacitor flash cache, the playing field changes entirely.
  • Why: A BBWC allows the host to immediately report a write operation as “complete” the millisecond it hits the controller’s physical onboard RAM cache, safely knowing that if the power cuts out, the battery will preserve the data until reboot.
  • The Impact: If I had a BBWC under my LVMthin or qcow2 tests, that terrible 2.4 MB/s random write bottleneck would have completely vanished, skyrocketing up into hundreds of megabytes per second. It essentially provides ZFS-like memory caching speeds without the software complexity or RAM overhead.

🏁 Final Verdict: Which One Should You Build?

  • Choose ZFS if you have plenty of host RAM to spare, you want elite enterprise data safety features, you rely on rapid multi-node backup replication, and your storage drives are connected to an HBA (Host Bus Adapter) in IT/Passthrough mode.
  • Choose LVMthin if you are using consumer-grade hardware or mini-PCs, you want to maximize the physical lifespan of your SSDs, you need every single drop of host RAM allocated to running your actual VMs, or you are deploying on top of a Hardware RAID card.
  • Choose qcow2 on EXT4 if absolute peak performance takes a back seat to sheer simplicity, storage portability, and easy file-level drive movement between isolated hypervisors.
I’m still trying to figureout the root cause of Veeam restore being unreal poor on a LVMThin target. Stay tuned!

Leave a Reply

Your email address will not be published. Required fields are marked *