Least-Privilege ZFS over iSCSI for Proxmox VE via TrueNAS

How I Configured Least-Privilege ZFS over iSCSI for Proxmox VE via TrueNAS

The built-in Proxmox VE “ZFS over iSCSI” storage provider relies on direct root SSH access to TrueNAS to dynamically spin up block devices. However, this legacy method is broken on modern TrueNAS SCALE versions (like Dragonfish and Electric Eel) because TrueNAS no longer uses the target engines (like Comstar/LIO) expected by PVE.
To solve this securely without enabling risky root SSH access, I installed the “official TrueNAS Proxmox VE Storage Plugin” aka plugin managed by “TheGreatWazoo”, worked around network anomalies, and configured a custom Role-Based Access Control (RBAC) policy for a true least-privilege service account.

Step 1: Setup TrueNAS

Virtualizing TrueNAS: CPU NUMA and Memory Logistics
    • Bypassing the CPU NUMA Trap: Since my TrueNAS virtual machine is small enough to fit comfortably inside a single physical CPU socket and its local memory pool, I left virtual NUMA (vNUMA) disabled in my hypervisor. Enabling it on a small VM would force the guest OS to waste processing cycles managing artificial boundaries. It would also restrict hypervisor scheduling, pinning my storage VM to specific sockets even if other sockets are completely idle. Leaving it disabled allows the hypervisor to handle optimization automatically via UMA (Uniform Memory Access) mode.
    • Fixing the RAM Allocation: TrueNAS was initially starved at a default 2048 MiB (2 GB), which is an instant recipe for an Out-Of-Memory (OOM) kernel crash under iSCSI loads. I scaled the allocation up to 8192 MiB (8 GB) (with a ceiling of 12288 MiB / 12 GB if needed). This provides enough overhead for the base OS middleware and the 4TB ZFS metadata.
    • Disabling Memory Ballooning: Because ZFS calculates its cache (ARC) based on a fixed expectation of available memory, I disabled the ballooning device and set fixed memory. If the host tried to dynamically reclaim RAM mid-test, the TrueNAS kernel would instantly panic and drop my storage network offline.

After successfully installing TrueNAS, I presses 1 at the console to configure the interface.

I noticed I could only toggle DHCPv4 and DHCPv6 on or off, but the actual IP address input field seemed completely missing.
Here is how I finally managed to configure my static IP:
    • I navigated down to the Aliases field. In TrueNAS SCALE, static IPs are assigned as aliases rather than a traditional static IP entry field.
    • I selected the empty list. I used my arrow keys to highlight aliases: <empty list> and pressed Enter to open the edit sub-menu.
    • I entered my IP with CIDR notation. I typed my desired address and subnet mask combined (for example, 192.168.1.100/24) instead of typing the netmask separately.
    • I saved the changes. I scrolled down to the <Save> button at the bottom of the screen and pressed Enter to apply the network settings.

Only then I could go to configure network settings and define the gateway and DNS servers. lil weird but whatever, easy enough.

Then just like before HBA to passthrough the hard drives to the VM

Just like before, created 2 isolated networks (one mgmt plane (HTTPS/SSH) and one data plane (iSCSI). *Note* I’m using VLANs (on the hypervisor for the VM guest) to simulate what you normally would do physically. So the configuration on TrueNAS looks the exact same as physical (two dedicated NICs for the two planes) AND… MPIO will not be covered in the scope of this blog. I’ll have to cover that in the future when I get some hosts with multiple NICs.

Step 2: TrueNAS Service Account & API Configuration

To avoid assigning global administrative rights, I built a dedicated service user and a custom privilege profile to authorize specific storage metadata calls.
1. Account Creation Constraints
I navigated to Credentials > Users to create my service account container:
    • Username: pve-iscsi
    • TrueNAS Access: Checked the box and selected Sharing Administrator rather than a generic read-only or full admin profile.

      • Create the account, Give it a Name, Check off TrueNAS Access, select read-only admin.
    • Shell/SSH Access: Kept unchecked. The plugin communicates purely via HTTPS REST API calls, so command-line execution is unnecessary.

🛠️ Caveat & Tweak: This creates a less restrictive account, it also creates a group with the same name as the account. We will be creating a custom “Privilege” Role with only the required “Roles” permissions, then bind that role to this group to complete the least privilege configuration.
User:
Group:
As you can see the group created here shows each permission already defined on this account (snippet here is after the needful was done), but you need to create the actual Role by clicking on the Privileges button next to add. Create a “Privilege” and name it after this setup.
Under “Roles” aka Permissions I granted this “Privilege”, what I call a Role,
1. Sharing (iSCSI Track)
    • Sharing iSCSI Extent Read
    • Sharing iSCSI Extent Write
    • Sharing iSCSI Initiator Read
    • Sharing iSCSI Initiator Write
    • Sharing iSCSI Portal Read
    • Sharing iSCSI Portal Write
    • Sharing iSCSI Target Read
    • Sharing iSCSI Target Write
    • Sharing iSCSI Target Extent Read
    • Sharing iSCSI Target Extent Write

2. Storage & Snapshots (ZFS Track)
    • Dataset Read
    • Dataset Write
    • Dataset Delete (Crucial so Proxmox can delete VM disks)
    • Snapshot Read
    • Snapshot Write
    • Snapshot Delete (Crucial so Proxmox can clear old snapshots/backups)
    • Pool Read (Allows the API to safely query the storage layout tree)

3. System Base Discovery
    • System General Read (Allows the API to check system info and uptime metrics)

4. The Breakthrough Secret (Filesystem Track)
    • Filesystem Read
    • Filesystem Write

(Note: Keeping Allow All Sudo Commands and Sudo without Password unchecked on the user account ensures this remains a strict least-privilege setup.)
As you can see in the snip once you bind this custom role the GUI only supports the built in “Privileges”, so if you edit this account it’ll show the TrueNAS Access checked off with the dropdown empty:

Step 3. Generating the API Token as Admin

Instead of logging out and accessing the UI under the new service account context, I provisioned the token from my master administrator account:
    1. Navigated to Credentials > Users and clicked on my new pve-iscsi user row.
    2. Inside the profile card’s Access widget, I clicked Add API Key.
    3. Named the token (e.g., PVE-iSCSI-Token) and copied the generated string immediately.


Step 4. Preparing the TrueNAS iSCSI Network Targets

Before connecting the hypervisors, I gathered configuration parameters and eased network access limits under Apps/Services > iSCSI.
    1. Bound iSCSI: Ensure iSCSI Service is bound to iSCSI NIC only.
    2. The Portal ID: Verified my active iSCSI listening group ID under the Portals tab (typically ID 1).
    3. The Target IQN: Copied my system base name string under Target Global Configuration (iqn.2005-10.org.freenas.ctl).
    4. The Initiator Workaround: Under the Initiators tab, I enabled Allow All Initiators and saved it as Initiator ID 1.

💡 Network Tweak: In a trusted home lab or private storage network, allowing all initiators simplifies initial setup. For strict isolation, this can be unchecked later by pasting the unique IQN strings of each individual Proxmox node directly into the allowed targets box.


Step 5. Driver Installation on Proxmox VE Nodes

Because storage modules run at the system level rather than the cluster level, I executed these steps directly on every node in my cluster via SSH/Shell.
# 0.  Firewall:
I had to configure a firewall rule to allow access to the DEVs githubsource.
URL: thegrandwazoo.github.io/
Application: Github-Pages
# 1. Download and register the official repository GPG key:
curl -fsSL https://thegrandwazoo.github.io/freenas-proxmox/public.gpg.key | gpg --dearmor | tee /etc/apt/keyrings/truenas-proxmox.gpg > /dev/null 
# 2. Inject the driver repository metadata block using the DEB822 standard format:
cat << 'SOURCES' | tee /etc/apt/sources.list.d/truenas-proxmox.sources
Types: deb URIs: https://thegrandwazoo.github.io/freenas-proxmox
Suites: v3
Components: main
Signed-By: /etc/apt/keyrings/truenas-proxmox.gpg
SOURCES
# 3. Synchronize package databases and install the kernel driver plugin:
apt update apt install truenas-proxmox -y
Installing this package successfully registers the truenas-iscsi framework and automatically recycles local management services (pvedaemon, pveproxy) to render the integration options in the UI.

Step 6. Mounting Storage and Navigating GUI Flaws

I refreshed my Proxmox browser interface, navigated to Datacenter > Storage > Add, and selected the newly exposed TrueNAS iSCSI storage provider.
The Connection Profile Layout
    • ID: TrueNAS-ZFS (or any friendly cluster name)
    • Nodes: Set to All
    • TrueNAS Host: 172.16.21.2 (Management network interface)
    • Portal IP: 69.69.69.1 (The isolated network interface bound to the iSCSI portal)
    • Pool: LexarPool (Case-sensitive master storage pool name)
    • API Token: [Pasted the custom service token generated in Phase 1]

Hitting Add automatically builds the connection link, turning my cluster resource status icon green.

Core Architecture Takeaway: The Proxmox Offline Migration “Flaw”

During testing, I stumbled on a major counterintuitive design choice in how Proxmox handles moving virtual machine storage: Proxmox allows you to move local storage blocks to another node effortlessly if the VM is powered ON, but throws a hard block if the VM is powered OFF.

Why this happens under the hood:
    • Online Migration (VM ON): Proxmox delegates operations to the QEMU/KVM virtualization engine. QEMU handles the network stream natively, actively duplicating running disk blocks to an entirely different storage name on the destination node on the fly.
    • Offline Migration (VM OFF): Proxmox bypasses QEMU and drops back to its legacy internal cluster scripts. This simple script reads the text configuration of the VM, checks the source storage name (e.g., msiNM620), and checks if an identical storage name exists on the destination host. If the destination node does not have a local storage engine with that exact name, the script panics and blocks the migration through the UI.

How TrueNAS fixes this behavior:
Because my new TrueNAS integration is defined as shared storage named identically (TrueNAS-ZFS) across all nodes in my datacenter, it completely satisfies Proxmox’s basic offline script constraints. When a VM’s disks live on the TrueNAS storage block, the offline migration wizard unblocks instantly, allowing lightning-fast configuration transfers between cluster nodes without moving any underlying data blocks.

BONUS

I had weird errors when moving from shared storage back to local (this was when I had the setup using full local administrator “Privileges”, it didnt’ cause the task to fail and AI said this about.

When copying virtual machine disk blocks off of the TrueNAS iSCSI pool back onto local storage, you may notice repeating logs stating:
qemu-img: iSCSI GET_LBA_STATUS failed... SENSE KEY:ILLEGAL_REQUEST(5)
Why it triggers:
During a file transfer copy block, qemu-img sequentially pings the TrueNAS target engine with GET_LBA_STATUS calls to query which specific disk block ranges are blank or sparse so it can bypass copying empty space. Because the TrueNAS storage framework handles these optimizations at a different filesystem descriptor tier, it rejects the application block format structure, generating the standard warning lines.
The Impact:
This error is entirely harmless. The script immediately falls back to basic raw block-by-block copying and finishes streaming the virtual machine disk data over the network safely with zero block corruption.
What happens on hosts that have no installed the plug in? I dunno haven’t tested, I’m assuming they won’t have access to the storage.
What happens if the iSCSI server has a hard fail? Dunno, I’ll find out soon in my lab as I need to move the server. Any new findings I’ll update this blog post.
Hope this helps someone, feel free to leave a comment.

Using FreeNAS as a VM to Test Storage Speeds

Using FreeNAS as a VM to Test Storage Speeds

OK sooo this is gonna seem kinda commical… I

  1. set up a PVE host. Most of which ended up being about storage, I didn’t even blog the actual installation process, I just got right into the storage after a completed installation lol.
  2. wrote a blog attempting to discuss managing a PVE host. A little better cover some basic host management stuff, but again end up talking about VM storage options and performance.. lol
  3. wrote a blog post above recovering a VM from ESXi to PVE using Veeam. Which was also poor, most a just a video reference to a Veeam tech who shows the technical steps, then just me wondering why I got such unreal poor performance from my setup.

I don’t suggest you read any of them cause they are some of the worst blogs I have ever written. They are, however, not entirely useless as they provide some bases to the tests I continue to complete.

So now I had another thought, and can you guess what it was around… yeah… storage… anyway, so if LVMthin is not the best choice for Random IO, and even using the Linux kernel as cache causes issues, what if I just throw the SSDs to a FreeNAS VM. Will it perform better?

My Storage Engineering Journey: Rebuilding Proxmox Local SSD Storage into an Isolated iSCSI San

🛠️ Phase 1: Deconstructing My Original Proxmox LVM Storage

Originally, my Proxmox host had three 240GB Kingston SSDs (sda, sdb, sdc) grouped into a striped LVM volume group called pve-fast. This group hosted an LVM-Thin data pool (fast-data) and an active striped system swap space (fast-swap).
To completely free these drives up for raw passthrough without throwing GUI errors, I opened my Proxmox host SSH shell and dismantled the entire architecture in reverse order:
    1. Remove GUI mapping: I commanded Proxmox to stop monitoring the LVM-Thin dashboard pool:
      bash
      pvesm remove Striped-SSDs
      
    2. Deactivate and scrub the swap space: I disabled the active swap space running on the SSDs:
      bash
      swapoff -v /dev/pve-fast/fast-swap
      

    3. Remove swap from the boot configuration: I edited the filesystem table:
      bash
      nano /etc/fstab
      

      I located the active line /dev/pve-fast/fast-swap none swap sw 0 0 and deleted it (or added a # at the beginning) to prevent my system from hanging or crashing on its next boot.

    4. Destroy the Logical Volumes: I permanently purged the inner allocation containers:
      bash
      lvremove /dev/pve-fast/fast-data -y
      lvremove /dev/pve-fast/fast-swap -y
      
    5. Destroy the Volume Group: I deleted the master pool layout itself:
      bash
      vgremove pve-fast
      

    6. Wipe LVM Labels from Raw Disks: I forced LVM to entirely release its ownership tags over the physical hardware:
      bash
      pvremove /dev/sda /dev/sdb /dev/sdc
      
    7. Verify Everything is Clean: I ran the validation checks to ensure only my main OS boot drives remained:
      bash
      pvs
      vgs
      lvs
      

💾 Phase 2: Passing Individual Raw Disks to TrueNAS

I spun up a brand-new TrueNAS SCALE virtual machine assigned with a flat, stable 16GB of RAM. Because memory ballooning breaks ZFS caching calculations, I disabled ballooning by keeping the Minimum Memory and Maximum Memory values identical inside the Proxmox UI.
    1. Locate My Persistent Disk IDs: I knew using changing identifiers like /dev/sdb would break my VM mapping if I added or removed hardware down the line. I ran this command to pull the unique hardware serial strings:
      bash
      ls -l /dev/disk/by-id/
      

      I copied down my three distinct Kingston identifiers:

        • ata-KINGSTON_SA400S37240G_50026B7785138139
        • ata-KINGSTON_SA400S37240G_50026B77851380CD
        • ata-KINGSTON_SA400S37240G_50026B7785137D21

    2. Map the Disks into the VM: Using the Proxmox host shell, I manually bound the raw block devices to my TrueNAS VM (VM ID 101) sequentially on the SCSI controller, appending critical flags to disable host-level caching and keep Proxmox backup routines from touching my storage array:
      bash
      qm set 101 -scsi1 /dev/disk/by-id/ata-KINGSTON_SA400S37240G_50026B7785138139,cache=none,backup=0
      qm set 101 -scsi2 /dev/disk/by-id/ata-KINGSTON_SA400S37240G_50026B77851380CD,cache=none,backup=0
      qm set 101 -scsi3 /dev/disk/by-id/ata-KINGSTON_SA400S37240G_50026B7785137D21,cache=none,backup=0
      
    3. Force Hardware Serial Numbers: When I first booted the VM, TrueNAS’s ZFS middleware threw validation errors. QEMU passes virtual disks as generic entities, meaning TrueNAS saw three separate paths sharing a blank or overlapping virtual serial identifier. I shut down the TrueNAS VM and edited the backend configuration file on my Proxmox host:
      bash
      nano /etc/pve/qemu-server/101.conf
      

      I navigated to the scsi1, scsi2, and scsi3 entry lines and appended ,serial= followed by their real hardware suffixes directly to the end of the config text:

      text
      scsi1: /dev/disk/by-id/ata-KINGSTON_SA400S37240G_50026B7785138139,backup=0,cache=none,size=234431064K,serial=50026B7785138139
      scsi2: /dev/disk/by-id/ata-KINGSTON_SA400S37240G_50026B77851380CD,backup=0,cache=none,size=234431064K,serial=50026B77851380CD
      scsi3: /dev/disk/by-id/ata-KINGSTON_SA400S37240G_50026B7785137D21,backup=0,cache=none,size=234431064K,serial=50026B7785137D21
      

      After saving and booting TrueNAS up, the duplicate ID issue disappeared completely.


🌐 Phase 3: Building an Isolated Layer 2 Virtual Network

To keep my heavy, high-throughput iSCSI storage data entirely off my flat subnet and isolated from my Management VLAN 21, I engineered an in-memory virtual switch pipeline.
    1. Create the Private Bridge in Proxmox: In the Proxmox Web GUI, I went to Node -> System -> Network -> Create -> Linux Bridge.
        • Name: vmbr1
        • IPv4/CIDR: 10.10.10.1/24
        • Gateway / Bridge Ports: Left completely blank.
        • This isolated the bridge entirely inside host memory, enabling packets to move at CPU speed without hitting a physical switch. I clicked Apply Configuration to spin it up live.

    2. Add a Second NIC to TrueNAS: In VM 101 -> Hardware -> Add -> Network Device.
        • Bridge: Selected vmbr1
        • Firewall Checkbox: Unchecked. This was a critical step. By turning off the Proxmox software firewall for this card, I bypassed heavy packet inspection overhead, saving CPU cycles and ensuring lower latency for my storage loop.

    3. Configure the Storage IP in TrueNAS: I logged into my TrueNAS SCALE web dashboard, opened Network -> Interfaces, and edited the newly populated unconfigured adapter (e.g., vtnet1). I unchecked DHCP and manually assigned a flat Layer 2 static IP configuration:
        • IP Address: 10.10.10.2
        • CIDR: 24
        • Gateway: Left completely blank to eliminate any potential multihoming asymmetric routing loops.

    4. Verify the Pipeline: I went into the Proxmox terminal and ran a quick check to make sure the internal memory link was intact:
      bash
      ping -c 3 10.10.10.2
      

      🏗️ Phase 4: Carving out Zvol and Provisioning the iSCSI SAN

    5. Create My Zvol Container: In TrueNAS SCALE, I went to the Datasets tab in the left panel. I selected my master pool (FastPool), clicked Add Zvol, and input the following configuration parameters:
        • Zvol Name: pve-zvol
        • Size: 450 GiB
        • Sparse Volume: Unchecked (Thick Provisioned). I did this to intentionally carve out and lock down this exact slice of my 639 GiB total raw pool upfront. This prevents Proxmox from accidentally over-allocating storage down the line and protects my ZFS array from hitting 100% capacity and freezing.
        • Compression: LZ4 (Lightning fast, low CPU overhead, reduces physical write amplification by compressing data blocks before they hit flash).
        • ZFS Deduplication: OFF. (I verified this was disabled because dedupe consumes roughly 5GB of system RAM per 1TB of data tracked, which would quickly starve and crash my 16GB TrueNAS VM).

    6. Execute the iSCSI Wizard: I navigated to Shares -> Block (iSCSI) -> Wizard:
        • Target Page: Named it pve-target and set Target Intent to Modern OS.
        • Extent Page: Set Extent Type to Device and selected my new volume FastPool/pve-zvol (450G). Under sharing platform, I selected Modern OS to guarantee proper 512-byte block alignment mapping.
        • Protocol Options Page: Switched the Portal dropdown to Create New, and selected my static storage IP address link 10.10.10.2. I hit save and ensured the global iSCSI service was flipped to Running and configured to Start Automatically.

    7. Map the Target in Proxmox: In the Proxmox Web GUI, I went to Datacenter -> Storage -> Add -> iSCSI:
        • ID: TrueNAS-iSCSI
        • Portal: 10.10.10.2
        • Target: I clicked the drop-down box, and Proxmox instantly queried the virtual switch, auto-populating my exact target IQN string: iqn.2005-10.org.freenas.ctl:pve-target.
        • Use LUNs Directly: Unchecked. By leaving this unchecked, I prevented Proxmox from locking the raw connection down to one exclusive VM.

    8. Layer LVM for Dynamic Multi-VM Support: To allow Proxmox to carve up that 450 GiB network block into multiple separate virtual machine hard drives, I added a management layer. Still under Datacenter -> Storage, I clicked Add -> LVM:
        • ID: iscsi-storage (I discovered this field requires alphanumeric characters/text and cannot be a plain integer like 69).
        • Base Storage: Selected TrueNAS-iSCSI
        • Base Volume: Selects the auto-discovered 450 GiB LUN block.
        • Volume Group: Named it tg-pool.
        • Content: Selected both Disk Image and Container.
        • Shared: Checked.
        • Wipe removed volumes / Allow snapshots as volume-chains: Left both Unchecked to save unnecessary write wear on my consumer SSDs and to bypass broken thin-provisioned snapshot metadata lookups over standard iSCSI blocks.


📊 Performance Testing, Caveats, & Troubleshooting Lessons

I deployed a Windows guest VM directly onto my new iscsi-storage pool and ran benchmarks using CrystalDiskMark. The numbers revealed exactly how complex, multi-layered storage virtualization behaves under the hood.
My Benchmark Results & The ZFS Sync Breakthrough
    • Sequential Reads (~6,193 MB/s): I hit incredibly high, near-PCIe numbers. This proved that TrueNAS’s ZFS ARC (Adaptive Replacement Cache) was working flawlessly—intercepting my test read requests and streaming them directly out of my VM RAM across the high-speed virtual memory switch.
    • The Initial Write Bottleneck (120 MB/s Seq / 1.07 MB/s Random 4K Q1T1): My write metrics initially hit a performance wall. Because Proxmox handles iSCSI network targets with strict write-safety guarantees, it flags every transaction as a Synchronous Write. My consumer-grade Kingston A400 SSDs do not have a physical onboard RAM battery protection module (PLP – Power Loss Protection). As a result, every time a sync request came down the pipe, the drive controllers were forced to freeze operations and perform a hard cache flush down to physical flash cells, tanking my speed.
    • The Performance Fix (283 MB/s Seq / 4.00 MB/s Random 4K Q1T1): To fix this, I adjusted my parameters in the TrueNAS dashboard by going to Datasets, selecting my pve-zvol, and modifying its advanced properties to change Sync from “Standard” to “Disabled”. This forced TrueNAS to treat incoming IO as Asynchronous, allowing my write commands to buffer safely in TrueNAS RAM first. My sequential speeds more than doubled, and my random 4K write speeds instantly surged by 400%.

⚠️ My Lab Warnings & Core Caveats To Remember

    1. The “Virtualization Tax” on RND4K Q1T1: Even with ZFS Sync disabled, my single-threaded, single-queue random writes max out at 4 MB/s (whereas a standalone, bare-metal Windows installation on this same single SSD can easily hit ~25 MB/s). I now understand that this is standard for virtual storage. Forcing a single 4K file over a deep chain of abstraction layers (Windows File System → VirtIO Driver → Proxmox LVM → iSCSI Network → TrueNAS Kernel → ZFS Allocation → Physical Storage Controller) introduces microscopic amounts of computational latency. Because a Q1T1 test forbids parallel actions, the system must wait for a full round-trip confirmation before sending the next block. My parallel performance is healthy, however, as shown by my high RND4K Q32T1 queue numbers.
    2. Why My Old Striped LVM-Thin Setup Died: This project helped me diagnose why my original configuration suffered from terrible 4K performance. Layering an LVM-Thin allocation pool on top of an LVM striped storage block caused severe sector misalignment and block-write amplification. A tiny 4K operating system write would get split across physical drive block boundaries, forcing the host controller to continuously run slow read-modify-write loops across multiple SSDs simultaneously just to update a single 4K data sector.
    3. The Data Integrity Tradeoff: Setting ZFS Sync to Disabled is perfect for my home test lab to achieve fast performance, but it carries a risk. Because TrueNAS is caching incoming writes inside its RAM buffer before they actually finish sinking onto the physical SSD flash chips, a sudden home power outage or a hard freeze of the physical Proxmox host will cause data loss for whatever was floating in memory. This can easily lead to a corrupted VM operating system filesystem.
    4. No Native GUI Snapshot Functionality: Because standard LVM sits on top of raw network blocks, Proxmox’s blue “Take Snapshot” button is greyed out/unsupported for these VMs. If I want to schedule automated backup states or snapshots for my testing, I must manage them directly through TrueNAS’s native ZFS snapshot tasks dashboard at the Zvol level.


Summary, was it faster? Well in terms of I/O performance, technically yes although at the cost of a lot of implementation steps, and at the cost of Server Memory, and CPU threads. Would I recommend this, even for a home lab… meh I mean for learning its cool, but the performance while the SEQ read is kind of insane, it doesn’t provide much practical use.
Here you can see the amount of memory the FreeNAS has to do its ZFS magic, and how much CPU it takes on a high SEQ operation:
no matter what the RAN4K Q1T1 always seem to perform poorly in my tests:
if You have the Memory to spare, and have a decent CPU server with a poor storage controller, this isn’t really that bad of an option, you can also tie it into other PVE bridges/networks and serve other storage needs.
Would I recommend this, over all probably not, not honestly this is more robost then the LVMthin on the striped LVM group of the same SSDs, and disabling the write protection, that caused the entire storage stack to come to a halt and made my one server become unresponsive. “So, I tried this, and with a 64M target saw speeds up to 5x to 10x better results. So, I figured really test it and pick 8GB target file. And the other VM I just migrated onto this host lost it pings. Apparently… Optimizing Proxmox storage using a VirtIO SCSI Single controller paired with io_uring,IO Thread, and Write Back caching can yield a massive 5x to 10x performance boost in small bursts. However, executing a massive storage stress test (like an 8GB CrystalDiskMark run) on budget, DRAM-less hardware (such as Kingston A400 SSDs) can cause a severe cascading system freeze.”
So, this not only performed better, it also did cause a storage kernel panic on the same PVE host. I still tore it down cause it was too much overhead. Still neat to see it work though.

Using Veeam to Migrate from ESXi to Proxmox

Step 1) Have Veeam with Backups from an ESXi Host.

Check.

Step 2) Have a PVE Host.

Check.

Step 3) Add PVE Host to Veeam.

Check. I had a whole bunch of images saved on Img but I lost all the links so the above is a YouTube video that I basically followed to get er done.

The only thing of note here that was annoying is I wanted the worker VM’s network to be in a certain VLAN and the wizard in Veeam didn’t have an option to set it, so I had to enter the network config, and when the wizard was at the testing stage, connect to the PVE host and apply the VLAN tag on the network of the worker VM.

This problem can either be resolved using VNets, SDN for PVE, but the real solution (having a simple text field and applying it into an API call) is “on the roadmap” for Veeam after 2 years knowing about this limitation, that the fix is so simple, it’s mind boggling it not in the initial offering… 

Another thing I found weird with my particular setup (step 2), is that for the snapshot storage I could only pick my EXT4 storage and not the LVMthin.

Step 4) Restore VM to PVE

Even with the worker VM on the host, Veeam wouldn’t give me recovery speed or estimate to recovery, I used the glances command on the PVE host and noticed it was indicating CPU-IOWAIT was the bottle neck,  and seeing the logical disk and each SSD in glances showing only roughly 3 MB/s. I believe this might be due to how the worker VM was configured for its storage settings and how it coded. Took 6 hours but it did work once I did these steps after Veeam said success.

It worked but it wouldn’t boot even with the SCSI controller set to VMware SCSI. I had to detach the HDD, and reattach using SATA and then under VM options pick it for the boot order.

Network wasn’t working had to apply VLAN manually, then install virtio drivers to see the NIC, then manually re-IP and it said IP on old phantom NIC, so remove from that? yes, and network back up.

*Note you should really uninstall VMware tools…. cause for some reason the UN-installer does a hardware check to see if it is a VMware VM, I remember this from when I did a V2P a while back, what does Broadcom have to say about it? “Fuck you, if you didn’t remove the application before converting… fuck you. Uninstaller won’t work, sit there like a tattoo, fuck you bruuuh.”

Quoted from this KB

Issue/Introduction

  • A Windows virtual machine was migrated from vSphere to a non-vSphere environment without uninstalling VMware Tools.
  • Microsoft installer fails to uninstall the VMware tools.
  • No errors are identified during the uninstallation process.

Cause

After migrating the VM to a different platform, the VMware Tools uninstallation process fails because the virtual machine is not running within a vSphere environment.

Resolution

  • This is an expected behavior. <- AKA: We coded this deliberately
  • VMware Tools should be uninstalled prior to migrating the virtual machines out of the vSphere environment. <- AKA: You should of been a perfect admin.
  • Once the migration is completed, the virtual machine is no longer under the support of VMware by Broadcom. <- AKA: Fuck you!

Maybe I’ll cover a blog doing that, but I forgot in my test example. So, make sure you have the latest backup of the VM after removing VMware tools.

Issue #1

Storage Speeds

If you check out, Managing a Proxmox Host – Zewwy’s Info Tech Talks, you can see on the Test VM my Crystal Disk mark speeds and it performed poor on the RAN-t1q1 R/W but the others were fine. I asked AI, it mentioned that LVMThin and the lack of the battery write cache, and the fact it’s DRAM less SSDs creates the issue.

1. The LVM-Thin “Allocation on Commit” Penalty
LVM-Thin allocates space dynamically on demand. When a brand new Windows VM runs CrystalDiskMark, it writes to sectors that have never been written to before.
Every single time a tiny 4KB write occurs, the host operating system has to pause, check the hidden LVM metadata tracker, carve out a raw block from the pool, update the metadata index, and then commit the write. Doing this chunk-by-chunk at Queue Depth 1 (one file at a time) destroys random I/O performance.
2. The Interleaved Stripe Stripe-Size Mismatch
When you created the volume, you explicitly declared a 64k stripe size (-I 64k) across 3 disks (-i 3). This means LVM expects data chunks to be written in 192KB sweeps (64KB x 3) to evenly split the load.
  • CrystalDiskMark is attempting to write a tiny 4KB packet.
  • 4KB is a fraction of a single 64KB stripe.
  • Because it doesn’t span all three drives, the kernel doesn’t gain parallel execution speed. Instead, the storage driver must execute a Read-Modify-Write (RMW) cycle, adding physical disk latency overhead to a minute transaction.
3. Missing VirtIO SCSI Drivers & Write Caching
If you are testing this inside a fresh Windows VM without optimizations, Windows defaults to a generic IDE/SATA controller emulation, which forces synchronous, non-cached writes. Coupled with your HPE B140i controller lacking a battery-backed physical cache, Windows is forcing every single 4KB block to wait until the flash chips on the Kingston A400 SSDs send an physical acknowledgment back up the chain.
For point 3, I didn’t tell AI about my test and second blog post where I actually did cover the difference in the different virtual controllers. However, as we saw we didn’t get much better performance in the RANIO results, OK double in the reads but nothing in the writes. It suggested to change the cache to write back

So I tried this, and with a 64M target saw speeds up to 5x to 10x better results. So I figured really test it and pick 8GB target file. And the other VM I just migrated onto this host lost it pings. Apparently…

Optimizing Proxmox storage using a VirtIO SCSI Single controller paired with io_uring,IO Thread, and Write Back caching can yield a massive 5x to 10x performance boost in small bursts. However, executing a massive storage stress test (like an 8GB CrystalDiskMark run) on budget, DRAM-less hardware (such as Kingston A400 SSDs) can cause a severe cascading system freeze.
Here is exactly what happens behind the scenes when a heavy synthetic workload breaks a virtualized storage layer:
  • The Host RAM Trap (Linux Dirty Throttling): When a VM uses Write back caching, the Proxmox host intercepts writes and absorbs them instantly into its own memory pool. However, once the cache volume hits Linux’s internal threshold (dirty_ratio), the kernel hits an emergency brake. It forcefully halts all concurrent disk I/O requests across the entire storage layer to flush the data down to the physical disks.
  • The DRAM-less Wall: Consumer-grade SSDs lack dedicated onboard DRAM to map where files live. Under a massive, continuous random 4K write assault, their internal Flash Translation Layer (FTL) becomes heavily bottlenecked. Once their small, temporary SLC burst cache fills up, write speeds plunge down to a crawl (1–2 MB/s), causing I/O latency to spike into full seconds.
  • The LVM Storage Deadlock: With the physical drives running at a snail’s pace, the Linux kernel thread managing the thin-LVM volume pool drops into an Uninterruptible Sleep (D state). While the main Proxmox Web GUI stays responsive, any management process trying to hook directly into the VM’s active hardware layer—such as the vncproxy console stream—locks up instantly. The target guest VM drops completely off the network because its virtual hard drive stops responding.
The Fix & Takeaway: To safely benchmark real-world storage limits without triggering a host-level queue lockup, bypass the host RAM cache entirely by setting the VM disk cache mode to Default (No Cache) or Write through.
More testing and learning to commence. interesting finds.

Managing a Proxmox Host

In my last post we went over installing a Proxymox host and we did a fair bit of managing already… ok mostly just storage but we had to manage the host after the initial install of the base OS. This should be pretty obvious, it’s a web interface, which is stated right on the Console output after you install. All the commands in the previous blog could have all been done from the direct system console, but also via remote SSH.

So, first act is to change the update repo, by removing the enterprise ones and adding the no-sub repo for updates. This alone won’t resolve the nagging pop up when you log in about having no subscription. to get rid of that:

Remove the annoying subscription pop-up

  1. Open the Shell terminal from your Proxmox web UI or connect via SSH as root.
  2. Navigate to the widget toolkit directory:
    cd /usr/share/javascript/proxmox-widget-toolkit/

    Make a backup copy of proxmoxlib.js:

    cp proxmoxlib.js proxmoxlib.js.bak
    
  3. Open the file in a text editor like nano:
    nano proxmoxlib.js
    
  4. Search for the text active (press Ctrl + W in nano).
  5. Locate the conditional check that looks for an active status, which typically contains !== 'active' or !res logic. Change the inequality exclamation mark ! to make it an equality check == 'active' (removing the ! so it evaluates positively instead of triggering the warning when inactive). Alternatively, comment out or bypass the function call according to your specific Proxmox minor version.
  6. Save the file (Ctrl + O, then Enter) and exit (Ctrl + X).
  7. Restart the Proxmox proxy service to apply the change:
    systemctl restart pveproxy.service
    
  8. Perform a hard refresh or clear your browser cache (Ctrl + F5)

System Resources

I basically just click on the host summary tab. or a VMs summary tab.

Or install “glances” on the terminal shell.

Networking

I know, I know, you’re probably screaming about authentication and user management, groups, permissions. probably yelling “RBAC!!” I’m gonna stick to using root for now and concentrate on infrastructure stuff for now.

When configuring Proxmox in a multi-subnet or VLAN environment, you quickly run into the limitations of the Linux kernel’s “Weak Host Model,” which handles routing very differently than enterprise firewalls like Palo Alto Networks (PAN-OS). Unlike zone-based firewalls that use policy-based forwarding to automatically reply out of the same interface a packet arrived on, Linux relies strictly on destination-based routing tables. This becomes a major trap if you assign identical or overlapping subnets to multiple network interfaces; even if you physically unplug a network cable, the Linux kernel holds onto that dead route at the top of its table. This causes traffic to be shoved down a disconnected interface, resulting in frustrating “No Route to Host” errors and dropped connections, even when your other live interface is perfectly configured.

Furthermore, setting up multi-homed access to the Proxmox Web UI introduces asymmetric routing challenges, as Linux only permits a single global default gateway by default. If traffic arrives on a secondary VLAN interface, Proxmox will mistakenly attempt to send the reply back out the primary management gateway, causing firewalls to log “aged out” or “incomplete” states due to the routing mismatch. To resolve this, administrators must either use the CLI to inject custom policy-based routing rules (ip rule and separate routing tables) into the network configuration file or cleanly isolate their subnets by stripping duplicate IP layers off disconnected bridges. Additionally, when testing these secondary access points, remember that the Proxmox web service strictly binds to its local hosts configuration and requires explicit HTTPS formatting over port 8006 (https://<IP>:8006) to successfully initialize a session.

That’s a long winded way to say that when I was trying to keep the flat home network (untagged) ip address on the Proxmox server, while also giving it a virtual interface attached to another VLAN tagged subnet. It wouldn’t connect (or it wouldn’t load) the web interface from my home untagged network, even though it was routed, and tagged properly on all network devices along the network path.

So from what I can tell:

a “Linux Bridge” is like a vmware vSwitch. You define the physical connection the host has to these bridges.

a “Linux VLAN” is like a VMK. This is where you define another IP address the host can use. You select which device by defining VLAN raw device, which seem you can pic the physical NIC or the bridge, I don’t know the implication if you pic the nic when its already configured for a bridge though… I’m still learning here.

When you edit a VMs NIC settings you pick a bridge, and you can define what vlan the traffic will be at the VMs NIC settings, this is like the VMPG on ESXi.

Did I break updates from this… yup looks like it.. DNS works.. but can’t reach out anywhere or yeah.. locked down subnet, that was easily fixed.

Edited a VM NIC settings, change bridge, added VLAN Tag. Disconnect, power on VM, apply static IP, change to connect, yup.. works just fine.

To move the MGMT IP of the PVE host from untagged, to tagged follow these steps.

Step 1) Bridge Needs to be VLAN aware

In my case using the base bridge vmbr0.

PVE host (left hand side) -> System -> Network -> vmbr0 -> edit -> Check off VLAN Aware.

Step 2) Create a VLAN

Under name give it vlan#, where # corresponds to the VLAN tag you need applied.

Set the IP address and the new Gateway (if it complains about gateway already set on the bridge network, make sure you remove the gateway from it, and it’s IP address else you’ll fall into the problem I described at the beginning of this networking section.)

Step 3) Apply the Config

Hit Apply at the top and watch the ping flip over…

Step 4) Change IP under /etc/hosts

nano /etc/hosts

Find your old IP and update, otherwise things like creating cluster info will bind to old IP in the cluster info.

Interacting with VMs

Virt-Viewer and Virt-Manager

PVE has the web console built right in, so you can just manage the VM directly that way. I like being able to have an app window for the connection much like VMRC for VMware. Which PVE has, called Virt-viewer, get it here: Virtual Machine Manager. I installed the Winx64 binaries.

One trick I like to do is connect a USB stick to my main mgmt machine, then copy files to it, then to a VM if I need to get files on to said VM if the VM is an offline only machine.

Since I deployed this VM from a generalized image I had created I needed to pick a storage controller that I knew would be natively available to the image I was using so I stuck with the LSI, I installed virt-viewer so SPICE as my GPU, and again a native supported NIC, so the E1000.

As you can also see, they are all generic drivers, but.. working:

So as you see, not terrible, but also not crazy, I know those 3 SSDs can perform better then these results since I did an I/O test on them via the host backend, so I’m assuming I have so loss in the virtual bus controller (the LSI 53C) or the standard Windows drivers. So, the first thing I want to test is installed the guest tools, will they change how devices show in the device manger, and will there be any performance increases?

Spice Guest Tools

So downloaded them on the guest VM from “www.spice-space.org/download.html”

not sure what was up with the serial driver, but I just accepted it:

Well…

Windows Main Device? No USB Trick for you!

  1. Even after all that, the video drivers showed up without basic drivers, and I can move in and out of the VM in the virt-viewer with having to press CTRL+ALT+R. That’s Good.
  2.  The Storage device in device manager still shows generic SATA ACHI so I don’t believe I’ll get any better I/O results.
  3. Attempting to add a SPICE USB port to the VM hardware worked but…

after shutting down the VM and power it back on, the device list wasn’t greyed out and showed one free channel. but picking any of my devices…

ok… this might be cause my mgmt machine is Windows?

I want to like PVE, but there are a lot of little niche things that are pissing me off about it. Then when you want to use SPICE with virt-viewer, it downloads a spice.vv file that you have to open, which auto deletes when the VM is shutdown or close (haven’t tested this). just feels like weird UX. anyway…

Storage Controller vs Virtual Hard Drives

I changed the SCSI controller from LSI 53C to VirtIO SCSI Single. But when I booted the VM back up I still saw the same generic SATA ACHI Controller. I felt like there was some ignorance on my part so I asked AI for any insights, it informed me to add a drive cause the type on the actual virtual disk could still be bound to the old type. So I temp added a disk (just for testing) and changed the connection from SATA to SCSI and checked the dev mgmt and ran a test and the performance was a fair bit better…

compared to

Performance Increase Overview

Benchmark Test Metric Type Performance Change Percentage Increase
Seq1M-Q8T1 Read 433 MB/s → 651 MB/s +50.3%
Write 56 MB/s → 95 MB/s +69.6%
Seq1M-Q1T1 Read 400 MB/s → 511 MB/s +27.8%
Write 51 MB/s → 54 MB/s +5.9%
Ran4K-Q32T1 Read 14 MB/s → 136 MB/s +871.4%
Write 7 MB/s → 7 MB/s 0.0% (No Change)
Ran4K-Q1T1 Read 6 MB/s → 13 MB/s +116.7%
Write 1.5 MB/s → 2 MB/s +33.3%

Key Takeaways
  • Massive Random Read Improvement: The biggest leap is in Ran4K-Q32T1 Read, sky-rocketing by 871.4%. This means heavy multi-threaded background random tasks will feel exponentially faster.
  • Solid Sequential Gains: Large file transfers (Seq1M) see a great bump, with reads improving by roughly 28% to 50%, and multi-queued writes jumping by nearly 70%.
  • Lagging Write Speeds: Random deep-queue writes (Ran4K-Q32T1) didn’t improve at all, and sequential single-thread writes (Seq1M-Q1T1) only crawled up by 5.9%.

That’s a bit improvment, I need to get the base OS HDD on this new type to gain the performance increase. Do to that:

Swap the Real Drive to SCSI

  1. In the Proxmox Hardware tab, select the 1 GB dummy disk you just made and click Detach. Then select the detached unused disk and click Remove.
  2. Select your main Windows boot disk (currently sitting on sata0 or ide0) and click Detach. It will immediately drop down to the bottom of the hardware list as an Unused Disk 0.
  3. Double-click that Unused Disk 0.  (I don’t know why double click seemed the only option I couldn’t see any action items at the top)
  4. In the pop-up window, change the Bus/Device dropdown to SCSI (it will likely assign scsi0). Click Add.

Fix the Boot Order

  1. Go to the VM’s Options tab in Proxmox.
  2. Double-click Boot Order.
  3. Check the box for your newly reattached scsi0 drive and drag/button it to the very top of the list so it is the primary boot device. Click OK.

Yeah for some reason it wasn’t checked off, so reattaching a vHDD has this implication something I didn’t instinctively had to do, in the snip above I unchecked net boot and checked off the scsi0.

Start your VM. before I ran the test I wanted to make sure the baseline VM was fine for it since now it was the Windows main OS drive that was running on the new virtual SCSI bus. however sure enough windows updates were alerady hitting the disk and the CPU.. I noticed it in task a manager, which was also showing me…
like what?! 84% active time constant, with a contant 800+ ms repsonse time and a measly 1.7MB/s … is windows doing insane I/o and bottle necking the I/O bus? was the theory all BS, or would this have happened on the settings I had before…? So many questions, so little answers… but the results are not good the Windows updates process is low CPU and high wait time on disk it seems the disk is slowing things down….
well system is back to idle windows updates completed.. lets see what diskmark has to say… shows the same results as “D:\” so we should have got the I/O performance increase, yet.. I remain skeptical….

Summary

So, we touched a bit on some basic management of a Proxmox host, like checking system resources, networking, storage, and managing VMs. Each of these are not covered in depth by any means, but just the simple fundamentals to getting a VM up and running and basic management of them.
These fundamentals are needed for the next stage, migrating VMs from ESXi to ProxMox. I know, I know, you’re saying I already did a basic pilot of that in the past here: Migrate ESXi VM to Proxmox – Zewwy’s Info Tech Talks but that was a bare metal, bare FS and using a linux VM with a convertion tool to just convert the  base HDD and it’s associated FS intact the version required by the hypervisor. It also took a lot of space, bandwidth, I didn’t explain what each step was really doing in detail. Anyway, long story short in the next blog post I’m gonna see how we can use Veeam to do a migration instead of a linux machine.

Rebuild FreeNAS/TrueNAS

So, the boot drive died in my FreeNAS server that was running 11.1u1. Good build. I figured since it died now would be a time to try TrueNAS Core. My SAN didn’t seem to be booting the installer, I tried different USB sticks and eventually just ended up using my IODD device, but all of which wouldn’t show the installer (albet on xterm/serial connection since the SAN is headless).

I decided to run the installer on a laptop and installing on …a….USB stick… wait a second….. when FreeNAS was first failing I tried moving the system log files off there since USB drives aren’t really meant for high R/W operations. These are set by the System -> System Dataset. What I think happed was the drive probably could’ve lasted a while longer but since the log files was still configured for the default “boot-pool” aka on the USB boot stick, and I had configured iSCSI with proxmox and at the end I discovered the log flooding problem… I bet it was this log flooding along with writing and log rotating on the USB boot device must have killed it…. Just wow man….

Anyway, I installed on a USB stick on a laptop picking the USB drive as a souce and picking BIOS boot option. I plugged the device into my SAN and boot up, now watching the boot operation (after verifying the boot order was correct) via the serial console window, after the system info the console gets stuck at “Press any key to continue”. I looked at the USB drive and the light indicated that something was happening, also looking at the NIC lights seemed something was going on, but also seem a bit pattern like and I wasn’t sure if something or anything was working. I decided to quickly check my DHCP server for any new entries and sure enough there it was….

I navigated to the address on my browser and sure enough, it booted!

There it is the dataset we need to change, now I don’t have any other pools defined so I can’t change it just yet, but lets make this our first priority so that if another bug from proxmox creeps in it won’t break my SAN. To do that I’m going to need to import the ZFS pools I had on this SAN when FreeNAS died.

Storage -> Pools -> Add

What a clean look. Import, No encryption. Looking for pools, this took a while.

There’s my pools! Sweet:

Next… exciting stuff…

Click Import.

Waiting….. waiting…. Let’s go!

What’s this notification…

I’ll leave that for now… can I change my system dataset? Oh no way it did it automatically….

well… even though I managed to import the pools successfully, since I can’t remember how the extents were made, I think they were file based, but no matter how I make it file based or device based the ESXi hosts see the drive but they aren’t automatically importing the datastore, probably cause they can’t see the FS that’s suppose to be on the iSCSI disk that’s being presented (the extents). I should have simply kept them device based and used the root pool. To add icing to the cake the one server I wanted to recover was the only one I didn’t run the backup copy job on after my USB drive that was hosting those files went belly up and I had to rerun all the jobs, AND the Veeam servers main backup repo was on one of the extents I can’t recover… so even though I had 3-2-1 rule in place I still ended up losing this server….

this is too much for me today, I’m goin to bed…

A new day, but end of it, so won’t finish this yet today either but there is hope.

I found this, which lead to this, which lead to this, and there’s sign of hope seeing the same thing in my own vmkernel log file.

So there it is, and ran the commands as specified in the threads and KB:

tail /var/log/vmkernel.log

Which showed me the volume not mounting due to snapshot.

esxcli storage vmfs snapshot list

Which showed me my old Datastore information.

esxcli storage vmfs snapshot resignature -l "iOSlow"

Which mounted my datastore under a new name:

Sure enough connected the backup drive to my Veeam server. And restored my missing VM. Now I just need to move the rest of the data and create a clean datastore. Or I guess I could simply rename it, but it didn’t auto mount on the other esxi server.

Share NTFS USB HDD via SMB on FreeNAS

I’m boiling down an entire night of knowledge as short as possible:

Is it possible? Yes, reference (this post)

Does the internet say it’s possible? No and More, No

Jeff “In the FreeNAS documentation it says using USB attached devices as shares is not allowed.”

Let’s do it anyway. Couple point notes:
*I created an account on FreeNAS “veeam” account ID 1001.

  1. Mounting The USB HDD to FreeNAS:
    Using the “Import Disk” option doesn’t work well:

    1. requires existing zpool aka volume, configured.
    2. when completed doesn’t show files properly.
    3. Mounts Disk in Read Only.
    4. Much like the link shared above we just mount it manually via the backend.
      1. ntfs-3g /dev/da6s1 /mnt/USBHDD/ -o rw,user_allow_other,uid=1001,gid=1000
      2. to make this stick after reboots have to edit fstab file. *I haven’t done this yet, when I have and tested it, I’ll update this area.
      3. The command mounts the NTFS using FUSE, and you can’t change ownership of files n folders after mounting only during.
  2. Sharing the Drive via SMB:
    1. Attempting to create a share via the Front End UI will show the path available in the path selector but it will simply state “This field is required” when trying to create the SMB Share. or you might get “The path None does not exist“.
    2. symlinking or mounting directly to existing zpool pool path that’s already shared via SMB, results in failure accessing the drive and Freenas Logs “smbd: dnssd_clientstub write_all(36) failed -1/53 57 Socket is not connected“
    3. The above line alone, I went through hell trying to solve, it’s what lead me to learning about FUSE and the chown issues and all that jazz, I went down so many rabbit holes I thought I was defeated, till I had one final idea: just like I manually mucked with the backend to get NTFS mounted in RW, maybe I can edit the backend Samba config to share the path since the front end python scripts were coded to prevent it.
      1. Find the config file: Samba config file:
        /usr/local/etc/smb4.conf
      2. Add a shared path entry:
        [usbhddd] 
            path = "/mnt/USBHDD"
            printable = no
            veto files = /.snapshot/.windows/.mac/.zfs/
            writeable = yes
            browseable = yes
            access based share enum = no
            hide dot files = yes
            guest ok = no
      3. Save the file and restart the Samba Service:
        service samba_server restart
        

When I saw that share path available, and when I double clicked it and I saw the files saved there show up, my jaw dropped!!! I couldn’t believe it worked.

Much like the manually having to edit the FSTAB to get the drive to mount automatic at boot, I have a feeling the smb4.conf file maybe overwritten at boot, which may require a cron job script to resolve. I again haven’t got to that point yet, I just finished this proof of concept that was, from my research, deemed to be impossible. Yet here I am blogging my success. See below for some info regarding Samba.

Samba options

Samba for FreeBSD

Key take away is that there’s a “link” between the Unix user and the “SMB” user. “FreeBSD user accounts must be mapped to the SambaSAMAccount database for Windows® clients to access the share. Map existing FreeBSD user accounts using pdbedit(8):”

pdbedit -a -u username

Final Note. I did this so I could have Backup Copy Jobs run, the Veeam server is a VM and this allows the VM to be migrated to other hosts while still being able to do both regular backup jobs and Backup copy jobs. and now that the USB drive on FreeNAS is NTFS based, I can just take the drive plug it into a windows machine and start restore operations. Having said that I’m doing this for my HomeLab and is for educational purposes only.

Here’s a snip of the repo in use via Veeam.

3TB Drive Shows up as 750GB

There’s a lot of stuff on this, so I’ll keep it short.

On windows, check Intle RST drivers (assuming there storage controller the hard drive is connected to is Intel based).

In my case it was behind a USB Enclosure. The drive showed properly as 3TB, but it didn’t recognize the File Systems.

Figuring I could see the files in linux that’s when the problem presented itself.

Lucky for me I had another machine that was 64 bit and had sata ports, plugged it into that and checked there (the storage controller was old nvidia nforce4, if anyone remembers that lol)

and it worked it saw the drive. When I went to mount the partition though it stated “unknown filesystem type ‘linux_raid_member’”

So I did the same thing and mounted it using mdadm, I also had to do “mdadm –stop /dev/md0” or else it always say the /dev/sdb3 was busy.  Strange.

This was cause the drive was from a RAID 1 member, so all files were accessible.

Never seen this one before, and yes I’m aware of 2TB limit of 32 bit systems, So I knew that was not the issue. This was good to know though in case of future file recovery attempts. 🙂

FreeNAS Single SSD as ZIL and L2ARC

Quick Story I remember I set this up on my FreeNAS server in hopes to get better performance, in reality, I don’t think it helped anything cause of my FreeNAS servers setup. Which was an old desktop with 3 Gigs of memory and a couple SATA drives, 2 spindle and 1 SSD.

Took me a while but I finally found the original source I followed.

Main Parts (assume SSD is ada0):

root@freenas1:~ % gpart create -s gpt ada0
ada0 created
root@freenas1:~ % gpart add -a 4k -b 128 -t freebsd-zfs -s 10G ada0
ada0p1 added
root@freenas1:~ % gpart add -a 4k -t freebsd-zfs ada0
ada0p2 added

List Disk to get GUIDs

root@freenas1:~ % gpart list

Add partitions as Zil and L2ARC on a logical disk (volume0)

root@freenas1:~ % zpool add volume0 log gptid/94a4bd28-aeb7-11e5-99ac-bc5ff42c6cb2
root@freenas1:~ % zpool add volume0 cache gptid/9a79622f-aeb7-11e5-99ac-bc5ff42c6cb2

Nice you can use the zpool command to verify their used as such:

If you are paying attention you’ll noticed the guid are different. Anyway you can use the GUI to see the results as well, if you click the main volume under Storage -> Volume, Then click the Show details button.

If you pick any of the partitions in this list, at the bottom you get a button labeled “Remove”. To undo the previous additions made via the back end SSH.

After this I removed the old Volume completely, including the old File based extent I was using on it.

I then created all new Volumes, one volume on each drive, then created 1 zVol on each volume, then used those zVols as Device based Extents on the iSCSI service…. and I couldn’t believe the performance increase, I couldn’t saturate the 1gbps link before with storage vMotions. Now every single Datastore maxes out the NIC and I hot 100 MB/s plus on every storage vMotion and I increased my storage capacity. W00t (of course I never had storage redundancy to begin with so nothing lost, all gains.

Summary

Don’t bother using a SSD to try and gain speeds on simple homelab FreeNAS servers. It’s useless… “Some more specifics: as a rule of thumb L2ARC is only really useful if you have lots of RAM (64GB+) and a ZIL is only useful if you’re performing lots of synchronous writes.” – anodos

FreeNAS Volume Down.

Quick Note, This is NOT a deep dive post into troubleshooting a downed volume, in this case I knew the drive was unavailable since boot and my goal was to re add the logical drive after correcting the physical connection issue.

This happened to me due to a Hardware issue. A power surge killed my UPS, like fully in that it wouldn’t turn on. SO had to rip it out and rebuild my DataCentre since I’m a poor man without proper servers, or server mounts. It’s a ghetto mans DataCenter.,.. anyway. The single USB enclosure housing a 2 TB HDD which was mounted and shared via SMB on the FreeNAS server didn’t power on. I decided to open the case to see if I could find the issue  (the PSU was fine as I was reading 12 v from the standard barrel connector. After I removed the case I was shocked find it was powering on… ok what gives. Put the case back on and nothing, it’s like the power barrel isn’t reaching the internal pins all of a sudden. I’m not sure if this was cause I swapped it with another 12v unit within the rack, either way I found an adapter to fit the same female and male ends and amazingly it worked lol, how useless but randomly came in use in my life.

So now back to FreeNAS with the USB drive powered on and connected.

First thing on the UI was the critical alert of the Volume being down. I wasn’t sure how to bring it back online with commands like lsusb being useless.

I found this FreeNAS form post with someone having a similar issue were the logs stated the simplest solution:

Recovery can be attempted by executing ‘zpool import -F vol1′

I SSH’d in and ran that command ageist the known volume that was down and lo and behold it appeared to have fixed my mounted USB drive…. but my SMB share just wasn’t available…

SO restart the SMB share… nothing… OK what gives… I dont’ remember documenting exactly how I set this up and it older FreeNAS 11.1-U1… so now I check the source server via SSH…

“zpool status” now shows the volume is there. checking “df -h” shows it’s mounted as /SMB… yet going to the Sharing -> Windows Shares and checking the shared volume states it should be /mnt/SMB but it’s not mounted as such hence why it’s not showing up…

Now 2 questions pop in my head 1) did I mis-configure something or 2) is the mount process different during boot in which it will mount the volume under /mnt instead of the root… not sure what happened here.. also not sure exactly how I should fix it. I want to avoid a reboot as it hosts iSCSI based VMFS volumes for my ESXI hosts.. what a pain…

ok… sigh mmmm I can either link or mount the volume accordingly at this time, but not sure how that will affect the server at boot….

So after talking to the “experts” apparently I did something wrong (how classic) due to a mix of my ignorance and … ahem… a system design in which the backend shouldn’t be touched outside the frontend… like lame SharePoint… anyway to read the details see this snippet:

Though have to give credit where it’s due and it’s nice to get clarification on things that piss me off so much it actually triggers my “flight or fight” response in my brain and I get like raged.

So taking a few minutes to cool down to hopefully resolve what should have, as usual, been a rather easy process became a royal pain in the fucking ass. But a “learning” experience none the less. Say that shit more than enough times in this stupid field of shit… ughhhh

OK now not pissed…. I went to Storage -> Volumes via the front end, and even though it showed green and healthy from the backend import command, I clicked the volume and selected “detach” from the bottom. I chose not to destroy my data (default, good stuff), and to not remove the share configuration (SMB service stopped anyway).

Then I clicked import volume (no encryption) and lucky for me the volume in question was the only one available in the dropdown list. The wizard successfully imported the volume, and sure enough doing a “df -h” on teh backend showed it mounted as /mnt/SMB ands retarting the SMB services worked and navigating the share also worked.

Yay well this sure was a learning experience…. don’t mess with the backend too much with FreeNAS (soon to be TrueNAS CORE).

Cheers

 

Windows MPIO to FreeNAS iSCSI Target

Intro

Well I made some mistake, the system worked but not utilizing its max capabilities..

I had been successfully using FreeNAS as a iSCSI target for  a disk mounted in Windows Server, but only one path being used at all times…

Windows Side

Source

I first needed the MPIO feature installed:

  1. Click Manage > Add Roles And Features.
  2. Click Next to get to the Features screen.
  3. Check the box for Multipath I/O (MPIO).
  4. Complete the wizard and wait for the installation to complete.

Noice.

Then we need to configure MPIO to use iSCSI

  1. Click Start and run MPIO.
  2. Navigate to the Discover Multi-Paths tab.
  3. Check the box to Add Support For iSCSI Devices.
  4. Click OK and reboot the server when prompted.

For me I didn’t get prompted for a reboot and reopening MPIO showed the checkbox unchecked, I had to click the add button then I got a prompt to reboot:

Now before I continue to get MPIO working on the source side, I need to fix some mistakes I made on the Target side. To ensure I was safe to make the required changes on the target side I first did the following:

  1. Completed any tasks that were using the disk for I/O
  2. Validated no I/O for disk via Resource manager
  3. Stopped any services that might use the disk for I/O
  4. Took the disk offline in Disk Manager
  5. Disconnected the Disc in iSCSI initiator

We are now safe to make the changes on the target before reconnecting the disk to this server, now on to FreeNAS.

FreeNAS Side

Source

I much like the source specified added an IP to the existing portal.. which I apparently shouldn’t have done.

Stop the iSCSI service for changes to be made.

Now delete the secondary IP from the one portal:

Now click add portal to create the secondary portal with the alternative IP.

There we go now just have to edit the target:

Now, that you have multiple portals/Group IDs configured with different IP addresses, these can be added to the targets.

Editing the existing targets to add iSCSI Group IDs

Once you have a target defined, you can click the Add extra iSCSI Group link to add the multiple Port Group ID backings.

Add extra iSCSI group IDs to each target in FreeNAS

Make sure you have the iSCSI service running. It does hurt at this point to bounce the service to ensure everything is reading the latest configuration, however with FreeNAS the configuration should take effect immediately.

Make sure iSCSI service is running in FreeNAS

Now we can go back to Windows to get the final configurations done. 🙂

Back on Windows

Configuring iSCSI

Launch iSCSI on the application server and select the iSCSI service to start automatically. Browse to the Discovery tab. Do the following for each iSCSI interface on the storage appliance:

  1. Click Discover Portal.
  2. Enter the IP address of the iSCSI appliance.
  3. Click OK.
  4. Repeat the above for each IP address on the iSCSI storage appliance.

Browse to Targets. An entry will appear for each available volume/LUN that the server can see on the storage appliance.

Configure Each Volume

For each volume, do the following:

  1. Click Connect to open the Connect To Target dialogue.
  2. Check the box to Enable Multi-Path.
  3. Click Advanced. This will allow us how to connect the first iSCSI session from the first NIC on the server. We can connect to the first interface on the iSCSI appliance.
  4. In the Advanced Settings box, select Microsoft iSCSI Initiator in Local Adapter, the first NIC of the server in Initiator IP, and the first NIC of the storage appliance in Target Portal IP.
  5. Click OK to close Advanced Settings.
  6. Click OK to close Connect To Target.

The volume is now connected. However, we only have 1 session between the first NIC of the server and the first NIC of the storage appliance. We do not have a fault-tolerant connection enabled:

  1. Click Properties in the Targets dialogue to edit the properties of the volume connection.
  2. Click Add Session.
  3. Check the box to Enable Multi-Path.
  4. Click Advanced.
  5. Select Microsoft iSCSI Initiator in Local Adapter. Select the second iSCSI NIC of the server in Initiator IP and the second NIC of the storage appliance in Target Portal IP.

Click OK a bunch of times.

If you open Disk Management, your new volume(s) should appear. You can right-click a disk or volume that you connected, select properties, and browse to MPIO. From there, you should see the paths and the MPIO customizable policies that are being used by this disk.

I left the load balancing algo to Round Robin, as Noted from here:

MCS

Fail Over Only – This policy utilizes one path as the active path and designates all other paths as standby. Upon failure of the active path the standby paths are enumerated in a round robin fashion until a suitable path is found.
Round Robin – This policy will attempt to balance incoming requests evenly against all paths.
Round Robin With Subset – This policy applies the round robin technique to the designated active paths. Upon failure standby paths are enumerated round robin style until a suitable path is found.
Least Queue Depth – This policy determines the load on each path and attempts to re direct I\O to paths that are lighter in load.
Weighted Paths – This policy allows the user to specify the path order by using weights. The larger the number assigned to the path the lower the priority.
MPIO

As above plus

Least Blocks – This policy sends requests to the path with the least number of pending I\O blocks.

Now did it actually work?

Seems like it.. performance is still not as good as I expected. must keep optimizing!

Hope this helps someone…