Virtualized TrueNAS Crashing

So, remember when I first build a test Frankenstein TrueNAS box to edge out some performance on otherwise shady consumer SSDs?  and then I ran with that Idea to test ZFS over iSCSI using TrueNAS?

Well since I was using a poormans setup, something wasn’t happy (PVE Virtio network stack? TrueNAS stack? I unno but I was using both VLAN segregation… Virtual NIC setup via Virtio, not sure but the host would become unresponsive to any network requests (pings or otherwise) but the host and VM would be otherwise fine (if I went to the console, I could issue commands without problem, but if I restarted the networking service only the host PVE would become responsive, TrueNAS would stay dead I had to reboot the whole host. AI assumes its the saturation, I have my doubts as I throttled the VLANs which didn’t help. But we’ll try anyway, I’ll move the iSCSI NICs to a PCI network card I apparently forgot I already had in this host. via hardware passthrough.. ok Let’s see how this goes…

I didn’t realize this one active VM is still on the ZFS over iSCSI datastore… move it and… host crash, so doing some testing interesting finds:

  1. Just as I stated above, restoring the network stack on the PVE host brings it back up fine, but
  2. Restarting the nework stack or even restarting the TrueNAS VM does not bring it back
  3. I have to reboot the PVE host as a whole
  4. The GUI task viewer, showing the task will stall on the transfer but the hosts stay perfectly stable and running fine, the transfer resumes where it left off after the storage PVE host and TrueNAS VM boot back up (takes roughly 20 seconds for PVE to boot, and 1 minute 30 seconds for TrueNAS) 2 min.
  5. Even when testing had a 6 minute outage window and the transfer resumed just fine when the backend storage came back. Is there a timeout out that checks actual transfer and if its stalled? I don’t know.
  6. The ZFS storage status in the left hand nav doesn’t reflect unknown status till roughly about 5 minutes of the storage becoming unavailable. But showing unknown back to good is almost instant.
The Solution: Hardware Isolation via PCIe Passthrough
The definitive fix is to physically separate the traffic by cutting Proxmox completely out of the storage loop. My motherboard has a spare PCIe slot, added a dedicated network card (a cheap, enterprise-stable multi-port Intel Gigabit adapter) and passed it directly to the storage VM.
Find the dedicated network card’s IOMMU group and add it as a Raw PCI Device to the TrueNAS VM hardware tab in Proxmox. Ensure All Functions and PCI-Express are checked.
This completely bypasses the Proxmox network stack, the virtual Linux bridge (vmbr0), and the VirtIO translation layers. TrueNAS now talks directly to bare-metal Intel hardware using native, highly optimized FreeBSD drivers.
After doing this I haven’t had the hosting PVE server or TrueNAS VM stop responding to pings, or services. It’s a lil less Frankensteiny now.
… or so I thought that separating the iSCSI traffic resolved it, everything far more stable, my usual instant test (Booting Windows VM against the SSD backed ZFS over iSCSI target) was working just fine… So, while I thought this resolved it, I jumped back to my AI chat window to state that this fix didn’t in fact work. I took a snip of my intuitional moment…
Which it then gave some examples of what it could be, I was kinda shocked when one of it’s hunches, was backed by hard evidence in the logs after the fact.. it stated and I quote..
“Clear out the e1000e Driver Bugs on your Management NIC

Your onboard management NIC (00:19.0) uses the Intel Linux e1000e driver, which is notorious for dropping off the network during heavy CPU or PCIe bus loads due to power management and interrupt coalescing.”

then told me to check the log….

journalctl -b -1 -n 200
and low and behold…
Oct 05 21:29:09 desktop-pve kernel: e1000e 0000:00:19.0 nic0: Detected Hardware Unit Hang:
TDH <9c>
TDT <c4>
next_to_use <c4>
next_to_clean <9a>
buffer_info[next_to_clean]:
time_stamp <100904184>
next_to_watch <9c>
jiffies <1009ac040>
next_to_watch.status <0>
MAC Status <80083>
PHY Status <796d>
PHY 1000BASE-T Status <7800>
PHY Extended Status <3000>
PCI Status <10>
Oct 05 21:29:10 desktop-pve systemd[1]: pve-ha-crm.service: Deactivated successfully.
Oct 05 21:29:10 desktop-pve systemd[1]: Stopped pve-ha-crm.service - PVE Cluster HA Resource Manager Daemon.

Ain’t that but a bitch…. and it turns out this is a known problem, now only was the first AI source noting this, here’s a PVE thread started in 2019 still on going, this guy exact same signs and symptoms from April this year, this Garret guy who for some reason used AI to created a service script that is way over kill when you can just disable the offloading feature via a simple edit to the interfaces file…just add

post-up ethtool -K {interface} tso off

under the specific interface you want to disable. verify with:

ethtool -k nic0 | grep tcp-segmentation-offload
tcp-segmentation-offload: off

This time its fixed…. 11PM at night, I swear… fucking tech man…

Leave a Reply

Your email address will not be published. Required fields are marked *