So, remember when I first build a test Frankenstein TrueNAS box to edge out some performance on otherwise shady consumer SSDs? and then I ran with that Idea to test ZFS over iSCSI using TrueNAS?
Well since I was using a poormans setup, something wasn’t happy (PVE Virtio network stack? TrueNAS stack? I unno but I was using both VLAN segregation… Virtual NIC setup via Virtio, not sure but the host would become unresponsive to any network requests (pings or otherwise) but the host and VM would be otherwise fine (if I went to the console, I could issue commands without problem, but if I restarted the networking service only the host PVE would become responsive, TrueNAS would stay dead I had to reboot the whole host. AI assumes its the saturation, I have my doubts as I throttled the VLANs which didn’t help. But we’ll try anyway, I’ll move the iSCSI NICs to a PCI network card I apparently forgot I already had in this host. via hardware passthrough.. ok Let’s see how this goes…
I didn’t realize this one active VM is still on the ZFS over iSCSI datastore… move it and… host crash, so doing some testing interesting finds:
- Just as I stated above, restoring the network stack on the PVE host brings it back up fine, but
- Restarting the nework stack or even restarting the TrueNAS VM does not bring it back
- I have to reboot the PVE host as a whole
- The GUI task viewer, showing the task will stall on the transfer but the hosts stay perfectly stable and running fine, the transfer resumes where it left off after the storage PVE host and TrueNAS VM boot back up (takes roughly 20 seconds for PVE to boot, and 1 minute 30 seconds for TrueNAS) 2 min.
- Even when testing had a 6 minute outage window and the transfer resumed just fine when the backend storage came back. Is there a timeout out that checks actual transfer and if its stalled? I don’t know.
- The ZFS storage status in the left hand nav doesn’t reflect unknown status till roughly about 5 minutes of the storage becoming unavailable. But showing unknown back to good is almost instant.
vmbr0), and the VirtIO translation layers. TrueNAS now talks directly to bare-metal Intel hardware using native, highly optimized FreeBSD drivers.e1000e Driver Bugs on your Management NIC
Your onboard management NIC (00:19.0) uses the Intel Linux e1000e driver, which is notorious for dropping off the network during heavy CPU or PCIe bus loads due to power management and interrupt coalescing.”
then told me to check the log….
journalctl -b -1 -n 200
Oct 05 21:29:09 desktop-pve kernel: e1000e 0000:00:19.0 nic0: Detected Hardware Unit Hang: TDH <9c> TDT <c4> next_to_use <c4> next_to_clean <9a> buffer_info[next_to_clean]: time_stamp <100904184> next_to_watch <9c> jiffies <1009ac040> next_to_watch.status <0> MAC Status <80083> PHY Status <796d> PHY 1000BASE-T Status <7800> PHY Extended Status <3000> PCI Status <10> Oct 05 21:29:10 desktop-pve systemd[1]: pve-ha-crm.service: Deactivated successfully. Oct 05 21:29:10 desktop-pve systemd[1]: Stopped pve-ha-crm.service - PVE Cluster HA Resource Manager Daemon.
Ain’t that but a bitch…. and it turns out this is a known problem, now only was the first AI source noting this, here’s a PVE thread started in 2019 still on going, this guy exact same signs and symptoms from April this year, this Garret guy who for some reason used AI to created a service script that is way over kill when you can just disable the offloading feature via a simple edit to the interfaces file…just add
post-up ethtool -K {interface} tso off
under the specific interface you want to disable. verify with:
ethtool -k nic0 | grep tcp-segmentation-offload tcp-segmentation-offload: off
This time its fixed…. 11PM at night, I swear… fucking tech man…
