{"id":1888,"date":"2026-10-05T20:28:28","date_gmt":"2026-10-06T01:28:28","guid":{"rendered":"https:\/\/zewwy.ca\/?p=1888"},"modified":"2026-10-05T22:43:58","modified_gmt":"2026-10-06T03:43:58","slug":"virtualized-truenas-crashing","status":"publish","type":"post","link":"https:\/\/zewwy.ca\/index.php\/2026\/10\/05\/virtualized-truenas-crashing\/","title":{"rendered":"Virtualized TrueNAS Crashing"},"content":{"rendered":"<p>So, remember when I <a href=\"https:\/\/zewwy.ca\/index.php\/2026\/09\/13\/using-freenas-as-a-vm-to-test-storage-speeds\/\">first build a test Frankenstein TrueNAS box to edge out some performance on otherwise shady consumer SSD<\/a>s?\u00a0 and <a href=\"https:\/\/zewwy.ca\/index.php\/2026\/10\/01\/least-privilege-zfs-over-iscsi-for-proxmox-ve-via-truenas\/\">then I ran with that Idea to test ZFS over iSCSI using TrueNAS<\/a>?<\/p>\n<p>Well since I was using a poormans setup, something wasn&#8217;t happy (PVE Virtio network stack? TrueNAS stack? I unno but I was using both VLAN segregation&#8230; Virtual NIC setup via Virtio, not sure but the host would become unresponsive to any network requests (pings or otherwise) but the host and VM would be otherwise fine (if I went to the console, I could issue commands without problem, but if I restarted the networking service only the host PVE would become responsive, TrueNAS would stay dead I had to reboot the whole host. AI assumes its the saturation, I have my doubts as I throttled the VLANs which didn&#8217;t help. But we&#8217;ll try anyway, I&#8217;ll move the iSCSI NICs to a PCI network card I apparently forgot I already had in this host. via hardware passthrough.. ok Let&#8217;s see how this goes&#8230;<\/p>\n<p>I didn&#8217;t realize this one active VM is still on the ZFS over iSCSI datastore&#8230; move it and&#8230; host crash, so doing some testing interesting finds:<\/p>\n<ol>\n<li>Just as I stated above, restoring the network stack on the PVE host brings it back up fine, but<\/li>\n<li>Restarting the nework stack or even restarting the TrueNAS VM does not bring it back<\/li>\n<li>I have to reboot the PVE host as a whole<\/li>\n<li>The GUI task viewer, showing the task will stall on the transfer but the hosts stay perfectly stable and running fine, the transfer resumes where it left off after the storage PVE host and TrueNAS VM boot back up (takes roughly 20 seconds for PVE to boot, and 1 minute 30 seconds for TrueNAS) 2 min.<\/li>\n<li>Even when testing had a 6 minute outage window and the transfer resumed just fine when the backend storage came back. Is there a timeout out that checks actual transfer and if its stalled? I don&#8217;t know.<\/li>\n<li>The ZFS storage status in the left hand nav doesn&#8217;t reflect unknown status till roughly about 5 minutes of the storage becoming unavailable. But showing unknown back to good is almost instant.<\/li>\n<\/ol>\n<div class=\"otQkpb\" role=\"heading\" aria-level=\"3\" data-animation-nesting=\"\" data-sfc-cp=\"\" data-wiz-attrbind=\"aria-level=EDE91b_1o\/fLk2Md\" data-sfc-root=\"ep\" data-epip=\"\" data-complete=\"true\" data-processed=\"true\" data-sae=\"\">The Solution: Hardware Isolation via PCIe Passthrough<!--TgQPHd|||[]--><\/div>\n<div class=\"n6owBd awi2gc\" data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIFBAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">The definitive fix is to physically separate the traffic by cutting Proxmox completely out of the storage loop. My motherboard has a spare PCIe slot, added a dedicated network card (a cheap, enterprise-stable multi-port Intel Gigabit adapter) and passed it directly to the storage VM.<\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIFBAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\"><\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIFBAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">\n<div class=\"n6owBd awi2gc\" data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGhAA\" data-complete=\"true\" data-processed=\"true\">Find the dedicated network card&#8217;s IOMMU group and add it as a <strong class=\"rQesXe MPyX\" data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-complete=\"true\" data-processed=\"true\">Raw PCI Device<\/strong> to the TrueNAS VM hardware tab in Proxmox. Ensure <strong class=\"rQesXe MPyX\" data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-complete=\"true\" data-processed=\"true\">All Functions<\/strong> and <strong class=\"rQesXe MPyX\" data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-complete=\"true\" data-processed=\"true\">PCI-Express<\/strong> are checked.<\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGhAA\" data-complete=\"true\" data-processed=\"true\"><\/div>\n<div class=\"n6owBd awi2gc\" data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">This completely bypasses the Proxmox network stack, the virtual Linux bridge (<code class=\"KDcb0c\" dir=\"ltr\" data-sfc-root=\"ep\" data-epip=\"\" data-complete=\"true\" data-sae=\"\">vmbr0<\/code>), and the VirtIO translation layers. TrueNAS now talks directly to bare-metal Intel hardware using native, highly optimized FreeBSD drivers.<\/div>\n<\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\"><\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">After doing this I haven&#8217;t had the hosting PVE server or TrueNAS VM stop responding to pings, or services. It&#8217;s a lil less Frankensteiny now.<\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\"><\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">&#8230; or so I thought that separating the iSCSI traffic resolved it, everything far more stable, my usual instant test (Booting Windows VM against the SSD backed ZFS over iSCSI target) was working just fine&#8230; So, while I thought this resolved it, I jumped back to my AI chat window to state that this fix didn&#8217;t in fact work. I took a snip of my intuitional moment&#8230;<\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\"><\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\"><a href=\"https:\/\/i.imgur.com\/KKhOTT2.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter\" src=\"https:\/\/i.imgur.com\/KKhOTT2.png\" alt=\"\" width=\"1015\" height=\"634\" \/><\/a><\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\"><\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">Which it then gave some examples of what it could be, I was kinda shocked when one of it&#8217;s hunches, was backed by hard evidence in the logs after the fact.. it stated and I quote..<\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\"><\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">&#8220;<span style=\"font-size: 1rem;\">Clear out the <\/span><code class=\"KDcb0c\" dir=\"ltr\" data-sfc-root=\"ep\" data-epip=\"\" data-complete=\"true\" data-sae=\"\">e1000e<\/code><span style=\"font-size: 1rem;\"> Driver Bugs on your Management NIC<\/span><\/p>\n<div class=\"n6owBd awi2gc\" data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAEIKhAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">\n<p>Your onboard management NIC (<code class=\"KDcb0c\" dir=\"ltr\" data-sfc-root=\"ep\" data-epip=\"\" data-complete=\"true\" data-sae=\"\">00:19.0<!--TgQPHd|||[]--><\/code>) uses the Intel Linux <code class=\"KDcb0c\" dir=\"ltr\" data-sfc-root=\"ep\" data-epip=\"\" data-complete=\"true\" data-sae=\"\">e1000e<!--TgQPHd|||[]--><\/code> driver, which is <a href=\"https:\/\/forum.proxmox.com\/threads\/proxmox-host-loses-connection-on-management-nic-reboot-fixes-it.182314\/\">notorious for dropping off the network during heavy CPU or PCIe bus loads due to power management and interrupt coalescing.<\/a>&#8221;<\/p>\n<p>then told me to check the log&#8230;.<\/p>\n<\/div>\n<pre data-sfc-root=\"ep\" data-hveid=\"CAEIKhAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">journalctl -b -1 -n 200<\/pre>\n<\/div>\n<div data-sfc-cp=\"\" data-sfc-root=\"ep\" data-epip=\"\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">and low and behold&#8230;<\/div>\n<pre data-sfc-root=\"ep\" data-hveid=\"CAAIGxAA\" data-complete=\"true\" data-processed=\"true\" aria-owns=\"action-menu-parent-container\">Oct 05 21:29:09 desktop-pve kernel: e1000e 0000:00:19.0 nic0: Detected Hardware Unit Hang:\r\nTDH &lt;9c&gt;\r\nTDT &lt;c4&gt;\r\nnext_to_use &lt;c4&gt;\r\nnext_to_clean &lt;9a&gt;\r\nbuffer_info[next_to_clean]:\r\ntime_stamp &lt;100904184&gt;\r\nnext_to_watch &lt;9c&gt;\r\njiffies &lt;1009ac040&gt;\r\nnext_to_watch.status &lt;0&gt;\r\nMAC Status &lt;80083&gt;\r\nPHY Status &lt;796d&gt;\r\nPHY 1000BASE-T Status &lt;7800&gt;\r\nPHY Extended Status &lt;3000&gt;\r\nPCI Status &lt;10&gt;\r\nOct 05 21:29:10 desktop-pve systemd[1]: pve-ha-crm.service: Deactivated successfully.\r\nOct 05 21:29:10 desktop-pve systemd[1]: Stopped pve-ha-crm.service - PVE Cluster HA Resource Manager Daemon.<\/pre>\n<p>Ain&#8217;t that but a bitch&#8230;. and it turns out this is a known problem, now only was the first AI source noting this, <a href=\"https:\/\/forum.proxmox.com\/threads\/e1000e-eno1-detected-hardware-unit-hang.59928\/\">here&#8217;s a PVE thread started in 2019 still on going<\/a>, <a href=\"https:\/\/forum.proxmox.com\/threads\/proxmox-host-loses-connection-on-management-nic-reboot-fixes-it.182314\/\">this guy exact same signs and symptoms from April this year,<\/a> <a href=\"https:\/\/www.garrettlaman.com\/Homelab\/Fixing-Intel-e1000e-NIC-hangs-on-Proxmox-nodes\">this Garret guy who for some reason used AI to created a service script that is way over kill<\/a> when you <a href=\"https:\/\/forum.proxmox.com\/threads\/intel-nic-e1000e-hardware-unit-hang.106001\/page-6\">can just disable the offloading feature via a simple edit to the interfaces file&#8230;<\/a>just add<\/p>\n<pre class=\"bbCodeCode\" dir=\"ltr\" data-xf-init=\"code-block\" data-lang=\"\"><code>post-up ethtool -K {interface} tso off<\/code><\/pre>\n<p>under the specific interface you want to disable. verify with:<\/p>\n<pre>ethtool -k nic0 | grep tcp-segmentation-offload\r\ntcp-segmentation-offload: off<\/pre>\n<p>This time its fixed&#8230;. 11PM at night, I swear&#8230; fucking tech man&#8230;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>So, remember when I first build a test Frankenstein TrueNAS box to edge out some performance on otherwise shady consumer SSDs?\u00a0 and then I ran with that Idea to test ZFS over iSCSI using TrueNAS? Well since I was using a poormans setup, something wasn&#8217;t happy (PVE Virtio network stack? TrueNAS stack? I unno but &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/zewwy.ca\/index.php\/2026\/10\/05\/virtualized-truenas-crashing\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Virtualized TrueNAS Crashing&#8221;<\/span><\/a><\/p>\n","protected":false},"author":3,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6,7],"tags":[461,465],"class_list":["post-1888","post","type-post","status-publish","format-standard","hentry","category-networking","category-storage","tag-pve","tag-truenas"],"_links":{"self":[{"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/posts\/1888","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/comments?post=1888"}],"version-history":[{"count":4,"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/posts\/1888\/revisions"}],"predecessor-version":[{"id":1893,"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/posts\/1888\/revisions\/1893"}],"wp:attachment":[{"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/media?parent=1888"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/categories?post=1888"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/zewwy.ca\/index.php\/wp-json\/wp\/v2\/tags?post=1888"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}