ZFS dedupe on v2.4.1 will silently corrupt data
What?
The short version is that due to this bug introduced somewhere between
2.3.4 and 2.4.0, ZFS versions probably up to the current latest of v2.4.3 when running with dedup=on can silently
corrupt data.
The patch is not currently released and can be found on the staging branch for zfs-2.4.4.
If you're currently running ZFS with dedupe on, stop immediately if you can or you'll need to upgrade to a staging branch.
Syptoms
Under load the system will silently zero out files it's writing. A sure sign that it's happened is a file which when queried with zdb -bbb -vvv -O $zfs_filesystem_name path/from/root/of/pool shows only an Indirect Block, and not a data block. The file itself will return all zeros, but be the correct size.
The server I found this on is a cache and I'd been chasing some type of an issue for a few weeks to the point I had a separate daemon watching for file corruption and flagging when it was found and repaired. It hadn't occurred to me that this could be coming from ZFS itself.
Important
REALLY IMPORTANT
This issue seems to be caused by any type of object writing on a filesystem with dedupe turned on. I tracked this down because I had two servers syncing with snapshots, and found a zeroed out file on the destination filesystem but not the source.
This is particularly insidious because as far as ZFS is concerned, the snapshots are all perfectly in sync - the corruption has happened during snapshot receive and manifested the same way (zeroing out a random file) despite the same snapshot being fully intact on the source.
Discussion
At this point I'd basically say do not use the ZFS deduplication code for production work if you can avoid it. This is essentially a catastrophic silent data corruption bug. I've used ZFS for over 15 years at this point, it was last on my list of suspects here.
I was using dedupe here because the code had recently gotten an improvement in memory use, so before I found the actual issue on Github disabling dedupe was on the action list (though I really didn't want to because in the specific application it's achieving a vital 30% 50% reduction in overall data size).
Conclusions
In my specific setup I'll be doing a fairly painstaking manual revalidation of all the files on the second server and the primary - this was hitting both. On my non-dedupe systems I've had no issues. It is however extremely disappointing this problem has not been more widely circulated - it is a nasty one and once it starts happening is very frequent (the machine I'm upgrading is still collecting errors, they're just being fixed by a checker script which is how I'll validate the patches.)