Silent data corruption (bit rot) is a real problem that affects every storage system: a bit flips without the operating system noticing, and when you discover the damage weeks or months later, the backup already contains the corrupted version. Traditional filesystems such as ext4 or XFS do not verify data integrity once it has been written to disk, leaving production servers exposed to invisible losses.
ZFS was designed from scratch to solve exactly this problem. Created by Sun Microsystems in 2005 and continued today as OpenZFS, it is far more than a filesystem: it is a complete storage management platform that integrates a filesystem, volume manager and data protection into a single coherent solution.
In this article we explore the architecture of ZFS, its data protection mechanisms, how it compares with other filesystems and why it is a key component of professional storage server infrastructure.
What Is ZFS: From Sun Microsystems to OpenZFS
ZFS (Zettabyte File System) was born in 2005 at Sun Microsystems as part of Solaris 10. Its creator, Jeff Bonwick, designed it to eliminate the limitations of existing filesystems and standalone volume managers. Instead of stacking separate layers (partitions, hardware RAID, filesystem), ZFS unifies them into a single stack that controls the entire path from physical disk to file.
After Oracle acquired Sun in 2010, the open-source community forked the project under the name OpenZFS, which is today's reference implementation. OpenZFS runs natively on FreeBSD, Linux (through the ZFS on Linux kernel module), macOS and even Windows experimentally. Distributions such as Ubuntu, Proxmox and TrueNAS include it as a first-class filesystem option.
Copy-on-Write: The Foundation of Integrity
The fundamental mechanism that sets ZFS apart from conventional filesystems is copy-on-write (CoW). In a traditional system like ext4, when you modify a file, the new data is written over the old data. If the server loses power during the write, the block is left in an inconsistent state and the data is lost or corrupted.
ZFS never overwrites existing data. Every write goes to a new block, and only when the write has been completed and verified with a SHA-256 checksum does the metadata tree pointer atomically update to point to the new block. If the system loses power mid-operation, the original data remains intact because the pointer was never updated.
Furthermore, ZFS verifies checksums on every read. If it detects that a block does not match its checksum (bit rot, controller error, bad sector), it automatically reads the redundant copy from the mirror or reconstructs the block from RAIDZ parity. This self-healing is transparent to applications and guarantees that the data you read is always the data you wrote.
Key concept:
Copy-on-write + end-to-end checksums = ZFS automatically detects and corrects silent data corruption that other filesystems do not even detect. No external tools like fsck are needed.
Vdevs and Zpools: Mirrors, RAIDZ, SLOG and L2ARC
ZFS organises storage into two levels. A vdev (virtual device) is a group of physical disks that work together with a given redundancy configuration. A zpool is a container that groups one or more vdevs and distributes data among them through automatic striping.
- content_copy Mirror: two or more disks with identical copies. Equivalent to RAID 1. Offers the best read performance and fastest recovery from failure, at the cost of higher space usage.
- shield RAIDZ1/Z2/Z3: equivalents of RAID 5/6 with single, double or triple parity. RAIDZ2 tolerates two simultaneous disk failures and is the recommended configuration for storage servers with four or more disks.
- speed SLOG (Separate Log): a dedicated NVMe device that accelerates synchronous writes (ZIL). Critical for databases, NFS and any workload with heavy fsync usage.
- cached L2ARC: an SSD or NVMe that acts as a second-level read cache, extending the RAM-based ARC cache with flash storage for datasets that do not fit in memory.
The advantage of this model is flexibility: you can mix vdevs of different configurations within the same zpool, add vdevs to expand capacity without downtime and hot-swap disks. ZFS manages data distribution transparently, without the need for hardware RAID controllers.
Datasets, Snapshots and Clones
Within a zpool, ZFS organises data into datasets, which are logical units with independent properties (compression, quotas, permissions, mount point). Each dataset can have its own policies without affecting others, which greatly simplifies management on servers with multiple workloads.
Snapshots are instant point-in-time copies of a dataset. Thanks to copy-on-write, creating a snapshot is practically instantaneous and consumes no additional space until the original data changes. They are ideal for consistent backups, restore points before upgrades and protection against human errors (deleting a file, modifying a database).
Clones are writable datasets created from a snapshot. They share common blocks with the original snapshot and only consume space for data that differs. They are perfect for development and testing environments: clone a production database in seconds, test against the copy and destroy it when you no longer need it, without impacting the performance of the original dataset.
Compression and Deduplication
ZFS offers inline compression that operates transparently: data is compressed before being written to disk and decompressed on read. The recommended algorithm for general use is LZ4, which delivers typical compression ratios of 1.5x-2x with virtually zero CPU impact. For archival, ZSTD provides ratios above 3x with higher CPU usage.
Deduplication eliminates identical blocks by storing a single copy and maintaining a reference table (DDT). While it can save space for workloads with highly repetitive data (such as VDI or backups of similar VMs), it requires a significant amount of RAM (approximately 5 GB per TB of deduplicated data). For most production servers, LZ4 compression offers a better cost-benefit ratio than deduplication.
ARC and L2ARC: Intelligent Caching
The ARC (Adaptive Replacement Cache) is the RAM-based read cache in ZFS. Unlike the generic Linux page cache, ARC uses an adaptive algorithm that balances recently accessed blocks with frequently accessed blocks, optimising the hit rate for mixed workloads. ARC dynamically consumes available RAM and releases it when other applications need it.
The L2ARC extends the ARC cache to an SSD or NVMe device. When a block is evicted from the RAM-based ARC, it can be promoted to the flash-based L2ARC before having to be read from mechanical disk. This is especially useful on servers with HDD pools that handle datasets larger than available RAM: the L2ARC provides an intermediate performance tier between RAM and disk.
Comparison Table: ZFS vs ext4 vs XFS vs Btrfs
The following table compares ZFS with the most widely used filesystems on Linux servers:
| Criterion | ZFS | ext4 | XFS | Btrfs |
|---|---|---|---|---|
| Copy-on-Write | Yes | No | No | Yes |
| Checksums | SHA-256/Fletcher4 on data and metadata | Metadata only (journal) | Metadata only | CRC32C on data and metadata |
| Built-in RAID | Mirror, RAIDZ1/Z2/Z3 | No (requires mdadm) | No (requires mdadm) | RAID 0/1/10, RAID 5/6 (unstable) |
| Snapshots | Instant, unlimited | Not native | Not native | Instant, native |
| Compression | LZ4, ZSTD, GZIP inline | No | No | LZO, ZLIB, ZSTD inline |
| Production maturity | 20+ years, Solaris/FreeBSD/Linux | Very high, Linux default | Very high, RHEL default | Maturing, SUSE default |
ZFS Send/Receive: Efficient Replication
One of the most powerful features of ZFS is zfs send and zfs receive, which allow you to serialise a complete snapshot or an incremental delta between two snapshots and send it to another pool, another server or even a file. The receiver recreates the dataset with all its data, properties and metadata intact.
This enables extremely efficient offsite replication strategies: the first time you send the complete snapshot, and from then on you only send incremental deltas (the blocks that changed). A 10 TB dataset with 50 GB of daily changes only needs to transfer those 50 GB, not the full 10 TB. Tools such as syncoid (from Sanoid) automate this process with periodic snapshots and scheduled incremental replication.
ZFS Use Cases
ZFS is especially valuable in scenarios where data integrity and efficient storage management are critical:
- storage NAS/SAN servers: TrueNAS, a widely popular storage operating system, is built on OpenZFS. Ideal for sharing files via SMB/NFS with integrated checksums, snapshots and replication.
- backup Backup repositories: instant snapshots and ZFS send/receive make ZFS a native backup platform with incremental replication, without relying on external backup software for the primary copy.
- database Databases: with SLOG on NVMe for fast synchronous writes and snapshots for consistent backups without stopping the database. PostgreSQL, MySQL and MariaDB run excellently on ZFS.
- dns Virtualisation (Proxmox): ZFS is a natively supported storage backend in Proxmox VE, with VM snapshots, thin provisioning and built-in replication between cluster nodes.
- inventory_2 Long-term archival: the combination of checksums, ZSTD compression and RAIDZ3 makes ZFS the safest option for cold storage where integrity over years is critical.
EasyDataHost and ZFS
EasyDataHost uses ZFS on its dedicated storage servers and on the backup nodes of its infrastructure. ZFS pools with RAIDZ2 and LZ4 compression provide the foundation for storing client backups with bit rot protection, automatic snapshots and offsite replication via ZFS send/receive.
For clients who need a dedicated server with reliable storage, EasyDataHost offers configurations with pre-installed ZFS, including SLOG on NVMe, custom-sized RAIDZ2 pools and snapshot and replication scripts configured out of the box.
The managed services team handles ongoing maintenance: pool health monitoring, proactive alerts for degraded disks, periodic scrubs to detect bit rot, capacity expansion and performance optimisation.
- check_circle RAIDZ2 + LZ4 compression: double parity with transparent compression to maximise capacity and security.
- check_circle Automatic snapshots: retention policies configured to protect against human errors and ransomware.
- check_circle Offsite replication: incremental ZFS send/receive to a second data centre for disaster recovery.
- check_circle Data in Spain: Tier III+ data centre in Madrid with guaranteed data sovereignty.
Conclusion
ZFS is not just a filesystem: it is a complete storage management platform that integrates data protection, redundancy, compression and replication into a single coherent solution. Its copy-on-write architecture with end-to-end checksums solves the silent corruption problem that plagues traditional filesystems, while its instant snapshots and ZFS send/receive provide the most efficient backup and disaster recovery tools available.
- arrow_right ZFS unifies a filesystem and volume manager with copy-on-write and end-to-end checksums.
- arrow_right RAIDZ provides software redundancy without the need for hardware RAID controllers.
- arrow_right Snapshots and clones enable instant backups and testing environments at no space cost.
- arrow_right ARC + L2ARC provide intelligent caching that automatically optimises read performance.
- arrow_right EasyDataHost uses ZFS on its storage and backup servers with RAIDZ2, snapshots and replication.
If you need a professionally managed storage server with ZFS, contact our team to design the configuration that best fits your capacity, performance and data protection requirements.