Storage

Deduplication and Compression: How Much Space You Really Save

Vendors promise spectacular reduction ratios, but the real savings depend on the data type, the algorithm and the order of operations. We explain what each technology does, which ratios are realistic and the hidden price you pay in CPU, RAM and restore times.

business EasyDataHost calendar_today August 19, 2026 schedule 8 min read

Few marketing figures raise as many expectations — and as many disappointments — as data reduction ratios. "Up to 10:1", "cut your storage by 90%"… Reality is more nuanced: compression and deduplication are different technologies, with different costs, and their effectiveness depends almost entirely on the kind of data you store.

Understanding the difference is not an academic exercise: it determines how many TB you need to buy for your backup repositories, how much RAM your storage server must have, and how long a restore will take when something breaks. It also explains why two providers advertising "the same price per TB" can cost you very different amounts at the end of the month.

In this article we take both technologies apart, give realistic ratios per data type, analyse the hidden cost in CPU and RAM, and review how Veeam, Proxmox Backup Server and storage appliances apply them in practice.

Compression and Deduplication: What Each One Does

Compression works inside each data block: it looks for local redundancy (repeated sequences, frequent symbols) and encodes it more compactly. Each block is reduced independently, with no knowledge of the rest of the storage. It is a pure CPU process, with no global state, which is why it scales well and is easy to reason about.

Deduplication works between blocks: it computes a cryptographic fingerprint (hash) of each block and, if an identical one already exists in the repository, stores only a reference instead of a second copy. It does not shrink each block; it eliminates repeated blocks. To do so it must maintain an index of every known fingerprint, and that index is precisely where its hidden cost comes from.

They are complementary and are almost always applied together: deduplication first discards the repeated blocks, and compression then shrinks the unique blocks that actually have to be written. A backup of 20 Windows virtual machines is the perfect example: deduplication removes the 19 redundant copies of the operating system files, and compression reduces what remains.

Deduplication Types: Inline, Post-Process, File and Block Level

Not all deduplication is created equal. Three axes define how each implementation works:

  • check_circle Inline vs post-process: inline dedup compares and discards blocks before writing them to disk (less space required, more CPU and RAM during the backup). Post-process dedup writes everything first and deduplicates later in the background (it needs extra staging space, but does not penalise the backup window).
  • check_circle File level vs block level: file-level dedup only detects entirely identical files (useful on file servers, useless on VM images). Block-level dedup splits data into chunks and finds redundancy even when the containing files differ, which is the usual case in backup.
  • check_circle Fixed vs variable chunking: with fixed-size blocks, inserting a single byte at the start of a file shifts every following block and breaks dedup. Variable (content-defined) chunking, using techniques such as Rabin fingerprinting, cuts blocks at natural boundaries in the data and survives those shifts, at the cost of more CPU.

The most effective combination for backup is usually inline, block-level dedup with variable chunking. It is also the most computationally expensive, which is why many products settle for middle-ground designs: large fixed blocks, dedup only within each job, or dedup delegated to the repository filesystem.

LZ4, zstd and gzip: Choosing a Compression Algorithm

In backup and storage, three algorithm families cover practically every case:

  • arrow_right LZ4: extreme speed (several GB/s per core) with a moderate ratio. It is the choice when CPU is the bottleneck or when you compress on the production host itself. It decompresses so fast that its impact on restores is negligible.
  • arrow_right zstd: the modern balance. At low levels it approaches LZ4 speed with a noticeably better ratio, and its high levels compete with the heavy compressors. It is today's de facto standard: ZFS, Proxmox Backup Server, btrfs and a growing number of backup tools use it.
  • arrow_right gzip (DEFLATE): the universal veteran. Decent ratios but far slower than zstd at the same compression level. Today it only makes sense for compatibility; for new storage, zstd beats it on almost every front.

The practical rule: LZ4 when speed matters most, zstd for almost everything else, and high zstd levels only for cold data written once and rarely read.

Realistic Ratios per Data Type

The reduction ratio is not decided by the software: it is decided by the redundancy that actually exists in your data. These are the ranges we consistently see in real environments:

Data type Compression Deduplication Typical total reduction
VMs with similar OS 1.5–2x High (repeated OS blocks) 2–4x
Active databases 1.5–2.5x Low–medium 1.5–2.5x
Modern office files (docx, xlsx, PDF) 1.2–1.5x (already compressed) Medium (repeated versions) 1.3–2x
Logs and plain text 3–8x Medium 3–8x
Media: video, images, ZIP ~1x ~1x (except exact copies) ~1x
Data encrypted at source ~1x ~1x ~1x

Two important nuances. First: video, JPEG images and compressed archives have already been through a compressor; compressing them again burns CPU for practically no gain, so do not insist. Second: where deduplication truly shines is in the time dimension — incremental copies of the same VM, day after day, share the vast majority of their blocks, and there the cumulative ratios across the retention chain can indeed reach double digits.

The Hidden Cost: CPU, RAM and Restore Times

Compression costs CPU, and with LZ4 or low-level zstd that cost is nowadays almost irrelevant compared to the savings. Deduplication is a different story: its fingerprint index must be consulted on every single write, and if it does not fit in RAM, every written block triggers additional random reads from disk.

The canonical example is ZFS. Its deduplication table (DDT) stores an entry of several hundred bytes for every unique block in the pool; the classic rule of thumb is around 5 GB of RAM per deduplicated TB. When the DDT overflows memory, write performance collapses dramatically. That is why the general recommendation — reflected in the OpenZFS documentation itself — is that ZFS dedup almost never pays off outside very specific cases. Compression with zstd, on the other hand, almost always does: minimal cost, no global state, and it even benefits read performance because fewer bytes are read from disk. If you want to dig into how ZFS protects the integrity of your data, we cover it in our article on the ZFS filesystem.

The other cost that comparisons tend to forget is the restore. A heavily deduplicated repository turns a sequential read into thousands of scattered reads: your VM's blocks are spread across the entire store, and "rehydrating" the data is especially punishing on repositories backed by spinning disks. An excellent reduction ratio is worthless if your RTO explodes on the day you actually need to restore.

Encryption and Order of Operations: the Mistake That Destroys Ratios

Good encryption produces output that is indistinguishable from random data: no patterns, no repetitions, no redundancy. If you encrypt before deduplicating or compressing, there is nothing left to reduce — two identical blocks encrypted with different keys or initialization vectors no longer resemble each other at all, and the ratio drops to ~1x instantly.

The correct order:

1) Deduplicate → 2) Compress → 3) Encrypt. Let the backup software be the one that encrypts after reducing. If you send data to the repository that is already encrypted at source (BitLocker/LUKS disks dumped raw, application-encrypted files), assume a ~1x ratio and size your storage accordingly.

This is why serious tools integrate encryption into their own pipeline: Veeam encrypts blocks after compressing them, and Proxmox Backup Server encrypts chunks that have already been deduplicated. Security is not compromised and the ratios survive.

Where It Happens in Practice: Veeam, Proxmox Backup Server and Appliances

Veeam applies compression and deduplication per job: blocks are deduplicated within each job's backup chain and compressed with a configurable algorithm (the "optimal" level, based on LZ4, is the recommended balance). On top of that, on ReFS or XFS repositories it takes advantage of the filesystem's block cloning: synthetic fulls are built as references to already existing blocks, taking a fraction of the space and without moving data. It is one of the reasons a well-designed repository matters as much as the software; in our Veeam Cloud Connect offsite backup service, repositories are sized precisely with this in mind.

Proxmox Backup Server puts deduplication at the core of its design: every backup is split into chunks identified by their hash and stored only once for the whole datastore, with per-chunk zstd compression. Each new copy of a VM only contributes the chunks that did not exist yet, so dedup is global and across machines, not just within a job.

Deduplication appliances and storage arrays (Data Domain, StoreOnce and the like) apply inline dedup with variable chunking to everything they receive, and that is where the double-digit ratios on their datasheets come from — measured over weeks of repeated fulls, not over the first backup. Object storage systems, such as our S3 storage, usually bill for the already-reduced data you actually occupy, which brings us to the final point.

Advice when comparing providers:

When you compare prices per TB, always ask whether the billed TB is measured before or after data reduction (front-end vs back-end). A "more expensive" price per TB on already deduplicated and compressed data can end up considerably cheaper than a "budget" one on raw data. Run the numbers for your real case with our Veeam backup calculator.

Frequently Asked Questions

What reduction ratio can I expect for my backups?

It depends on the data type: VMs with similar operating systems reach 2–4x combining dedup and compression; active databases, 1.5–2.5x; logs and plain text, 3–8x with compression alone; and video, images or already-compressed files barely shrink (~1x).

Should I enable ZFS deduplication?

Almost never. The deduplication table must live in RAM (on the order of several GB per deduplicated TB) and, if it does not fit, write performance collapses. Compression with zstd or LZ4, on the other hand, has a minimal cost and pays off in almost every case: always turn it on.

Why don't my encrypted backups deduplicate?

Encryption turns data into statistically random blocks: two identical blocks encrypted with different keys or initialization vectors no longer resemble each other, and neither dedup nor compression can find any redundancy. Reduce first and encrypt afterwards, leaving encryption to the backup software.

Conclusion

Deduplication and compression save real space, but neither as much as the marketing promises nor for free. What is worth remembering:

  • arrow_right Compression shrinks each block; deduplication eliminates repeated blocks. They complement each other and the order matters: dedup, then compression, then encryption.
  • arrow_right Ratios depend on the data: 2–4x for VMs with similar OS, moderate for databases and office files, ~1x for media, compressed archives and data encrypted at source.
  • arrow_right zstd almost always pays off; ZFS dedup almost never does. Dedup demands RAM for its index and penalises restores by scattering blocks.
  • arrow_right When comparing providers, ask whether the billed TB is before or after reduction: it completely changes the price comparison.

If you want to size your backup repositories with realistic ratios instead of brochure figures, contact our team: we will analyse your case and prepare an estimate with no obligation.

Deduplication Compression Backup Storage ZFS Veeam
compress

Backup that saves space without surprises

EasyDataHost: offsite backup with Veeam Cloud Connect, S3 storage and repositories sized with realistic ratios. Infrastructure in Spain, 24/7 support.