Storage

Ceph: Distributed Storage for High Availability

How Ceph eliminates single points of failure in storage: RADOS architecture, CRUSH algorithm, automatic replication, erasure coding and native Proxmox integration to build truly resilient cloud infrastructures.

business EasyDataHost calendar_today March 28, 2026 schedule 9 min read

Traditional storage based on SAN arrays or local RAID systems has worked for decades, but it carries a fundamental architectural problem: the single point of failure. When a SAN controller fails, when a NAS server runs out of space, or when a RAID array suffers a double disk failure, data becomes inaccessible and the business stops. In environments where downtime has a direct cost (cloud IaaS, SaaS platforms, hospitals, finance), depending on a single centralised storage point is a risk that can no longer be accepted.

The industry's answer to this problem is software-defined storage (SDS) with a distributed architecture, and within this paradigm, Ceph has established itself as the reference open-source solution. Ceph spreads data across tens or hundreds of nodes, eliminates any single component whose failure could cause service loss, and scales horizontally without practical limits.

In this article we explain what Ceph is, how its architecture works, what access types it offers, how it ensures high availability, how it integrates with Proxmox, and what role it plays in EasyDataHost's cloud infrastructure.

What Is Ceph

Ceph is a distributed, open-source, software-defined storage system that provides object, block and file storage on a single unified platform. It was created by Sage Weil in 2004 as part of his doctoral thesis and is now backed by the open-source community and companies such as Red Hat (IBM), SUSE, Canonical and Proxmox.

At the core of Ceph lies RADOS (Reliable Autonomic Distributed Object Store), a distributed object storage layer that manages replication, failure recovery and data redistribution autonomously. All the access protocols that Ceph exposes to the outside world are built on top of RADOS.

What sets Ceph apart from other distributed solutions is the CRUSH algorithm (Controlled Replication Under Scalable Hashing). Instead of maintaining a centralised table indicating where each piece of data resides, CRUSH deterministically calculates the location of each object from its name and a cluster map. This eliminates the need for a centralised metadata service for block and object data, which in turn removes a classic performance and availability bottleneck.

Ceph Architecture: Key Components

A Ceph cluster is composed of several types of daemons that cooperate to store, replicate and serve data. Understanding each component is essential for properly sizing and operating a cluster:

  • storage OSD (Object Storage Daemon): each physical disk in the cluster runs an OSD daemon. It is the basic unit of storage. A typical production cluster has tens or hundreds of OSDs. Each OSD stores data, manages replication with other OSDs and reports its status to the monitor.
  • monitor_heart MON (Monitor): maintains the cluster map (the CRUSH map, the OSD map, the pool map) and manages consensus among nodes using the Paxos protocol. They are deployed in odd numbers (3 or 5) to guarantee quorum in the event of failures.
  • dashboard MGR (Manager): collects performance, status and usage metrics from the cluster. It provides the web dashboard, Prometheus/Grafana integration and management modules such as the automatic balancer.
  • folder_open MDS (Metadata Server): only required when using CephFS (the file system). It manages file system metadata (directories, permissions, names) independently of the data, allowing metadata and data to scale separately.

Data is organised into pools, which define replication or erasure coding rules, and within each pool into placement groups (PGs), which are the logical unit for distributing data among OSDs. The CRUSH map defines the physical topology of the cluster (racks, hosts, disks) and the placement rules that determine how replicas are spread to maximise fault tolerance.

Access Types: Block, File and Object

Ceph is a unified storage platform that exposes three access interfaces on top of the same RADOS infrastructure:

  • view_in_ar RBD (RADOS Block Device): provides block volumes that behave like virtual physical disks. It is the most commonly used mode for virtual machine disks on Proxmox, OpenStack and Kubernetes. It supports thin provisioning, snapshots, clones and live migration.
  • folder_shared CephFS (Ceph File System): a distributed POSIX file system that is mounted on clients as a network volume. Ideal for workloads requiring shared file access across multiple servers, such as HPC, rendering or data repositories.
  • cloud_upload RGW (RADOS Gateway): an S3 and Swift-compatible gateway that exposes the cluster as an object storage service. It allows Ceph to be used as a destination for S3 backup, data lakes or multimedia repositories with the same API as AWS S3.
  • code librados: a native library that allows applications to access RADOS directly without going through the abstraction layers. It offers maximum performance for applications that can integrate at the code level.

Key concept:

Ceph unifies block, file and object storage on a single distributed platform. This simplifies management, reduces the infrastructure required and allows the same cluster to serve VM disks, shared volumes and S3 buckets simultaneously.

High Availability: No Downtime on Failure

High availability in Ceph does not depend on special hardware or redundant controllers: it is an inherent property of its distributed architecture. Ceph offers two data protection mechanisms that are configured at the pool level:

Replication (2x or 3x) writes identical copies of each placement group to different OSDs, located on different hosts or racks according to CRUSH rules. With triple replication, the cluster tolerates the simultaneous failure of two OSDs (or two entire hosts) without losing a single piece of data and without interrupting service. Writes are acknowledged to the client only when all replicas have been written, guaranteeing strong consistency.

Erasure coding fragments each object into k data chunks and m parity chunks. A typical 4+2 configuration tolerates the loss of any 2 fragments with a 50% overhead, compared to 200% for triple replication. It is ideal for cold storage or high-volume pools where space efficiency takes priority over write latency.

When an OSD fails, Ceph automatically initiates the self-healing process: it detects the failure within seconds, marks the OSD as out and begins rebuilding the lost data on the remaining healthy OSDs. The process is completely transparent to applications that are reading and writing data. There is no maintenance window, no manual intervention and no downtime.

Comparison Table: Ceph vs Traditional RAID vs SAN

The following table compares Ceph with traditional storage architectures across the aspects that matter most in production environments:

Criterion Ceph Traditional RAID SAN (FC/iSCSI)
Scalability Horizontal, no practical limit Limited to the server chassis Vertical, limited to the array
Redundancy 2x/3x replication or erasure coding across nodes Parity across local disks (RAID 5/6/10) Dual controllers, internal RAID
Cost Commodity hardware, no licences Low (controller + disks) High (array + licences + FC switches)
Flexibility Block + File + Object on a single platform Local block only Block (some with NAS add-on)
Self-healing Automatic, rebuilds across the entire cluster Slow rebuild on a single server Rebuild within the array
Protocol RBD, CephFS, S3/Swift, librados Local access (SATA/SAS/NVMe) FC, iSCSI, NVMe-oF

Ceph + Proxmox: True Hyperconvergence

One of Ceph's most powerful integrations is with Proxmox VE, the open-source hypervisor that EasyDataHost uses as the foundation for its Cloud IaaS platform. Proxmox includes Ceph natively: a complete Ceph cluster can be deployed and managed from the Proxmox interface itself, without external tools or complex configurations.

This integration enables the construction of hyperconverged (HCI) architectures where the same physical nodes run virtual machines and store Ceph data simultaneously. The result is a simpler infrastructure with fewer components, less cabling and no dependency on external SAN arrays.

VM disks are stored as RBD volumes in the Ceph cluster, which enables critical capabilities: live migration (moving VMs between nodes without downtime, because the storage is shared and accessible from any node), instant block-level snapshots, fast VM cloning and automatic high availability (if a Proxmox node fails, VMs restart on another node with immediate access to the same Ceph disks).

Practical advantage:

With Ceph + Proxmox you do not need a dedicated SAN for shared storage between nodes. Data is automatically replicated across all servers in the cluster, enabling live migration and high availability without additional hardware.

Performance: NVMe, BlueStore and Tuning

Historically, Ceph had a reputation for being slow compared with dedicated SAN arrays. That perception has become completely outdated with the latest Ceph releases and the adoption of NVMe disks. A modern Ceph cluster with NVMe OSDs, 25 GbE or faster networking and BlueStore as the backend can deliver hundreds of thousands of IOPS with sub-millisecond latencies.

BlueStore is the default storage backend in Ceph since Luminous (2017). Unlike the old FileStore that relied on an underlying file system (XFS), BlueStore writes directly to the block device without a filesystem intermediary. This eliminates the double write (write-ahead journal + filesystem), reduces latency and improves performance significantly. In addition, BlueStore supports inline compression and per-block checksums to detect bit-rot.

To squeeze maximum performance, production clusters separate the WAL (Write-Ahead Log) and the DB (metadata database) of BlueStore onto fast NVMe devices, while data can reside on SSDs or HDDs depending on the tier. This architecture combines NVMe latency for writes with HDD capacity for bulk storage.

Ceph Use Cases

Ceph's versatility makes it the reference distributed storage platform for a wide range of production scenarios:

  • cloud Cloud IaaS: VM disk storage with triple NVMe replication, live migration and automatic high availability. This is the primary use case for EasyDataHost Cloud.
  • desktop_windows VDI (Virtual Desktop Infrastructure): thousands of virtual desktops simultaneously accessing RBD disks with low latency. Ceph distributes IOPS across all OSDs, avoiding the bottlenecks typical of a centralised SAN.
  • backup Backup repository: pools with erasure coding to maximise net capacity, combined with CephFS or RGW as a target for Veeam or any backup tool compatible with NFS or S3.
  • inventory_2 Native S3 with RGW: S3-compatible object storage for data lakes, offsite backup, archiving and any application that consumes the S3 protocol. It is the foundation of the EasyDataHost S3 service.
  • videocam Media streaming: video platforms that need massive storage with high read throughput. Ceph with erasure coding delivers large net capacity with elevated sequential performance.

EasyDataHost and Ceph

EasyDataHost's cloud infrastructure is built on Proxmox VE with NVMe Ceph storage in triple replication. Every piece of data written by a virtual machine is simultaneously replicated to three OSDs located on three different physical servers, ensuring that the failure of any disk, server or even rack does not cause data loss or service interruption.

All EasyDataHost Ceph clusters use enterprise NVMe disks with BlueStore, a dedicated 25 Gbps storage network and monitors/managers in high availability with a three-node quorum. The result is storage with sub-millisecond latencies, hundreds of thousands of distributed IOPS and the ability to scale horizontally simply by adding nodes to the cluster.

Beyond VM storage, EasyDataHost offers managed services for monitoring, maintenance and optimisation of Ceph clusters, including version upgrades, performance tuning, capacity expansion and disaster recovery planning.

  • check_circle Triple NVMe replication: every piece of data is replicated across three different physical servers.
  • check_circle Dedicated 25 Gbps network: a separate storage network for Ceph traffic.
  • check_circle Automatic self-healing: transparent reconstruction following any disk or node failure.
  • check_circle Data in Spain: Tier III+ data centre in Madrid with guaranteed data sovereignty.

Conclusion

Ceph has redefined what high-availability storage means. By distributing data across multiple nodes with automatic replication and self-healing, it eliminates the single points of failure that characterise traditional RAID and SAN architectures. Its ability to deliver block, file and object storage on a single platform makes it the ideal foundation for modern cloud infrastructures.

  • arrow_right Ceph is open-source distributed storage that unifies block, file and object storage on top of RADOS.
  • arrow_right The CRUSH algorithm eliminates centralised routing tables, scaling without bottlenecks.
  • arrow_right Triple replication and erasure coding ensure availability despite disk, node or rack failures.
  • arrow_right Integration with Proxmox enables hyperconvergence, live migration and HA without an external SAN.
  • arrow_right EasyDataHost uses NVMe Ceph with triple replication as the foundation of its Cloud IaaS platform.

If you need high-availability distributed storage for your cloud infrastructure, contact our team to design the Ceph architecture that best fits your requirements.

Ceph Storage High Availability Proxmox Cloud
hub

Distributed Ceph storage with triple NVMe replication

EasyDataHost Cloud: Proxmox + NVMe Ceph, triple replication, automatic self-healing, live migration, data in Spain. No single points of failure.