Network

Load Balancing: Techniques and Architectures

How to distribute traffic across servers intelligently: balancing algorithms, solution comparison, health checks, SSL offloading, load balancer high availability, and session persistence for production architectures.

business EasyDataHost calendar_today June 20, 2026 schedule 9 min read

Any web application that receives real traffic faces, sooner or later, a fundamental problem: a single server has a finite limit of simultaneous connections, CPU, memory and bandwidth. When that limit is reached, users experience slowness, timeouts or outright 503 errors. Scaling vertically (more CPU, more RAM) has a physical and economic ceiling. The architectural solution is to scale horizontally, distributing traffic across multiple servers, and the component that orchestrates that distribution is the load balancer.

A load balancer is not simply a proxy that distributes requests: it is a critical piece of infrastructure that decides which server should handle each connection, verifies that backends are healthy, terminates SSL connections, maintains session persistence and, in many cases, is the first point of contact between the user and the platform. If the load balancer fails, the entire platform stops responding. That is why understanding its techniques, algorithms and high-availability patterns is essential for any systems architect.

In this article we analyse the fundamentals of load balancing: L4 and L7 types, the main distribution algorithms, a comparison of solutions, health checks, SSL offloading, load balancer high availability, sticky sessions, DNS-based balancing and how EasyDataHost implements these techniques in its network infrastructure.

What Is Load Balancing

Load balancing is the technique of distributing incoming network traffic across a group of backend servers (also called a pool or upstream) so that no single server receives more load than it can handle. The component that performs this function is called a load balancer and can be implemented as dedicated hardware, a virtual appliance or software running on commodity servers.

The fundamental objectives of load balancing are threefold: performance (distributing the load so that response times remain optimal), availability (if a backend fails, traffic is automatically redirected to the others) and scalability (adding more servers to the pool without modifying the client configuration). In production environments with demanding SLAs, the load balancer is the component that enables 99.99% uptime commitments.

L4 vs L7 Load Balancing

The most important distinction in load balancing is between layer 4 (transport) and layer 7 (application) of the OSI model. An L4 load balancer operates at the TCP/UDP level: it sees IP addresses, ports and connection flags, but does not inspect the content of the request. It makes routing decisions based solely on the source and destination IP:port tuple. It is extremely fast because it does not need to parse application-layer protocols.

An L7 load balancer, on the other hand, understands the application protocol (HTTP, HTTPS, gRPC, WebSocket). It can inspect HTTP headers, URLs, cookies, virtual hosts and make routing decisions based on content. This enables advanced capabilities such as path-based routing (/api to one pool, /static to another), header rewriting, header injection, per-URL rate limiting and SSL termination with SNI.

Practical rule:

Use L4 when you need maximum performance and do not require content inspection (databases, DNS, mail). Use L7 when you need intelligent HTTP-based routing, centralised SSL termination or header manipulation.

Load Balancing Algorithms

The balancing algorithm determines how the load balancer selects the backend server for each new connection or request. Each algorithm has advantages and disadvantages depending on the type of workload:

  • sync Round Robin: distributes requests sequentially across backends. It is the simplest algorithm and works well when all servers have similar capacity and requests have uniform computational cost.
  • compress Least Connections: sends the request to the server with the fewest active connections at that moment. Ideal when requests have variable duration (slow APIs, downloads, WebSockets) because it avoids overloading servers that are already busy.
  • balance Weighted (Round Robin/Least Conn): assigns a weight to each backend so that more powerful servers receive proportionally more traffic. A server with weight 3 receives three times as many requests as one with weight 1.
  • fingerprint IP Hash: computes a hash of the source IP to always assign the same backend to the same client. It guarantees session affinity without cookies, but can create imbalances if one IP range generates more traffic than others.
  • speed Least Response Time: combines active connections and backend response time. It sends traffic to the server that not only has fewer connections but also responds faster. It requires continuous latency measurement.

Comparison: HAProxy vs Nginx vs Traefik vs F5 vs Cloud LB

The following table compares the most widely used load balancing solutions in production environments:

Criterion HAProxy Nginx Traefik F5 BIG-IP Cloud LB
Type Software OSS Software OSS Software OSS Hardware/Virtual Managed service
Layers L4 + L7 L4 + L7 L7 L4 + L7 L4 + L7
Performance Very high (millions of conn/s) High Medium-high Very high (hardware) Auto-scaling
Configuration Flat file Flat file Auto-discovery (Docker/K8s) GUI + CLI API / Console
Cost Free Free (Plus is paid) Free High (licence) Pay-as-you-go
Best for High traffic, bare metal Web + reverse proxy Microservices, Kubernetes Enterprise, compliance Cloud-native workloads

Health Checks: TCP, HTTP and Custom

A load balancer is only useful if it knows which backends are healthy and which are not. Health checks are periodic probes that the load balancer sends to each server to verify its status. If a backend does not respond or responds with an error, it is automatically removed from the pool until it recovers.

  • lan TCP check: the load balancer attempts to open a TCP connection to the backend port. If the handshake completes, the server is considered healthy. This is the most basic check and does not validate that the application is functioning correctly.
  • http HTTP check: sends an HTTP GET request to a specific endpoint (typically /health or /status) and verifies that the response has a 2xx status code. It validates that the application is responding, not just that the port is open.
  • tune Custom check: executes a script or verifies specific conditions (database connectivity, disk space, CPU load). It allows defining application-specific health criteria that go beyond simple connectivity.

The critical parameters of a health check are the interval (how often the check runs), the timeout (how long to wait for a response) and the threshold (how many consecutive failed checks are needed to mark a backend as down). An interval that is too short generates unnecessary load; one that is too long delays failure detection.

SSL Termination and Offloading

SSL termination (or TLS termination) means that the load balancer is the point where HTTPS traffic is decrypted. The client establishes the encrypted connection with the load balancer, which decrypts the request and forwards it to the backend in plain text (HTTP) or re-encrypted. This is known as SSL offloading because it offloads the computational cost of encryption from the backends.

The advantages are significant: SSL certificates are managed at a single point (the load balancer) instead of on each backend, server performance improves by eliminating the cryptographic workload, and the load balancer can inspect HTTP traffic to make L7 routing decisions. In environments with strict compliance requirements, SSL re-encryption can be configured: the load balancer decrypts, inspects and re-encrypts before sending to the backend, maintaining end-to-end encryption.

Load Balancer High Availability: VRRP and Keepalived

If the load balancer is the single entry point to the platform, it becomes a potential single point of failure. The standard solution is to deploy two or more load balancers in an active-passive (or active-active) configuration using VRRP (Virtual Router Redundancy Protocol).

Keepalived is the most widely used open-source implementation of VRRP on Linux. Two HAProxy servers share a virtual IP (VIP): the master node handles traffic normally and sends periodic VRRP messages to the backup. If the backup stops receiving those messages (because the master has failed), it assumes the VIP within milliseconds and begins handling traffic. The failover is transparent to clients because the IP address does not change.

In active-active configurations, both load balancers receive traffic simultaneously (each with its own VIP or sharing traffic via DNS round robin), doubling processing capacity and eliminating idle resources.

Key concept:

A load balancer without redundancy is an anti-pattern. Always deploy at least two instances with VRRP/Keepalived so that failover is automatic and imperceptible to users.

Session Persistence: Sticky Sessions

Some applications store session state in the server's local memory (shopping carts, login sessions, temporary configurations). If the load balancer sends requests from the same user to different backends, the session is lost. Session persistence (sticky sessions) ensures that all requests from the same user are directed to the same backend.

  • cookie Cookie-based: the load balancer inserts a cookie with the identifier of the assigned backend. On subsequent requests, it reads the cookie and routes to the same server. This is the most reliable method because it works correctly even with NAT and proxies.
  • fingerprint Source IP: uses the client's IP as the affinity key. Simple but problematic when multiple users share the same public IP (corporate networks, carrier-grade NAT).

The general recommendation is to avoid sticky sessions whenever possible by designing stateless applications that store session state in an external service (Redis, database). This allows the load balancer to distribute traffic freely and simplifies horizontal scaling.

DNS-Based Load Balancing

DNS-based load balancing is the simplest form of traffic distribution: multiple A (or AAAA) records are configured for the same domain, each pointing to a different server or load balancer. DNS resolvers naturally rotate responses (DNS round robin), distributing clients across the available IPs.

The main limitation is that DNS has no knowledge of backend health: if a server goes down, DNS will continue returning its IP until the record is manually updated or the TTL expires. Intelligent DNS services (such as Route 53, Cloudflare or NS1) solve this with integrated health checks that automatically remove the IPs of failed servers. They also enable geographic balancing (sending users to the nearest data centre) and weighted DNS to distribute traffic proportionally across regions.

Load Balancing at EasyDataHost

EasyDataHost Cloud infrastructure uses load balancing at multiple levels. At the network layer, HAProxy in high availability with Keepalived distributes incoming traffic across compute nodes. HTTP health checks verify the state of each backend every 5 seconds, automatically removing nodes that fail to respond.

For clients who need dedicated load balancing for their applications, EasyDataHost offers managed services for HAProxy and Nginx configuration with SSL offloading, custom health checks, sticky sessions and active-active setups with VRRP. All deployed on enterprise servers with redundant connectivity in our data centre in Spain.

  • check_circle HAProxy HA with Keepalived: L4/L7 load balancing with sub-second automatic failover.
  • check_circle Centralised SSL offloading: certificate management at a single point with automatic renewal.
  • check_circle Advanced health checks: TCP, HTTP and custom with configurable thresholds.
  • check_circle Data centre in Spain: low latency and guaranteed data sovereignty.

Conclusion

Load balancing is a fundamental component of any production architecture that requires performance, availability and scalability. It is not simply about distributing requests: it involves choosing the right algorithm for your traffic type, implementing health checks that detect real failures, managing SSL centrally, ensuring that the load balancer itself is not a single point of failure and deciding whether you need session persistence or can design your application as stateless.

  • arrow_right L4 for maximum performance without inspection; L7 for intelligent HTTP-based routing.
  • arrow_right Round Robin for uniform workloads; Least Connections for variable-duration requests.
  • arrow_right HAProxy is the open-source reference for high traffic; Traefik for microservices and Kubernetes.
  • arrow_right VRRP + Keepalived eliminate the load balancer's single point of failure with sub-second failover.
  • arrow_right EasyDataHost implements HAProxy HA load balancing with SSL offloading, health checks and data in Spain.

If you need to implement load balancing for your infrastructure or applications, contact our team to design the architecture that best suits your traffic, availability and performance requirements.

Load Balancing HAProxy Network High Availability SSL
balance

Professional load balancing with high availability

EasyDataHost: HAProxy HA with Keepalived, SSL offloading, advanced health checks, sticky sessions and sub-second failover. Data in Spain.