Failover, load balancing, and bonding, in plain English

A technician labels two internet cables connected to a multi-WAN router in a field vehicle.

OPERATE · 10 MIN · LAST FACT-CHECKED JULY 16, 2026

A router with two internet connections can use them in several ways. The labels sound similar, but the results are different.

  • Failover keeps a backup ready and uses it when the preferred path fails.
  • Load balancing spreads different sessions or devices across available paths.
  • Bonding uses a tunnel to distribute packets across more than one path, potentially allowing one protected session to use multiple links.
  • Smoothing is a related technique that can send redundant information across paths to reduce the effect of loss, usually at the cost of more data.

Most small systems do not need every feature. The right question is what should happen to the essential application when one path becomes unusable.

At a glance

Mode Main job Can one session normally use multiple links? Does an active session necessarily survive a path change? Typical cost
Failover Move traffic to a backup No Not necessarily Additional path and a multi-WAN router
Load balancing Share many flows across links Usually no Not by itself Configuration and multiple active paths
Packet-level bonding Combine or protect paths through a tunnel Yes, for traffic inside the tunnel Can, when the design maintains the remote endpoint Tunnel endpoint or cloud service, overhead, configuration
Smoothing / duplication Reduce the effect of loss or jitter Traffic is duplicated rather than simply added Can protect real-time traffic within the supported tunnel Increased data use and tunnel service

The details vary by platform. Treat a product’s feature name as a prompt to read its current documentation, not as proof of a particular behavior.

First, understand the health check

Failover begins with a decision: is a connection healthy enough to use?

A router may check one or more of these:

  • Physical link state.
  • Ability to reach a gateway.
  • DNS response.
  • Ping or another probe to one or more internet targets.
  • Packet loss or latency thresholds.
  • Application-specific reachability.

A cable can remain connected while the path beyond it is unusable. That is why link state alone is often insufficient.

The health-check policy also creates a trade-off:

  • A slow or conservative check avoids unnecessary switching but takes longer to react.
  • An aggressive check reacts faster but may move traffic during a brief, recoverable event.
  • A single test target can create a false failure if that target has a problem.

Write down the detection time, the threshold, and the recovery behavior. Otherwise “automatic failover” remains too vague to test.

Failover: a preferred path and a backup

In a basic failover design, the router sends traffic through the preferred connection. When its health check declares that path unusable, the router changes the route to a backup.

This is often enough for:

  • Web browsing.
  • Email.
  • Cloud applications that reconnect cleanly.
  • Unattended devices that retry automatically.
  • Work where a brief interruption is acceptable.

What failover does not promise

Ordinary failover does not automatically preserve every active session. When traffic leaves through a different provider, the public internet address and network path may change. A meeting, VPN, remote desktop session, upload, or payment connection may need to reconnect.

That does not make failover ineffective. It means the acceptance test should match the requirement:

If the primary cable is disconnected during the real application, what does the user experience, and how long until normal work resumes?

If a short reconnect is acceptable, ordinary failover may be the cleanest answer.

Load balancing: spread the work

Load balancing uses more than one active path and assigns traffic according to a policy.

Common policies include:

  • Send new sessions to the least-used path.
  • Divide traffic by a configured weight.
  • Keep a device, application, or destination on a specific WAN.
  • Prefer one path for work traffic and another for guest or bulk traffic.
  • Use the lower-cost or higher-capacity path until a threshold is reached.

The important practical point is that session-based load balancing normally keeps one session on one path. Ten people may use two connections efficiently, while one large upload still uses only the path assigned to that session.

Load balancing is useful when the goal is aggregate capacity across many devices or flows. It is not the same as adding two speed-test results for every user.

Persistence matters

Some websites, VPNs, payment services, and authenticated applications expect a session to keep the same source address. A good policy keeps related traffic on one path long enough to avoid needless reauthentication or interruption.

This is sometimes called persistence, stickiness, or session affinity. The exact control depends on the router.

Bonding: multiple paths inside a tunnel

Packet-level bonding places traffic inside a tunnel to a remote endpoint. The system can then distribute packets from one protected session across multiple WAN links and present a consistent endpoint to the wider internet.

Peplink’s current SpeedFusion documentation is one vendor-specific example. It describes bandwidth bonding as distributing data at the packet level so one session can use combined links. It separately describes Hot Failover as keeping active sessions online through a path failure. (Peplink SpeedFusion overview)

This can help when the requirement is:

  • More throughput for a single transfer than one available path can provide.
  • Better continuity for an active call, VPN, broadcast, or remote-control session.
  • A consistent public endpoint while local paths change.
  • Protection against short loss events on an unstable path.

Bonding has a remote half

A router cannot generally perform internet bonding by itself. The tunnel needs a compatible endpoint: another appliance, a hosted server, or a vendor cloud service. That endpoint reorders or reconstructs traffic and sends it to the public internet.

The remote half affects:

  • Added latency from the tunnel route.
  • Geographic endpoint selection.
  • Service cost and licensing.
  • Throughput limits.
  • Data usage.
  • What happens if the tunnel service itself is unavailable.

Treat the tunnel as part of the architecture and support plan.

Two paths do not always equal their sum

Bonding can increase useful throughput, but the result depends on link capacity, latency, loss, ordering, overhead, endpoint limits, and the bonding algorithm. A slow or unstable link may contribute less than its headline speed suggests.

If raw speed is the only goal, measure a representative transfer through the intended tunnel. Do not promise arithmetic addition from two isolated speed tests.

Smoothing: spend data to reduce gaps

Some systems can duplicate packets or add recovery information across more than one path. Peplink describes its Smoothing mode as sending redundant packets through multiple channels to reduce the effect of packet loss, and notes elsewhere that smoothing consumes more service usage than basic Hot Failover. (Peplink SpeedFusion overview; Peplink app and service modes)

That trade-off can make sense for live audio, video, or control traffic. It may be wasteful for ordinary browsing or a large download that can retry.

Real-time media is sensitive to more than bandwidth. Microsoft’s current call-quality guidance identifies packet loss, round-trip time, and jitter as important factors in video quality. (Microsoft network-quality guidance)

Use protection selectively. Sending every guest stream or software update through a duplicated tunnel can consume data without improving the work that matters.

Common misunderstandings

“Dual-WAN means bonded”

No. Dual-WAN only means the router can connect to two upstream paths. The configured policy determines whether they are failover, balanced, bonded, or used for different traffic.

“Dual SIM means two active cellular connections”

Not necessarily. Some modems hold two SIMs but use only one at a time. Concurrent cellular paths require the appropriate number of active modems or radios and a router architecture that can use them as intended.

“Failover means a call cannot drop”

Not necessarily. Basic route failover and session-preserving tunnel failover are different behaviors.

“Bonding always makes the connection faster”

No. It adds capability, but also overhead and a remote path. Measure the specific application.

“More paths always make a better system”

No. Each path adds a plan, cable, power draw, configuration, alert, and support question. Add one when it covers a defined failure or capacity need.

Choose by the consequence of failure

A brief reconnect is acceptable

Use ordinary failover. Keep the policy understandable and test the detection and recovery time.

Many users need more aggregate capacity

Use load balancing or traffic steering. Keep important applications on appropriate paths and preserve session affinity where needed.

One transfer needs more throughput

Evaluate packet-level bonding with a suitable remote endpoint. Test the actual transfer and account for tunnel limits and overhead.

An active call or control session should survive a path loss

Evaluate a session-preserving tunnel or hot-failover design. Test it under the same application, motion, and link conditions expected in the field.

Loss and jitter are the main problem

Evaluate selective smoothing or packet duplication for the sensitive traffic. Confirm the extra data use is acceptable.

A practical acceptance test

Do this before calling the system finished.

  1. Record the design. Label WAN 1 and WAN 2, their providers, plans, physical paths, and intended priority.
  2. Record the policy. Note health-check targets, failure threshold, recovery threshold, traffic rules, and tunnel endpoint.
  3. Establish a baseline. Run the real meeting, upload, VPN, payment, or control session on the normal path.
  4. Fail the preferred path. Disconnect it in a controlled and safe way.
  5. Observe the application. Record whether it continues, pauses, reconnects, or fails, and how long recovery takes.
  6. Restore the path. Confirm whether traffic moves back immediately, after a delay, or not until a new session begins.
  7. Repeat. One successful changeover is not enough to understand inconsistent behavior.
  8. Test the other failure. If both paths share power or a router, test or document that common dependency too.

For a moving system, repeat under representative motion and coverage. For an unattended site, include a power-cycle and remote-recovery test.

What we would do

We would choose the least complex mode that meets the application’s recovery requirement:

  • Failover for a tolerable reconnect.
  • Load balancing for many simultaneous flows.
  • Bonding or hot failover for a tested need to protect one session.
  • Smoothing only for traffic that benefits from the additional data.

Then we would document the behavior in seconds and user impact, not just feature names.

Start with the broader design method in Plan the connection before you buy the gear.

Sources

Editorial note: Peplink and SpeedFusion are Peplink trademarks. Peplink documentation is cited for the behavior of its own implementation; other platforms use different feature names and policies. Verify current licensing, service limits, and firmware documentation for the selected system.

Use what you learned. Shop a known component, or bring us the system that still needs a plan.