Scaling in a CUPS Architecture: How to Grow Without Disrupting Subscribers

October 5, 2026
Mobile Networks
5 out of 5
Scaling in a CUPS Architecture: How to Grow Without Disrupting Subscribers
When scaling the User Plane horizontally, new traffic-processing nodes are added to the cluster. But the entire traffic flow cannot simply be redistributed from scratch. Active sessions must retain their state and continue to be processed on the same node. Therefore, new connections and existing sessions are distributed according to different rules, while active sessions are migrated gradually. In this article, we explain why disruptions occur and how User Plane scaling works when increasing capacity does not come at the expense of inconvenience for individual subscribers.

Cluster Architecture: Control and User Planes

Inside the cluster, components are divided into the Control Plane and User Plane.

The User Plane performs traffic-handling functions: receiving, processing, and forwarding subscriber traffic.

One of the key User Plane functions is DPI — a system for classifying traffic by sessions and applications. DPI nodes receive packets, process connections, apply the required policies, and collect subscriber usage statistics. These are the nodes that are added when capacity needs to be increased or taken out of service for maintenance.

VAS Experts cluster architecture with Control and User Planes

Figure 1 — Cluster architecture

The Control Plane stores subscriber session state and assigns the User Plane. Based on the number of active nodes, it determines which node serves a particular subscriber, which parameters and services are associated with that subscriber, and performs session migration when this assignment changes.

The Requirement Everything Follows From

Let’s start not with scaling itself, but with the constraint that makes it more complex.

Traffic classification and per-service accounting require both directions of a connection to be handled by the same DPI (User Plane) node.

Moreover, the unit of service is not an individual IP address or even a single connection. A subscriber can use several services simultaneously, have multiple connections, use IPv4 and IPv6, or share internet access with other devices. Therefore, one subscriber may have multiple connections and addresses, but from a service perspective, they belong to a single service session. This session becomes the unit of scaling: all connections associated with the subscriber must be handled together by the same traffic-processing node and the same control node.

This leads to the key rule

The unit of scaling is not a flow, an address, or an individual connection, but the service session itself, together with all its parameters: connections, IPv4 and IPv6 addresses, tariff plan, quota, and accounting data. This is the unit of distribution and cannot be split between nodes.

Service session diagram

Figure 2 — Service session: all connections and addresses of one subscriber form a single service unit

If packets belonging to the same TCP connection reach different nodes, this can cause the connection to break, the classification context to be lost, and a single session to be split across multiple nodes for traffic accounting.

This is particularly clear in the case of NAT. If the address and port translation is stored on the old node while the next packet reaches a new node, the new node will not find the corresponding entry and the traffic will be dropped. For an established connection, this results in a disruption.

Two Ways to Distribute Subscribers Across Nodes and How They Differ

How can new sessions be distributed across nodes so that packets belonging to each individual session always reach the same node?

Hashing by Subscriber Identifier

One common and effective approach is to calculate a hash based on the IMSI/APN pair of the subscriber’s PDN session. We take the subscriber identifier, IMSI+APN, calculate its hash, match the result against the list of active nodes, and obtain the number of the node to which the session should be sent. Since a subscriber has one IMSI, all of their sessions within the same APN, regardless of the IP address used (IPv4 or IPv6), are directed to the same processing node. This approach does not require state synchronization between nodes: as long as each node has the same information about the cluster composition, each independently calculates the same hash and node number.

When a new node is added, an algorithm such as Maglev or Randevu rebuilds the distribution so that routing changes for only a minimal portion of subscribers — approximately 1/(N+1) of the total when the cluster grows from N to N+1 nodes. The rest continue to be served by the same nodes.

But there is also a fundamental limitation: the algorithm does not take into account the current state of a session on the node.

As soon as a new node is added to the cluster, the hash result changes for some subscribers. The algorithm directs their new connections to another node. At the same time, the previous node may still contain the service context for these subscribers: information about connections, usage counters, and other data required to process current sessions. The new node does not yet have this state. If the new distribution is applied immediately, active sessions will end up on nodes that do not have their context. Therefore, it is not enough to ensure that each individual connection reaches the correct node. The subscriber-related context must also remain intact.

Initial Placement and Controlled Migration

Hashing solves only one part of the problem: initial distribution. The next question is: “What should be done with sessions that are already active?” To address this, it is useful to separate two processes: the initial placement of new sessions and the controlled migration of existing ones.

Controlled Migration

Hashing is still used for new sessions — it distributes the load quickly and evenly and also serves as a fallback mode if the Control Plane is temporarily unavailable. Active sessions, however, require a separate mechanism that allows them to be redistributed gradually.

This is handled by PCEF, a Control Plane component. When the number of DPI nodes changes, PCEF goes through all of its sessions. For each one, it compares the current node assignment with the target assignment calculated using consistent hashing based on the new number of nodes. If they differ, PCEF switches the session to the target node. This is how a smooth User Plane migration is implemented.

Therefore, when a new traffic-processing node is added, the Control Plane does not need to recalculate the entire system from scratch. Instead, it can select specific sessions for migration.

The scaling of the Control Plane itself — PCEF instances — works in a similar way. When their number increases, each active PCEF goes through the sessions it currently serves and releases those that, according to the new hash, should now belong to the newly added instance. If the number of PCEF instances decreases — for example, because one instance fails — the remaining instances go through not only their own sessions but all sessions in the shared storage. They find those that have been left without an owner and take over the ones that, according to the current hash, now belong to them.

Comparison of two methods for distributing subscriber sessions across nodes

Figure 3 — Two ways to distribute subscriber sessions across nodes

For example, an operator adds a new DPI node. New connections immediately start being assigned to it. The Control Plane then gradually selects some active sessions and moves them away from overloaded nodes. Once the state has been prepared, the assignment is changed and subsequent traffic follows the new route.

For the subscriber, nothing fundamentally changes: the video call continues, a file download does not have to start over, and accumulated accounting data is preserved.

The UP load balancer itself does not make the redistribution decision. It receives the command, prepares the session state, and continues processing after the switch. This is a fundamental difference from an architecture where each node independently calculates its assignment using a hash.

This creates a clear separation of roles: the Control Plane calculates the target assignment and performs the migration, while the User Plane processes traffic according to that assignment.

More about the Control and User Planes can be found in the article — CUPS for PCEF/PGW: Why Separate the Control Plane and User Plane in the Mobile Core.

The Cost of Control

For controlled migration to work correctly, it is necessary to ensure consistency of the data used by PCEF to make decisions, control the migration rate, and account for sessions whose state cannot be safely migrated, such as NAT.

Consistency of Node Information

User Plane nodes do not calculate assignments themselves — they receive them from the Control Plane and apply them when processing traffic. Therefore, it is important to ensure that every node is actually working with the current version.

Whenever the number of nodes changes, PCEF calculates the target node for each session based on the new number of DPI nodes, which it retrieves from Consul. If there are several PCEF instances in the cluster, it is important that all of them see the same number of nodes at the same time. Otherwise, different PCEF instances may calculate different target nodes for the same session, and the distribution will no longer be consistent.

Migration Takes Time

Migrating a single session takes a fraction of a second, but there may be millions of sessions, and they have to be migrated sequentially rather than all at once. Migration takes time, but active sessions remain uninterrupted.

Not Everything Can Be Migrated

Finally, not all sessions are equally portable. Some parts of their state can be safely copied to another node, while other parts are tightly bound to the specific hardware on which the session was created. NAT is one such case: the address and port translation entry exists only in the memory of a specific node and cannot simply be recreated on another node without risking the loss of already established connections.

Such sessions are handled differently — not through migration, but through natural termination: the node stops accepting new sessions, while existing ones continue to be served there until they end naturally.

How to Verify That It Works: Three Measurable Criteria

Everything described above — IMSI-based hashing and controlled migration — exists for the sake of three specific properties that can be measured on real traffic rather than simply declared.

Property How it is verified If the property is violated
Both directions of a session on the same node Verify that incoming and outgoing traffic for the same session are processed on the same node Per-service accounting and classification may be incomplete
Migration without packet loss Verify that traffic continues without interruption while the session is being migrated The subscriber may experience a connection drop
Migration without accounting loss Verify that the total accounting volume on the original and new nodes matches actual consumption An accounting error may only become apparent when the final report is generated

The third property deserves special attention. During controlled migration, the original node sends the final usage data before the session is completed. The Control Plane combines it with the data from the new node and adjusts the usage threshold. This prevents data from being lost or counted twice.

At the same time, there is no need to close and reopen the charging session. The subscriber continues using the service without interruption, while usage accounting continues based on data from both nodes.

What This Scaling Approach Provides

The main advantage of this approach is that an operator can change the infrastructure composition without massive redistribution of active sessions. A new node can immediately be used for new connections, while existing sessions can be migrated gradually as needed.

This produces several practical effects:

  • Capacity can be increased without taking the network offline. The new node starts accepting connections immediately, while active sessions are migrated in the background.
  • Load is rebalanced selectively. An overloaded node is relieved by exactly the amount required rather than through a complete recalculation of connections.
  • Subscriber state remains intact. Connections, addresses, tariff plan, quota, and per-service accounting remain associated with the subscriber rather than being split across multiple nodes.
  • Planned maintenance does not turn into emergency failover. When a node is taken out of service, its sessions can be migrated to other nodes in advance instead of terminating all connections at once.

Network growth is measured in terabits, while the result of scaling is experienced at the level of an individual connection. An architecture in which the unit of distribution is not a flow or even an individual session, but the subscriber together with everything associated with them, simply makes it impossible to overlook this requirement at the design stage.