The same address can mean different things
Two tenants can now use the same private address range in our Cilium-based stack. For a company moving workloads into a private cloud, keeping its addresses avoids changes to application settings, firewall rules and connections to other systems.
A Virtual Private Cloud should make that overlap unremarkable. Each tenant gets its own network, even on shared physical servers, and an address resolves only within the network it belongs to.
Our earlier stack chained a tenant-networking system with Cilium for policy and traffic visibility. Chained Cilium keys identity by IP address, so tenant ranges could not overlap, and startup depended on two networking systems coordinating. We chose to extend Cilium itself and maintain the fork.
Isolation has to hold along the whole path
The central change is that network decisions retain the tenant context. An address in one tenant’s network cannot be treated as the same destination in another, even when the addresses match.
We keep tenant membership separate from security identity. Cilium’s label-based policy still answers which connections are allowed. Tenant context answers which network an address belongs to. That separation lets us add overlapping networks without making tenant membership stand in for access control.
Tenant context extends past routing into address resolution and connection state. With separate routes alone, two tenants could still collide later in the traffic path. Carrying the context that far was the main engineering work in bringing the VPC model into an existing networking stack.
Shared physical servers
Tenant A network
Tenant B network
A route must be intentional
A tenant network shares hardware with the services that run the platform. Sharing that hardware must not create an accidental route to those services.
A new tenant network has no internet path by default. Access requires an explicit route, and a missing route drops the traffic. We also keep tenant routes from reaching the platform’s own service network.
Tenant network
Platform service networkNot reachable from tenant routes
A restart should not change the network
A workload can restart or move while its network configuration is meant to stay the same. Tying that configuration too closely to a running process makes routine maintenance harder.
We give network interfaces a lifecycle separate from the workloads using them. An interface holds the address and access-group membership, so a workload restart does not have to create a new network identity. This also gives migration and recovery a stable configuration to restore.
Access policy follows the interface. Only the Cilium agent turns an interface’s group list into identity labels, so a workload cannot add itself to a group. Policy the platform rejects fails closed.
Recovery is part of the network design
After a failure, the platform has to know which addresses are still allocated and which resources they belong to. Rebuilding from an incomplete view can create conflicting allocations or report a network as ready too soon.
The platform keeps a durable record of network allocations and reconciles the running network against it. That gives recovery an intended state to rebuild. Readiness requires confirmation from the network, rather than just a successfully created resource record.
Model checking tests that recovery logic against partial failures. It covers recovery that stops halfway through and an operation that runs again before the previous attempt has settled.
Test results and the bare-metal plan
The latest cluster run passed 468 of 476 checks on three nodes, with two tenant networks on the same 10.64.0.0/16 range. All eight failures trace to one known endpoint-claim race. Public addresses have also run on our multi-node lab hardware, announced by the node agent answering ARP, with no BGP required.
Next, we’re running the datapath on bare metal. Cross-server delivery, Security Groups including ICMP and NAT gateways come first. The network load balancer, public address pools on customer networks and public address failover times follow.
Work that informed ours
Cilium provides the policy and traffic-processing foundation for this work. Kube-OVN and OpenStack Neutron offer useful examples of tenant networking in private clouds.
Andromeda at Google and VFP at Microsoft put the virtual network in the host datapath at much larger scale. Both shaped how we approached isolation and host networking.
References
- Cilium multi-pool IPAM
The overlapping-CIDR limit this work starts from.
- Cilium security identity
How Cilium derives identity from labels.
- Kube-OVN VPC
Kube-OVN documentation.
- OpenStack Neutron networking
OpenStack documentation.
- Andromeda
Dalton et al., NSDI 2018.
- VFP
Firestone, NSDI 2017.
