Proteus: a hypervisor-agnostic VPC for Triton Cloud
Proteus implements Triton Cloud VPCs in the kernel of each compute node. It attaches below the hypervisor, so the same code runs under bhyve on illumos and on FreeBSD.
By Triton DataCenter ·
Proteus is the VPC dataplane in Triton Cloud. It runs in the kernel of each compute node, between each VM's virtual NIC and the physical network, and applies the firewall, routing, NAT, and Geneve encapsulation to every packet. On two illumos compute nodes connected by 2x10G LACP, VM-to-VM traffic through Proteus reaches 19.4 Gbit/s, which is line rate after encapsulation overhead. Proteus attaches below the hypervisor, so the same policy and packet engine run on illumos and FreeBSD, with a small driver for each kernel.
The VPC is implemented on the compute nodes
A Triton Cloud VPC has no configuration in the physical network. The switches forward IPv6 between compute nodes and know nothing about tenants. Each tenant network, with its subnets, firewall rules, routes, NAT, and floating IPs, is implemented by Proteus on the compute nodes that run the tenant's VMs.
When a VM sends a packet, Proteus on that node checks the packet against policy, rewrites it if needed, adds a Geneve header, and sends it over the physical network (the underlay) to the node that runs the destination VM. Proteus on that node removes the Geneve header, checks policy again, and delivers the packet to the VM.
We chose this model so that creating a tenant, a subnet, or a firewall rule never requires a switch change. The cost is that every packet needs a policy decision in the host kernel, so the dataplane has to be fast.
Proteus follows the VFP design
The design of Proteus follows VFP, the virtual switch that Daniel Firestone of Microsoft described in VFP: A Virtual Switch Platform for Host SDN in the Public Cloud at NSDI 2017. The paper reports that VFP had run on more than a million Azure hosts. We use three of its ideas.
Policy is a sequence of layers. Each layer is a table of match-action rules for one concern. Proteus has six: gateway services, egress posture, firewall, routing, NAT, and encapsulation. A packet passes through the layers in order, and each matching rule adds a header transformation.
The layer walk happens once per flow. Proteus combines the transformations from all six layers into one and stores it in a flow table keyed by the flow. Later packets in the same flow need one hash lookup and one rewrite.
Flows are stateful. When the firewall allows an outbound connection, it records the flow, so the replies are allowed without an inbound rule. Tenants expect this behavior from a security group.
For the packet engine we use OPTE, an open-source Rust implementation of this design by Oxide Computer, licensed MPL-2.0. We include it unmodified as a git submodule and use it as a library. The engine supplies the flow table, layer and rule evaluation, header transforms, and TCP state tracking. It does not define a network. Everything that defines a Triton VPC is Proteus code.
Proteus is four components
triton-vpcProteusproteus-coreProteusPlatform trait.triton-vpc holds the product decisions: one VNI per VPC, an IPv6 underlay, floating IPs, the metadata service, and the operator's egress posture. They change as the product changes, so they live in one crate we own and can change without touching the engine. The flow key is the VNI plus the inner 5-tuple, so two VPCs that use the same addresses never share a flow table entry.
proteus-core holds the kernel code where a bug leaks traffic between tenants or panics the node. We want one copy of that code, not one per OS. Because it is written against the Platform trait, it also runs in tests on a laptop, against a fake platform that records every call it makes to the OS.
The OS driver is the only part that is specific to a kernel. It implements Platform: how a VM port attaches, how Geneve is taken off the underlay link, how the next-hop MAC is looked up, and how counters and the control device are exposed.
How a packet crosses the VPC
The example uses the two acme-prod VMs from Figure 1. web is 10.0.0.3 on node A, db is 10.0.0.4 on node B, and web opens a TCP connection to db on port 5432.
web first needs the MAC address of its gateway, 10.0.0.1, so it sends an ARP request. The request does not leave node A. The gateway layer answers it, and also answers DHCP and IPv6 neighbor discovery. No router exists; the gateway layer on the local compute node generates these replies.
The SYN then arrives at web's Proteus port. The engine builds a flow key from the VNI and the inner 5-tuple and looks it up in the flow table. The lookup misses, so the packet goes through the six layers in order.
The overlay layer looks up 10.0.0.4 in the virtual-to-physical (V2P) table. The result is node B's underlay address and db's MAC address. The layer adds outer Ethernet, IPv6, UDP, and Geneve headers, writes the VPC's VNI, and changes the inner destination MAC from the gateway's MAC to db's MAC.
The driver then looks up the MAC address of the next hop toward node B in the host's IPv6 neighbor table (ip2mac on illumos, nd6 on FreeBSD) and caches it. The NIC computes the outer UDP checksum when it supports that offload. The frame goes out on the underlay link.
Proteus on node B inspects frames on its underlay link. Frames that are not Geneve to node B's own underlay address, such as SSH, NTP, or storage traffic, go to the host IP stack unchanged. For a Geneve frame, Proteus puts the VNI into the flow key before the engine sees the packet, so the packet can only match state that belongs to its own VPC.
web to db. Steps 1 to 4 happen on node A. Step 5 is the only traffic the physical network carries. If web and db ran on the same node, the overlay layer would detect the local destination in step 3, and Proteus would deliver the frame to db's port without encapsulation.Geneve carries the VM's frame over IPv6
Proteus uses Geneve (RFC 8926) to carry VM traffic between compute nodes. Geneve puts the VM's complete Ethernet frame in a UDP datagram to port 6081, after an 8-byte Geneve header. The main field in that header is the 24-bit Virtual Network Identifier (VNI). Each Triton VPC has its own VNI. Two VPCs can both use 10.0.0.0/16, and the receiving node uses the VNI to keep their packets apart.
Figure 5 shows the SYN from step 5 as it crosses the underlay.
db's MAC.Geneve
Added by Proteus · bytes 62–69 · 8 bytes
| Field | Bytes | Value | Note |
|---|---|---|---|
| 62 | 0 / 0 | Version 0, no options. | |
| 63 | 0 / 0 / 0 | Not a control packet, no critical options. | |
| 64–65 | 0x6558 | Transparent Ethernet Bridging: the payload is an Ethernet frame. | |
| 66–68 | 4097 (0x001001) | Identifies the VPC. The receiving node adds it to the flow key before any lookup. | |
| 69 | 0x00 |
All seven packet headers and field details
Outer Ethernet
Added by Proteus · bytes 0–13 · 14 bytes
| Field | Bytes | Value | Note |
|---|---|---|---|
| 0–5 | 3c:fd:fe:0b:00:01 | Next-hop MAC toward fd00:cabe::b, from the host's IPv6 neighbor table. | |
| 6–11 | 3c:fd:fe:0a:00:01 | This node's underlay NIC. | |
| 12–13 | 0x86dd IPv6 | The underlay carries IPv6 only. |
Outer IPv6
Added by Proteus · bytes 14–53 · 40 bytes
| Field | Bytes | Value | Note |
|---|---|---|---|
| 14–17 | 6 / 0 / 0 | ||
| 18–19 | 70 bytes | UDP, Geneve, and the VM's frame. | |
| 20 | 17 (UDP) | ||
| 21 | 128 | Engine default. | |
| 22–37 | fd00:cabe::a | Node A's tunnel endpoint. | |
| 38–53 | fd00:cabe::b | Node B's tunnel endpoint, from the V2P lookup of 10.0.0.4. |
Outer UDP
Added by Proteus · bytes 54–61 · 8 bytes
| Field | Bytes | Value | Note |
|---|---|---|---|
| 54–55 | 49623 | Hash of the inner flow, so LACP and ECMP spread different connections across links. Example value. | |
| 56–57 | 6081 (Geneve) | IANA-assigned Geneve port. | |
| 58–59 | 70 bytes | ||
| 60–61 | 0xc854 | Computed over the IPv6 pseudo-header. Proteus offloads it to the NIC when the NIC supports it; this page computes it. |
Geneve
Added by Proteus · bytes 62–69 · 8 bytes
| Field | Bytes | Value | Note |
|---|---|---|---|
| 62 | 0 / 0 | Version 0, no options. | |
| 63 | 0 / 0 / 0 | Not a control packet, no critical options. | |
| 64–65 | 0x6558 | Transparent Ethernet Bridging: the payload is an Ethernet frame. | |
| 66–68 | 4097 (0x001001) | Identifies the VPC. The receiving node adds it to the flow key before any lookup. | |
| 69 | 0x00 |
Inner Ethernet
From the VM · bytes 70–83 · 14 bytes
| Field | Bytes | Value | Note |
|---|---|---|---|
| 70–75 | 90:b8:d0:c1:3a:04 | db's MAC. The VM addressed the frame to the gateway's MAC, and the overlay layer replaced it. | |
| 76–81 | 90:b8:d0:5e:77:03 | web's MAC. | |
| 82–83 | 0x0800 IPv4 |
Inner IPv4
From the VM · bytes 84–103 · 20 bytes
| Field | Bytes | Value | Note |
|---|---|---|---|
| 84 | 4 / 5 | ||
| 86–87 | 40 bytes | ||
| 90–91 | DF | ||
| 92 | 64 | Set by the VM. Proteus does not decrement it. | |
| 93 | 6 (TCP) | ||
| 94–95 | 0x0a84 | ||
| 96–99 | 10.0.0.3 | Unchanged inside the VPC. For floating IP traffic, the nat layer rewrites it. | |
| 100–103 | 10.0.0.4 |
Inner TCP
From the VM · bytes 104–123 · 20 bytes
| Field | Bytes | Value | Note |
|---|---|---|---|
| 104–105 | 49152 | ||
| 106–107 | 5432 | PostgreSQL on db. | |
| 108–111 | 0x9a3b0c11 | ||
| 117 | SYN | The firewall layer starts tracking the flow on this SYN. | |
| 118–119 | 64240 | ||
| 120–121 | 0x2566 |
Proteus sends the Geneve base header with no options. Figure 7 shows its fields.
Proteus adds 70 bytes to every packet, plus 4 on a VLAN-tagged underlay. Counting the VM's own headers too, headers are 8.8% of each full-size frame at a VM MTU of 1500 and about 1.5% at 8900.
Why Geneve. We chose Geneve over VXLAN (RFC 7348). Both carry a 24-bit VNI. VXLAN's header has a fixed format, and Geneve's has typed, variable-length options. Proteus sends no options today. They give us a place for per-packet metadata later without changing the encapsulation. Geneve is also in wide use: AWS uses it for its Gateway Load Balancer, and Taras Gritsenko describes its use for traffic between VPCs at AWS.
Why an IPv6 underlay. Each compute node's tunnel endpoint is an address in a private fd00:cabe:: range. Tenant VPCs are IPv4 today, so underlay addresses and tenant addresses never overlap, and the range has room for every node in a fleet.
Why a flow hash in the source port. The outer UDP source port is a hash of the inner flow. Switches and LACP bonds pick a link by hashing the outer headers, so different VM connections spread across links, and all packets of one connection use the same link and stay in order.
Proteus attaches below the hypervisor
The hypervisor has no part in the steps above. bhyve moves frames between the VM's virtio queues and a network interface on the host, and Proteus provides that interface. On illumos, each Proteus port registers as a GLDv3 MAC provider, a VNIC is created on it, and bhyve's in-kernel virtio backend (viona) uses the VNIC. Native illumos zones attach the same way. bhyve needs no changes.
On FreeBSD, each Proteus port is a netgraph node, usually joined to an ng_eiface interface that bhyve's virtio-net device uses. Geneve comes off the underlay through a link-layer pfil hook, and next-hop MACs come from nd6. The driver sets each outgoing frame's flow ID from the Geneve source port, so lagg spreads VM connections across the underlay links.
Every OS that runs a hypervisor has a kernel interface for connecting a VM to the network. To support a new OS, we write a driver for that interface and for the underlay hook. The policy, the engine, and proteus-core do not change. We moved the OS-independent code into proteus-core when we wrote the FreeBSD driver. The illumos driver implements the same Platform trait and imports the same kernel symbols as before the split.
FreeBSD compute-node diagram
| Piece | illumos | FreeBSD |
|---|---|---|
| Hypervisor | bhyve | bhyve |
| VM port | GLDv3 MAC provider + VNIC | netgraph node + ng_eiface |
| Underlay receive | MAC siphon | link-layer pfil |
| Underlay links | aggr (LACP) | lagg |
| Neighbor lookup | ip2mac |
nd6 |
| Control | /dev/proteus ioctl |
/dev/proteus ioctl |
Engine, triton-vpc, proteus-core, blueprint format |
Same on both. proteusadm works unchanged. |
Same on both. proteusadm works unchanged. |
Both drivers send the same Geneve frames, so one VPC can include VMs on illumos and FreeBSD nodes.
How Triton Cloud uses Proteus
Tenants and operators do not configure Proteus directly. tritond, the Triton Cloud control plane, stores the desired state of each VPC: subnets, security groups, routes, floating IPs, and load balancers. For each VM NIC, tritond compiles that state into a blueprint, which is a complete description of the port's policy. tritonagent on the compute node applies the blueprint through the Proteus control device and reports which blueprint generation the kernel accepted. We made the blueprint declarative so the agent can reapply it after any failure, and so the control plane can show whether each change has reached each node.
The V2P table changes separately. When a VM starts on another node, tritond sends its new location as a peer entry, and existing blueprints stay the same.
Tenant features map to Proteus as follows:
- A security group change produces a new blueprint for each affected port.
- A floating IP is a rule in the
natlayer and a hook on the node's external link. - The metadata service at 169.254.169.254 is answered on the node that runs the VM.
- A load balancer is a VPC member, and Proteus carries its traffic like any other VM's.
An operator adds a FreeBSD node by assigning the FreeBSD platform image from the same booter (tritonadm cn boot assign <mac> --pi freebsd-<stamp>). The image includes proteus.ko and loads it at boot.
Measured throughput
We measured VM-to-VM throughput with iperf3 between two illumos compute nodes. Each node has two 10GbE ports in an LACP bond, and the VMs use an MTU of 8900.
| iperf3 streams | Gbit/s |
|---|---|
| 1 | 7.8 to 9.8 |
| 4 | 10.3 |
| 8 | 19.4 |
| 16 | 19.4 |
Two illumos nodes, 2x10GbE LACP per node, VM MTU 8900. The single-stream result varies between runs.
19.4 Gbit/s is line rate for two 10GbE links after encapsulation overhead, with the traffic split evenly across both links in both directions. A single TCP connection stays on one link, because the bond places each flow on one link. We will run the same tests on 100GbE next. The next section shows where the CPU time goes.
Where the time goes
To see what the dataplane costs per packet, we sampled kernel stacks on two compute nodes while one VM streamed TCP to a VM on the other node. The VMs used an MTU of 8900, and every packet in the run was a flow-table hit. Figure 9 shows the cost of one 8.9 KB frame on each node, in the order the steps run.
Transmit · sending node
| Step, in order | per frame | share |
|---|---|---|
| LSO + checksum emulationmac_hw_emul | 926 ns | 23% |
| Flow-table hit + transformparse, flow-table get, apply transform | 130 ns | 3% |
| Proteus gluewrap, VLAN tag, neighbor, inlined code | 311 ns | 8% |
| Copy frame into one segmentlinearize for mlxcx; the NIC does the checksum | 467 ns | 11% |
| mac tx soft ring / dldDlsTxLease::emit | 277 ns | 7% |
| Promiscuous copy of every frameidle DLPI handle for the mlxcx VLAN multicast workaround | 890 ns | 22% |
| NIC driver pathaggr, mlxcx DMA bind | 1102 ns | 27% |
| Total | 4103 ns | 100% |
Receive · receiving node
| Step, in order | per frame | share |
|---|---|---|
| msgpullupcopies every received frame | 1171 ns | 26% |
| Proteus gluelock, validate, lookup, inlined code | 473 ns | 11% |
| Flow-table hit + transformparse, flow-table get, apply transform | 102 ns | 2% |
| Delivery into guestmac_rx, viona copy | 2689 ns | 61% |
| Total | 4435 ns | 100% |
The flow-table hit is about 130 ns on transmit and 100 ns on receive, roughly 3% of a frame's cost. Most of the rest belongs to the host: the illumos mac layer, the NIC driver, and on receive the copy into guest memory. At 130 ns, the engine could keep up with about 90 Gbit/s of 1500-byte traffic on one core.
Three steps in the path are overhead we can remove, shown hatched in Figure 9. Each copies every full frame:
- On transmit, an idle DLPI handle in promiscuous mode, held for the mlxcx VLAN multicast workaround, makes the mac layer copy every sent frame (890 ns).
- On transmit, Proteus copies each frame into a single segment before it goes to mlxcx (467 ns).
- On receive,
msgpullupcopies each frame (1,171 ns).
Together they cost 2.5 µs, about 30% of the 8.5 µs the two nodes spend on a frame. Figure 10 converts each per-frame cost to the bandwidth one core can carry at 8900-byte frames.
One core can carry 16 to 17 Gbit/s through the full path, and 22 to 26 Gbit/s without the three copies. A single flow is limited by the link before it is limited by the CPU: the bond places a flow on one 10G member, so one stream tops out near 9.8 Gbit/s, and eight streams reach 19.4 Gbit/s.
How this was measured
- Two Dell R640 compute nodes with Xeon Cascade Lake 2.6 GHz CPUs, a 2x10G LACP underlay on mlxcx NICs, underlay MTU 9000, Geneve over IPv6. Platform image 20260928, Proteus build g7d353c9. One iperf3 TCP stream between two bhyve VMs on different nodes.
- DTrace kernel stack sampling at 4999 Hz for 10 seconds, with no probes in the datapath. We attributed samples by call path and divided by the frame counts from the Proteus port counters. An earlier run with probes overstated each enabled probe by about 300 ns and is not used.
- The engine's
Port::processis inlined into Proteus, so samples inside engine functions count as the engine (130 ns transmit, 102 ns receive). Another 180 ns on transmit and 80 ns on receive landed in inlined Proteus code that may include part of the engine, so the upper bound for the engine step is about 310 ns. - Costs are per 8.9 KB frame. The guest copy,
msgpullup, and checksum steps scale with frame size. The engine step does not. - The throughput results come from earlier runs on the same nodes (20 and 25 September). Sampling lowered throughput to about 6 Gbit/s during this run.
Next steps
Next we will run the throughput tests on 100GbE. On illumos, we will pass packets to the MAC layer in chains instead of one at a time, and add hardware offload for the encapsulation. We will also remove the three full-frame copies in Figure 9, which account for 2.5 µs of the 8.5 µs the two nodes spend on each frame.
Proteus is MPL-2.0 and uses the OPTE engine from Oxide Computer under the same license. Triton Cloud is the next generation of Triton DataCenter, tritoncloud.io.