Edge Telco Infrastructure

Production-grade guide to edge telco infrastructure covering architecture patterns, implementation strategies, testing approaches, and operational best practices for enterprise engineering teams.

Edge Telco Infrastructure enables low-latency, high-reliability connectivity between mobile and fixed networks and distributed edge compute nodes. It matters when real-time control of cellular access, network slicing, and service orchestration must span from core to periphery — such as in industrial automation, autonomous vehicles, and multi-tenant edge data centers. It is the physical and logical fabric that binds base stations, core services, and edge servers into a unified, coordinated network.

Deploying a New Edge Site

To bring up a new edge site, you configure a network edge gateway (NEG) using a standardized template. The NEG is a Linux-based appliance running systemd, netplan, containerd, and k3s (Kubernetes minimal). You begin with a base image built via buildroot:

# buildroot/configs/edge-telco-site.defconfig
BR2_TARGET_GENERIC_GETTY_PORT="ttyS0"
BR2_TARGET_GENERIC_HOSTNAME="edge-site-01"
BR2_PACKAGE_SYSTEMD="y"
BR2_PACKAGE_NETPLAN="y"
BR2_PACKAGE_CONTAINERD="y"
BR2_PACKAGE_K3S="y"
BR2_PACKAGE_OPENSSL="y"
BR2_PACKAGE_OPENSSH="y"

After flashing the image to an Intel NUC or Rockchip RK3399 board, you configure the network via netplan:

# /etc/netplan/01-eth0.yaml
network:
  version: 2
  renderer: networkd
  ethernets:
    eth0:
      match:
        macaddress: 00:1a:2b:3c:4d:5e
      set-name: eth0
      addresses:
        - 192.168.10.10/24
      routes:
        - to: default
          via: 192.168.10.1
          table: 100
      nameservers:
        addresses:
          - 8.8.8.8
          - 1.1.1.1
      dhcp6: true
      link-local:
        - ipv6
    eth1:
      match:
        macaddress: 00:1a:2b:3c:4d:5f
      set-name: eth1
      addresses:
        - 10.20.30.1/24
      routes:
        - to: 10.20.30.0/24
          via: 10.20.30.1
          table: 100
      dhcp6: true
      link-local:
        - ipv6
  ip:
    routing-tables:
      100:
        - from: 10.20.30.0/24
          table: 100

You then bootstrap k3s with a --cluster-cidr and --service-cidr explicitly set:

sudo k3s server \
  --cluster-cidr=10.200.0.0/16 \
  --service-cidr=10.201.0.0/16 \
  --advertise-address=192.168.10.10 \
  --node-taints="edge-site=true:NoSchedule" \
  --kube-apiserver-args="--enable-admission-plugins=NodeRestriction,PodSecurity" \
  --flannel-backend=host-gw \
  --disable=traefik \
  --disable=servicelb

The key failure here is assuming flannel defaults are sufficient. In reality, host-gw performs better than vxlan on high-bandwidth, low-latency sites, but misconfiguring flannel's backend leads to unexpected MTU issues. You must set flannel.conf explicitly:

{
  "Network": "10.200.0.0/16",
  "Backend": {
    "Type": "host-gw"
  },
  "Delegate": {
    "HairpinMode": true,
    "MTU": 1450
  }
}

When flannel runs in vxlan, packets are fragmented across the 1500-byte MTU boundary. A packet from a 5G UE to an edge service may exceed 1500 bytes, and without setting MTU: 1450, you see IP: fragmentation needed, MTU = 1500 in tcpdump output on the eth1 interface.

Connecting to the Core Network

The edge site must connect to the core via IPsec or MPLS. For IPsec, use strongSwan with charon as the daemon:

# /etc/ipsec.conf
conn edge-core
    left=192.168.10.10
    leftid=edge-site-01.example.com
    leftcert=server.cert.pem
    leftsubnet=10.200.0.0/16
    right=10.20.30.100
    rightid=core-telco.example.com
    rightsubnet=10.20.30.0/24
    ike=aes256-sha256-modp2048
    esp=aes256-sha256
    keyexchange=ikev2
    auto=start
    dpdaction=clear
    dpddelay=30s
    dpdtimeout=120s
    rekey=no
    compress=yes
    fragmentation=yes
# /etc/ipsec.d/edge-core.secrets
: PSK "super-secret-psk-for-edge-site-01"
edge-site-01.example.com : RSA server.key.pem

You must enable fragmentation=yes in ipsec.conf. Without it, packets larger than 1400 bytes — common in 5G uplink traffic — are dropped. The error appears in journalctl:

charon[1234]: IKE_SA edge-core[1]: received fragmented packet (total: 1840 bytes, fragment #1/2)
charon[1234]: IKE_SA edge-core[1]: received fragment #2/2, complete packet received
charon[1234]: IKE_SA edge-core[1]: sending fragmented packet (total: 1840 bytes, fragment #1/2)

If you use MPLS instead, you must configure mpls and vrf on the edge router. Use bird as the BGP daemon:

# /etc/bird/bird.conf
protocol device {
    ipv4;
    ipv6;
}

protocol bgp core {
    local as 65000;
    neighbor 10.20.30.100 as 65001;
    ipv4;
    ipv6;
    import all;
    export all;
    next hop self;
    mesh all;
    hold time 30;
    keepalive 10;
    connect delay 5;
    startup 10;
    graceful restart;
    fast decay 5;
    route refresh;
    vrf core;
}

protocol mpls {
    label 1000;
    import all;
    export all;
    use vrf core;
    bgp;
    interface eth1;
}

The silent failure is assuming that mpls and vrf are automatically synchronized. But vrf interfaces must be manually created and bound to mpls instances. A missing vrf configuration leads to traffic from the core being routed into the wrong VRF, causing packet loss and looped routes.

Implementing Network Slicing

Each network slice requires a dedicated slice gateway (SGW) and service function chain (SFC). Use Open vSwitch with ovs-vsctl to create slice-specific bridges and flows:

# Create slice bridge for slice ID 100
sudo ovs-vsctl add-br br-slice-100
sudo ovs-vsctl add-port br-slice-100 eth1
sudo ovs-vsctl set port br-slice-100 tag=100

# Assign VLAN to slice
sudo ovs-vsctl add-br br-slice-100-vlan100
sudo ovs-vsctl add-port br-slice-100-vlan100 eth1
sudo ovs-vsctl set port br-slice-100-vlan100 tag=100

# Set up OpenFlow rules
sudo ovs-ofctl add-flow br-slice-100 \
    "priority=100,ip,nw_dst=10.200.10.10,actions=mod_vlan_vid:100,output:1"

sudo ovs-ofctl add-flow br-slice-100 \
    "priority=200,ip,nw_src=10.200.10.10,actions=mod_vlan_vid:100,output:2"

The confusion point is between VLAN tagging and VLAN stacking. You assume that mod_vlan_vid:100 sets the outer VLAN, but it only adds one layer. For multi-slice traffic, you must use mod_vlan_vid:100,mod_vlan_pcp:3 to set both VLAN ID and priority. Without pcp, the core network cannot classify traffic by slice quality.

You must also configure DSCP and QoS on the edge gateway. Use tc (traffic control) to set up hierarchical queues:

# Set up root qdisc on eth1
sudo tc qdisc add dev eth1 root handle 1: htb default 10
sudo tc class add dev eth1 parent 1: classid 1:1 htb rate 1000mbit ceil 1000mbit
sudo tc class add dev eth1 parent 1:1 classid 1:10 htb rate 100mbit ceil 200mbit prio 1
sudo tc class add dev eth1 parent 1:1 classid 1:20 htb rate 50mbit ceil 100mbit prio 2
sudo tc class add dev eth1 parent 1:1 classid 1:30 htb rate 20mbit ceil 50mbit prio 3

# Map DSCP to traffic class
sudo tc filter add dev eth1 protocol ip prio 1 u32 \
    match ip dscp 0x2e/0x30 action simple \
    set classid 1:10
sudo tc filter add dev eth1 protocol ip prio 2 u32 \
    match ip dscp 0x1e/0x30 action simple \
    set classid 1:20
sudo tc filter add dev eth1 protocol ip prio 3 u32 \
    match ip dscp 0x10/0x30 action simple \
    set classid 1:30

The sharp edge is that tc filters are not cumulative. A packet with dscp=0x2e triggers only the first filter — it does not fall through to the next. You must use actions=classify and chain to build a cascading QoS stack.

Orchestrating Edge Services

Service orchestration uses Kubernetes with k3s, but the edge site must run edge-specific operators. Deploy edge-service-operator via Helm:

helm install edge-operator \
  --set config.image=registry.example.com/edge-operator:v1.3.0 \
  --set config.slices=100,101,102 \
  --set config.sliceGatewayIP=10.20.30.100 \
  --set config.ingressClass=edge-telco \
  --set config.networkPolicyEnforcement=strict \
  --set config.logLevel=debug

The edge-service-operator watches for EdgeService CRDs:

apiVersion: edge.example.com/v1
kind: EdgeService
metadata:
  name: vehicle-traffic-control
  namespace: edge-ns-01
spec:
  sliceID: 100
  replicas: 3
  image: registry.example.com/traffic-control:v2.1.0
  ports:
    - name: http
      port: 80
      targetPort: 8080
    - name: metrics
      port: 9090
      targetPort: 9090
  resources:
    requests:
      memory: "512Mi"
      cpu: "250m"
    limits:
      memory: "1Gi"
      cpu: "1"
  nodeSelector:
    edge-type: vehicle
  affinity:
    podAntiAffinity:
      requiredDuringScheduling:
        - topologicalKey: kubernetes.io/hostname
          labelSelector:
            matchLabels:
              app: vehicle-traffic-control
  service:
    type: ClusterIP
    externalIPs:
      - 10.200.10.10
    loadBalancerIP: 10.200.10.11

When deploying services, the critical misstep is assuming that LoadBalancer services automatically configure Nginx or MetalLB. But in a k3s-based edge site, LoadBalancer is not automatically enabled unless you start metallb with a ConfigMap:

apiVersion: v1
kind: ConfigMap
metadata:
  namespace: metallb-system
  name: config
data:
  config: |
    address-pools:
      - name: default
        protocol: layer2
        auto-assign: true
        addresses:
          - 10.200.10.10-10.200.10.20

Without this, LoadBalancer services remain in Pending state, with no external IP assigned. You see:

kubectl get svc
NAME                      TYPE           CLUSTER-IP     EXTERNAL-IP   PORT(S)     AGE
vehicle-traffic-control   LoadBalancer   10.201.10.10   <pending>     80:30000/TCP   2m

The silent failure is that k3s ships with Traefik, but it does not auto-configure Ingress for EdgeService CRDs. You must enable Ingress in k3s and set ingressClass in the EdgeService spec:

spec:
  ingress:
    class: edge-telco
    hosts:
      - traffic-control.edge.example.com
    rules:
      - http:
          paths:
            - path: /api
              backend:
                serviceName: vehicle-traffic-control
                servicePort: 80

You must also deploy cert-manager with letsencrypt to issue TLS certificates:

apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-prod
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: [email protected]
    privateKeySecretRef:
      name: letsencrypt-prod-account-key
    solvers:
      - http01:
          ingress:
            class: edge-telco

Without ingress.class, cert-manager fails to issue certificates because Ingress objects are not recognized by the edge-telco class.

Monitoring and Troubleshooting

Edge telco infrastructure relies on Prometheus, Grafana, and Loki for observability. Configure node-exporter with a custom textfile collector for edge-specific metrics:

# /etc/prometheus/node-exporter.d/edge-metrics.text
# HELP edge_uplink_packet_loss Percentage of packets lost in 5G uplink
# TYPE edge_uplink_packet_loss gauge
edge_uplink_packet_loss 0.023

# HELP edge_core_latency_ms Round-trip time to core network
# TYPE edge_core_latency_ms gauge
edge_core_latency_ms 18.4

# HELP edge_slice_status Slice ID, status as 1=active, 0=inactive
# TYPE edge_slice_status gauge
edge_slice_status{slice_id="100"} 1
edge_slice_status{slice_id="101"} 0

The key failure is that node-exporter does not automatically pick up textfile directories. You must start it with:

node_exporter \
  --collector.textfile.directory=/etc/prometheus/node-exporter.d \
  --web.listen-address=:9100 \
  --log.level=info

You must add textfile directory to systemd:

[Unit]
Description=Node Exporter
After=network.target

[Service]
ExecStart=/usr/local/bin/node_exporter \
  --collector.textfile.directory=/etc/prometheus/node-exporter.d \
  --web.listen-address=:9100 \
  --log.level=info
Restart=always
User=root

[Install]
WantedBy=multi-user.target

When Prometheus scrapes node_exporter, you see textfile metrics in /metrics, but if node_exporter fails to write to the directory, the metrics are stale. You see edge_uplink_packet_loss at 0.0, even though the service reports 0.023.

Use journalctl to troubleshoot:

journalctl -u node_exporter -f

You see:

node_exporter[1234]: Error writing to /etc/prometheus/node-exporter.d/edge-metrics.text: Permission denied
node_exporter[1234]: Failed to write textfile metrics: write /etc/prometheus/node-exporter.d/edge-metrics.text: no space left on device

The sharp edge is that Prometheus scrapes node_exporter every 15 seconds, but node_exporter writes textfile metrics every 30 seconds. If textfile updates are not atomic, you can see a 30-second gap between edge_uplink_packet_loss updates, leading to misleading graphs.

Maintaining Edge Consistency

To maintain consistency across edge sites, use Ansible with a site.yml playbook:

---
- name: Deploy edge telco infrastructure
  hosts: all
  become: yes
  vars:
    network_config_dir: /etc/netplan
    k3s_config_dir: /etc/rancher/k3s
    slice

This page was rewritten on 10 October 2026. It replaced a templated version whose text was largely shared with other pages in this section and was not specific to its own title. The new text was drafted with a locally run language model, checked by a separate reviewer model for specificity and for invented figures, and measured against its sibling pages for duplication before publication. If anything here is wrong, tell us at [email protected] and we will correct it.