Production-grade guide to edge telco infrastructure covering architecture patterns, implementation strategies, testing approaches, and operational best practices for enterprise engineering teams.
Edge Telco Infrastructure enables low-latency, high-reliability connectivity between mobile and fixed networks and distributed edge compute nodes. It matters when real-time control of cellular access, network slicing, and service orchestration must span from core to periphery — such as in industrial automation, autonomous vehicles, and multi-tenant edge data centers. It is the physical and logical fabric that binds base stations, core services, and edge servers into a unified, coordinated network.
To bring up a new edge site, you configure a network edge gateway (NEG) using a standardized template. The NEG is a Linux-based appliance running systemd, netplan, containerd, and k3s (Kubernetes minimal). You begin with a base image built via buildroot:
# buildroot/configs/edge-telco-site.defconfig
BR2_TARGET_GENERIC_GETTY_PORT="ttyS0"
BR2_TARGET_GENERIC_HOSTNAME="edge-site-01"
BR2_PACKAGE_SYSTEMD="y"
BR2_PACKAGE_NETPLAN="y"
BR2_PACKAGE_CONTAINERD="y"
BR2_PACKAGE_K3S="y"
BR2_PACKAGE_OPENSSL="y"
BR2_PACKAGE_OPENSSH="y"
After flashing the image to an Intel NUC or Rockchip RK3399 board, you configure the network via netplan:
# /etc/netplan/01-eth0.yaml
network:
version: 2
renderer: networkd
ethernets:
eth0:
match:
macaddress: 00:1a:2b:3c:4d:5e
set-name: eth0
addresses:
- 192.168.10.10/24
routes:
- to: default
via: 192.168.10.1
table: 100
nameservers:
addresses:
- 8.8.8.8
- 1.1.1.1
dhcp6: true
link-local:
- ipv6
eth1:
match:
macaddress: 00:1a:2b:3c:4d:5f
set-name: eth1
addresses:
- 10.20.30.1/24
routes:
- to: 10.20.30.0/24
via: 10.20.30.1
table: 100
dhcp6: true
link-local:
- ipv6
ip:
routing-tables:
100:
- from: 10.20.30.0/24
table: 100
You then bootstrap k3s with a --cluster-cidr and --service-cidr explicitly set:
sudo k3s server \
--cluster-cidr=10.200.0.0/16 \
--service-cidr=10.201.0.0/16 \
--advertise-address=192.168.10.10 \
--node-taints="edge-site=true:NoSchedule" \
--kube-apiserver-args="--enable-admission-plugins=NodeRestriction,PodSecurity" \
--flannel-backend=host-gw \
--disable=traefik \
--disable=servicelb
The key failure here is assuming flannel defaults are sufficient. In reality, host-gw performs better than vxlan on high-bandwidth, low-latency sites, but misconfiguring flannel's backend leads to unexpected MTU issues. You must set flannel.conf explicitly:
{
"Network": "10.200.0.0/16",
"Backend": {
"Type": "host-gw"
},
"Delegate": {
"HairpinMode": true,
"MTU": 1450
}
}
When flannel runs in vxlan, packets are fragmented across the 1500-byte MTU boundary. A packet from a 5G UE to an edge service may exceed 1500 bytes, and without setting MTU: 1450, you see IP: fragmentation needed, MTU = 1500 in tcpdump output on the eth1 interface.
The edge site must connect to the core via IPsec or MPLS. For IPsec, use strongSwan with charon as the daemon:
# /etc/ipsec.conf
conn edge-core
left=192.168.10.10
leftid=edge-site-01.example.com
leftcert=server.cert.pem
leftsubnet=10.200.0.0/16
right=10.20.30.100
rightid=core-telco.example.com
rightsubnet=10.20.30.0/24
ike=aes256-sha256-modp2048
esp=aes256-sha256
keyexchange=ikev2
auto=start
dpdaction=clear
dpddelay=30s
dpdtimeout=120s
rekey=no
compress=yes
fragmentation=yes
# /etc/ipsec.d/edge-core.secrets
: PSK "super-secret-psk-for-edge-site-01"
edge-site-01.example.com : RSA server.key.pem
You must enable fragmentation=yes in ipsec.conf. Without it, packets larger than 1400 bytes — common in 5G uplink traffic — are dropped. The error appears in journalctl:
charon[1234]: IKE_SA edge-core[1]: received fragmented packet (total: 1840 bytes, fragment #1/2)
charon[1234]: IKE_SA edge-core[1]: received fragment #2/2, complete packet received
charon[1234]: IKE_SA edge-core[1]: sending fragmented packet (total: 1840 bytes, fragment #1/2)
If you use MPLS instead, you must configure mpls and vrf on the edge router. Use bird as the BGP daemon:
# /etc/bird/bird.conf
protocol device {
ipv4;
ipv6;
}
protocol bgp core {
local as 65000;
neighbor 10.20.30.100 as 65001;
ipv4;
ipv6;
import all;
export all;
next hop self;
mesh all;
hold time 30;
keepalive 10;
connect delay 5;
startup 10;
graceful restart;
fast decay 5;
route refresh;
vrf core;
}
protocol mpls {
label 1000;
import all;
export all;
use vrf core;
bgp;
interface eth1;
}
The silent failure is assuming that mpls and vrf are automatically synchronized. But vrf interfaces must be manually created and bound to mpls instances. A missing vrf configuration leads to traffic from the core being routed into the wrong VRF, causing packet loss and looped routes.
Each network slice requires a dedicated slice gateway (SGW) and service function chain (SFC). Use Open vSwitch with ovs-vsctl to create slice-specific bridges and flows:
# Create slice bridge for slice ID 100
sudo ovs-vsctl add-br br-slice-100
sudo ovs-vsctl add-port br-slice-100 eth1
sudo ovs-vsctl set port br-slice-100 tag=100
# Assign VLAN to slice
sudo ovs-vsctl add-br br-slice-100-vlan100
sudo ovs-vsctl add-port br-slice-100-vlan100 eth1
sudo ovs-vsctl set port br-slice-100-vlan100 tag=100
# Set up OpenFlow rules
sudo ovs-ofctl add-flow br-slice-100 \
"priority=100,ip,nw_dst=10.200.10.10,actions=mod_vlan_vid:100,output:1"
sudo ovs-ofctl add-flow br-slice-100 \
"priority=200,ip,nw_src=10.200.10.10,actions=mod_vlan_vid:100,output:2"
The confusion point is between VLAN tagging and VLAN stacking. You assume that mod_vlan_vid:100 sets the outer VLAN, but it only adds one layer. For multi-slice traffic, you must use mod_vlan_vid:100,mod_vlan_pcp:3 to set both VLAN ID and priority. Without pcp, the core network cannot classify traffic by slice quality.
You must also configure DSCP and QoS on the edge gateway. Use tc (traffic control) to set up hierarchical queues:
# Set up root qdisc on eth1
sudo tc qdisc add dev eth1 root handle 1: htb default 10
sudo tc class add dev eth1 parent 1: classid 1:1 htb rate 1000mbit ceil 1000mbit
sudo tc class add dev eth1 parent 1:1 classid 1:10 htb rate 100mbit ceil 200mbit prio 1
sudo tc class add dev eth1 parent 1:1 classid 1:20 htb rate 50mbit ceil 100mbit prio 2
sudo tc class add dev eth1 parent 1:1 classid 1:30 htb rate 20mbit ceil 50mbit prio 3
# Map DSCP to traffic class
sudo tc filter add dev eth1 protocol ip prio 1 u32 \
match ip dscp 0x2e/0x30 action simple \
set classid 1:10
sudo tc filter add dev eth1 protocol ip prio 2 u32 \
match ip dscp 0x1e/0x30 action simple \
set classid 1:20
sudo tc filter add dev eth1 protocol ip prio 3 u32 \
match ip dscp 0x10/0x30 action simple \
set classid 1:30
The sharp edge is that tc filters are not cumulative. A packet with dscp=0x2e triggers only the first filter — it does not fall through to the next. You must use actions=classify and chain to build a cascading QoS stack.
Service orchestration uses Kubernetes with k3s, but the edge site must run edge-specific operators. Deploy edge-service-operator via Helm:
helm install edge-operator \
--set config.image=registry.example.com/edge-operator:v1.3.0 \
--set config.slices=100,101,102 \
--set config.sliceGatewayIP=10.20.30.100 \
--set config.ingressClass=edge-telco \
--set config.networkPolicyEnforcement=strict \
--set config.logLevel=debug
The edge-service-operator watches for EdgeService CRDs:
apiVersion: edge.example.com/v1
kind: EdgeService
metadata:
name: vehicle-traffic-control
namespace: edge-ns-01
spec:
sliceID: 100
replicas: 3
image: registry.example.com/traffic-control:v2.1.0
ports:
- name: http
port: 80
targetPort: 8080
- name: metrics
port: 9090
targetPort: 9090
resources:
requests:
memory: "512Mi"
cpu: "250m"
limits:
memory: "1Gi"
cpu: "1"
nodeSelector:
edge-type: vehicle
affinity:
podAntiAffinity:
requiredDuringScheduling:
- topologicalKey: kubernetes.io/hostname
labelSelector:
matchLabels:
app: vehicle-traffic-control
service:
type: ClusterIP
externalIPs:
- 10.200.10.10
loadBalancerIP: 10.200.10.11
When deploying services, the critical misstep is assuming that LoadBalancer services automatically configure Nginx or MetalLB. But in a k3s-based edge site, LoadBalancer is not automatically enabled unless you start metallb with a ConfigMap:
apiVersion: v1
kind: ConfigMap
metadata:
namespace: metallb-system
name: config
data:
config: |
address-pools:
- name: default
protocol: layer2
auto-assign: true
addresses:
- 10.200.10.10-10.200.10.20
Without this, LoadBalancer services remain in Pending state, with no external IP assigned. You see:
kubectl get svc
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
vehicle-traffic-control LoadBalancer 10.201.10.10 <pending> 80:30000/TCP 2m
The silent failure is that k3s ships with Traefik, but it does not auto-configure Ingress for EdgeService CRDs. You must enable Ingress in k3s and set ingressClass in the EdgeService spec:
spec:
ingress:
class: edge-telco
hosts:
- traffic-control.edge.example.com
rules:
- http:
paths:
- path: /api
backend:
serviceName: vehicle-traffic-control
servicePort: 80
You must also deploy cert-manager with letsencrypt to issue TLS certificates:
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
email: [email protected]
privateKeySecretRef:
name: letsencrypt-prod-account-key
solvers:
- http01:
ingress:
class: edge-telco
Without ingress.class, cert-manager fails to issue certificates because Ingress objects are not recognized by the edge-telco class.
Edge telco infrastructure relies on Prometheus, Grafana, and Loki for observability. Configure node-exporter with a custom textfile collector for edge-specific metrics:
# /etc/prometheus/node-exporter.d/edge-metrics.text
# HELP edge_uplink_packet_loss Percentage of packets lost in 5G uplink
# TYPE edge_uplink_packet_loss gauge
edge_uplink_packet_loss 0.023
# HELP edge_core_latency_ms Round-trip time to core network
# TYPE edge_core_latency_ms gauge
edge_core_latency_ms 18.4
# HELP edge_slice_status Slice ID, status as 1=active, 0=inactive
# TYPE edge_slice_status gauge
edge_slice_status{slice_id="100"} 1
edge_slice_status{slice_id="101"} 0
The key failure is that node-exporter does not automatically pick up textfile directories. You must start it with:
node_exporter \
--collector.textfile.directory=/etc/prometheus/node-exporter.d \
--web.listen-address=:9100 \
--log.level=info
You must add textfile directory to systemd:
[Unit]
Description=Node Exporter
After=network.target
[Service]
ExecStart=/usr/local/bin/node_exporter \
--collector.textfile.directory=/etc/prometheus/node-exporter.d \
--web.listen-address=:9100 \
--log.level=info
Restart=always
User=root
[Install]
WantedBy=multi-user.target
When Prometheus scrapes node_exporter, you see textfile metrics in /metrics, but if node_exporter fails to write to the directory, the metrics are stale. You see edge_uplink_packet_loss at 0.0, even though the service reports 0.023.
Use journalctl to troubleshoot:
journalctl -u node_exporter -f
You see:
node_exporter[1234]: Error writing to /etc/prometheus/node-exporter.d/edge-metrics.text: Permission denied
node_exporter[1234]: Failed to write textfile metrics: write /etc/prometheus/node-exporter.d/edge-metrics.text: no space left on device
The sharp edge is that Prometheus scrapes node_exporter every 15 seconds, but node_exporter writes textfile metrics every 30 seconds. If textfile updates are not atomic, you can see a 30-second gap between edge_uplink_packet_loss updates, leading to misleading graphs.
To maintain consistency across edge sites, use Ansible with a site.yml playbook:
---
- name: Deploy edge telco infrastructure
hosts: all
become: yes
vars:
network_config_dir: /etc/netplan
k3s_config_dir: /etc/rancher/k3s
slice
This page was rewritten on 10 October 2026. It replaced a templated version whose text was largely shared with other pages in this section and was not specific to its own title. The new text was drafted with a locally run language model, checked by a separate reviewer model for specificity and for invented figures, and measured against its sibling pages for duplication before publication. If anything here is wrong, tell us at [email protected] and we will correct it.
We use cookies for analytics (Google Analytics) and advertising (Google AdSense) to improve your experience and support free content. Privacy Policy