Production-grade guide to edge fog computing covering architecture patterns, implementation strategies, testing approaches, and operational best practices for enterprise engineering teams.
Edge fog computing is a distributed computing model where data processing and storage occur at the network edge, with multiple layers of intermediate nodes—fog nodes—between the cloud and end devices. It matters when latency, bandwidth, and real-time decision-making are critical: in industrial automation, autonomous vehicles, smart cities, and real-time video analytics, where a single cloud round-trip can break a control loop.
Configure fog nodes to run containerized microservices with minimal overhead. Use dockerd with --containerd=/run/containerd/containerd.sock and --exec-opt native.cgroupdriver=systemd. Set up a static IP for each fog node via /etc/systemd/network/eth0.network:
[Match]
Name=eth0
[Network]
DHCP=yes
Address=192.168.10.100/24
Gateway=192.168.10.1
DNS=192.168.10.1
Use systemd-networkd and systemd-resolved for consistent naming and DNS resolution. Install k3s (lightweight Kubernetes) for orchestration:
curl -sfL https://get.k3s.io | K3S_KUBELET_EXTRA_ARGS="--node-ip=192.168.10.100" sh -
Deploy a fog node as a worker node with --node-label fog=true and --node-taint fog=true:NoSchedule.
A fog node silently fails when k3s fails to start containerd due to missing cgroup v2 support. The error Failed to start containerd.service: Unit not found appears in journalctl -u k3s when systemd expects cgroupv2 but /sys/fs/cgroup/cgroup.subtree_control is absent. Fix by adding cgroup_enable=cpuset,memory,io to the kernel command line.
When a fog node boots but fails to join the cluster, check k3s/server/logs/etcd.log. A common cause is etcd not binding to the correct interface: failed to bind to 0.0.0.0:2379: listen tcp 0.0.0.0:2379: bind: cannot assign requested address. Ensure --bind-address=192.168.10.100 is passed to k3s and that the k3s service file includes:
[Service]
ExecStartPre=/bin/sh -c 'mkdir -p /var/lib/rancher/k3s/etcd && chown -R k3s:k3s /var/lib/rancher/k3s'
ExecStart=/usr/local/bin/k3s server --bind-address=192.168.10.100 --node-ip=192.168.10.100
Deploy microservices across fog nodes using k3s's built-in Helm and kubectl. Define a service with replicas: 3, tolerations, and affinity to ensure it runs on fog nodes:
apiVersion: apps/v1
kind: Deployment
metadata:
name: sensor-processor
namespace: fog
spec:
replicas: 3
selector:
matchLabels:
app: sensor-processor
template:
metadata:
labels:
app: sensor-processor
spec:
tolerations:
- key: fog
operator: Equal
value: "true"
effect: NoSchedule
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/hostname
operator: In
values:
- fog-node-01
- fog-node-02
- fog-node-03
containers:
- name: processor
image: sensor-processor:v1.2.0
ports:
- containerPort: 8080
env:
- name: FOG_NODE_ID
valueFrom:
fieldRef:
fieldPath: spec.nodeName
Use kubectl rollout status deployment/sensor-processor -n fog and kubectl get pods -n fog -o wide to verify deployment across nodes.
Fog services use k3s’s built-in traefik ingress controller. Configure ingress.yaml to route traffic to fog services:
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: sensor-ingress
namespace: fog
annotations:
traefik.ingress.kubernetes.io/router.entrypoints: web
traefik.ingress.kubernetes.io/router.middlewares: "fog-auth@kubernetescrd"
spec:
rules:
- host: sensors.fog.local
http:
paths:
- path: /data
pathType: Prefix
backend:
service:
name: sensor-processor
port:
number: 8080
Ensure traefik.toml includes:
[entryPoints]
[entryPoints.web]
address = ":80"
[providers]
[providers.kubernetesCRD]
namespaces = ["fog"]
[api]
dashboard = true
A common failure: traffic to sensors.fog.local resolves to 192.168.10.100, but the service is not exposed via LoadBalancer. The root cause is k3s not enabling traefik in config.yaml:
# /etc/rancher/k3s/config.yaml
traefik: true
traefik-args:
- "--entryPoints.web.address=:80"
- "--providers.kubernetesCRD"
- "--api.dashboard=true"
Data flows from edge devices to fog nodes and up to the cloud. Use kafka for streaming data between layers. Deploy kafka on a dedicated fog node with docker-compose.yml:
version: '3.8'
services:
zookeeper:
image: confluentinc/cp-zookeeper:7.0.0
environment:
ZOOKEEPER_CLIENT_PORT: 2181
ZOOKEEPER_TICK_TIME: 2000
ZOOKEEPER_SYNC_LIMIT: 2
ports:
- "2181:2181"
volumes:
- zookeeper-data:/var/lib/zookeeper
kafka:
image: confluentinc/cp-kafka:7.0.0
depends_on:
- zookeeper
environment:
KAFKA_BROKER_ID: 1
KAFKA_ZOOKEEPER_CONNECT: zookeeper:2181
KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: PLAIN:PLAIN,SSL:SSL
KAFKA_ADVERTISED_LISTENERS: PLAIN://192.168.10.105:9092,SSL://192.168.10.105:9093
KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 2
KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR: 2
ports:
- "9092:9092"
- "9093:9093"
volumes:
- kafka-data:/var/lib/kafka/data
volumes:
zookeeper-data:
kafka-data:
Edge devices publish sensor data to kafka using kafka-python:
from kafka import KafkaProducer
import json
producer = KafkaProducer(
bootstrap_servers=['192.168.10.105:9092'],
value_serializer=lambda v: json.dumps(v).encode('utf-8'),
acks='all',
retries=3,
batch_size=16384,
linger_ms=50
)
for sensor_data in sensor_stream():
producer.send('sensor-stream', sensor_data)
producer.flush()
producer.close()
Fog nodes process data in real time using Apache Flink or Spark Streaming. Deploy flink on k3s with flink.yaml:
apiVersion: apps/v1
kind: Deployment
metadata:
name: flink-jobmanager
namespace: fog
spec:
replicas: 1
selector:
matchLabels:
app: flink
template:
metadata:
labels:
app: flink
spec:
containers:
- name: jobmanager
image: flink:1.16.0
args:
- jobmanager
- --host
- flink-jobmanager
- --port
- "6123"
- --high-availability
- "zookeeper"
- --high-availability.storageDir
- "hdfs:///flink/ha/"
- --high-availability.zookeeper.quorum
- "zookeeper:2181"
- --high-availability.zookeeper.path.root
- "/flink"
ports:
- containerPort: 6123
name: jobmanager
- containerPort: 8081
name: webui
volumeMounts:
- name: flink-config
mountPath: /opt/flink/conf
- name: flink-logs
mountPath: /opt/flink/log
- name: taskmanager
image: flink:1.16.0
args:
- taskmanager
- --host
- flink-jobmanager
- --tm.heap.size
- "2g"
- --tm.numberOfTaskSlots
- "4"
ports:
- containerPort: 6122
name: taskmanager
volumeMounts:
- name: flink-config
mountPath: /opt/flink/conf
- name: flink-logs
mountPath: /opt/flink/log
volumes:
- name: flink-config
configMap:
name: flink-config
- name: flink-logs
persistentVolumeClaim:
claimName: flink-logs-pvc
The critical error: JobManager not reachable at flink-jobmanager:6123. This happens when k3s doesn't resolve flink-jobmanager within the cluster. The fix is to set clusterDNS and dnsConfig in k3s's config.yaml:
dns:
clusterIP: 10.43.0.10
enable: true
service: true
node: true
enableNodeLocalDNS: true
nodeLocalDNS: 192.168.10.100
Fog nodes must maintain consistent state across layers. Use etcd for distributed configuration and coordination. Configure etcd with etcd.conf.yml:
name: fog-etcd
data-dir: /var/lib/etcd
wal-dir: /var/lib/etcd/wal
snapshot-count: 10000
heartbeat-interval: 100
election-timeout: 500
listen-client-urls: http://0.0.0.0:2379
advertise-client-urls: http://192.168.10.100:2379
listen-peer-urls: http://0.0.0.0:2380
initial-advertise-peer-urls: http://192.168.10.100:2380
initial-cluster: fog-node-01=http://192.168.10.100:2380,fog-node-02=http://192.168.10.101:2380,fog-node-03=http://192.168.10.102:2380
initial-cluster-state: new
initial-cluster-token: fog-cluster
Use etcdctl to monitor cluster health:
etcdctl --endpoints=192.168.10.100:2379,192.168.10.101:2379,192.168.10.102:2379 endpoint health --cluster
A silent failure: etcd returns {"name":"fog-etcd","health":"true","version":"3.5.0"} but k3s fails to join the cluster. The cause is etcd not writing to data-dir due to missing fsync in journalctl:
journalctl -u etcd --since "2023-10-05 10:00:00"
Look for failed to create snapshot: open /var/lib/etcd/snap/db: no such file or directory. Fix by mounting a persistent volume with fsync enabled and setting sync=1 in fstab:
/dev/sdb1 /var/lib/etcd ext4 defaults,noatime,nodiratime,fsync=1 0 2
Edge devices send data to fog nodes via MQTT. Use mosquitto as the MQTT broker on the fog node:
apiVersion: apps/v1
kind: Deployment
metadata:
name: mosquitto
namespace: fog
spec:
replicas: 1
selector:
matchLabels:
app: mosquitto
template:
metadata:
labels:
app: mosquitto
spec:
containers:
- name: mosquitto
image: eclipse-mosquitto:2.0
ports:
- containerPort: 1883
name: mqtt
- containerPort: 9001
name: web
volumeMounts:
- name: mosquitto-config
mountPath: /mosquitto/config
- name: mosquitto-data
mountPath: /mosquitto/data
volumes:
- name: mosquitto-config
configMap:
name: mosquitto-config
- name: mosquitto-data
persistentVolumeClaim:
claimName: mosquitto-pvc
mosquitto.conf:
listener 1883
protocol mqtt
log_destinations stdout
log_type all
persistence true
persistence_location /mosquitto/data
persistence_interval_ms 30000
allow_anonymous true
# TLS
listener 8883
protocol mqtt
cafile /mosquitto/config/ca.crt
certfile /mosquitto/config/server.crt
keyfile /mosquitto/config/server.key
require_certificate true
Edge devices use paho-mqtt to publish to fog:
import paho.mqtt.client as mqtt
client = mqtt.Client("sensor-node-01")
client.tls_set(ca_certs="/etc/ssl/certs/ca.crt", certfile="/etc/ssl/certs/client.crt", keyfile="/etc/ssl/certs/client.key")
client.connect("192.168.10.100", 8883, 60)
client.publish("sensors/temperature", '{"node": "sensor-01", "value": 23.4, "timestamp": 1696500000}')
client.loop_start()
The edge device publishes sensors/temperature but the fog node receives no message. The error is Connection failed: 1 (Connection refused) from mosquitto logs. The cause: mosquitto not binding to 0.0.0.0 in listener 883. Fix by adding bind_address 0.0.0.0 to mosquitto.conf.
Another failure: edge devices lose connection and reconnect, but mosquitto does not resume QoS 1 deliveries. The root cause is missing persistence in mosquitto.conf. Without persistence true, messages are lost on restart.
A subtle issue: mosquitto uses 192.168.10.100 as clientid but the edge device uses sensor-node-01. The fog service consumes sensors/temperature but misses updates from sensor-node-01. The fix is to set client_id in the mqtt.Client call and use clientid in mosquitto:
client = mqtt.Client("sensor-node-01")
client.connect("192.168.10.100", 8883, 60, client_id="sensor-node-01")
Use Prometheus and Grafana to monitor fog nodes and services. Deploy prometheus with prometheus-config.yaml:
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
- job_name: 'k3s'
static_configs:
- targets: ['192.168.10.100:9100', '192.168.10.101:9100', '192.168.10.102:9100']
- job_name: 'flink'
static_configs:
- targets: ['flink-jobmanager:8081']
metrics_path: /prometheus
- job_name: 'etcd'
static_configs:
- targets: ['192.168.10.100:2379']
metrics_path: /metrics
Use node-exporter on each fog node:
docker run -d \
--name node-exporter \
--restart=always \
-v /proc:/host/proc:ro \
-v /sys:/host/sys:ro \
-v /:/host/root:ro \
-p 9100:9100 \
prom/node-exporter \
--path.procfs /host/proc \
--path.sysfs
This page was rewritten on 10 October 2026. It replaced a templated version whose text was largely shared with other pages in this section and was not specific to its own title. The new text was drafted with a locally run language model, checked by a separate reviewer model for specificity and for invented figures, and measured against its sibling pages for duplication before publication. If anything here is wrong, tell us at [email protected] and we will correct it.
We use cookies for analytics (Google Analytics) and advertising (Google AdSense) to improve your experience and support free content. Privacy Policy