Edge Fog Computing

Production-grade guide to edge fog computing covering architecture patterns, implementation strategies, testing approaches, and operational best practices for enterprise engineering teams.

Edge fog computing is a distributed computing model where data processing and storage occur at the network edge, with multiple layers of intermediate nodes—fog nodes—between the cloud and end devices. It matters when latency, bandwidth, and real-time decision-making are critical: in industrial automation, autonomous vehicles, smart cities, and real-time video analytics, where a single cloud round-trip can break a control loop.

Deploying Fog Nodes

Configure fog nodes to run containerized microservices with minimal overhead. Use dockerd with --containerd=/run/containerd/containerd.sock and --exec-opt native.cgroupdriver=systemd. Set up a static IP for each fog node via /etc/systemd/network/eth0.network:

[Match]
Name=eth0

[Network]
DHCP=yes
Address=192.168.10.100/24
Gateway=192.168.10.1
DNS=192.168.10.1

Use systemd-networkd and systemd-resolved for consistent naming and DNS resolution. Install k3s (lightweight Kubernetes) for orchestration:

curl -sfL https://get.k3s.io | K3S_KUBELET_EXTRA_ARGS="--node-ip=192.168.10.100" sh -

Deploy a fog node as a worker node with --node-label fog=true and --node-taint fog=true:NoSchedule.

Fog Node Failure Modes

A fog node silently fails when k3s fails to start containerd due to missing cgroup v2 support. The error Failed to start containerd.service: Unit not found appears in journalctl -u k3s when systemd expects cgroupv2 but /sys/fs/cgroup/cgroup.subtree_control is absent. Fix by adding cgroup_enable=cpuset,memory,io to the kernel command line.

When a fog node boots but fails to join the cluster, check k3s/server/logs/etcd.log. A common cause is etcd not binding to the correct interface: failed to bind to 0.0.0.0:2379: listen tcp 0.0.0.0:2379: bind: cannot assign requested address. Ensure --bind-address=192.168.10.100 is passed to k3s and that the k3s service file includes:

[Service]
ExecStartPre=/bin/sh -c 'mkdir -p /var/lib/rancher/k3s/etcd && chown -R k3s:k3s /var/lib/rancher/k3s'
ExecStart=/usr/local/bin/k3s server --bind-address=192.168.10.100 --node-ip=192.168.10.100

Orchestrating Fog Services

Deploy microservices across fog nodes using k3s's built-in Helm and kubectl. Define a service with replicas: 3, tolerations, and affinity to ensure it runs on fog nodes:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: sensor-processor
  namespace: fog
spec:
  replicas: 3
  selector:
    matchLabels:
      app: sensor-processor
  template:
    metadata:
      labels:
        app: sensor-processor
    spec:
      tolerations:
      - key: fog
        operator: Equal
        value: "true"
        effect: NoSchedule
      affinity:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
            - matchExpressions:
              - key: kubernetes.io/hostname
                operator: In
                values:
                - fog-node-01
                - fog-node-02
                - fog-node-03
      containers:
      - name: processor
        image: sensor-processor:v1.2.0
        ports:
        - containerPort: 8080
        env:
        - name: FOG_NODE_ID
          valueFrom:
            fieldRef:
              fieldPath: spec.nodeName

Use kubectl rollout status deployment/sensor-processor -n fog and kubectl get pods -n fog -o wide to verify deployment across nodes.

Fog Service Discovery and Load Balancing

Fog services use k3s’s built-in traefik ingress controller. Configure ingress.yaml to route traffic to fog services:

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: sensor-ingress
  namespace: fog
  annotations:
    traefik.ingress.kubernetes.io/router.entrypoints: web
    traefik.ingress.kubernetes.io/router.middlewares: "fog-auth@kubernetescrd"
spec:
  rules:
  - host: sensors.fog.local
    http:
      paths:
      - path: /data
        pathType: Prefix
        backend:
          service:
            name: sensor-processor
            port:
              number: 8080

Ensure traefik.toml includes:

[entryPoints]
  [entryPoints.web]
  address = ":80"

[providers]
  [providers.kubernetesCRD]
  namespaces = ["fog"]

[api]
  dashboard = true

A common failure: traffic to sensors.fog.local resolves to 192.168.10.100, but the service is not exposed via LoadBalancer. The root cause is k3s not enabling traefik in config.yaml:

# /etc/rancher/k3s/config.yaml
traefik: true
traefik-args:
  - "--entryPoints.web.address=:80"
  - "--providers.kubernetesCRD"
  - "--api.dashboard=true"

Managing Data Flow Across Fog Layers

Data flows from edge devices to fog nodes and up to the cloud. Use kafka for streaming data between layers. Deploy kafka on a dedicated fog node with docker-compose.yml:

version: '3.8'
services:
  zookeeper:
    image: confluentinc/cp-zookeeper:7.0.0
    environment:
      ZOOKEEPER_CLIENT_PORT: 2181
      ZOOKEEPER_TICK_TIME: 2000
      ZOOKEEPER_SYNC_LIMIT: 2
    ports:
      - "2181:2181"
    volumes:
      - zookeeper-data:/var/lib/zookeeper

  kafka:
    image: confluentinc/cp-kafka:7.0.0
    depends_on:
      - zookeeper
    environment:
      KAFKA_BROKER_ID: 1
      KAFKA_ZOOKEEPER_CONNECT: zookeeper:2181
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: PLAIN:PLAIN,SSL:SSL
      KAFKA_ADVERTISED_LISTENERS: PLAIN://192.168.10.105:9092,SSL://192.168.10.105:9093
      KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 2
      KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR: 2
    ports:
      - "9092:9092"
      - "9093:9093"
    volumes:
      - kafka-data:/var/lib/kafka/data

volumes:
  zookeeper-data:
  kafka-data:

Edge devices publish sensor data to kafka using kafka-python:

from kafka import KafkaProducer
import json

producer = KafkaProducer(
    bootstrap_servers=['192.168.10.105:9092'],
    value_serializer=lambda v: json.dumps(v).encode('utf-8'),
    acks='all',
    retries=3,
    batch_size=16384,
    linger_ms=50
)

for sensor_data in sensor_stream():
    producer.send('sensor-stream', sensor_data)
    producer.flush()

producer.close()

Fog Data Processing and State Management

Fog nodes process data in real time using Apache Flink or Spark Streaming. Deploy flink on k3s with flink.yaml:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: flink-jobmanager
  namespace: fog
spec:
  replicas: 1
  selector:
    matchLabels:
      app: flink
  template:
    metadata:
      labels:
        app: flink
    spec:
      containers:
      - name: jobmanager
        image: flink:1.16.0
        args:
        - jobmanager
        - --host
        - flink-jobmanager
        - --port
        - "6123"
        - --high-availability
        - "zookeeper"
        - --high-availability.storageDir
        - "hdfs:///flink/ha/"
        - --high-availability.zookeeper.quorum
        - "zookeeper:2181"
        - --high-availability.zookeeper.path.root
        - "/flink"
        ports:
        - containerPort: 6123
          name: jobmanager
        - containerPort: 8081
          name: webui
        volumeMounts:
        - name: flink-config
          mountPath: /opt/flink/conf
        - name: flink-logs
          mountPath: /opt/flink/log
      - name: taskmanager
        image: flink:1.16.0
        args:
        - taskmanager
        - --host
        - flink-jobmanager
        - --tm.heap.size
        - "2g"
        - --tm.numberOfTaskSlots
        - "4"
        ports:
        - containerPort: 6122
          name: taskmanager
        volumeMounts:
        - name: flink-config
          mountPath: /opt/flink/conf
        - name: flink-logs
          mountPath: /opt/flink/log
      volumes:
      - name: flink-config
        configMap:
          name: flink-config
      - name: flink-logs
        persistentVolumeClaim:
          claimName: flink-logs-pvc

The critical error: JobManager not reachable at flink-jobmanager:6123. This happens when k3s doesn't resolve flink-jobmanager within the cluster. The fix is to set clusterDNS and dnsConfig in k3s's config.yaml:

dns:
  clusterIP: 10.43.0.10
  enable: true
  service: true
  node: true
  enableNodeLocalDNS: true
  nodeLocalDNS: 192.168.10.100

Synchronizing State and Consistency

Fog nodes must maintain consistent state across layers. Use etcd for distributed configuration and coordination. Configure etcd with etcd.conf.yml:

name: fog-etcd
data-dir: /var/lib/etcd
wal-dir: /var/lib/etcd/wal
snapshot-count: 10000
heartbeat-interval: 100
election-timeout: 500
listen-client-urls: http://0.0.0.0:2379
advertise-client-urls: http://192.168.10.100:2379
listen-peer-urls: http://0.0.0.0:2380
initial-advertise-peer-urls: http://192.168.10.100:2380
initial-cluster: fog-node-01=http://192.168.10.100:2380,fog-node-02=http://192.168.10.101:2380,fog-node-03=http://192.168.10.102:2380
initial-cluster-state: new
initial-cluster-token: fog-cluster

Use etcdctl to monitor cluster health:

etcdctl --endpoints=192.168.10.100:2379,192.168.10.101:2379,192.168.10.102:2379 endpoint health --cluster

A silent failure: etcd returns {"name":"fog-etcd","health":"true","version":"3.5.0"} but k3s fails to join the cluster. The cause is etcd not writing to data-dir due to missing fsync in journalctl:

journalctl -u etcd --since "2023-10-05 10:00:00"

Look for failed to create snapshot: open /var/lib/etcd/snap/db: no such file or directory. Fix by mounting a persistent volume with fsync enabled and setting sync=1 in fstab:

/dev/sdb1  /var/lib/etcd  ext4  defaults,noatime,nodiratime,fsync=1  0  2

Coordinating Edge and Fog

Edge devices send data to fog nodes via MQTT. Use mosquitto as the MQTT broker on the fog node:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: mosquitto
  namespace: fog
spec:
  replicas: 1
  selector:
    matchLabels:
      app: mosquitto
  template:
    metadata:
      labels:
        app: mosquitto
    spec:
      containers:
      - name: mosquitto
        image: eclipse-mosquitto:2.0
        ports:
        - containerPort: 1883
          name: mqtt
        - containerPort: 9001
          name: web
        volumeMounts:
        - name: mosquitto-config
          mountPath: /mosquitto/config
        - name: mosquitto-data
          mountPath: /mosquitto/data
      volumes:
      - name: mosquitto-config
        configMap:
          name: mosquitto-config
      - name: mosquitto-data
        persistentVolumeClaim:
          claimName: mosquitto-pvc

mosquitto.conf:

listener 1883
protocol mqtt

log_destinations stdout
log_type all

persistence true
persistence_location /mosquitto/data
persistence_interval_ms 30000

allow_anonymous true

# TLS
listener 8883
protocol mqtt
cafile /mosquitto/config/ca.crt
certfile /mosquitto/config/server.crt
keyfile /mosquitto/config/server.key
require_certificate true

Edge devices use paho-mqtt to publish to fog:

import paho.mqtt.client as mqtt

client = mqtt.Client("sensor-node-01")
client.tls_set(ca_certs="/etc/ssl/certs/ca.crt", certfile="/etc/ssl/certs/client.crt", keyfile="/etc/ssl/certs/client.key")
client.connect("192.168.10.100", 8883, 60)

client.publish("sensors/temperature", '{"node": "sensor-01", "value": 23.4, "timestamp": 1696500000}')
client.loop_start()

Fog-Edge Synchronization Pitfalls

The edge device publishes sensors/temperature but the fog node receives no message. The error is Connection failed: 1 (Connection refused) from mosquitto logs. The cause: mosquitto not binding to 0.0.0.0 in listener 883. Fix by adding bind_address 0.0.0.0 to mosquitto.conf.

Another failure: edge devices lose connection and reconnect, but mosquitto does not resume QoS 1 deliveries. The root cause is missing persistence in mosquitto.conf. Without persistence true, messages are lost on restart.

A subtle issue: mosquitto uses 192.168.10.100 as clientid but the edge device uses sensor-node-01. The fog service consumes sensors/temperature but misses updates from sensor-node-01. The fix is to set client_id in the mqtt.Client call and use clientid in mosquitto:

client = mqtt.Client("sensor-node-01")
client.connect("192.168.10.100", 8883, 60, client_id="sensor-node-01")

Monitoring and Observability

Use Prometheus and Grafana to monitor fog nodes and services. Deploy prometheus with prometheus-config.yaml:

global:
  scrape_interval: 15s
  evaluation_interval: 15s

scrape_configs:
  - job_name: 'k3s'
    static_configs:
      - targets: ['192.168.10.100:9100', '192.168.10.101:9100', '192.168.10.102:9100']

  - job_name: 'flink'
    static_configs:
      - targets: ['flink-jobmanager:8081']
    metrics_path: /prometheus

  - job_name: 'etcd'
    static_configs:
      - targets: ['192.168.10.100:2379']
    metrics_path: /metrics

Use node-exporter on each fog node:

docker run -d \
  --name node-exporter \
  --restart=always \
  -v /proc:/host/proc:ro \
  -v /sys:/host/sys:ro \
  -v /:/host/root:ro \
  -p 9100:9100 \
  prom/node-exporter \
  --path.procfs /host/proc \
  --path.sysfs

This page was rewritten on 10 October 2026. It replaced a templated version whose text was largely shared with other pages in this section and was not specific to its own title. The new text was drafted with a locally run language model, checked by a separate reviewer model for specificity and for invented figures, and measured against its sibling pages for duplication before publication. If anything here is wrong, tell us at [email protected] and we will correct it.