Field Notes

Back

This is Part 1 of a series on building a scalable, fault-tolerant log pipeline. I’m studying log pipelines at scale and building one to test the ideas hands-on. Over the coming posts, I’ll cover each layer in detail: agents, forwarders, buffers, hot storage, and long-term storage.


Layer 1: Log Agents & Forwarders#

Your log agent is the entry point of the entire pipeline. Everything downstream depends on what leaves here.

Agent vs Forwarder#

An agent runs close to your workload, on each node, VM, or pod. Its job is to collect and ship. Nothing more. It should not consume the CPU and memory that your business workloads need.

A forwarder sits between agents and storage. It handles the heavy work: aggregation, transformation, enrichment, and routing.


What Each Layer Is Responsible For#

Agent must#

  • Filesystem buffer: no data loss when downstream is unreachable
  • Batch and compress: saves bandwidth at scale
  • Alert on failure: know when it stops shipping
  • Stay lightweight: no TLS needed on the local agent-to-forwarder path

Agent must not#

  • Enrich or transform logs
  • Handle complex routing
  • Run any expensive processing

Keeping the agent minimal protects your application. A DaemonSet agent on a database node or a sidecar next to a latency-sensitive service cannot afford to spike CPU for log parsing.

Forwarder must#

  • Parse & structure: convert raw stdout/stderr into structured JSON before reaching storage
  • Transform: rename keys, remap values, drop noisy fields, apply regex, type-cast strings to numbers
  • Enrich: add service_name, host_name, environment, Kubernetes metadata, infer log_level from message content
  • Handle high-volume streams: aggregate from hundreds of agents simultaneously without dropping records
  • Backpressure + large disk buffer: absorb traffic spikes when downstream slows; the agent buffer is small, the forwarder buffer is large
  • Multi-stream routing: route different log streams to different destinations based on content or label
  • TLS: from forwarder to Kafka or storage; security adds latency and CPU overhead, plan accordingly
  • Multi-tenancy + AuthN / AuthZ: isolate logs per team or namespace
  • Multiple output plugins: Kafka, S3, ClickHouse, VictoriaLogs, all simultaneously if needed
  • PII data masking

When the agent goes direct (no forwarder)#

If you skip the forwarder tier, the agent must also cover:

  • TLS in transit to storage or Kafka
  • Basic enrichment: service_name, host_name, namespace
  • Multi-tenancy support
  • Predictable resource ceiling under load

Going direct simplifies the deployment but pushes more responsibility onto every agent instance running on your production nodes.


Deployment Patterns#

The checklist above tells you what each component needs. The pattern you choose determines where those responsibilities fall.

Pattern 1 — Single-Hop#

Agents → Forwarder → ClickHouse

Fewer moving parts. The forwarder handles all parsing, transformation, enrichment, and routing. ClickHouse handles the analytics layer directly — no buffer needed between collection and query. Works well when you have a single storage destination and predictable ingest volume.

Pattern 2 — Multi-Tier#

Agents → Forwarder → Kafka → VictoriaLogs / S3

Agents stay minimal. The forwarder handles all transformation. Kafka adds durability and fan-out. Enterprise Kubernetes logging operators implement this exact pattern: Fluent Bit as a DaemonSet agent on every node, Fluentd as the centralised forwarder.

Kubernetes logging operator flow: Fluent Bit DaemonSet agents on each node forwarding to a centralised Fluentd aggregator
Credit: Kube Logging Operator

This is the right choice when:

  • You need multiple consumers reading the same log stream
  • You want to decouple ingest rate from storage write speed
  • You need replay capability for backfill or reprocessing

Pattern 3 — Kafka-Native#

Agents → Kafka → VictoriaLogs / S3

Skip the forwarder. The agent sends directly to Kafka with a filesystem buffer. Kafka becomes the single source of truth. A separate consumer (Flink, a custom processor, or a simple Kafka consumer) handles transformation and routing downstream.

This trades forwarder simplicity for Kafka-native durability. You lose the centralised transformation tier but gain a clean separation between collection and processing.


Agents Under Review#

In this series I am exploring four agents:

AgentFootprintStrengths
Fluent BitVery lowLightweight, fast, wide plugin ecosystem
VLAgentLowNative VictoriaLogs integration, simple config
OpenTelemetry CollectorMediumUnified telemetry, OTLP native
Grafana AlloyMediumGrafana ecosystem, River config language

Each handles buffering, backpressure, and routing differently. I will benchmark them as we go.

Designing a Scalable & Fault-Tolerant Log Pipeline [Part 1]: Agents and Forwarders
https://blogs.thedevopsguy.biz/blog/log-pipeline-at-scale-p1
Author Akash Rajvanshi
Published at September 6, 2026