Designing a Scalable & Fault-Tolerant Log Pipeline [Part 2]: The Buffer Layer
Whether you actually need Kafka, what a buffer layer makes possible, and how to configure it correctly for log pipelines at scale.
This is Part 2 of a series on building a scalable, fault-tolerant log pipeline. Part 1 covered agents and forwarders — the collection tier. This part covers what sits between them and storage: the buffer layer.
Agents → Forwarders → Kafka → Storage & AnalyticsLayer 2: The Buffer Layer#
Do you actually need Kafka?#
Many teams avoid Kafka in their telemetry pipeline. And honestly, they’re not wrong. It adds operational overhead, storage costs, and complexity that feel unnecessary when all you want is queryable logs.
Newer patterns are making this easier to skip. ClickHouse is increasingly used as a unified backend for all three telemetry signals logs, metrics, and traces in a single storage layer. Tools like SigNoz, Uptrace, and Vector write directly to ClickHouse, bypassing the buffer entirely. Fewer things to manage. One place to query. SQL on all your signals from day one. Hard to argue with that.
So the question I kept coming back to: what do I need that ClickHouse alone can’t give me?
If your pipeline has one consumer and fast queries are the goal, ClickHouse-native is a strong choice.
If you need multiple independent consumers, real-time compute on the stream, replay, or cross-signal correlation at the ingestion layer, that’s where a buffer changes what’s possible.
What a buffer layer actually enables#
One Kafka topic. Multiple consumers. Each reads independently at its own speed.
In my homelab, a single topic o11y.logs.v1 feeds four independent consumers:
- VictoriaLogs — immediate search and Grafana dashboards
- Flink (archival) — continuous write to Parquet on S3 via Iceberg
- Flink (correlation engine) — real-time anomaly detection across metrics and logs
- ClickHouse — analytics without touching the main search pipeline
Each consumer has its own offset. If ClickHouse goes down, VictoriaLogs keeps ingesting. If the Flink job restarts, it replays from its last checkpoint. Nothing knows about anything else.
Without a buffer layer, every new consumer is another output config on your agent or forwarder. Every change happens at the edge, across every node.
Replay#
Kafka retains data for a configurable window. Point a new consumer at old offsets and replay everything within that window, no changes to the source.
New storage system? New analytics pipeline? Missed events during an outage? Replay from the beginning of the retention window. I haven’t found anything else in the stack that does this without touching the source.
Key concepts for log pipelines#
Partitions#
More partitions equals more parallelism. Partition by service or host for ordered logs per source. Partition count is fixed at topic creation — changing it later is painful, so decide upfront.
A reasonable starting point: one partition per consumer thread you expect to run. Under-partition and your consumers bottleneck on a single thread. Over-partition and you waste broker resources managing empty partitions.
Consumer groups#
Each group tracks its own offset independently. This is what makes multiple independent consumers possible.
My stack runs four groups off the same topic:
| Group | Consumer | Purpose |
|---|---|---|
o11y.logs.search | VictoriaLogs | Hot search and Grafana dashboards |
o11y.logs.archive | Flink LTS | Iceberg/Parquet archival to S3 |
o11y.logs.correlate | Flink | Real-time anomaly detection |
o11y.logs.analytics | ClickHouse | Analytics queries |
The first time I misconfigured group IDs, two jobs raced for the same offsets and one silently received half the data. Took a while to track down. Consumer group IDs are the one config you do not want to copy-paste without thinking.
Retention vs compaction#
For log topics, use time-based retention. Set retention.ms to how long you want replay capability, 24 hours for operational pipelines, 7 days if you’re using Kafka as a short-term archive.
Compaction is for config or metrics topics where you want the latest value per key. Not for logs where every event matters. Compacted topics silently drop duplicate keys — the wrong behaviour for an event stream.
Replication#
Replication factor 3 on a 3-node cluster means one broker can fail without data loss. Kafka is not a single point of failure if you deploy it correctly.
Do not run a single-node Kafka cluster for anything you care about. Replication is the one place where the minimal homelab setup differs meaningfully from production.
Tuning levers#
| Setting | Trade-off |
|---|---|
acks=1 | Higher throughput, risk losing messages if the leader fails before replication |
acks=all | Full durability, lower throughput |
linger.ms + batch.size | Higher values = better throughput, higher latency, tune for your ingest rate |
compression.type=snappy or zstd | Reduces storage and network overhead; zstd compresses better, snappy is faster |
retention.ms | Balance replay window against disk cost |
Alternatives#
The buffer layer pattern works without Kafka:
- Redpanda — Kafka-compatible API, single binary, no ZooKeeper or KRaft. If Kafka’s ops burden is the reason you’re hesitating, this is the closest drop-in. Same API, much less to manage. My homelab uses Redpanda for exactly this reason.
- Apache Pulsar — multi-tenancy built in, tiered storage native, separates compute and storage. Worth considering at scale, harder to operate than Redpanda.
- NATS JetStream — fast and lightweight, but not what I’d use for a mission-critical log pipeline. The ecosystem around consumers and exactly-once semantics isn’t there yet.
- AWS Kinesis / GCP Pub/Sub — managed, no ops burden, but you’re locked in to the cloud provider and the pricing model scales with volume.
For a homelab or small production environment, Redpanda is the pragmatic choice. Run it on a single node to start, expand to three nodes when you care about durability.
Drawbacks#
Kafka is not free. The trade-offs are real:
- Ops complexity — cluster management, broker health, partition rebalancing, and upgrade paths are non-trivial. Someone has to own this.
- Storage at multiple stages — logs live on Kafka and downstream simultaneously during the retention window. Budget for it.
- Cost — disk, compute, and network for a 3-node cluster adds up at scale.
- Security surface — mTLS, ACLs, and inter-broker encryption need to be configured, not assumed. The defaults are not secure.
- Overkill for small setups — one consumer, under 10 nodes? You probably don’t need it. A forwarder writing directly to VictoriaLogs with a large disk buffer covers most of what Kafka adds.
The point is not that Kafka is wrong. It’s that the decision should be deliberate. If you know why you need it, the trade-offs are manageable.
What to look for in a consumer#
A consumer that can’t keep up or handle failures correctly undermines everything the buffer layer gives you.
Consumer must:
- Track offsets independently — resume exactly where it left off after a restart, no manual reset required
- Handle backpressure — don’t overwhelm downstream storage under burst traffic
- Monitor lag — alert when falling behind the producer; a consumer that’s silently hours behind is worse than one that’s visibly down
- Handle schema changes — graceful on new or missing fields without crashing the job
- Dead letter handling — isolate malformed records without blocking the pipeline; bad events should land somewhere queryable, not disappear
- Delivery guarantee — exactly-once or at-least-once — pick one and understand the trade-off; exactly-once is more expensive but eliminates duplicates in analytics
- Parallelism — one consumer thread per partition maximum; plan partition count around the parallelism you need
- Fault tolerance — crash and resume without manual offset reset or data loss
Lag monitoring is the one I see skipped most often. A consumer can be running, healthy by every process metric, and silently hours behind. By the time you notice, the retention window may have passed and the data is gone.
Next: Part 3 will cover the hot storage layer. It looks at VictoriaLogs, Loki, and Elasticsearch for recent logs, and ClickHouse and Iceberg on S3 for historical analytics.