Field Notes

Back

An observability stack usually starts with a few applications and infrastructure components. They send logs, metrics, and traces to a telemetry backend. At first, you may not know how much data these signals produce. As traffic and retention increase, the stack needs more storage and compute.

At that point, tweaking a few values or scaling the pods is not enough. You need to know how much data arrives, how long it stays, how many copies are stored, and whether the cluster can handle a busy hour or a failed pod.

Capacity planning answers four questions:

  • How much data arrives?
  • How much data must be stored?
  • How much CPU, RAM, network, and disk throughput are needed?
  • Can the stack keep running when a pod fails?

The calculator compares three stacks with the same workload:

  1. LGTM with Loki, Tempo, Mimir, and Grafana
  2. Victoria with VictoriaLogs, VictoriaTraces, VictoriaMetrics, and Grafana
  3. ClickHouse OSS with Grafana, or ClickStack OSS with HyperDX

You can use one stack as it is. You can also create a hybrid plan. Loki for logs, VictoriaMetrics for metrics, Tempo for traces, and Grafana for visualization is a valid choice when it fits your operational needs.

Open the calculator and test it with your workload.


Why Capacity Planning Matters#

Telemetry grows as you add applications and infrastructure. Incidents produce more logs and traces. New labels and workloads can also create more metric series.

Without a capacity plan, teams usually discover limits in production:

  • An ingester runs out of memory during a traffic burst
  • A WAL or local data disk fills before retention removes old data
  • Object storage capacity is affordable, but API requests and data transfer can increase the final cost
  • A query scans too much data and competes with ingestion
  • A worker node has enough CPU but not enough RAM or disk throughput
  • Replication protects data but multiplies network and storage demand
  • One unavailable pod leaves the remaining pods below the required capacity

A capacity plan is a starting point. It is not a performance guarantee. You still need to test ingestion and real queries before production.


Inputs You Need#

Use measured data when possible. If you do not have it, start with an estimate and replace it later.

Logs and Traces#

Collect these values:

  • Raw ingestion in GB per day
  • Retention in days
  • Peak to average traffic ratio
  • Compression ratio
  • Replication factor
  • Required outage buffer

Do not plan only for the average. A workload that sends 100 GB per day may still send twice its normal rate during a busy hour.

Metrics#

Metrics are better described by series and samples than by GB per day.

  • Active series
  • Scrape interval
  • Expected series growth
  • Label cardinality
  • Retention
  • Recording rule and alert load

A shorter scrape interval creates more samples without adding more series. High cardinality can increase memory and query cost even when the raw sample rate looks manageable.

Infrastructure#

You also need:

  • Worker node size
  • CPU and RAM reserved for Kubernetes
  • Target node utilization
  • Availability zones
  • Storage type and disk performance
  • Object storage price
  • Cross zone traffic

Minimum is the smallest capacity estimate for the selected workload. Recommended adds availability, peak traffic, and planning headroom.


Calculate Your Telemetry Volume#

The best source is your collector or backend. Measure how many bytes it sends during a known period. Use the size before backend compression because the calculator applies compression later.

Logs#

If your collector exposes exported bytes:

logs GB/day = exported bytes ÷ measured seconds × 86,400 ÷ 1,000,000,000

If you know the event rate and average event size:

logs GB/day = events/s × average bytes/event × 86,400 ÷ 1,000,000,000

Add logs from applications, Kubernetes, hosts, databases, network devices, and cloud services. Measure after filtering and before storage compression.

Traces#

Trace volume depends on request rate, spans per request, sampling, and span size.

exported spans/s = requests/s × average spans/request × sampling rate
traces GB/day = exported spans/s × average bytes/span × 86,400 ÷ 1,000,000,000

A 10 percent sampling rate is 0.10. Include span attributes, resource attributes, and events when you measure the average span size.

If you have spans per minute, divide by 60 first.

Metrics#

Metrics start with active series and scrape interval.

samples/s = active series ÷ scrape interval in seconds
samples/day = samples/s × 86,400
monthly samples in millions = samples/s × 86,400 × 30 ÷ 1,000,000

Some backends also need a raw byte estimate:

raw metrics GB/day = samples/day × raw bytes/sample ÷ 1,000,000,000

Every unique label set is a separate series. Read the active series count from your collector or metrics backend when possible.

Worked Example#

Assume this workload:

  • 1,000 log events per second at 1,200 bytes per event
  • 500 requests per second with 8 spans per request
  • 10 percent trace sampling and 1,000 bytes per exported span
  • 1 million active metric series with a 60 second scrape interval

Logs:

1,000 events/s × 1,200 bytes × 86,400 ÷ 1,000,000,000
= 103.68 GB/day

Traces:

exported spans/s = 500 requests/s × 8 spans × 0.10 sampling
                 = 400 spans/s

400 spans/s × 1,000 bytes × 86,400 ÷ 1,000,000,000
= 34.56 GB/day

Metrics:

samples/s = 1,000,000 active series ÷ 60 seconds
          = 16,667 samples/s

monthly samples = 16,667 × 86,400 × 30
                = 43.2 billion samples

Enter 104 GB/day for logs, 35 GB/day for traces, 1,000,000 active series, and a 60s scrape interval. These are raw inputs. The selected backend applies compression, retention, replication, and storage overhead.


How SignalCost and CloudRaft Estimate Telemetry#

You may know the number of clusters, applications, databases, and virtual machines without knowing their telemetry volume.

Use the CloudRaft Observability Pricing Calculator to turn infrastructure counts into an initial telemetry estimate.

Use the SignalCost Observability Cost Calculator to estimate telemetry and compare managed and self-hosted prices.

Both tools use fixed rates when measurements are missing. Their results are starting estimates. Our calculator takes that telemetry estimate and sizes the components in LGTM, Victoria, and ClickHouse OSS.

CloudRaft#

CloudRaft uses virtual machines or nodes, Kubernetes clusters, databases, services, and requests per second.

Its log estimate follows this model:

logs GB/day = 0.85 × VMs or nodes
            + 0.28 × services
            + 1.6 × databases
            + 8 × Kubernetes clusters

Its trace estimate applies an 8 percent sampling rate and about 1.8 KiB per sampled request:

traces GB/day = requests/s × 0.08 × 1.8 KiB × 86,400 ÷ 1,048,576

CloudRaft also estimates metric sample volume from infrastructure counts. Use it as a starting point, then validate it against your active series and scrape interval.

SignalCost#

SignalCost estimates telemetry from your infrastructure. It uses the number of virtual machines, Kubernetes clusters, pods, services, and custom metrics:

total pods = Kubernetes clusters × pods per cluster

logs GB/day = VMs × log GB per VM/day
            + total pods × log GB per pod/day

metric series = VMs × 300
              + total pods × 100
              + Kubernetes clusters × 8,000
              + custom metric series

spans/min = traced services × 200

The log values depend on the selected verbosity:

VerbosityGB per VM/dayGB per pod/day
Low0.50.05
Medium30.3
High101

SignalCost uses 0.5 GB per million spans and 10 bytes per metric sample. These are useful starting values. They are not measurements from your environment.

CloudRaft, SignalCost, and our calculator solve different problems:

AreaCloudRaftSignalCostOur Calculator
Main goalEstimate telemetry and compare pricesEstimate telemetry and compare vendor costSize self-hosted components
Starting pointInfrastructure and request rateInfrastructure or exact telemetryMeasured or estimated telemetry
ComputeStack level estimatesProduct level workload tiersCPU, RAM, and replicas per component
StorageStack level estimatesFixed compression valuesCompression, retention, replication, PVC, and object storage
PerformanceCost estimateCost estimateNetwork and disk write checks

Use CloudRaft or SignalCost to create an initial estimate. Then enter that workload into our calculator. Replace it with measured data when it becomes available.

Measure one normal week before production. Include a deployment and a busy period. Estimate development, staging, and production separately because they rarely produce the same volume.


How Our Calculator Works#

Ingestion#

Logs and traces are converted from daily volume to average and peak rates.

average MB/s = daily GB × 1,000 ÷ 86,400
peak MB/s = average MB/s × peak factor

The peak rate drives the Recommended compute and replica count.

Metrics start with samples per second:

samples/s = active series ÷ scrape interval in seconds

Each metrics backend uses its own bytes per sample. Mimir calculates compacted object storage and Kafka buffering. VictoriaMetrics calculates local storage, replication, and merge space. ClickHouse calculates a raw row size before column compression.

Storage#

The base storage equation is:

compressed storage = daily ingestion × retention ÷ compression ratio

Object storage adds planning headroom:

planned object storage = compressed storage × (1 + overhead)

Local disks keep free space for merges and uneven data placement:

planned local storage = compressed storage ÷ (1 - free disk fraction)

Replication creates physical copies. ClickHouse multiplies data by the replicas per shard. Loki, Tempo, and Mimir keep retained blocks in object storage. Their WAL, cache, index, and scratch paths use PVCs where required.

Kafka storage for Tempo and Mimir is visible in the tables. Kafka compute and service cost remain external.

Using the worked example, 104 GB/day of logs with 30 days of retention and 5× compression needs:

compressed logs = 104 GB/day × 30 days ÷ 5
                = 624 GB

with 10 percent overhead = 624 GB × 1.10
                         = 686.4 GB

This is the retained storage before any backend-specific local PVC, cache, WAL, or replication requirement.

Compute and Workers#

Each component has Minimum and Recommended CPU, RAM, and replicas. Workload formulas increase them when the selected load needs more capacity.

The calculator adds every component request, then applies node reservation and target utilization:

required CPU = component CPU ÷ (1 - CPU reservation) ÷ target utilization
required RAM = component RAM ÷ (1 - RAM reservation) ÷ target utilization

It calculates workers by CPU and RAM. The largest of those results and the configured worker floor becomes the worker count.

Network and Disk#

The network model includes raw ingestion, peak traffic, replication, and durable writes. It also calculates peak traffic per worker.

The disk model checks peak writes for Loki WAL, Victoria data PVCs, VictoriaMetrics storage, and ClickHouse data PVCs.

IOPS is checked only when you enter an average I/O size. Ingestion alone cannot predict query reads, cache misses, compaction, or ClickHouse merge load.

Cross zone cost is included only when you enter a cross zone traffic share.


Choosing a Stack#

There is no single best stack. Choose based on the query language, storage model, and operational work your team can support.

StackGood FitMain Work
LGTMGrafana users who need LogQL, TraceQL, PromQL, and object storageMore distributed services, caches, and external Kafka
VictoriaTeams that want an efficient metrics TSDB and fast operational searchPVC planning, backups, and Victoria specific query behavior
ClickHouse OSSTeams that want tiered storage and SQL across all signalsSchemas, shards, replicas, Keeper, and merge operations
HybridTeams with different needs for each signalMore integrations and operational boundaries

LGTM#

LGTM fits teams that want the full Grafana ecosystem. Loki, Tempo, and Mimir use familiar Grafana query languages and integrate directly with Grafana. Their distributed modes scale each write, read, cache, and scheduling path separately.

Loki, Tempo, and Mimir can keep retained data in S3-compatible object storage. This reduces dependence on large local PVCs, but the stack has more services to deploy and tune.

Victoria#

VictoriaMetrics is strong as a metrics TSDB. VictoriaLogs provides fast search for recent operational logs. The distributed components are easier to follow because their names describe their roles, such as vminsert, vmstorage, vmselect, vlinsert, and vlstorage.

The Victoria plan uses PVCs for primary data. VictoriaMetrics provides vmbackup for copying snapshots to S3-compatible storage. VictoriaLogs and VictoriaTraces need volume snapshots or a separate backup and archive path until object storage support fits your production requirements.

ClickHouse OSS#

ClickHouse provides columnar compression, fast analytical search, and SQL. Structured logs can reach about 14× to 33× compression, depending on the schema. Storage policies can keep recent data on local SSD and move older parts to S3. Keeping all three signals in one database also makes correlation easier when they share service, trace, and resource attributes.

The tradeoff is ownership. You must manage schemas, indexes, materialized views, shards, replicas, Keeper, and background merges.

The log storage comparison in Part 3 covers VictoriaLogs search, Loki on S3, and ClickHouse storage tiering in more detail. Part 4 compares their cold storage and query paths.

The default ClickHouse plan uses Grafana. Selecting HyperDX adds the ClickStack application and MongoDB.


Cloud Data Costs#

Cloud collection can add costs that are not included in the calculator.

  • Check whether API calls, exports, and managed connectors are billed
  • Check whether data crosses an availability zone, region, or cloud provider
  • Price private connectivity and public egress separately
  • Filter unwanted data before it leaves the source
  • Buffer data when the destination is unavailable

Filter close to the source. There is no value in paying to transfer, store, and index data that nobody uses.

The calculator does not estimate Internet egress, load balancer, API request, or managed connector charges.


Practical Workflow#

  1. Measure or estimate logs, traces, active series, and scrape interval
  2. Set retention and peak factor
  3. Select a backend for each signal
  4. Review Minimum as the starting plan
  5. Review Recommended for peak load and availability
  6. Check object storage, PVCs, network, disk, and external Kafka
  7. Compare all three stacks with the same inputs
  8. Test real ingestion and queries
  9. Replace compression and I/O estimates with measurements
  10. Recalculate after a large change in traffic, retention, or labels

What Is Not Included#

The calculator does not predict exact query latency. It does not price every object storage request, backup, snapshot, load balancer, Internet transfer, collector, or Kafka broker.

It gives you a clear first plan. Use it to compare architectures and find the main capacity limits. Then validate the result with a representative load test.

Open the Self-Hosted Observability Capacity Calculator and test it with your workload.

If you find a missing component or a better source, email [email protected] or connect with me on LinkedIn.

Capacity Planning for a Self-Hosted Observability Stack
https://blogs.thedevopsguy.biz/blog/observability-capacity-planning
Author Akash Rajvanshi
Published at September 25, 2026