Prometheus as the data backend and Grafana as the visualization layer were the natural starting point. The combination made operational data visible quickly and worked while the number of series, dimensions and users remained bounded. Then product and system analysis expanded across users, links, pages, downtime and other behavioral and operational signals. Volume and cardinality grew beyond Prometheus's comfortable operating range.

During working hours, expensive queries competed with ingestion and retention. Grafana dashboards became slow or unavailable because the Prometheus backend was under pressure. Scaling the same storage design meant paying more cloud cost for a system that remained fragile under normal use.

The issue was workload fit, not bad tools

Prometheus is excellent for recent infrastructure metrics and bounded label sets. The problem appeared when we asked it to behave like a large analytical platform: long retention, high-cardinality dimensions, broad historical queries and many concurrent exploratory aggregations. Grafana itself was not the bottleneck, so there was no reason to replace the interface users already knew.

A tool can be excellent at its intended job and still be the wrong storage and compute model for a workload that has changed.

The same data followed two parallel paths

Product and system data was being delivered in parallel. Prometheus supplied the low-latency path used by Grafana, while Snowplow collected the event stream and landed it in Snowflake for analytical history. Both paths represented users, page and link activity, availability, downtime and other operating conditions.

Before: duplicated delivery
users + pages + links + system events
  ├─→ Prometheus ─────────────→ Grafana
  │       real-time path
  │
  └─→ Snowplow → Snowflake
          analytical history

This duplication delivered freshness, but it also meant operating an increasingly expensive Prometheus backend for data that already existed in Snowflake. The goal became consolidation: keep Grafana, make Snowflake fresh enough for operational use and remove the parallel Prometheus path.

First move: scheduled Snowflake Tasks

Because Snowplow data already landed in Snowflake, the first consolidation attempt used scheduled Tasks to build the user, page, link, availability and system views consumed by Grafana. Warehouse compute could scale independently and suspend when idle, avoiding the resource pressure of the Prometheus backend.

But scheduled Tasks were not real-time enough. Freshness remained tied to the schedule. Shorter intervals increased repeated work and warehouse starts; longer intervals left Grafana behind the current state. The architecture was stable, but the user experience could not yet replace the parallel Prometheus path.

Second move: incremental change with Snowflake Streams

Snowflake Streams changed the processing unit from “recompute a time window” to “consume the Snowplow rows committed since the last transactional offset.” A change-aware consumer could update only the user, page, link, downtime and system aggregates affected by new events.

Evolution of the processing model
# Scheduled recomputation
every interval:
    scan recent_window
    rebuild aggregates

# Incremental change processing
when stream_has_changes:
    delta = consume(stream)
    validate(delta)
    merge affected aggregates
    commit atomically

Streams did not execute the transformation themselves. They provided transactional CDC metadata and an offset. The processing layer consumed the stream with DML, advanced that offset on commit and could run only when change data was available.

After: one real-time analytical path
users + pages + links + system events
  → Snowplow
  → Snowflake event tables
  → Snowflake Streams
  → incremental operational models
  → Grafana

Prometheus path: removed

Why the design cost less

Less repeated scanning: processing touched changed rows rather than rescanning large recent windows.
Elastic compute: analytical work used Snowflake warehouses that could resize and auto-suspend.
Storage fit: large historical operational datasets moved to a platform designed for analytical retention and scans.
Separated concerns: Grafana remained the familiar interface while Snowflake took responsibility for scalable storage, transformation and analytical queries.

Why it became more responsive

A fixed schedule asks whether work should run because time passed. A stream-aware design asks whether committed data changed. That removed unnecessary executions and allowed relevant deltas to be processed much sooner, bringing operational analytics close to real time without continuously running the largest warehouse.

Reliability details that matter

Incremental designs move complexity rather than removing it. Each independent consumer needs its own stream offset. Streams must be consumed before they become stale. MERGE logic must be idempotent, updates and deletes need correct handling, schema changes require compatibility checks, and lag must be observable.

Operational contract
monitor:
  stream_lag
  stale_after
  rows_consumed
  merge_duration
  warehouse_credits
  data_freshness

alert when:
  lag > freshness_slo
  or stream_near_stale
  or merge_failed

The result

Snowflake Streams made the Snowplow-to-Snowflake path real-time enough to replace the parallel Prometheus delivery. Grafana remained the visualization surface, but Snowflake became its operational and analytical backend. Cloud cost fell because the platform stopped operating an increasingly oversized Prometheus system and stopped recomputing unchanged data.

The broader lesson is that visualization and data infrastructure are separable decisions. A familiar interface can remain while its storage and processing model evolves. Put high-cardinality operational history and exploratory analysis on a system designed for analytical scale, then connect the visualization layer with explicit freshness and cost contracts.

This case study is intentionally anonymized and simplified. It does not identify the company or expose internal data, capacity figures or schemas.