Prometheus as the data backend and Grafana as the visualization layer were the natural starting point. The combination made operational data visible quickly and worked while the number of series, dimensions and users remained bounded. Then product and system analysis expanded across users, links, pages, downtime and other behavioral and operational signals. Volume and cardinality grew beyond Prometheus's comfortable operating range.
During working hours, expensive queries competed with ingestion and retention. Grafana dashboards became slow or unavailable because the Prometheus backend was under pressure. Scaling the same storage design meant paying more cloud cost for a system that remained fragile under normal use.
The issue was workload fit, not bad tools
Prometheus is excellent for recent infrastructure metrics and bounded label sets. The problem appeared when we asked it to behave like a large analytical platform: long retention, high-cardinality dimensions, broad historical queries and many concurrent exploratory aggregations. Grafana itself was not the bottleneck, so there was no reason to replace the interface users already knew.
A tool can be excellent at its intended job and still be the wrong storage and compute model for a workload that has changed.
The same data followed two parallel paths
Product and system data was being delivered in parallel. Prometheus supplied the low-latency path used by Grafana, while Snowplow collected the event stream and landed it in Snowflake for analytical history. Both paths represented users, page and link activity, availability, downtime and other operating conditions.
users + pages + links + system events
├─→ Prometheus ─────────────→ Grafana
│ real-time path
│
└─→ Snowplow → Snowflake
analytical history
This duplication delivered freshness, but it also meant operating an increasingly expensive Prometheus backend for data that already existed in Snowflake. The goal became consolidation: keep Grafana, make Snowflake fresh enough for operational use and remove the parallel Prometheus path.
First move: scheduled Snowflake Tasks
Because Snowplow data already landed in Snowflake, the first consolidation attempt used scheduled Tasks to build the user, page, link, availability and system views consumed by Grafana. Warehouse compute could scale independently and suspend when idle, avoiding the resource pressure of the Prometheus backend.
But scheduled Tasks were not real-time enough. Freshness remained tied to the schedule. Shorter intervals increased repeated work and warehouse starts; longer intervals left Grafana behind the current state. The architecture was stable, but the user experience could not yet replace the parallel Prometheus path.
Second move: incremental change with Snowflake Streams
Snowflake Streams changed the processing unit from “recompute a time window” to “consume the Snowplow rows committed since the last transactional offset.” A change-aware consumer could update only the user, page, link, downtime and system aggregates affected by new events.
# Scheduled recomputation
every interval:
scan recent_window
rebuild aggregates
# Incremental change processing
when stream_has_changes:
delta = consume(stream)
validate(delta)
merge affected aggregates
commit atomically
Streams did not execute the transformation themselves. They provided transactional CDC metadata and an offset. The processing layer consumed the stream with DML, advanced that offset on commit and could run only when change data was available.
users + pages + links + system events
→ Snowplow
→ Snowflake event tables
→ Snowflake Streams
→ incremental operational models
→ Grafana
Prometheus path: removed
Why the design cost less
Why it became more responsive
A fixed schedule asks whether work should run because time passed. A stream-aware design asks whether committed data changed. That removed unnecessary executions and allowed relevant deltas to be processed much sooner, bringing operational analytics close to real time without continuously running the largest warehouse.
Reliability details that matter
Incremental designs move complexity rather than removing it. Each independent consumer needs its own stream offset. Streams must be consumed before they become stale. MERGE logic must be idempotent, updates and deletes need correct handling, schema changes require compatibility checks, and lag must be observable.
monitor:
stream_lag
stale_after
rows_consumed
merge_duration
warehouse_credits
data_freshness
alert when:
lag > freshness_slo
or stream_near_stale
or merge_failed
The result
Snowflake Streams made the Snowplow-to-Snowflake path real-time enough to replace the parallel Prometheus delivery. Grafana remained the visualization surface, but Snowflake became its operational and analytical backend. Cloud cost fell because the platform stopped operating an increasingly oversized Prometheus system and stopped recomputing unchanged data.
The broader lesson is that visualization and data infrastructure are separable decisions. A familiar interface can remain while its storage and processing model evolves. Put high-cardinality operational history and exploratory analysis on a system designed for analytical scale, then connect the visualization layer with explicit freshness and cost contracts.
This case study is intentionally anonymized and simplified. It does not identify the company or expose internal data, capacity figures or schemas.