Back to Blog
Engineering

Building a Global Monitoring Mesh: Architecture Deep Dive

How we designed SteadyStack's distributed check infrastructure for reliability, low latency, and horizontal scalability across 7 pinned global edge regions.

A
Alex GutscherSteadyStack Engineering
June 5, 202612 min read

SteadyStack's monitoring infrastructure is built as a distributed mesh — a network of lightweight check nodes pinned across 7 global edge regions via Cloudflare Durable Object location hints that coordinate to verify endpoint availability and response times. Here's how it works under the hood.


High-Level Architecture Overview

At a high level, the system decouples into three distinct operational layers: Regional Probe DOs, Quorum Consensus Engine, and Alert Dispatch Pipelines:

CODE
[Target Application / API]
       ^             ^             ^
       |             |             |  (Parallel 60s Edge Checks)
+------+-------------+-------------+------+
| Regional Probe DO  Regional Probe DO    |  (7 Pinned Cloudflare DO Regions)
| (weur: London)     (enam: Virginia)     |
+------+-------------+-------------+------+
       |             |             |  (Sub-30ms RPC Consensus)
       v             v             v
+-----------------------------------------+
| Quorum Consensus Engine (Durable Actor) |  (Consensus & Flapping Manager)
| - Aggregates 4-of-7 multi-region voting |
| - Computes rolling P50/P95/P99 latency  |
+--------------------+--------------------+
                     |
                     v
  [Alert Dispatch: Slack, Webhooks, PagerDuty]

Each layer is horizontally scalable, geographically pinned at the edge, and coordinates via stateful actors to maintain high-throughput check aggregation.


1. Regional Probe Durable Objects

Each check probe is executed by a geographically pinned Cloudflare Durable Object located at one of 7 verified regions (wnam, enam, weur, eeur, apac, apac-ne, apac-se). When scheduled:

  • Probes pull target configurations cached locally in memory or fast KV stores.
  • Probes issue non-blocking HTTP, DNS, SSL, or TCP probes with strict timeout budgets.
  • Latency and TLS handshake metrics are measured with sub-millisecond precision using native Web APIs (performance.now()).

Because edge nodes are distributed worldwide, checks accurately reflect real-world user latency rather than synthetic measurements taken inside a single cloud provider's private network.


2. The Stateful Consensus Coordinator

When an edge node detects a non-2xx status code or a connection timeout, it does not page your team immediately. Instead, it engages the MonitorChannel Durable Object:

TYPESCRIPT
// Quorum Consensus Dispatch Logic
export async function verifyFailure(
  targetId: string,
  originFailure: ProbeResult,
  env: Env,
): Promise<ConsensusResult> {
  // Dispatch parallel verification to two geographically distinct edge nodes
  const [probeA, probeB] = await Promise.all([
    dispatchSecondaryProbe(targetId, "fra-eu-central", env),
    dispatchSecondaryProbe(targetId, "sin-ap-southeast", env),
  ]);

  const votes = [originFailure, probeA, probeB];
  const confirmedDown = votes.filter((v) => !v.success).length;

  return {
    isDown: confirmedDown >= 2,
    voteCount: confirmedDown,
    totalProbes: 3,
    timestamp: Date.now(),
  };
}

This 2/3 consensus protocol eliminates 99.9% of transient false positives caused by single-node transit drops or BGP route flapping.


3. Streaming Telemetry Pipeline

Check results flow through a streaming ingestion pipeline that normalizes latency, status codes, and DNS timings:

Metric StageResolutionRetention (Free)Retention (Pro)
Real-Time Pings60s30 Days365 Days
Latency Rollups1hr / 24hr90 Days3 Years
Incident LogsExact logsUnlimitedUnlimited

The pipeline feeds real-time dashboard graphs, incident timelines, and public status pages with under 300ms propagation delay.


Key Lessons Learned

Building a global monitoring mesh taught us that network reliability is inherently non-deterministic. Even with redundant nodes in every continent, regional ISP drops and CDN cold-starts happen daily. Multi-region quorum consensus isn't an optional optimization — it's the prerequisite for building an alerting system that on-call engineers can trust with their sleep.

Tags
#monitoring architecture#distributed systems#Cloudflare Workers#Durable Objects#infrastructure
A

Alex Gutscher

Author

Core engineer and distributed systems enthusiast at SteadyStack. Building global edge monitoring mesh networks and 4-of-7 quorum incident alert pipelines.

Found this article helpful?
Quorum-Verified Monitoring

Stop 3 AM false alarms with SteadyStack

Get multi-region edge quorum consensus verification, zero false alarms, and custom branded status pages — completely free for up to 50 monitors.