Building a global monitoring platform that checks thousands of endpoints every 60 seconds presents a classic distributed systems problem: How do you execute synthetic pings from geographically pinned edge regions without introducing multi-second consensus delays or false positives?
Traditional monitoring platforms rely on dedicated fleets of virtual machines in AWS or GCP regions. Centralized pollers execute cron loops, write raw results to centralized Postgres or TimeScale instances, and process alerting queues asynchronously.
At scale, this centralized architecture suffers from three major flaws:
- High Infrastructure Costs: Running idle EC2/Compute instances in 20+ regions just to issue HTTP pings incurs massive fixed compute and egress costs.
- Alert Latency: Serialized database writes and queue polling add 10 to 30 seconds of latency between failure detection and dispatch.
- Rotating Origin False Positives: Unpinned cron executions jump continents and trigger false degradation alarms on geo-routed targets.
To solve this, we built SteadyStack directly on top of Cloudflare's geographically pinned Durable Objects and our 4-of-7 Quorum Consensus Engine. Here is how the architecture works under the hood.
The System Architecture Overview
SteadyStack's monitoring engine consists of three decoupled layers:
[Target Application / Server]
^ ^ ^
| | | (Parallel pings)
+------+-------------+-------------+------+
| Regional Probe DO Regional Probe DO | (7 Pinned Cloudflare Edge Regions)
| (weur: London) (enam: Virginia) |
+------+-------------+-------------+------+
| | | (Internal Edge RPC)
v v v
+-----------------------------------------+
| Quorum Consensus Engine (Durable Actor) | (Stateful Quorum Coordinator)
| - Aggregates 4-of-7 consensus |
| - Manages alert deduplication & flapping|
+--------------------+--------------------+
|
v
[Alert Dispatches: Slack / Webhooks / PagerDuty]- Regional Probe Durable Objects: Geographically pinned Cloudflare Durable Objects deployed across 7 verified regions (
wnam,enam,weur,eeur,apac,apac-ne,apac-se). - Quorum Consensus Engine: Stateful actor that maintains rolling latency state, excludes slow/flapping probes, and coordinates 4-of-7 consensus voting.
- Dispatch Engine: Asynchronous WebCrypto-signed alert pipelines that push notifications to Slack, Discord, PagerDuty, and custom HTTP webhooks in under 300ms.
1. Executing 60-Second Checks at the Edge
Instead of running long-lived timer processes on virtual machines, SteadyStack uses Cloudflare Scheduled Triggers (Cron Triggers) synchronized down to the minute.
Each minute, scheduled edge workers pull target configurations cached in Cloudflare KV with a 60-second TTL:
// Edge Worker Scheduled Handler
export default {
async scheduled(
event: ScheduledEvent,
env: Env,
ctx: ExecutionContext,
): Promise<void> {
// 1. Fetch cached target batch assigned to this edge POP
const targets = await env.DNS_CACHE.get<MonitorTarget[]>(
"targets:batch:pop",
"json",
);
if (!targets || targets.length === 0) return;
// 2. Execute parallel pings with strict AbortController timeouts
const pingPromises = targets.map((target) => executeEdgePing(target, env));
// 3. Process results without blocking scheduled handler lifecycle
ctx.waitUntil(Promise.all(pingPromises));
},
};
async function executeEdgePing(target: MonitorTarget, env: Env): Promise<void> {
const controller = new AbortController();
const timeoutId = setTimeout(
() => controller.abort(),
target.timeoutMs || 5000,
);
const startTime = performance.now();
try {
const response = await fetch(target.url, {
method: target.method || "GET",
headers: target.headers,
signal: controller.signal,
});
const latencyMs = Math.round(performance.now() - startTime);
const isSuccess = response.status >= 200 && response.status < 400;
if (!isSuccess) {
// Initiate RPC quorum check to Durable Object Coordinator
await triggerQuorumVote(
target,
{ success: false, statusCode: response.status, latencyMs },
env,
);
} else {
// Report success telemetry to state coordinator
await reportTelemetry(
target.id,
{ success: true, statusCode: response.status, latencyMs },
env,
);
}
} catch (err: any) {
const latencyMs = Math.round(performance.now() - startTime);
// Timeout or network error — trigger consensus validation
await triggerQuorumVote(
target,
{ success: false, error: err.name, latencyMs },
env,
);
} finally {
clearTimeout(timeoutId);
}
}2. Low-Latency Consensus via Durable Object Actors
When an edge node observes a failed check, it needs to verify whether the target is down globally or if the failure was localized to its own network path.
We use Durable Objects as stateful actors. Each monitor target has a designated MonitorChannel Durable Object actor that lives in the primary region closest to the target infrastructure.
// Durable Object Quorum Coordinator
export class MonitorChannel implements DurableObject {
private state: DurableObjectState;
private activeVotes: Map<string, ProbeVote[]> = new Map();
constructor(state: DurableObjectState, env: Env) {
this.state = state;
}
async fetch(request: Request): Promise<Response> {
const url = new URL(request.url);
if (url.pathname === "/vote") {
const vote: ProbeVote = await request.json();
return this.handleQuorumVote(vote);
}
return new Response("Not Found", { status: 404 });
}
private async handleQuorumVote(vote: ProbeVote): Promise<Response> {
const voteId = vote.batchId;
const votes = this.activeVotes.get(voteId) || [];
votes.push(vote);
this.activeVotes.set(voteId, votes);
// If this is the first failed vote, immediately request pings from 2 partner POPs
if (votes.length === 1) {
await this.requestPartnerProbes(vote.targetUrl, vote.originRegion);
// Wait up to 1200ms for partner votes to arrive
await new Promise((resolve) => setTimeout(resolve, 1200));
}
const currentVotes = this.activeVotes.get(voteId) || [];
const failures = currentVotes.filter((v) => !v.success).length;
// 2 out of 3 (or more) failures = Confirmed Hard Outage
if (failures >= 2 && !this.isCurrentlyFailing) {
this.isCurrentlyFailing = true;
await this.dispatchIncidentAlert(vote.targetId, currentVotes);
}
return new Response(JSON.stringify({ confirmed: failures >= 2 }), {
headers: { "Content-Type": "application/json" },
});
}
}Why Durable Objects?
Unlike traditional databases that require locking and complex SQL transactions to coordinate state, a Durable Object guarantees single-threaded, in-memory execution with automatic persistence. This allows consensus evaluation to execute in under 15 milliseconds once partner votes arrive.
3. The Trade-Offs & Known Limitations
No architecture is without trade-offs. Here is where our edge consensus approach requires explicit engineering choices:
- Non-HTTP Protocol Constraints at the Edge: Cloudflare Workers excel at HTTP/HTTPS, SSL, and DNS checks. For raw ICMP PING or custom TCP port checks, edge workers cannot issue raw IP sockets. To handle TCP/Port checks, we deploy lightweight private probe binaries (
steadystack-probe) that connect back to Durable Objects via WebSockets. - Cold-Start Jitter on Low-Traffic Workers: If an edge POP hasn't executed a scheduled job in several hours, the worker script initialization can add 10-20ms of execution latency. We mitigate this by pre-warming worker isolates across major POP hubs.
Benchmark Comparison: Detection Speed & Cost
| Metric | Traditional Cloud VM Architecture | SteadyStack Edge Consensus Mesh |
|---|---|---|
| Check Resolution | 5 minutes (Standard) | 60 seconds (1 minute) |
| Consensus Latency | 15–30 seconds | < 1.2 seconds total |
| Sovereign Probe Nodes | 3–8 Datacenters | 7 Pinned Edge DOs |
| Compute Overhead / Mo | $400–$1,200 (VM Fleets) | $5 (Serverless Edge) |
| False Positive Rate | ~3.2% of alerts | < 0.02% (Quorum Locked) |
Conclusion
By combining Cloudflare's 7 pinned global edge regions with Durable Object state coordination, SteadyStack achieves sub-second consensus verification without paying the multi-region cloud VM tax or suffering from rotating origin jitter.
If you are interested in exploring the codebase or deploying private edge probes in your own homelab or VPC, check out SteadyStack on GitHub or start monitoring for free.
Alex Gutscher
AuthorCore engineer and distributed systems enthusiast at SteadyStack. Building global edge monitoring mesh networks and 4-of-7 quorum incident alert pipelines.
Stop 3 AM false alarms with SteadyStack
Get multi-region edge quorum consensus verification, zero false alarms, and custom branded status pages — completely free for up to 50 monitors.