Back to Blog
Guides

Uptime Monitoring 101: Synthetic vs RUM Guide

A complete beginner-to-intermediate guide to uptime monitoring: probe types, critical metrics, check intervals, alert routing, and multi-region consensus.

A
Alex GutscherSteadyStack Engineering
July 28, 20269 min read

If your production API or website drops right now, would your engineering team know within 60 seconds — or would you find out from an angry customer on Twitter?

Uptime monitoring is the practice of continuously testing your digital endpoints from outside your network perimeter to detect outages, latency spikes, and SSL issues before your users encounter them. This guide covers the essential principles of modern synthetic monitoring.


The 6 Types of Synthetic Checks

Different parts of your architecture require different inspection protocols:

Check TypeProtocol / TargetWhat It VerifiesBest Used For
HTTP / HTTPSGET, POST, HEADStatus codes, body assertions, headersWeb apps & REST APIs
TCP PortRaw TCP handshakeSocket connectivity & listener healthRedis, Postgres, SSH
ICMP PingPacket round-tripNetwork reachability & packet dropBare-metal hosts & VPNs
SSL / TLSX.509 CertificateCertificate validity & expiration daysPublic domains & origins
DNS RecordA, CNAME, TXTGlobal DNS propagation & consistencyNameservers & CDN edges
Heartbeat / CronIncoming HTTP pingBackground job completion deadlinesQueue workers & cron jobs

4 Core Metrics That Tell the Real Story

Availability is more than a binary "up or down" flag:

  1. Uptime Percentage: Overall percentage of successful checks over a rolling 30/90-day window (e.g. 99.95%).
  2. Mean Time to Detect (MTTD): How many seconds elapse between when a service crashes and when an alert triggers.
  3. P95 / P99 Latency: 95th-percentile response time. A service that takes 10 seconds to respond is functionally unavailable for users.
  4. SSL Expiration Horizon: Number of days remaining before an SSL/TLS certificate expires.

Why Check Frequency Dictates Detection Speed

Your check interval sets the mathematical limit for how quickly you can discover an incident:

  • 5-Minute Interval (300s): Worst-case detection latency is 5 minutes; completely misses transient micro-outages lasting 2–4 minutes.
  • 60-Second Interval (60s): Real-time awareness; worst-case detection time is 60 seconds.
  • 30-Second Interval: Used for high-frequency payment gateways and critical financial APIs.
CODE
5-Minute Poller (UptimeRobot Free):
00:00 [200 OK] ------------------ 5 mins ------------------ 00:05 [200 OK]
                     ^-- 3-min outage occurs here --^
                     (Customer checkout fails; zero alerts)

SteadyStack 60-Second Poller:
00:00 [200 OK] - 1m - 00:01 [200 OK] - 1m - 00:02 [500 DOWN] ==> Paged in 60s

The Golden Rule: Avoid Single-Probe False Alarms

Never configure your alerts to page an engineer on a single failed ping from a single location. Internet routing anomalies, local ISP packet loss, and CDN edge refreshes happen continuously.

Require multi-region quorum verification (such as SteadyStack's 2/3 edge consensus) to confirm failures before triggering high-priority notification channels.

Tags
#uptime monitoring#uptime monitoring guide#website monitoring#check interval#DevOps
A

Alex Gutscher

Author

Core engineer and distributed systems enthusiast at SteadyStack. Building global edge monitoring mesh networks and 4-of-7 quorum incident alert pipelines.

Found this article helpful?
Quorum-Verified Monitoring

Stop 3 AM false alarms with SteadyStack

Get multi-region edge quorum consensus verification, zero false alarms, and custom branded status pages — completely free for up to 50 monitors.