If your production API or website drops right now, would your engineering team know within 60 seconds — or would you find out from an angry customer on Twitter?
Uptime monitoring is the practice of continuously testing your digital endpoints from outside your network perimeter to detect outages, latency spikes, and SSL issues before your users encounter them. This guide covers the essential principles of modern synthetic monitoring.
The 6 Types of Synthetic Checks
Different parts of your architecture require different inspection protocols:
| Check Type | Protocol / Target | What It Verifies | Best Used For |
|---|---|---|---|
| HTTP / HTTPS | GET, POST, HEAD | Status codes, body assertions, headers | Web apps & REST APIs |
| TCP Port | Raw TCP handshake | Socket connectivity & listener health | Redis, Postgres, SSH |
| ICMP Ping | Packet round-trip | Network reachability & packet drop | Bare-metal hosts & VPNs |
| SSL / TLS | X.509 Certificate | Certificate validity & expiration days | Public domains & origins |
| DNS Record | A, CNAME, TXT | Global DNS propagation & consistency | Nameservers & CDN edges |
| Heartbeat / Cron | Incoming HTTP ping | Background job completion deadlines | Queue workers & cron jobs |
4 Core Metrics That Tell the Real Story
Availability is more than a binary "up or down" flag:
- Uptime Percentage: Overall percentage of successful checks over a rolling 30/90-day window (e.g. 99.95%).
- Mean Time to Detect (MTTD): How many seconds elapse between when a service crashes and when an alert triggers.
- P95 / P99 Latency: 95th-percentile response time. A service that takes 10 seconds to respond is functionally unavailable for users.
- SSL Expiration Horizon: Number of days remaining before an SSL/TLS certificate expires.
Why Check Frequency Dictates Detection Speed
Your check interval sets the mathematical limit for how quickly you can discover an incident:
- 5-Minute Interval (300s): Worst-case detection latency is 5 minutes; completely misses transient micro-outages lasting 2–4 minutes.
- 60-Second Interval (60s): Real-time awareness; worst-case detection time is 60 seconds.
- 30-Second Interval: Used for high-frequency payment gateways and critical financial APIs.
5-Minute Poller (UptimeRobot Free):
00:00 [200 OK] ------------------ 5 mins ------------------ 00:05 [200 OK]
^-- 3-min outage occurs here --^
(Customer checkout fails; zero alerts)
SteadyStack 60-Second Poller:
00:00 [200 OK] - 1m - 00:01 [200 OK] - 1m - 00:02 [500 DOWN] ==> Paged in 60sThe Golden Rule: Avoid Single-Probe False Alarms
Never configure your alerts to page an engineer on a single failed ping from a single location. Internet routing anomalies, local ISP packet loss, and CDN edge refreshes happen continuously.
Require multi-region quorum verification (such as SteadyStack's 2/3 edge consensus) to confirm failures before triggering high-priority notification channels.
Alex Gutscher
AuthorCore engineer and distributed systems enthusiast at SteadyStack. Building global edge monitoring mesh networks and 4-of-7 quorum incident alert pipelines.
Stop 3 AM false alarms with SteadyStack
Get multi-region edge quorum consensus verification, zero false alarms, and custom branded status pages — completely free for up to 50 monitors.