Skip to content
EnterraHost
Products
Tools
Hosting Email Support Client Area Find a domain

How to understand a downtime event in EnterraMon

EnterraMon reported your site down for four minutes. Whether that was a real outage, a network blip, or a deliberate restart is answerable, and the answer is in how the event was recorded rather than in how long it lasted.

An event needs confirmation

A single failed check does not raise an alert. EnterraMon waits for a run of consecutive failures before it decides the site is down, and the count is why a momentary blip does not wake anybody.

This is the single most useful thing to understand about the alerting, because it explains both directions of the surprise. A four-minute outage that appears in the dashboard may produce no email if the checks that failed were not consecutive. A site that is genuinely down produces an alert after the confirmation threshold rather than on the first failed request.

So the recorded duration is a measurement of the confirmed period, not of the whole time the server was unavailable.

The incident types

Not every incident is an outage. EnterraMon records several kinds, and they call for different responses.

Incident types EnterraMon records against a monitored site.
TypeWhat triggered it
Hard downThe site stopped answering
HTTP errorThe server answered with an error status rather than serving the page
ContentThe page content changed beyond the expected range
TitleThe page title changed
RedirectThe domain started redirecting somewhere unexpected

HTTP error and hard down are the pair that get confused. A site returning a 500 is up, in the sense that something is listening and answering, and completely broken from a visitor’s point of view. EnterraMon records the status code, which is what distinguishes the two in the incident list.

Content and title changes use a baseline

The content incident types work differently from the availability ones. EnterraMon records what the page looked like when it was healthy, then compares later checks against that.

The baseline is the page size and the page title from a known-good check. A change outside the expected range raises a warning, and a run of consecutive warnings escalates it.

Two consequences follow. A deliberate redesign will raise a content incident, which is expected rather than a fault. And a page that changes gradually, a little on each check, may stay inside the range the whole way and never raise one at all.

Where a content change was intended, the baseline is what needs updating, not the alert.

Recovery is its own event

When a site comes back, EnterraMon records a recovery and sends it on the same contacts as the alert. The pair is what makes the incident list readable, because a reader can see the end of an event rather than inferring it from the next line.

Recoveries are sent to everybody who received the alert. Nobody is told a site is down and then left to discover it came back by refreshing a dashboard.

Why a short outage may not appear

Three reasons, and all of them are ordinary.

The failed checks were not consecutive, so the confirmation never completed. A restart that takes forty seconds can land between two checks and show up as one failure rather than a run.

The check interval. A site that is down for less than the gap between checks may not be tested while it is broken.

The failure was on the path rather than at the server, and it did not affect every region. A single region monitoring from Cape Town will report a Cape Town routing problem as a downtime event, because from where it stands that is what it is.

That last one is the reason a second region changes the reading of an incident rather than only the latency. Adding a second region covers it.

What the alert does and does not mean

An alert means EnterraMon could not reach your site, repeatedly, from the region it checks from. It is a statement about reachability from that vantage point.

Whether your visitors saw the same thing depends on where they were and how they were routed. Where the incident matters to a client or a customer, the report is the better artefact than the email, because it carries the period, the incident list and the recovery together. Scheduling a weekly report is how that reaches them without anybody assembling it by hand.

Every guide in this category is on the EnterraMon page.