Get the time, duration, and stated cause in writing, keep your own log of every outage, and judge the host on the pattern across a year rather than on any single incident.

The explanation you will get

Something failed overnight, and by the time anybody noticed it was working again.

The status page says resolved. The support reply says a network issue affecting a subset of servers, now corrected.

That is not much of an explanation and it is usually all you will receive, because the detail belongs to somebody else's infrastructure and is not going to be shared with a small account.

Which means the useful response is not extracting a better answer. It is recording what happened so that the pattern becomes visible over a year.

Why it happens at three in the morning

Not coincidence, and worth knowing because it explains a great deal.

Maintenance, updates, and migrations are scheduled overnight deliberately, since that is when the fewest people are affected.

Automated processes also run then: backups, log rotation, database maintenance, and certificate renewals, any of which can fail or consume enough resource to take a site down briefly.

And nobody is watching, so a two-minute problem and a two-hour problem look identical the following morning unless something was monitoring.

An outage at three does not indicate anything sinister. It indicates that the things most likely to fail are scheduled when you are asleep.

What to ask for

Ask in writing and keep the reply, not because you expect much from any single answer, but because five of them across a year is evidence and one is an anecdote.

Keep your own record

The practice that makes this worth any effort at all.

A simple list: date, time, duration, what you were told, and whether you found out from monitoring or from a customer.

Four lines per incident, kept in the same place, for a year.

That record answers the question you actually have, which is not what happened on one night but whether this host is reliable enough.

Without it you are relying on impression, and impression is dominated by whichever outage was most inconvenient rather than by frequency.

Most businesses that keep this record for a year find either that outages are rarer than they felt, or that they cluster in a way that makes the conversation with the host straightforward.

Distinguish the site from the server

A diagnostic point, because these get conflated and have different remedies.

The server being unreachable is the host's problem: nothing responds, and other sites on the same machine are also down.

The site failing while the server is fine is usually yours: a plugin update, a script consuming memory, a database problem, or a certificate expiry.

The distinction is usually visible in the error, and in whether the hosting panel itself was reachable during the outage.

Recording which type each incident was matters, since a year of your own failures is a different conversation from a year of theirs.

A worked example

A business experienced several short outages over about six months and grew convinced their host was poor.

They started recording each one properly.

Over the following year there were eight incidents. Five were brief and overnight, two were during the day, and one was several hours.

Reading the record, three had followed a plugin update they had performed, and two coincided with a scheduled maintenance window the host published in advance and nobody had subscribed to.

So half were their own doing and two had been announced.

They subscribed to the status notifications, moved updates to a weekday morning, and stopped attributing everything to the host.

The remaining three genuine incidents in a year was a figure they decided they could accept, which they could not have concluded from impression alone.

Subscribe to the status page

A small step that changes the experience considerably.

Most hosts publish a status page with notifications for incidents and planned maintenance, and almost no small business customer subscribes to it.

That means announced maintenance arrives as an unexplained outage, which is where a large share of the frustration comes from.

Subscribe, and route the notifications somewhere a person will see them rather than into a general inbox.

Knowing at eleven the previous evening that a window is scheduled turns an incident into an expectation.

When the pattern justifies moving

The decision the record exists to inform.

Repeated unannounced outages during business hours, a support function that does not respond within a stated time, or the same stated cause recurring across months are the signals that matter.

A few brief overnight incidents in a year is normal and moving host over them is likely to buy a different set of the same.

Where you do decide to move, the record is what makes the conversation with the next host specific: this is what we experienced, what is your position on it.

And migration has its own risk of downtime, so it should be a response to a pattern rather than to an incident.

The counter-case

This can become disproportionate scrutiny of something that barely matters.

For a brochure site producing a handful of enquiries a month, a forty-minute outage at three in the morning has no measurable cost, and building a monitoring and record-keeping practice around it is effort spent on the wrong thing.

There is also a version where a business chases a technical explanation nobody is going to provide, spending hours on support tickets to establish a cause that would not change any decision.

Where a site takes orders or bookings, the record is worth keeping. Where it does not, note the date in a calendar and move on.

The point is proportion: know how often, and stop there unless the number is uncomfortable.

What to do

  1. Ask for start time, duration and cause, in writing.
  2. Establish whether it was the site or the server.
  3. Record four lines per incident.
  4. Note how you found out.
  5. Subscribe to the status page.
  6. Move your own updates to a weekday morning.
  7. Judge on the year, not the incident.

Step three is the whole practice, and a year of it usually produces a calmer conclusion than the impression it replaces.

What a guarantee actually promises is covered in uptime percentages in hours.


Frequently asked questions

Why do outages happen overnight?

Because maintenance, updates, migrations, backups, and certificate renewals are scheduled then deliberately. Nobody is watching, so a two-minute and a two-hour problem look identical afterwards.

What explanation should I expect?

Usually very little. The detail belongs to somebody else's infrastructure and is not shared with small accounts. The useful response is recording the pattern rather than extracting a better answer.

What should I record?

Date, time, duration, what you were told, and whether you found out from monitoring or a customer. Four lines per incident, kept in one place.

How do I tell whose fault it was?

If nothing responds and other sites on the same machine are down, it is the server. If the site fails while the hosting panel works, it is usually a plugin, script, or certificate.

Should I subscribe to the status page?

Yes, and route it to a person. Almost no small business does, which is why announced maintenance arrives as an unexplained outage.

When should I change host?

On a pattern: repeated unannounced outages in business hours, unresponsive support, or the same cause recurring. A few brief overnight incidents a year is normal.

West Coast Media Solutions Inc. provides web design, web development, hosting, digital marketing, and business consulting to organisations across Canada, drawing on more than twenty-five years in the field.

Site went down overnight?

Write four lines about it somewhere permanent. A year of those answers the question the incident cannot.

Start a Conversation