Change one thing, wait for enough data, then change the next. If you must ship several at once, accept that you are testing the bundle and cannot attribute the result to any part of it.

The afternoon that teaches nothing

Somebody decides the enquiry page is underperforming.

They rewrite the headline, move the form higher, change the button from grey to green, shorten it by two fields, and add a line about response times.

Enquiries rise by a third. Everybody is pleased, and nobody knows anything.

Five changes went out together. One of them may have produced the entire improvement while another was quietly making things worse, and the net result was positive.

There is no way to find out afterwards, and the lesson that could have been reused on nine other pages was not learned.

Why it matters more than it seems

The immediate result is the same either way, so the argument for isolating changes looks academic.

It is not, for two reasons.

A finding that transfers is worth many times a single improvement. Knowing that shortening the form was what worked means you can shorten every form you own.

And a bundle that fails leaves you nowhere. If the five changes had reduced enquiries by a fifth, the only available response is to revert all of them, including whichever one was helping.

The sequence that works

The fifth item is the one most often skipped, and a change that produced nothing is a genuine finding worth writing down.

The obvious objection

This is slow, and the objection is fair.

Five changes tested sequentially, with a sensible waiting period, could take most of a year on a small site.

There are two honest answers.

For anything you are confident about and which is clearly an improvement in itself, such as fixing a broken form or making a phone number tappable, do not test it. Just do it. Testing a fix is a waste of a testing slot.

Reserve sequential testing for genuine uncertainty, where reasonable people would disagree about the outcome. That is usually two or three questions per year, not fifteen, and at that rate the pace is entirely workable.

Running two versions at once

The alternative to waiting, where you have the traffic for it.

A split test shows two versions simultaneously to different visitors, which removes the seasonality and timing problems that make sequential testing unreliable.

It is the better method when it is available and it has a requirement most small sites cannot meet, which is enough visitors for the difference to separate from chance.

It also still requires one change per test. Two versions differing in four ways has exactly the same attribution problem as changing four things on a Tuesday.

A worked example

A service business tested three changes to their quote request page over five months, one at a time.

Removing two of the seven fields produced a clear rise in submissions.

Adding a line about typical response time produced no measurable change at all.

Moving the form above the explanatory text reduced submissions noticeably, which nobody had expected, and was reverted.

Had all three shipped together, the net effect would have been a modest improvement and they would have concluded that the bundle worked.

Instead they learned that field count mattered on their site and position did not, and applied the field reduction to four other forms over the following month, which produced more total improvement than the original test had.

Note the date, and everything else that changed

The discipline that makes sequential testing trustworthy.

Keep a simple log: date, what changed, what you expected, what happened.

Also record anything else that moved in the same period, because sequential testing is vulnerable to exactly this. A change made the week a campaign started, or the week a competitor closed, will take credit for something else entirely.

Where something significant coincides with your change, the honest response is to note it and treat the result as inconclusive rather than to claim it.

When shipping a bundle is right

Sometimes you cannot test one thing at a time and should not pretend otherwise.

A redesign is a bundle by definition. So is a page rewritten from nothing, and so is any change where the parts only make sense together.

In those cases test the bundle honestly: record a baseline, ship the whole thing, and evaluate the whole thing.

What matters is knowing which you are doing. A bundle evaluated as a bundle is a legitimate test. A bundle described afterwards as though one element caused the result is not.

The test somebody already ran

Without a log, the same ideas cycle back every eighteen months or so.

Somebody new joins, or an agency changes, or the season comes round again, and a suggestion arrives that was tried in 2019 and did nothing.

Nobody remembers, because the person who ran it has left or has simply forgotten, and there is no reason it would have stuck in anybody's memory. It produced no effect, which is precisely the kind of result that leaves no trace.

So the change gets made again, occupies another testing slot, and produces nothing again.

The remedy is that the log lives with the business rather than with whoever ran the test, in the same place as the enquiry records and the baseline documents.

Reading it before proposing anything takes two minutes and is the cheapest part of the whole discipline. It also makes the log worth keeping, since a record nobody consults stops being maintained within a year.

The counter-case

Rigid one-change discipline can cost more than it returns.

Small independent fixes on different pages do not interfere with each other, and holding back four unrelated corrections so they can be tested in sequence is pointless caution.

Changes that are obviously right by other standards, such as accessibility corrections or removing something broken, should not wait for a testing slot either.

And on a very low traffic site, sequential testing may never accumulate enough data to conclude anything, in which case the honest approach is to make changes on judgement, record what you did, and stop pretending it is measurement.

The rule applies to genuine uncertainty on pages that matter. Everywhere else, use judgement and get on with it.

The method

  1. Fix anything broken without testing it.
  2. Identify the two or three real questions.
  3. Change one thing and note the date.
  4. Record what else changed that period.
  5. Wait, then record the result including no effect.
  6. Apply anything you learned to comparable pages.
  7. Evaluate bundles as bundles, and say so.

Step six is where the value is, and it only exists if step three was done properly.

Learning nothing from a flawed method is covered in searching for yourself and learning nothing.


Frequently asked questions

Why change only one thing at a time?

Because a result from several simultaneous changes cannot be attributed to any of them. One may have helped while another quietly hurt, and the net looked positive.

Is sequential testing too slow for a small business?

It would be if you tested everything. Fix broken things without testing, and reserve sequential testing for the two or three genuine uncertainties a year.

Should I use a split test instead?

Where you have the traffic, yes, since it removes timing and seasonality problems. It still requires one change per test: two versions differing in four ways has the same attribution problem.

What if I have to ship several changes together?

Then test the bundle honestly. Record a baseline, ship it all, evaluate it all. The error is describing the result afterwards as though one element caused it.

What should I record?

Date, what changed, what you expected, what happened, and anything else that moved in the same period. A campaign starting the same week will take credit for your change.

Is a change that did nothing worth recording?

Yes. No effect is a genuine finding, and it stops the same idea being proposed again in eight months as though it were untried.

West Coast Media Solutions Inc. provides web design, web development, hosting, digital marketing, and business consulting to organisations across Canada, drawing on more than twenty-five years in the field.

Five changes planned for one afternoon?

Ship the fixes now, then test the two you are genuinely unsure about, one at a time.

Start a Conversation