Check how many pages your site actually has indexed against how many you wrote. Archives, tags, and attachment pages account for most of the difference.

The count that surprises people

Ask how many pages your site has and you will get a number based on what somebody wrote.

Check what is actually indexed and the number is frequently several times larger.

The difference is pages your platform generated: an archive for every tag, every category, every month, every author, and sometimes one for every image.

None of those was written by anybody, and most of them contain fragments of pages that already exist.

Where they come from

The fifth is the worst and the least known. Some platforms create a separate page for every uploaded image, containing that image and nothing else, which on a site with four hundred photographs is four hundred pages of nothing.

Why it matters

Since a page nobody visits sounds harmless.

They compete with your real pages, since an archive listing extracts from three service pages can rank instead of the service page itself.

They consume crawling, so your genuine updates are found more slowly.

They add a large volume of thin content to a site being assessed as a whole, which is exactly the pattern that has been discouraged.

And they produce a poor experience for anybody who does land on one, since an attachment page answers nothing.

None of that is catastrophic and all of it is unnecessary.

Find out what you have

Which takes about five minutes.

Open the indexing report in your search console and look at the number of indexed pages.

Compare it against the number of pages you believe you have.

Then look at the list and sort by address, which groups the generated ones together and makes the pattern obvious.

You are looking for repeated address fragments: tag, category, author, a year, or an image name.

The proportion is frequently three or four to one against your real pages.

A worked example

A business with about forty written pages had over nine hundred indexed.

The bulk were attachment pages for photographs, tag archives from tags applied once each, and monthly date archives going back six years.

They disabled attachment pages, removed the tags nobody used, and set date and author archives not to be indexed.

The indexed count fell to around sixty over about two months.

Their service pages improved modestly and, more usefully, the search console reports became readable for the first time.

The whole change was three settings and an afternoon of tidying.

What to do with each kind

Since they need different treatment.

Attachment pages should be disabled entirely or redirected to the image itself, since they serve nobody.

Date and author archives on a small site should be set not to be indexed, since they duplicate the blog index.

Tag archives should be kept only where the tag genuinely groups several useful pages, and removed where it was applied once.

Category archives are frequently worth keeping and indexing, since they are a real route for a reader.

Paginated pages beyond the first should generally be left indexable but are rarely worth attention either way.

How to make the change

Practically, since this is a settings question rather than a rebuild.

Most platforms have a settings screen or a plugin controlling exactly which of these are generated and indexed.

Ask whoever maintains the site to set archives you do not want to be excluded from indexing, which is one option per type.

Do not block them from being crawled instead, since a blocked page cannot be read and therefore cannot be seen to be excluded.

That distinction catches people out and leaves the pages in the index indefinitely.

Expect the count to fall over weeks rather than immediately.

Prevent it recurring

Since tags in particular accumulate.

Agree that tags are only used where three or more pages share one, which prevents a tag per post.

Check the indexed count quarterly alongside your other figures, since a rising number with no new pages written is the signal.

And check it after installing anything, since plugins add generated pages of their own.

That is a one-line check on a number you should be watching anyway.

The same applies to a rebuild, since a new theme frequently reintroduces every archive type that was switched off on the old one.

Check the search results for your own site

A quicker version of the diagnosis that needs no tools at all.

Search for your domain with the site operator and page through what comes back, which shows roughly what is indexed in the order it is judged.

Generated pages appear as entries with no description, a title that is a tag name or a date, or a filename from a photograph.

Two minutes of that gives the same answer as the indexing report and is easier to show somebody who does not use search console.

It is also a reasonable check to run on a competitor, which occasionally explains a site that looks large and is mostly archives.

The counter-case

Some of these earn their place.

Category archives on a content-heavy site are genuine navigation and can rank usefully for broad terms.

Tag archives around a real topic with a dozen posts behind it are a legitimate page.

And on a small site with forty pages, none of this is the reason you are not ranking, so it should not be the first thing addressed.

Treat it as tidying that makes your own reports readable rather than as a fix for anything commercial.

Compare indexed against written, disable attachment pages, exclude date and author archives, and keep the categories.

The five minutes

  1. Check the indexed count.
  2. Compare against pages you wrote.
  3. Sort the list by address.
  4. Disable attachment pages.
  5. Exclude date and author archives.
  6. Prune single-use tags.
  7. Recheck the count quarterly.

Step one is the whole diagnosis, since a business that has never compared those two numbers has no idea how much of its site it did not write.

The other common cause of the same symptom is covered in two versions of your site in the index.


Frequently asked questions

Where do these pages come from?

Your platform generating an archive for every tag, category, month, and author, plus a page for every uploaded image, and paginated versions of all of them.

Which is the worst?

Attachment pages. Some platforms create a separate page per uploaded image containing that image and nothing else, so four hundred photographs means four hundred empty pages.

Why does it matter?

They compete with your real pages, consume crawling so genuine updates are found slower, and add a large volume of thin content to a site assessed as a whole.

How do I find out?

Compare the indexed count in your search console against the number of pages you believe you have, then sort the list by address to group the generated ones.

What should I do with each kind?

Disable attachment pages, exclude date and author archives from indexing, keep only tags grouping several useful pages, and generally keep category archives.

What is the common mistake?

Blocking them from being crawled rather than excluding them from indexing. A blocked page cannot be read, so it cannot be seen to be excluded, and stays in the index.

West Coast Media Solutions Inc. provides web design, web development, hosting, digital marketing, and business consulting to organisations across Canada, drawing on more than twenty-five years in the field.

Never compared indexed pages against pages you wrote?

Open the indexing report. The ratio is frequently three or four to one.

Start a Conversation