Check whether text in your PDFs can be selected. If it cannot, the file is an image and is unreadable to assistive technology, search engines, and anybody who needs to enlarge it. Publish the content as a web page instead.

The thirty-second test

Open a PDF on your site and try to select a line of text with your cursor.

If you can highlight it, the file contains real text.

If nothing highlights, or the whole page selects as one block, it is an image of a document rather than a document.

That single check tells you which of your files are usable and which are, functionally, photographs that happen to be of words.

What a scanned file excludes

More people than the obvious group, which is why this is worth attention beyond an accessibility obligation.

The third and fourth are the ones that get a business's attention, and they are true regardless of anybody's view about obligations.

How these files get there

Never deliberately, and usually through one of a few routes.

Somebody scans a printed price list because that is the only copy anybody can find.

A supplier sends a specification sheet that was itself scanned.

A brochure is exported from design software with the text converted to outlines, which looks identical and contains no text.

Or a document is photographed on a phone and saved as a PDF, which is increasingly common and produces the worst version of all of these.

The better answer is usually a web page

Before fixing a PDF, ask whether it should be a PDF.

A price list, an opening hours notice, a service description, a set of frequently asked questions, or a specification table is all better as a page on your site.

A page can be read aloud, resized, searched, found, linked to a specific section, and updated without re-exporting anything.

PDFs earn their place where the layout genuinely matters and the document is meant to be printed or filed: a form somebody completes by hand, a certificate, a technical drawing, or something that has to look identical everywhere.

For everything else, the file exists because somebody had it in that format already, which is not a reason.

Fixing the ones that must stay

Where a PDF is genuinely the right format, the scanned version can be corrected.

Text recognition converts the image into selectable text, and it is built into most PDF software and available in free tools.

Run it, then check the result, because recognition makes predictable errors on numbers, part codes, and anything in a table, which are frequently the parts people need.

Beyond text recognition, a properly accessible document also needs a reading order, headings marked as headings, tables identified as tables, and descriptions on images.

That is real work, and it is another argument for publishing the content as a page instead, where all of that is easier.

A worked example

A supplier had a product catalogue on their site as a scanned PDF, twenty-eight pages, updated annually.

It was their most downloaded file and had never appeared in any search result.

The scan meant the entire catalogue was invisible: no part number, product name, or specification in it could be found by anybody searching, including customers using the site's own search.

They rebuilt the twelve most-referenced pages as ordinary web pages with tables, and kept the full catalogue as a PDF for customers who wanted to print it, this time exported properly with real text.

Search traffic to the new pages was substantial within four months, from part numbers that had existed on their site for six years without ever being findable.

The accessibility improvement was real and was not what convinced anybody internally.

Check what you already have

Most businesses do not know how many documents are on their site.

Files accumulate: an old price list, a form from 2017, a supplier sheet, a policy nobody has read, a newsletter.

A crawl of your site will list every linked document, and your server files will show ones that are still reachable but no longer linked from anywhere.

Open each one, do the selection test, and note which are images.

You will also find files that should not be public at all, which is a separate benefit of the same exercise.

Say what a link will do

A small courtesy that is frequently missed and costs nothing.

A link that opens a document should say so, and say how large it is: the specification sheet as a PDF, two megabytes.

On a phone, on mobile data, an unexpected fifteen megabyte download is a genuine cost to somebody, and a link that gives no warning is why people stop clicking them.

It also matters for anybody using assistive technology, since being moved into a different application without warning is disorienting.

Format and size, in the link text. It takes seconds and it is the difference between a considerate site and a careless one.

The counter-case

There are documents where none of this applies and effort would be wasted.

An archived document from a decade ago, kept for reference, downloaded twice a year, is not worth converting.

Legal and historical records frequently have to be preserved as they are, including their layout and any signatures, and altering them would defeat the purpose.

Technical drawings and plans are genuinely visual, and text recognition on a drawing produces nonsense.

In those cases the right approach is to leave the file, describe what it contains in the surrounding page, and offer an alternative format on request. That serves somebody who cannot use the file without pretending the file itself can be fixed.

The audit

  1. Crawl the site and list every linked document.
  2. Try to select text in each one.
  3. Ask whether it should be a page rather than a file.
  4. Rebuild the ones people actually use as pages.
  5. Run text recognition on files that must stay.
  6. Check the recognised numbers and part codes.
  7. Add format and size to every document link.

Step three removes more work than the other six combined, because most of these files were never meant to be documents.

Why print material does not transfer is covered in a print brochure is not a web page.


Frequently asked questions

How do I tell if a PDF is readable?

Try to select a line of text with your cursor. If nothing highlights, or the whole page selects as one block, it is an image of a document rather than a document.

Who is excluded by a scanned document?

Screen reader users receive nothing, anybody enlarging it gets a blurrier image, search engines and your own site search cannot read it, and nobody can copy a part number from it.

Should the content be a PDF at all?

Usually not. Price lists, service descriptions, and specification tables belong on a page. PDFs earn their place where layout matters and the document is meant to be printed or filed.

Can a scanned PDF be fixed?

Text recognition makes it selectable, and it makes predictable errors on numbers and part codes, so check those. A fully accessible document also needs reading order, headings, and table structure.

How do I find the documents on my site?

Crawl the site to list every linked file, and check server files for ones still reachable but no longer linked. You will usually find some that should not be public.

What should a document link say?

The format and the size. An unexpected fifteen megabyte download on mobile data is a real cost, and being moved into another application without warning is disorienting.

West Coast Media Solutions Inc. provides web design, web development, hosting, digital marketing, and business consulting to organisations across Canada, drawing on more than twenty-five years in the field.

Price list on your site as a PDF?

Try to select the text. If you cannot, nobody searching has ever found a single thing in it.

Start a Conversation