It reasons better, follows instructions more closely, and handles longer material. It still invents facts, still knows nothing recent, and still knows nothing about your business.

What was announced

OpenAI released GPT-4 today, a new version of the model behind the chatbot everybody has been discussing since December.

It is available through the paid tier of ChatGPT and, for developers, through a waiting list.

The free version continues to use the previous model, which is worth knowing, since most people describing their experience are describing that one.

So there are now two rather different things in circulation under the same name.

What genuinely improved

The second matters most for ordinary work. Asking for eight bullet points of under twelve words each in Canadian spelling now produces that, where before it produced something approximately like it.

What did not change

Which is the part worth dwelling on.

It still produces confident statements that are untrue, and OpenAI said so plainly in the announcement rather than leaving it to be discovered.

It still has no knowledge of events after its training period, so anything current is outside it.

It still knows nothing about your prices, your area, your customers, or your standards.

And it is still weighted toward American material, so the jurisdiction problem is unchanged.

Every checking habit described here over the past three months applies exactly as before.

Better makes verification harder

An uncomfortable consequence worth stating.

A tool that is wrong twenty per cent of the time in obvious ways trains you to check.

A tool that is wrong five per cent of the time in subtle ways trains you to trust it, and the five per cent is now harder to spot.

So the practical risk of an unchecked error may have gone up rather than down, even though the error rate went down.

Which means the discipline of extracting factual claims and checking each one matters more with the better version, not less.

That is the opposite of what most people will conclude this month.

What it is worth trying again

Since some things that failed before may now work.

Tasks with several constraints at once, where the previous version dropped one of them.

Longer documents that previously had to be broken up.

Anything requiring it to hold a structure across a long output.

And the questions from your own trade that it got wrong when you tested it, which is the calibration worth rerunning.

If you kept those five questions, this is the week to ask them again and see what moved.

A worked example

A business had tried using the earlier version to turn a long specification into a customer-facing summary and abandoned it, because the output kept losing detail and inventing structure.

They tried again with the new version on the same document.

The result held the structure, kept the detail, and required editing rather than rewriting.

It also contained two figures that were not in the source document, which the check caught.

Better and still requiring the same verification is the accurate summary of the week.

They now use it for that task and check every number.

The paid tier question

Which is the practical decision in front of most people reading this.

The better version costs a monthly subscription, and whether that is worth it depends entirely on how much you actually use it.

Somebody drafting content weekly will notice the difference quickly.

Somebody who has tried it twice out of curiosity will not.

Try the free version on a real task first, and pay only if you find yourself hitting its limits rather than because a better one exists.

The subscription is easy to start and, like every subscription, easy to keep paying for after you stop using it.

Diarise a look at it in three months, which is long enough to know whether it became part of how you work.

Keep it in proportion

Since the coverage this week will be considerable.

Nothing about your business changed today.

Your customers still choose you on price, availability, and whether they trust you, none of which this affects.

The competitive picture described a fortnight ago is unchanged: everybody has the same tool, so the advantage sits with whoever supplies the specifics it cannot.

A better tool for producing the generic middle makes the specific parts more valuable rather than less, since the middle just became cheaper for everybody at once.

It can also read images now

Worth mentioning since it will be widely discussed, with the caveat that it is announced rather than generally available.

The new version is described as able to accept a picture as part of a question, so you could show it a photograph or a diagram and ask about it.

For a trade that would be genuinely useful: a photograph of a part, a control panel, or a fault.

None of that is available to most people this week, and when it is, everything above about invented answers applies to it as well.

Treat it as something to look at later rather than a reason to change anything now.

The counter-case

This is a real step rather than an incremental one.

The improvement on complex reasoning is large enough that tasks previously not worth attempting now are, and dismissing it as more of the same would be wrong.

For anybody doing substantial writing or analysis, the difference is noticeable within an afternoon.

And the pace of change means anything concluded about these tools has a short shelf life, including this.

Retest your own five questions, keep every checking habit, and remember that it still knows nothing about your business.

This week

  1. Rerun your five trade questions.
  2. Note what improved and what did not.
  3. Retry a task you abandoned.
  4. Keep checking every factual claim.
  5. Assume American on anything legal.
  6. Try the free version before paying.
  7. Change nothing else yet.

Step four is the one that gets quietly dropped this month, since a tool that is right more often is a tool people check less often, which is exactly backwards.

Where this started is covered in everybody is talking about a chatbot.


Frequently asked questions

What was announced?

OpenAI released GPT-4, a new version of the model behind ChatGPT, available through the paid tier and a developer waiting list. The free version still uses the previous model.

What improved?

Reasoning through multi-step problems, following instructions precisely, handling longer material, and performance on exams. Instruction-following matters most for ordinary work.

What did not change?

It still states untrue things confidently, still knows nothing after its training period, still knows nothing about your business, and is still weighted toward American material.

Does better mean less checking?

No, the opposite. A tool wrong twenty per cent of the time in obvious ways trains you to check. One wrong five per cent of the time in subtle ways trains you to trust it.

What is worth trying again?

Tasks with several constraints at once, longer documents that had to be broken up, anything holding a structure across a long output, and the trade questions it got wrong before.

Should I pay for it?

Try the free version on a real task first and pay only if you hit its limits. Somebody drafting weekly will notice; somebody who has tried it twice will not.

West Coast Media Solutions Inc. provides web design, web development, hosting, digital marketing, and business consulting to organisations across Canada, drawing on more than twenty-five years in the field.

Tested it on your own trade in the autumn?

Ask the same five questions again this week and note what moved. That is the only comparison worth having.

Start a Conversation