Three stacked folders of client work on a desk beside a laptop showing an AI chat window.
AI adoptionLeadershipAI tools

Fable 5 is free on paid Claude plans until 19 July. Test it properly

Fable 5 is free on paid Claude plans until 19 July. Here is a half-day trial on real client work to decide whether the top AI model is worth paying for.

Good Transformer8 min read

Anthropic's Fable 5 is the most capable AI model the public can use, and until Sunday 19 July it is included in paid Claude plans at no extra cost.

We think it is powerful enough that a leader should take half a day this week and run it on real client work, to see whether it makes a genuine difference to what your firm can do.

This piece sets out how: pick three pieces of work, brief the model properly, judge the results as a client would, and finish with a written decision on whether the top model is worth paying for once the free window closes.

If your firm pays for Claude Pro, Max, Team or premium Enterprise seats, Fable 5 is already in your account. From 20 July, using it costs extra. An idle poke at a chatbot will not tell you whether it is worth that money. A structured trial will.

Update, 13 July 2026: Anthropic has twice extended the free window since this piece was published, most recently to 19 July, with metered usage credits from 20 July. The dates in this piece have been updated to match. The trial works exactly the same, and there are more free days left than the original deadline suggested.

Why this model deserves the time

Fable 5 has had the strangest arrival of any AI model to date. It launched on 9 June, was switched off worldwide on 12 June under a US export-control directive, and returned on 1 July with new safeguards once the restriction was lifted. Anthropic's account of the redeployment includes the detail that matters here: the model is free on paid plans for a limited window, then moves to metered usage credits. Anthropic has since extended that window twice, currently to 19 July.

Its power is not a marketing claim. The best measure of AI doing real work is the Remote Labor Index, run by the Center for AI Safety with Scale AI. The test is easy to describe. AI models are given 240 real freelance projects, work that clients commissioned and paid for, across 23 fields. A model scores only when the finished work is good enough that the client would accept it. Fable 5 now leads every public model at 16.1 per cent, roughly double the next-best model in the same round of testing.

Two things follow from that score, and both matter to a leader. First, the best AI in the world still completes only about one real project in six to a standard a client would pay for, so it is nowhere near replacing professional work wholesale. Second, the leading score is roughly four times the best published result from the round before, so tasks that were safely out of reach last year may not be out of reach now. A benchmark cannot tell you where your firm's work sits between those two facts. Half a day with your own tasks can.

Pick three pieces of real work

A trial is only as good as the work you feed it, and demo tasks teach you nothing. Choose three real pieces of work from your own firm.

1. A routine deliverable your firm produces every week. Think of the management-accounts commentary, a candidate shortlist summary for a client, or a first-pass mark-up of a standard contract. Routine work has a known standard, and you can put the output next to what your team actually produced last time and compare line by line.

2. A judgement-heavy deliverable you would normally give to a senior person. That might be an advisory letter on an awkward client question, the strategy section of a pitch, or a due-diligence summary that has to end in a recommendation. This is where the difference between the top model and the standard ones is claimed to live, so this is where the claim gets tested.

3. One task you privately believe AI cannot do. This is the most useful of the three. If the model fails it the way you expected, you have seen the ceiling with your own eyes. If it does not fail, you have learned something worth far more than the half-day it cost you.

The usual confidentiality rules apply: strip client identities from anything sensitive before it goes in, exactly as your existing AI policy already requires for any tool handling client material.

Brief it like an outside contractor

Most disappointing AI output is the result of a lazy brief. You would not give a new contractor one sentence of instructions and expect a usable deliverable back, so do not do it here. Give the model what you would give a capable outsider on their first day: the background, the source documents, the house style, the standard you expect, and a plain description of what a finished piece looks like. Then let it work. Fable 5 can plan, use the material and run a long task on its own; the brief is what points that effort at your standard rather than a generic one.

Have your most experienced people run the trial, for the same reason you would not send a trainee to assess a lateral hire. Experience is what tells you, quickly and with confidence, whether an output is genuinely good or merely fluent.

Judge it as a paying client would

Borrow the benchmark's scoring rule: an output passes only if you would send it to a paying client with your firm's name on it. An output that "just needs a polish" before it could go out is a fail. A confident paragraph resting on an invented figure is a fail. Hold that line, because a generous marker learns nothing.

Write down where each task fell short, because the failures are the most valuable thing the trial produces. A model that writes a fluent management-accounts commentary but invents one variance explanation has shown you exactly where the boundary sits: drafting can be delegated, unchecked analysis cannot. A contract mark-up that catches nine issues in ten has shown you it is a first-pass tool that needs a senior review behind it, and that is still worth real money. That written record of failures, made on your own work, is the most commercially useful thing this free window can produce.

Count your own time honestly too. If checking and correcting an output took longer than doing the work the old way, that task fails the trial no matter how good the prose looked. We published a short method for testing whether an AI use pays for itself, and the same arithmetic applies at this tier.

Decide before the meter starts

On 19 July the free window closes. End the trial with a decision in writing, however short. Three outcomes cover most firms.

1. The upgrade case is proven. The top model did something your current tier demonstrably cannot, on work that matters commercially. Budget for it deliberately, for the specific people and tasks where the difference showed, and nowhere else.

2. Your current tier is enough. The gap was small on your actual work. That is a happy result: you now have evidence, with a date on it, to resist upgrade pressure, and a record of what the top end could and could not do in your hands.

3. The capability is there but your firm is not ready to use it. That is common, and useful to know. The bottleneck is briefing, checking and process, and none of those is solved by a bigger model.

One caution belongs in the file next to the decision. The model you have been testing was unavailable worldwide for almost three weeks because a government directive required its withdrawal. Whatever the trial shows, do not let any single vendor's top model become something your firm cannot work without; the full argument, and the fallback plan, are in our note on vendor continuity.

Running exactly this kind of structured trial on a leader's own work, with someone experienced alongside, is what our 1-to-1 AI lessons for leaders are built around.

Questions leaders ask

Is Fable 5 actually better than the model a firm already uses?

On the benchmarks, clearly yes: its Remote Labor Index score is roughly double the next-best model's. Whether it is better on your work is what the trial exists to establish. The difference tends to show most on long, judgement-heavy tasks with a lot of source material, and least on short everyday drafting, which the standard tiers already handle well.

What does it cost after the free window?

From 20 July, Fable 5 use on paid Claude plans is metered through usage credits rather than included in the subscription. The practical point is simple: once the window closes every experiment has a price, which is why the free days deserve deliberate use.

Should the whole team get access during the window?

We would keep the trial small and senior. The free access draws on your existing usage limits, and the goal is a firm-level judgement, which needs experienced eyes more than volume. Share the written findings with the team afterwards; wider access can follow the decision rather than precede it.

What if the window closes before the trial is finished?

A half-day is enough for three tasks if the work is chosen in advance. If you do run out of time, the same trial still works at the metered price, and the cost of a few experiments is small against the cost of a wrong platform decision. The window makes the trial free this week; it does not make it impossible next week.

The next step takes ten minutes: choose the three pieces of work today, book the half-day before the window closes on 19 July, and decide which senior person runs it. If you would like help turning what the trial teaches into a working plan for your firm, book a discovery call and we will go through it with you.

Work with Good Transformer

Turn this thinking into working practice.

Explore team advisory

Newsletter

Get new Insights by email

Practical notes on using AI with judgement, and the AI news leaders actually need. No hype, no spam, unsubscribe anytime.

Choose how often you want the digest

Keep reading