‘The Invisible 66’: why I spiked my accountancy AI white paper
I tried to turn 100 homepage audits and 540 AI answers into a white paper that would win me accountancy clients. Then the evidence broke the argument I intended to sell, so I spiked it. What I found instead is more useful if you are thinking of buying GEO services.
I’ve spent the past few days trying to kill my own white paper.
This wasn’t the plan.
I wanted to understand why some accountancy firms appeared in AI recommendations while others didn’t. If the research found something firms could improve, it might also give them a reason to call me about consulting or copywriting.
Instead, I got a strong title and evidence that didn’t support my argument.
The findings
I examined every firm in the Accountancy Age Top 50+50 and collected 540 answers to 36 buyer questions from Claude, Perplexity and OpenAI.
- The larger firms appeared far more often. Clearer, more specific homepage titles made no detectable difference.
- The study found no evidence of a structured-data advantage. Seventy-one per cent of firms with JSON-LD appeared at least once, compared with 76 per cent of those without it. That does not mean structured data was harmful. The second group was small and skewed towards larger firms.
- The three assistants produced markedly different shortlists. For an average question, only 17 per cent of the firms named by any assistant appeared on all three.
- A firm’s name in an answer did not necessarily mean it was being recommended. Some firms appeared in passages explaining why the buyer might not need them.
- If somebody wants to charge a premium for schema or on-page changes because they will make an AI recommend your firm, ask to see the evidence. This study found none.
I started with 100 firms and 36 questions
I focused on the Accountancy Age Top 50+50, an annual ranking of 100 UK accountancy firms by fee income.
My first job was to examine how every firm presented itself online. I manually captured the code behind every homepage, then recorded its page title, meta description, main heading, positioning language and structured data.
Structured data, often called schema, is code that describes an organisation, its locations and its services in a format machines can read. My initial theory was that clearer pages and better structured information might make a firm easier for AI assistants to understand and recommend.
I then wrote 36 questions that a prospective client might ask. They covered mainstream accountancy work, specialist services and locations around the UK. One asked which firms a £2 million owner-managed business should consider. Others concerned R&D tax, selling a business or finding an accountant in a particular region.
The assistants were not shown the Accountancy Age list, and it was never mentioned in a question. They were simply asked to suggest suitable firms.
I sent each question to Claude and Perplexity five times through their APIs. This allowed me to use the same wording and settings each time, then save every answer.
Why five times? Simply because AI assistants can answer the same question differently from one query to the next. Asking five times helped me separate recurring appearances from the one-offs.
This produced 360 answers: 36 questions, multiplied by two assistants, multiplied by five runs.
Most people use public interfaces, not APIs. I therefore ran five questions once through each of Claude, ChatGPT, Google AI Mode and Perplexity. This small spot check did not overturn the broad picture, although the firms named varied considerably between interfaces.
To count the results, I compiled accepted versions of each firm’s name. This allowed a shorter trading name, for example, to be matched to its full name. I checked uncertain matches by reading the surrounding answer.
The first explanation was convenient
The first matching run said that 45 of the firms ranked 1 to 50 appeared at least once. Only 17 of those ranked 51 to 100 did.
The homepage review had also found plenty of generic language. Firms frequently described themselves as trusted, commercial, proactive and partner-led. Some page titles did little more than state the firm’s name.
When I captured these homepages in August 2026, the difference between generic and specific titles was easy to see. Grant Thornton UK’s title was simply “Grant Thornton UK”; Crowe UK’s was “Home | Crowe UK”. CBTax used “Your Global R&D Tax Specialists | CBTax”. Dow Schofield Watts used “Corporate Finance – Business Advisory – Due Diligence | Dow Schofield Watts”.
It was tempting to connect the results. Perhaps smaller firms were missing because their websites didn’t explain clearly enough who they helped, what they did, or why a client should choose them? Perhaps better titles, clearer positioning and structured data made a firm easier to recommend.
That became the argument behind a white paper I planned to call The Invisible 66.
The title referred to the 66 per cent of firms in the lower half of the ranking that did not appear in the AI answers. It sounded like a finding about the whole market. In reality, it was a percentage from one half of one study.
I commissioned an AI to break the argument
Before writing the white paper, I commissioned an adversarial review of the complete evidence pack. I gave a fresh Codex task the raw responses, homepage data, matching rules and calculations. The instruction was not to proofread. It was to rebuild the numbers and challenge every conclusion I had reached.
Most of the arithmetic survived. My argument didn’t.
Firm size was much more closely associated with AI mentions than the specificity of a firm’s homepage title. Large, famous firms often used generic titles because their names already carried meaning. Smaller firms tended to explain more, but they also had less market recognition and appeared in fewer results.
The title groups therefore contained different kinds of firms. When firms were compared within the two halves of the ranking, the study could detect no useful title effect in either direction.
Structured data did no better. I found JSON-LD, the most common format for structured data, in 83 captured homepages. Of those firms, 59 appeared at least once in queries (71 per cent). Among the 17 without it, 13 appeared (76 per cent).
That doesn’t mean structured data reduced visibility. The group without it contained only 17 firms and a disproportionate number of large names. It means the study found no evidence that structured data in and of itself improved visibility.
The review found ordinary errors too. Two apparent Andersen recommendations were references to Arthur Andersen, the historical firm, rather than the Andersen business in the ranking. A stability count was 178 when it should have been 174 because rejected matches had not been removed. A check for JSON-LD had been described as a check for all structured data.
At the end of the review, all I could safely say was that larger firms appeared more often in AI recommendations. The research didn’t show that changing a homepage title or adding structured data led to more mentions.
AI made the unsupported version sound finished
I used AI throughout the project. It helped refine questions, process response files, inspect uncertain names, check calculations and draft different versions of the paper. It was also used to challenge the work.
The problem was that it could turn several plausible observations into a smooth general explanation before the evidence had earned one.
One draft organised the results into five ways an “evidence chain” could fail, followed by a six-step improvement programme. The framework made sense. Much of the advice was probably sound.
But “this is sensible advice” and “my research has demonstrated this” are different claims.
That judgement remained mine.
What I found after the simple explanation failed
I expanded the study rather than abandoning it. I sent the same 36 questions to OpenAI five times through its API, with web search switched on. This added 180 answers and took the main dataset to 540.
The extra answers didn’t rescue the white paper. Instead, they showed me why a single visibility score was the wrong thing to build.
Different assistants produced different shortlists
After the final matching corrections, 72 of the 100 firms appeared somewhere. Thirty-nine appeared on all three assistants. Nineteen appeared on one assistant alone.
That left 28 firms which did not appear at all. This was the honest replacement for The Invisible 66.
For an average question, just 17 per cent of the firms named by any assistant appeared on all three. The same buyer could therefore receive a different competitive set depending on which assistant answered the question.
There was no single, stable AI ranking for an accountancy firm.
Even that analysis contained a trap. The reported overlap between Claude and OpenAI was initially 39 per cent. It should have been 36 per cent. On one question, neither assistant named a firm. The code treated the two empty lists as perfect agreement.
Names gave me another problem
The final review recovered 11 appearances missed because assistants used previous or alternative trading names – five for CT under “Chiene + Tait”, four for Cottons Group under “Cottons Chartered Accountants” and two for AMS Group. One further apparent appearance turned out to be the firm’s web address in the source list rather than its name in the answer, so I didn’t count it. In other cases, a subsidiary appeared more often than the firm ranked by Accountancy Age.
A group could therefore look almost invisible while one of its businesses was being recommended. A name count might be technically accurate and still give the wrong commercial picture.
Being named was not the same as being recommended
I later ran a smaller test focused solely on audit appointments. I asked OpenAI two questions five times each on 18 August.
The Big Four appeared in all ten answers. A simple name count would give each firm a perfect result. But they were often mentioned in passages explaining why a £50 million private company might not need them.
Those answers contained 88 visible citations. Seventy-six came from the Financial Reporting Council, five from GOV.UK, three from ICAEW, three from Forvis Mazars and one from Accountancy Age.
Most of the cited evidence therefore came from regulatory and public sources, not firms’ sales pages.
This was a small test, so I cannot assume every question or AI assistant would behave the same way. But it showed why a website-only audit can mislead. A firm may be named because the assistant is warning the buyer against using it. The assistant may also base that judgement mainly on guidance from regulators, not anything published on the firm’s website.
The factors I did not measure matter
A homepage is only one part of the public evidence an assistant may use. I did not measure enough of that wider evidence across 100 firms to isolate why one firm appeared and another did not. Nor could I honestly prescribe one technical fix.
This matters if somebody is trying to sell you “generative engine optimisation”, or GEO, based on a schema check, a few page edits and a handful of prompts.
My research does not show that schema or on-page work is worthless. Clear pages still help people understand a firm, and structured data helps machines interpret information about it.
What the research found was no evidence that those changes, by themselves, caused firms to appear more often in AI recommendations.
That is a poor basis for an expensive promise.
Six questions to ask before buying GEO
If a supplier promises to improve your firm’s AI visibility, I recommend that you ask:
- Which buyer questions will you test, and why do they matter commercially?
- How often will you repeat them, across which assistants and dates?
- Will you read each mention in context, rather than simply count names?
- How will you handle parent firms, subsidiaries, former names and acquired brands?
- What evidence beyond the firm’s own website will you examine?
- How will you show that any improvement resulted from the work, rather than normal variation between answers?
These questions force the supplier to explain what is being measured before selling you the remedy.
What I can offer
The assistants built these answers mostly from evidence that wasn’t on the firms’ own websites. Some firms were named only for the assistant to explain why a buyer didn’t need them. Neither of those is a technical fault. Both are questions about what a firm can prove, and where.
I audit content and I write it. For firms in the Accountancy Age Top 50+50, I can show what this study found about yours: which questions produced your name, who appeared instead, and which sources the assistants used. For any other firm, I can run the same test for your market. Just ask me.
The useful work is what follows from that – establishing which claims you can evidence, which you can’t, and what to publish about it.
I won’t promise you a place in tomorrow’s answer. Nobody can.
A credible study would compare similar firms, examine evidence beyond their websites, make controlled changes and repeat the tests across several assistants and dates. Mine didn’t. It produced useful observations, but no proof that changing a website caused an AI recommendation.
Research that is required to produce a marketable conclusion is not research. It is content production with an unusually expensive preamble.
So I spiked the paper.
Know someone who wants to buy GEO?
Send them this article first