Skip to content

I blocked two accountancy listicles from Claude. Its recommendations changed

I found a dramatic pattern in 540 AI answers, then discovered my comparison was invalid. So I blocked two accountancy listicles from Claude and collected 200 answers in a randomised experiment. Recommendations for three competitors fell sharply. Neither publisher was named.

This experiment began with a result I wanted to – but couldn’t – defend. I found a dramatic pattern in 540 AI answers. At first glance, it appeared that content from several small accountancy firms supplied the sources behind the AI recommendations, while larger competitors collected the mentions.

Then I discovered that my comparison was invalid. So I started again, blocked two individual articles from Claude’s web searches and asked the same buyer questions 200 times.

This time, the recommendations really did change.

What I tested and what happened

I collected 200 answers through the Claude API. They were divided into two tests:

Within each test, the question and settings stayed the same. Every result below therefore compares 50 answers with the article available against 50 with it blocked.

  • Saffery: 42 to 3. Claude recommended Saffery in 42 of 50 charity-audit answers when the Acumon article was available. When I blocked the article, Saffery appeared in only three.
  • Dixon Wilson: 14 to 0. Claude recommended Dixon Wilson in 14 of 50 private-client answers when the Nephos article was available. It wasn’t recommended once when I blocked the article.
  • Rawlinson & Hunter: 15 to 1. Rawlinson & Hunter appeared in 15 of 50 answers when the Nephos article was available, compared with one when it was blocked.
  • The publishers: 0. Acumon and Nephos weren’t named at all in any of the 200 answers. That includes the 63 answers that cited one of their articles.

The effect wasn’t the same for every firm. Price Bailey appeared in 38 of 50 answers when the Acumon article was available, but in 45 when it was blocked.

I therefore can’t claim that appearing in a listicle makes every named firm more likely to be recommended. 

The narrower finding is still striking: in this test, changing whether Claude could use two articles substantially changed its recommendations for three firms, while the firms that published the articles received no recommendations at all.

This started with a white paper I abandoned

My original project concerned the Accountancy Age Top 50+50, an annual ranking of the UK’s largest accountancy firms by fee income.

I captured the code from behind all 100 homepages, then recorded their titles, descriptions, positioning language and structured data. Structured data is code that helps machines identify what an organisation or page represents.

I also wrote 36 questions that a possible accountancy client might ask. They covered audit, tax, corporate finance, locations and different types of business. Each question went to Claude, Perplexity and OpenAI five times, producing 540 answers.

My theory was that firms with clearer websites and better structured information would appear more often. The evidence didn’t support it.

Large and well-known firms appeared much more frequently, regardless of whether their homepage titles were specific or vague. When I compared firms of similar size, I couldn’t detect a useful relationship between homepage titles and AI recommendations.

I’d built the research for a white paper about improving AI visibility, but it hadn’t shown that the improvements I wanted to discuss caused anything. So I spiked the white paper.

The failed project left me with a better question

Although the original argument failed, I still had 540 answers and their cited sources. One pattern stood out.

Acumon’s website appeared among the sources in 92 answers, while Acumon itself was named in two. Whiz Consulting’s site appeared in 21 answers, but Whiz wasn’t named at all.

Rise Accounting and Nephos showed a similar gap between their websites being used and the firms themselves being mentioned. The pages attracting attention weren’t conventional service pages. They were lists of other accountancy firms:

These firms had followed familiar content-marketing advice: answer a popular search question, publish a useful comparison and put your own firm at the top.

The pages were being found. The awkward possibility was that they were helping AI assistants recommend everybody else.

The first version of this article was wrong

I initially compared answers that cited one of these listicles with all the other answers in the study. The numbers looked spectacular.

Where the Nephos article was cited, Dixon Wilson and Rawlinson & Hunter appeared in 54 per cent of answers. Across the other 512 answers, they appeared in almost none.

The Acumon results looked equally impressive. Firms described on its charity-audit page appeared far more often in answers citing the page than in the rest of the study.

I had a strong line:

The list works. Just not for the publisher.

Unfortunately, it was better copy than research.

The other answers concerned entirely different problems. They included questions about cryptocurrency, R&D tax, landlords, selling businesses and finding accountants in different regions.

A private-client tax specialist appearing more often in private-client tax answers than in cryptocurrency answers tells us very little about the effect of a Nephos article.

The Acumon comparison was worse. Its charity page appeared in all 20 relevant Claude and Perplexity answers, but none of the ten OpenAI answers. Whether the page was cited and which engine produced the answer were effectively the same thing, so I was comparing assistants as much as pages.

The evidence still showed that small firms’ articles were supplying information in answers that recommended larger competitors. It didn’t show that the articles had caused those recommendations.

I’d allowed a bigger denominator to produce a more impressive result. It did, but it also made the comparison invalid.

One small signal survived

Within Claude alone, the original study contained ten answers to the two private-client questions.

The Nephos page appeared among the sources in seven. Dixon Wilson and Rawlinson & Hunter were each recommended in five of those seven answers, but neither appeared in the other three.

This was a comparison within one assistant and one tightly defined subject. It was still only ten answers, and Claude had chosen both the sources and the firms. It showed an association worth investigating, not that the Nephos article had caused anything.

I changed one thing

I used two articles from the original research:

  • The Acumon guide to non-profit auditors
  • The Nephos guide to UK tax advisers.

The Acumon article places Acumon first, then profiles firms including Price Bailey, Saffery and Grant Thornton. The Nephos article ranks Nephos first, followed by firms including Dixon Wilson and Rawlinson & Hunter.

I used one buyer question for each page:

Which UK accountancy firms would you recommend for charity and not-for-profit audit? Please name specific firms and say why.

And:

I have substantial personal wealth and need advisers for tax planning and estate matters in the UK. Which firms are good?

Each question went to Claude Sonnet 5 one hundred times through Anthropic’s API on 19 August 2026.

An API lets software send the same question directly to an AI service and save the complete response. It meant I could keep the wording, model and search settings unchanged.

For 50 answers, Claude searched the web normally. For the other 50, I used Anthropic’s web-search controls to block only the relevant listicle. The rest of the publisher’s website and the wider web remained available, and Claude wasn’t told that anything had been removed.

The two conditions were mixed into a random order, which prevented all the available-page answers being run first and all the blocked answers later. The only deliberate difference was whether Claude could find that exact article.

The order came from a fixed random seed recorded before the experiment. The independent reviewer later regenerated the schedule and obtained the same file, byte for byte, confirming that the stored order matched the method I had specified.

I chose the cases and firms before the main test

Before the controlled experiment, I ran a 60-answer pilot covering six combinations of articles and buyer questions.

I set the selection rule before seeing the results. A combination could proceed only if Claude cited the target article between two and eight times in ten answers. If the page appeared every time, there wouldn’t be a useful natural comparison. If it never appeared, blocking it would be unlikely to tell me anything.

Only two combinations qualified. The Acumon article appeared in eight of ten pilot answers to the charity-audit question. The Nephos article appeared in six of ten answers to the private-client question. The other four combinations stopped under the original rule.

I then ran a separate 100-answer observational study: 50 fresh answers for each of the two qualifying combinations. This helped me choose the four competitor results I would measure in the experiment. Price Bailey, Saffery, Dixon Wilson and Rawlinson & Hunter showed the clearest commercially relevant relationships between the target article being cited and the firm appearing in the answer.

I recorded those four firms before running the controlled experiment. The 60 pilot answers and 100 observational answers were kept separate from the 200 experimental answers and weren’t used in the final comparisons.

These were deliberately selected cases in which the articles had a realistic chance of mattering. They weren’t a random sample of accountancy content.

The Acumon page changed Saffery’s result

When the Acumon article was available, Claude recommended Saffery in 42 of 50 answers. When I blocked it, Saffery appeared in three.

That is a difference of 78 percentage points. The statistical range around the result still leaves a large effect, approximately 62 to 87 points.

Price Bailey didn’t follow the same pattern. It appeared in 38 answers with the page and 45 without it. That difference could reasonably be chance variation.

Acumon itself wasn’t named in either group. The article therefore produced a striking difference for one competitor selected before the results were opened, but not the other. It didn’t secure a place for its publisher.

The Nephos page changed two results

The Nephos article produced a similar result for two firms selected in advance.

With the page available:

  • Dixon Wilson appeared in 14 of 50 answers;
  • Rawlinson & Hunter appeared in 15.

With it blocked:

  • Dixon Wilson appeared in none;
  • Rawlinson & Hunter appeared in one.

Nephos didn’t appear in either group.

The differences for both competitors were 28 percentage points. The statistical ranges were approximately 15 to 42 points for Dixon Wilson and 14 to 42 for Rawlinson & Hunter.

The recommendation check exposed another error

The original Accountancy Age research taught me not to treat every firm name as a recommendation. An AI might mention a large firm to explain why it’s unsuitable. A name might appear only in a citation title, or refer to a different business with a similar name.

An automated screen found 158 visible-name appearances across the 600 primary firm-and-answer combinations. My first evidence pack labelled every one as a recommendation and claimed that I had read each appearance in context.

That wasn’t true.

The script had automatically converted every visible-name hit into a recommendation without recording an individual judgement. The independent reviewer found the problem by inspecting the code.

It then did the missing work. Without knowing whether each article had been available or blocked, it relabelled all 600 combinations and read every one of the 158 visible-name contexts. All 158 were genuine recommendations. None was a neutral mention, an exclusion or a name appearing only in a citation.

The results survived unchanged. My description of how I had produced them didn’t.

In these particular answers, visible names and recommendations happened to be the same because Claude was producing shortlists. That doesn’t make automated name counting a safe substitute for reading answers in context.

The review also found one response in the blocked Nephos group that had stopped at the maximum answer length. It hadn’t named either preselected competitor before it ended.

I kept it in the results. Replacing an inconvenient answer after opening the data would have caused a larger problem than retaining and disclosing the truncation.

The page didn’t behave like a simple vote

The tempting explanation is that Claude read a list of firms and repeated the names. The evidence is more interesting.

Blocking the Acumon page changed the other sources Claude found. Answers citing Azets’ website rose from six in 50 to 42 in 50. Price Bailey’s site rose from 36 to 45, while Lovewell Blake and PKF Francis Clark also appeared more often.

Claude hadn’t removed one page from a fixed reading list. It had assembled a different reading list.

A separate Acumon article about non-profit auditors in Bristol remained available. It also profiles Saffery and appeared among the search results for 17 of the 50 blocked answers, yet Saffery recommendations still fell from 42 to three.

The presence of Saffery’s name somewhere on the same website wasn’t enough. The particular page and the other sources assembled around it mattered.

Something similar happened in the Nephos test. When the tax-adviser article was blocked, Claude cited a different Nephos article about wealth-management companies in 20 of 50 answers. That replacement page doesn’t profile Dixon Wilson or Rawlinson & Hunter.

Removing one page changed the evidence Claude used to build its answer. It didn’t simply subtract one mention from a score.

What the experiment does and doesn’t show

Within these 200 API answers on 19 August 2026, blocking each listicle changed Claude’s recommendation rate for at least one competitor chosen before the results were opened.

The question and model settings remained the same, while page availability was assigned at random. That makes this a causal result within the experiment rather than another correlation between citations and recommendations.

It’s still a narrow result. I tested two articles, two buyer questions and one Claude model through an API on one afternoon.

I didn’t test ChatGPT, Perplexity, Gemini, Google’s AI results or Claude’s public interface. Nor did I measure enquiries, appointments or fee income, so I can’t say that either publisher lost work.

I also can’t claim that “best firms” articles generally help competitors. Price Bailey moved in the opposite direction, and four pilot cases never reached the controlled stage.

The result isn’t a universal warning against comparison content. It’s evidence that one such page can alter an AI shortlist in ways the publisher may not expect.

What accountancy firms should do before publishing a listicle

You don’t need to delete every article that names a competitor, but it might be worth looking at them afresh.

Before commissioning another “best accountancy firms” page, ask:

  • Which buyer question is this page supposed to answer?
  • Is my firm genuinely a strong answer to that question?
  • Does the page make a specific, evidenced case for choosing us?
  • Am I giving a competitor a clearer and more credible profile than I give myself?
  • Have I tested the exact buyer question repeatedly?
  • Have I checked what happens across several assistants and dates?
  • Do I know which sources appear when my page doesn’t?

One screenshot won’t answer those questions, and neither will a generic “AI visibility score”.

What this means for your content

This experiment doesn’t give accountancy firms a formula for getting into AI answers, and I don’t sell one. My work starts with more useful questions: what does your firm need to say, who needs to hear it, what evidence would make it credible and where should it appear?

I help accountancy and other professional services firms answer those questions, then turn the answers into content. That might begin with a proper content audit or clearer website copy. The aim is to make the firm easier for serious buyers to understand and trust, while giving search engines and AI assistants better public material to work with.

If your firm knows more than its website says, or has years of content without a clear purpose, tell me where you’re stuck. I’ll help you work out what to keep, what to change and what to write next.

Know someone who should read this?

Send it to them now

Share on LinkedIn
Ben Locker

The writer

Ben Locker

Ben Locker is a content auditor and copywriter for specialist businesses. He works for professional, technical and specialist businesses across the UK and Europe.

Back to the blog