In the same week, USA Today published a roundup of top SEO and AEO tools for building visibility across search and AI, and a tech outlet published a piece announcing that they had tested the eleven best AI SEO tools of 2026.
I want to be careful here, because there is a sneering version of this observation and it is not the one I am making. I run a tools directory. I have spent an unreasonable number of hours this year clicking through SEO software so that other people do not have to. I am not against writing about tools. I am against a specific word, and the word is "tested."
Tested Against What
Here is the problem, stated as plainly as I can.
To test a tool you need a ground truth. You need to know what the right answer is, independently of the tool, so you can check whether the tool got there. When somebody tests a rank tracker, this is easy: you can go and look at the search results yourself and see whether the tool reported them correctly. When somebody tests a crawler, this is easy: the site has a finite number of URLs and either the crawler found them or it did not.
Now tell me the ground truth for an AI visibility tool.
The tool tells you that your brand appears in some percentage of answers for a set of prompts. Appears where? Generated when? By which model version, at what temperature, with what retrieval configuration, for a user with what history, in which country, on which of that vendor's several different surfaces? Ask the same model the same question twice and you may get two different answers with two different citations. Ask it next Tuesday, after a silent model update, and you may get a third. There is no stable object to be correct about.
So what is being measured when someone reports that a tool is accurate? Mostly, whether the dashboard loaded. I am only slightly exaggerating. In the absence of ground truth, a review of these tools collapses into a review of the user interface, the pricing, and how confident the marketing copy sounds. Those are real things and it is fine to write about them. They are just not what "we tested" is being used to imply.
I have made this argument at more length about the whole GEO category being premature optimization, and the tool roundups are the retail end of that same problem. You cannot build a reliable instrument for a phenomenon nobody has managed to define stably. You can absolutely build a product that sells to people who want the phenomenon to be measurable, which is a different and much easier engineering problem.
Why the National Papers Are Doing This
The presence of a general-interest national daily in this category is not a mystery and it is not a scandal. It is affiliate revenue.
Software roundups in high-value B2B categories monetise extremely well. A reader who clicks through to a trial of a marketing tool is worth many multiples of a reader who clicks an ad next to a news story. Newsrooms have watched display advertising collapse for fifteen years, and commerce content is one of the few things that reliably pays. So the incentive is straightforward, the desk producing it is usually structurally separate from the newsroom, and the resulting article ranks well because the domain is enormous and Google, for its own reasons, has spent recent years rewarding large established publishers in exactly these commercial query spaces.
Which produces the actual outcome: the highest-ranking evaluation of a category is written by the party with the least domain expertise and the most direct financial interest in you clicking through. That is not a conspiracy. It is just what the incentives build when nobody is steering.
The Shoeshine Signal
There is an old market story, probably apocryphal, about a financier deciding to sell everything after a shoeshine boy offered him stock tips, on the theory that when the last uninformed participant arrives the move is over. I have never fully trusted the story, partly because it is a bit snobbish about shoeshine boys, but the underlying mechanism is real: categories have a phase where the marketing runs so far ahead of the substance that the marketing itself becomes the signal.
An AEO tool listicle in a national newspaper is that signal. It does not mean the underlying technology is fake. Answer engines are real, people genuinely use them, and being cited in them will genuinely matter. It means the tooling layer has been financialised well ahead of the measurement layer, and money is now moving on the strength of a promise rather than a result.
I have watched this exact sequence before, with social media analytics around 2012, and with a stack of attribution products a few years after. Same shape every time. The tools that survive are the two or three that eventually build the boring infrastructure. The rest are a dashboard over somebody else's API, sold on urgency, gone within three years, having quietly taken a lot of retainers with them. When I look closely at this current crop, an uncomfortable number are a scheduled prompt runner, a database, and a chart, and there is nothing wrong with that except the price.
What a Real Evaluation Would Require
If somebody wanted to genuinely evaluate this category, here is roughly what it would take, and the reason nobody does it is that it is expensive and unexciting.
You would fix a prompt set and freeze it. You would run it across every tool and, independently, by hand, on the same day, in the same market. You would repeat that on a schedule for at least a quarter, so you could measure not accuracy but stability, which is the property that actually matters and the one nobody reports. You would check whether two tools looking at the same model on the same day agree with each other, and my strong suspicion, based on informal poking, is that frequently they do not. You would disclose every affiliate relationship. And you would publish the disagreements rather than a ranking, because the disagreements are the finding.
That is a research project, not a listicle. It would not rank for "best AI SEO tools" and it would not pay for itself.
Until somebody does it, treat the category the way you would treat any market with no audited numbers. Buy the cheapest thing that answers a question you actually have. Cancel it the moment it stops answering that question. And read the roundups for what they are.
A roundup is not evidence. It is inventory with a byline.