Seven things a comparison site publishes that you can check from outside
You are not in the room and cannot audit what was tested. These come from the methodologies of Wirecutter, Consumer Reports, G2 and Tom’s Hardware, and I graded my own deleted benchmark against them.

On this page
The reasonable objection to an article like this one goes: you cannot audit a comparison site from outside. You are not in the room. You do not know what was tested, who paid for what, or which vendor called which editor. Absent that, a disclosure page and an affiliate notice are about all anyone can ask for, and a reader is better off just judging the verdicts on their merits.
The first half of that is true. The conclusion does not follow.
You cannot audit the verdicts. What you can audit (from outside, without access, in the time it takes to open four tabs) is the process, and you can do it wherever the publisher has committed to enough specifics that a mismatch between what they said they do and what the pages show them doing would be visible to a stranger. That is the whole game. A small number of organisations play it well enough to steal a standard from.
Four publish enough detail to reverse-engineer one: Wirecutter, Consumer Reports, G2 and Tom’s Hardware. Below is the checklist I pulled out of their documents. Each check comes with the way to run it on somebody else’s site, and then with its result on a comparison benchmark of my own, which ranked 19 online registration products and which I deleted in August 2026. Running them together is deliberate. A checklist nobody has ever failed in public is a wish list.
I am grading myself with someone else’s ruler. That is why the confidence label on this page reads medium and not high: the checks are documented, the weighting I put on them is my opinion, and I am the only source for the grades, because I removed the property.
1. A published methodology that the site’s own process actually follows
The first check is not “is there a methodology page.” It is “does the methodology page describe what the site does.”
Wirecutter’s guide template is public. Every guide is supposed to contain “discussions of why you should trust us, who the guide is for, how we picked and tested the products, our picks, flaws (but not dealbreakers) that we encountered, other options worth considering, and the models we tested but don’t recommend” (The Anatomy of a Wirecutter Guide). Thirty seconds and three guides is enough to falsify that.
The same page says why they bother, and the reasoning is the reason I keep citing them: “The Wirecutter journalists who make picks and write guides are dedicated experts. But it’s not enough for us to simply make that assertion and expect readers to trust us. We have to explain why.”
Run it: read the methodology, open two or three actual comparisons, and try to trace one specific verdict back through the stated process. If the trail goes cold, the methodology is decoration.
Result on mine: failed. The benchmark published a scoring methodology. The
scoring function did not implement it. Two documents, one public and one
executing, disagreeing with each other. Worse: a Math.abs() call in the
comparison logic took the absolute value of a score difference, so every product
appeared to lead every rival it was set against. Every head-to-head page was a
win. A rounding error does not do that. That is a claim generator.
2. Inclusion thresholds stated as numbers, not adjectives
“We only cover the leading products” commits its author to nothing at all.
G2 publishes real thresholds. For a category to get a Grid report, “it must have at least six products with 10+ reviews, and 150+ reviews overall,” and “To qualify for inclusion, a product or service must have at least 10 reviews in the corresponding category” (Research Scoring Methodologies). The same page states an update cadence you could catch them missing: “Placement on live Grids® is updated daily.”
Numbers like that are a commitment with a downside attached, which is what makes them worth reading: they tell you which products were excluded and on what ground, they leave the publisher nowhere to retreat to when a category on the site plainly fails to clear the bar it set for itself, and they convert a claim about editorial rigour into something a stranger with a browser can falsify in an afternoon.
Run it: find the stated minimum, then go hunting for a product on the site that falls under it.
Result on mine: failed. The site claimed “10+ sources including G2 and Capterra.” Its own source-coverage panel (which I built and shipped, on the same page) showed those sources as skipped or errored. The site contradicted itself in full view of the reader.
3. Named testers, dated tests
Consumer Reports names the size of the operation: the Auto Test Center “requires a full-time staff of about 30 — engineers, editors, statisticians, technicians, photographers, videographers, and support staff” (How Consumer Reports Tests Cars). Tom’s Hardware’s storage methodology carries a byline and a date (Chris Ramseyer, published 14 March 2015), and the page still tells you the author “was a senior contributing editor for Tom’s Hardware. He tested and reviewed consumer storage” (How We Test HDDs And SSDs).
That date cuts both ways and it is worth being fair about it. A methodology page from 2015 is old. Tom’s Hardware does maintain separate methodology pages per category: power supplies and network switches each have their own, so the 2015 date is not the whole picture. Even so, an old date you can see beats a fresh-looking page with no date, because it lets you ask the right question. An undated methodology cannot be stale, in the same way an unlabelled jar cannot be past its expiry.
Run it: look for a human name and a date on the methodology, and on the individual comparisons. Then ask whether the date on a comparison squares with the products inside it.
Result on mine: failed. It claimed hourly recomputation. The inputs that recomputation ran against were months old. The timestamp was real and it was measuring the wrong event: when the job ran, not when the data changed. Freshness that cannot decay is not being measured; it is being displayed.
4. Conflicts disclosed at the point of the conflict
A reader deciding between two products on a comparison page is not going to detour through your About page first, so a conflict that lives only there arrives after the decision it should have informed. The FTC’s guidance refuses to let the burden slide somewhere convenient: “the ultimate responsibility for clearly and conspicuously disclosing a material connection rests with the influencer and the brand – not the platform” (FTC’s Endorsement Guides: What People Are Asking; the underlying rules are 16 CFR Part 255).
G2 implements the idea at the row level. Reviewers “who have a business relationship with a vendor (or their competitor) that could create bias (such as a reseller) can share insight, but their reviews do not count toward scoring,” and those reviews carry a flag (How G2 ensures authentic reviews). Its community guidelines add: “If a review is incentivized, G2 will clearly label the review as incentivized” (Community Guidelines).
Run it: find a product whose vendor has any commercial relationship with the publisher, then see whether that relationship is named on the page where the product is ranked.
Result on mine: failed, and this is the one that matters. The benchmark ranked Jumbula. I was Marketing Specialist at Jumbula from September 2019 to March 2026. That relationship appeared nowhere in the benchmark property, while the same property asserted that no vendor could pay to influence rankings. The second statement may well have been true in the narrow sense that no money changed hands. It still worked as misdirection, because the conflict that existed was not the conflict being denied.
5. Test units bought, or the alternative stated plainly
Consumer Reports: “we purchase every vehicle we test from a dealership, just like you do. (Last year we spent more than $2.2 million buying cars.)”
Wirecutter’s arrangement is mixed and the page says so: it requests units or buys them, spends “tens of thousands of dollars buying models for testing” every month, returns or donates what it receives, and is blunt about unsolicited gifts: “We simply do not accept these freebies” (Yes, I Work at Wirecutter. No, We Don’t Get a Bunch of Free Stuff.).
Those are two different policies, arrived at by two organisations with different economics, and the honest part is not that either one is ideal but that both are described precisely enough for a reader to weigh what each arrangement might do to a verdict. Set either beside a sentence like “we maintain independence,” which describes no arrangement, names no counterparty, forecloses no possibility and could sit unchanged on the About page of a publisher taking money from every vendor it covers. “We bought it and here is what it cost” can be wrong. That is what makes it worth something.
Result on mine: not applicable, in a way that is itself the finding. Nothing was purchased because nothing was tested. A benchmark ranking 19 products without running any of them is a scoring exercise over secondary data, and it should have said exactly that in its first sentence.
6. Affiliate revenue disclosed where the link is
Wirecutter puts it in the site’s standing header: “We independently review everything we recommend. We may make money from the links on our site,” and expands on the mechanism in About Wirecutter: “we may get paid commissions on products purchased through our links to retailer sites.” Tom’s Hardware carries the equivalent line on the review pages themselves: “When you purchase through links on our site, we may earn an affiliate commission.”
The FTC guidance makes the same point from the other end: the disclosure belongs in the content, not only beside the link.
Result on mine: passed. There was no affiliate revenue, so there was nothing to disclose. The only line that passes, and it passes for the least impressive reason available.
7. A real corrections log
Not a contact form. A dated, public, standing list of things the publisher got wrong. The New York Times publishes one continuously. A comparison site carrying hundreds of verdicts and zero published corrections is either newly launched or not looking.
Result on mine: failed. There was none. Instead of a correction I deleted the
whole property, which is the crudest available remedy and destroys the evidence
along with the error. Had I kept a corrections log from the start, I would have
had somewhere to put the Math.abs() bug on the day I found it, and you would be
able to check my account of it now instead of taking my word.
What the pattern was
The checks that caught me were not the sophisticated ones. Nobody had to audit my statistics.
Three of the seven failures were visible on the page. A source panel that contradicted the sentence above it. A freshness claim contradicted by its own inputs. A comparison table in which nothing ever lost.
Underneath all of it sits one ordinary sequence: I wrote the claims first, because claims are the part you write when you are excited about a project and they cost nothing; I built the system second, against a spec that had already been published to the world as a description of something finished; and then I never went back, not once in months, to run the boring comparison between the two documents that would have taken an afternoon and caught every failure above. Automation widens that gap and then hides it, because output keeps arriving and output keeps looking like output, which is also how the publishing side of this domain ran for six months.
Which reduces the checklist to a single question, asked of any claim you can see: what would I expect to find on this site if that sentence were false? Then go look for it. On my benchmark it took under a minute.
Sources
Every source below was opened and checked on the date shown. Links open in this tab.
- About Wirecutter Wirecutter, The New York Times www.nytimes.com Accessed 5 August 2026
- Yes, I Work at Wirecutter. No, We Don't Get a Bunch of Free Stuff. Wirecutter, The New York Times www.nytimes.com Accessed 5 August 2026
- The Anatomy of a Wirecutter Guide Wirecutter, The New York Times www.nytimes.com Accessed 5 August 2026
- How Consumer Reports Tests Cars Consumer Reports www.consumerreports.org Accessed 5 August 2026
- Who We Are Consumer Reports www.consumerreports.org Accessed 5 August 2026
- Research Scoring Methodologies G2 Documentation documentation.g2.com Accessed 5 August 2026
- How G2 ensures authentic reviews G2 Documentation documentation.g2.com Accessed 5 August 2026
- Community Guidelines G2 legal.g2.com Accessed 5 August 2026
- How We Test HDDs And SSDs Tom's Hardware www.tomshardware.com Accessed 5 August 2026
- How We Test Power Supply Units Tom's Hardware www.tomshardware.com Accessed 5 August 2026
- How We Test Network Switches Tom's Hardware www.tomshardware.com Accessed 5 August 2026
- FTC's Endorsement Guides: What People Are Asking Federal Trade Commission www.ftc.gov Accessed 5 August 2026
- 16 CFR Part 255 — Guides Concerning Use of Endorsements and Testimonials in Advertising Electronic Code of Federal Regulations www.ecfr.gov Accessed 5 August 2026
- Corrections The New York Times www.nytimes.com Accessed 5 August 2026