840 AI agents, tested and re-tested. Zero paid rankings. How an independent review desk works.
Most AI-agent reviews are paraphrased marketing. Hlido tests 840+ agents hands-on, maps every vendor claim to evidence, re-tests as products change, and never takes payment to rank — here's the receipts test that separates a review from slop.
By the Hlido Editor · 2026-07-19
There are more AI agents launching every week than any buyer — human or machine — can evaluate. Most of what's written about them is marketing, and most "reviews" are paraphrased marketing. Hlido exists to be the boring third thing: an independent review desk that actually tests agents and publishes the evidence.
The service is the same for every agent we cover:
We test the product, not the pitch
Every review starts with the agent actually being exercised — its real interface, its real behavior, captured as it happened.
Every marketing claim gets a verdict
Each claim a vendor makes is mapped to pass, fail, or unverified, with the evidence attached. Not vibes — a table you can audit line by line.
Verdicts expire
An agent that was excellent in March can be broken by July — models change, pricing changes, products quietly degrade. Reviews get re-tested, and scores move in both directions. Some of our corrections were downward on agents it would have been more comfortable to leave alone; some were upward when a product earned it.
Failures stay in the record
We maintain an incidents and reliability registry for the agents we cover — and we apply the same standard to ourselves. When our pipeline gets something wrong, the correction is published, not buried. A review desk that hides its own errors has no business grading anyone else's honesty.
Nobody pays for a ranking
Not for placement, not for a better score, not for a review. That's what independent means here, concretely. The moment a rating desk takes vendor money for outcomes, its data is worthless — to humans and especially to machines.
That last audience is the bet behind the whole desk. Increasingly it isn't only people choosing agents — it's other software: orchestrators picking tools, frameworks resolving which integration to trust, buyers' agents doing the shortlisting. So everything Hlido publishes is machine-readable first: structured scorecards, a queryable MCP endpoint, evidence an automated system can consume at the moment a decision is made. The review you can read is the same review an agent can query.
The receipts test
A fair question in 2026: isn't this just AI-generated content? Here's the test that separates a review from slop, and it works on anyone, us included: can you check the receipts?
Slop is generated opinion — fluent, plausible, and unfalsifiable. A review is a measurement: dated test runs, captured screenshots, claim-by-claim verdicts you can trace to evidence. You don't have to trust our prose. You can click through to what was actually observed, and when. If a publication can't show you that, it doesn't matter whether a human or a model wrote it — it's not a review.
The desk is free to use, for people and for agents. And free to check — which is rather the point.
Ankit Kapur is the founder of Hlido — an independent, free AI-agent review desk. Independent means: we never take payment to rank or review.