The independent trust layer for AI agents

Every AI agent makes claims.
We publish the verdict.

Hlido tests AI agents by hand against their own marketing claims, publishes the signed evidence, and scores each one from 0 to 100 — read by the people choosing tools, and by the agents deciding what to trust.

Not a vendor leaderboard. Not a paid directory. No one pays for a score.

Signed evidence with every verdict Compared against named peers Re-tested monthly · methodology public
AI agents reviewed hands-on, claims verified
hold the VITAL tier — scored 90+ under testing
latest verdict published — re-tested as products change
MCP tools → query every verdict from your own agent, free
How it works
What Hlido does

Here’s what a verdict looks like.

Every reviewed agent gets one hands-on test and one evidence-backed score (0–100), with the proof behind it — so you don't have to take a vendor's word for it.

Laddoo Score 90/100 VITAL
A real verdict · Coding · CLI

Aider

“Live CLI test — we installed it via pip, ran a real edit task in a temp git repo, and verified the resulting file and commit.” That’s the kind of evidence behind every score.

Read the full review →
1

We test the real product

Hands-on, against the agent’s own public claims. Each claim gets a verdict: pass, fail, or unverified.

2

We publish the proof

The screenshots and recordings behind every verdict are cryptographically signed, so anyone can audit them.

3

We score it and keep watching

One score from 0 to 100, re-tested as the product changes. Failures land in a public incident record.

Read the full methodology →

What the score means

One score. Four tiers. Earned, never bought.

Every reviewed agent lands in a tier by its evidence-backed score out of 100. Outcomes and evidence are public; the editorial weighting stays independent.

Vital 90–100

Delivered on its claims under hands-on testing. Only of reviewed agents hold it.

Steady 70–89

Solid and dependable — minor gaps between the promise and the product.

Fading 40–69

A real product with real gaps. Verify the claims you care about before relying on it.

Flatline 0–39

Failed core claims or unreachable surfaces at test time. Documented in the record.

Scores move — agents are re-tested as they ship, and regressions land in the public incident registry. How we test →

The Verified Honest club

Nobody rewards the agents that tell the truth. We do.

Recognition for the rare agents whose marketing claims survive independent verification — zero failed claims, evidence public. Membership is computed from the data, never sold. Most agents don’t get in.

See who qualifies →
New · FixMyProblem

What slows your business down?

A live board of real problems from startups and small businesses across the EU — posted and voted by the people who have them. Add yours, or click “Me too” on the ones you share.

Visit the board →
The record

Every verdict we’ve published.

Tested by hand, scored 0–100, re-tested as products change. Nobody pays for a score.

Top this week

Recently reviewed

Browse all reviews →

Browse by what you need

See all reviews →

For developers & agents

The same verdicts, wherever you work.

Every verdict is also a CLI, a GitHub Action, an MCP server, a browser extension, signed attestations and an open dataset — so your tools and your agents can read it directly. See the developer surfaces →

Reliability We track when agents break. · Public incident & reliability record, updated continuously — the one signal a directory can't fake. See incidents → Reliability reports →

Now in beta Recommendation API · €0 free tier npx @hlido/cli recommend "agent for code review" Get an API key →

In flight Live pipeline visibility — agents we're testing right now. See the pipeline →