AgentProbe is a map of the agent web — the servers, projects and on-chain identities that agents can actually use, across multiple protocols and 14 EVM-compatible chains. But a map you can trust is not the same as a map of everything. Most directories chase the biggest number they can print. We do the opposite: measure hard, publish little, and only list what passes.
The public directory is deliberately curated. We index a large, messy surface of projects — aggregator wrappers, dead endpoints, test records and impersonation-style entries all live in there. What we show is the small slice that measurably works. We would rather show fewer servers you can trust than a long list you cannot.
From the raw registry to verified trust
Every server starts in the raw registry surface and has to earn its way down the funnel. It has to resolve to a reachable endpoint, speak the MCP handshake correctly, and answer when we probe it — live, not once. Only the servers that clear every step reach the verified online tier that the public Discovery and Projects views show. These figures are live from the same database the front page reads:
One validator for the whole protocol stack
Agents don't speak one protocol; they speak several. AgentProbe assesses the range — the tool interface (MCP), the agent card (A2A), HTTP-native payments (x402), agent commerce (UCP) and on-chain identity (ERC-8004) — so a server is measured against what it actually claims to do. Here is how many projects in the curated set speak each one, live from the current snapshot:
What makes a server pass
Curation here isn't a vibe — it's a concrete, repeatable assessment. Every server runs the same gauntlet: one automated pass that probes it live and checks three things that have to hold together. Miss any one of them and it never reaches the public surface, no matter how polished it looks from the outside.
It implements the MCP handshake correctly and answers the protocol as specified — not a look-alike endpoint that returns plausible-looking noise.
It discovers and exposes real, callable tools. An endpoint that advertises tools it can't actually run doesn't pass — advertising is cheap, answering is not.
Its transport is sound — and wherever it speaks a payment or commerce rail like x402 or UCP, we read and validate that signal too, rather than taking it on faith.
Clear all three with no assessment errors, and the server earns the top mark:
Grade A means the server is online and works, end to end, right now. It measures technical quality — not endorsement. Trust and provenance (who is behind a server, how far you should rely on them) live on a separate axis; Grade is only the "does it actually run" axis. Nothing but Grade-A servers reaches the public directory.
If your server isn't listed, that isn't a public verdict. We don't publish a searchable wall of broken or misconfigured endpoints — so absence means "not yet measured as passing," never "named and shamed."
That's the whole stance: positive-only in public, honest on request. Run your own server through the same assessment privately, see exactly where it stands, and it shows up here the moment it's ready. Fewer servers — each one worth connecting to.