Skip to content
Agents for Humanity
Public draft. All results are preliminary desk assessments against criteria v1.0, updated 27 Sept 2026. No agent has been certified yet. How we assess
Independent · Open criteria · Open evidence

Is your agent
working for you?

AI agents are one of the most consequential technologies we have ever built. They will know us better than any system before them, and they will act on our behalf as extensions of ourselves rather than tools we use. We think it is essential to know whether an agent serves the person using it or someone else's interests, so we evaluate them in the open.

Leaderboard

How sovereign are today's agents?

Preliminary desk assessments of 26 agents against all 34 tests. Last updated 27 Sept 2026.

26
agents assessed
34
tests per agent
495
sources cited
0
certified so far
Breakdown

Where each agent holds up, and where it doesn't

Share of tests passed per criterion. Partial results count half, and unverified claims count zero until demonstrated.

P PortableT TransparentA AuditableV VerifiableM ModifiableC Controllable
OpenClaw9292708010083
ADF10092706010092
Hermes Agent92100705010083
ZeroClaw83834060100100
goose75100505010075
nanobot83100405010075
Open Interpreter679260409283
Letta6792605010058
Claude Cowork84220204250
Claude85010104250
ChatGPT8331003350
Gemini8331003342
Gemini Spark8331003342
Meta Muse825004250
Grok Bot0172003350
Perplexity Comet825003350
Siri01710103342
Microsoft Copilot Cowork0251002550
Microsoft Copilot0171002550
Manus825003333
ChatGPT Atlas08002558
Grok8251002517
Meta AI817100258
Genspark08003325
Google CC08001742
Instinct00002517
The standard · v1.0

Six criteria. All required.

The criteria are individually necessary and jointly sufficient. There is no partial certification: an agent that fails one criterion serves some other interest.

Process

Structure, not intent

  1. Step 1

    Public criteria

    Every test and its 'does not require' clause are published and versioned. Anyone can challenge them, and changes are logged.

  2. Step 2

    Evidence-backed assessment

    Each of the 34 tests is scored with a written finding and cited sources: docs, terms, privacy policies and source code. Vendors can dispute any finding.

  3. Step 3

    Certification, then re-certification

    Only hands-on AFH verification certifies, and only when every test passes. Major versions trigger re-certification.

Why this matters

An agent that can be ad-supported, remotely switched off, or quietly re-instructed isn't yours. We don't judge what an agent does. We check who it answers to.

About the project →