Is your agent
working for you?
AI agents are one of the most consequential technologies we have ever built. They will know us better than any system before them, and they will act on our behalf as extensions of ourselves rather than tools we use. We think it is essential to know whether an agent serves the person using it or someone else's interests, so we evaluate them in the open.
- PortableCan you leave, and take the whole agent with you?6 tests
- TransparentCan you see everything the agent is, with ordinary tools?6 tests
- AuditableCan you reconstruct exactly what the agent did?5 tests
- VerifiableCan you prove the agent runs what it claims?5 tests
- ModifiableCan you change anything, without asking?6 tests
- ControllableIs your word final?6 tests
An agent is certified only when every test passes. Each test is scored pass, partial, fail or unverified, with the evidence published. Read the criteria
How sovereign are today's agents?
Preliminary desk assessments of 26 agents against all 34 tests. Last updated 27 Sept 2026.
- 1OpenClaw86
- 2ADF86
- 3Hermes Agent83
- 4ZeroClaw78
- 5goose75
- 6nanobot75
- 7Open Interpreter72
- 8Letta71
- 9Claude Cowork30
- 10Claude28
- 11ChatGPT23
- 12Gemini21
- 13Gemini Spark21
- 14Meta Muse21
- 15Grok Bot20
- 16Perplexity Comet19
- 17Siri19
- 18Microsoft Copilot Cowork18
- 19Microsoft Copilot17
- 20Manus17
- 21ChatGPT Atlas15
- 22Grok14
- 23Meta AI11
- 24Genspark11
- 25Google CC11
- 26Instinct7
- 26
- agents assessed
- 34
- tests per agent
- 495
- sources cited
- 0
- certified so far
Where each agent holds up, and where it doesn't
Share of tests passed per criterion. Partial results count half, and unverified claims count zero until demonstrated.
| P Portable | T Transparent | A Auditable | V Verifiable | M Modifiable | C Controllable | |
|---|---|---|---|---|---|---|
| OpenClaw | 92 | 92 | 70 | 80 | 100 | 83 |
| ADF | 100 | 92 | 70 | 60 | 100 | 92 |
| Hermes Agent | 92 | 100 | 70 | 50 | 100 | 83 |
| ZeroClaw | 83 | 83 | 40 | 60 | 100 | 100 |
| goose | 75 | 100 | 50 | 50 | 100 | 75 |
| nanobot | 83 | 100 | 40 | 50 | 100 | 75 |
| Open Interpreter | 67 | 92 | 60 | 40 | 92 | 83 |
| Letta | 67 | 92 | 60 | 50 | 100 | 58 |
| Claude Cowork | 8 | 42 | 20 | 20 | 42 | 50 |
| Claude | 8 | 50 | 10 | 10 | 42 | 50 |
| ChatGPT | 8 | 33 | 10 | 0 | 33 | 50 |
| Gemini | 8 | 33 | 10 | 0 | 33 | 42 |
| Gemini Spark | 8 | 33 | 10 | 0 | 33 | 42 |
| Meta Muse | 8 | 25 | 0 | 0 | 42 | 50 |
| Grok Bot | 0 | 17 | 20 | 0 | 33 | 50 |
| Perplexity Comet | 8 | 25 | 0 | 0 | 33 | 50 |
| Siri | 0 | 17 | 10 | 10 | 33 | 42 |
| Microsoft Copilot Cowork | 0 | 25 | 10 | 0 | 25 | 50 |
| Microsoft Copilot | 0 | 17 | 10 | 0 | 25 | 50 |
| Manus | 8 | 25 | 0 | 0 | 33 | 33 |
| ChatGPT Atlas | 0 | 8 | 0 | 0 | 25 | 58 |
| Grok | 8 | 25 | 10 | 0 | 25 | 17 |
| Meta AI | 8 | 17 | 10 | 0 | 25 | 8 |
| Genspark | 0 | 8 | 0 | 0 | 33 | 25 |
| Google CC | 0 | 8 | 0 | 0 | 17 | 42 |
| Instinct | 0 | 0 | 0 | 0 | 25 | 17 |
Six criteria. All required.
The criteria are individually necessary and jointly sufficient. There is no partial certification: an agent that fails one criterion serves some other interest.
Portable
Can you leave, and take the whole agent with you?
The owner can export the complete agent and run it on their own machine with a free, open-source runtime, with no dependency on the original provider.
Transparent
Can you see everything the agent is, with ordinary tools?
The agent's full internal state can be inspected with general-purpose tools, with nothing hidden from the owner.
Auditable
Can you reconstruct exactly what the agent did?
The agent keeps a tamper-evident record that is enough to reconstruct its actions.
Verifiable
Can you prove the agent runs what it claims?
Anyone can confirm that the agent runs the code and configuration it claims, and that its history is complete.
Modifiable
Can you change anything, without asking?
The owner can change any part of the agent without the provider's permission.
Controllable
Is your word final?
The owner has authoritative, runtime-enforced control over the agent's actions, communication and lifecycle.
Structure, not intent
- Step 1
Public criteria
Every test and its 'does not require' clause are published and versioned. Anyone can challenge them, and changes are logged.
- Step 2
Evidence-backed assessment
Each of the 34 tests is scored with a written finding and cited sources: docs, terms, privacy policies and source code. Vendors can dispute any finding.
- Step 3
Certification, then re-certification
Only hands-on AFH verification certifies, and only when every test passes. Major versions trigger re-certification.
An agent that can be ad-supported, remotely switched off, or quietly re-instructed isn't yours. We don't judge what an agent does. We check who it answers to.