How we assess, and what we can't yet claim
A certification body is only as trustworthy as its process. Here is exactly how the current results were produced, including their limits.
Every result on this site is a preliminary desk assessment. None is a certification. Desk reviews read public documentation and source code. They do not run hands-on tests such as a full export and a model swap, network inspection for telemetry, or audit-log tampering.
These first assessments were drafted by AI research agents working from the published criteria and required to cite a source for each finding. Human review is in progress and is marked on each assessment. We say so plainly because hiding it would contradict everything we ask of agents.
The process
- 01
Scope the product
We pin down exactly what is assessed: product, plan, platform and date. Consumer-facing defaults are assessed, not enterprise add-ons, unless they are available to individuals.
- 02
Gather primary evidence
Official documentation, help centers, terms of service, privacy policies, export tools and, for open-source projects, the source code itself. Press coverage is used only when no primary source exists, and is labeled as such.
- 03
Score all 34 tests
Each test gets a status, a written finding of one or two sentences, and its citations. 'Pass' needs concrete evidence. When evidence is absent, the test is 'Unverified' and scores zero.
- 04
Review and publish
Findings are checked for consistency across agents, then published in full: every finding, every source, the version assessed and the date.
- 05
Dispute and revise
Vendors and the public can dispute any finding with evidence. Accepted disputes change the record, and the change is logged.
Principles we hold ourselves to
Missing evidence is scored as unmet. Opacity is not neutral when you are asking people to trust an agent with their life.
A hosted product can use only its own models while you stay. What matters is what you can take, and run, when you leave.
We never score what a company promises to do, only what the architecture makes possible or impossible.
Every agent and format, from any company, is scored with the same criteria and the same rules.
Known limitations
- Some vendor help centers block automated access; where findings rely on excerpts or secondary sources, the citation shows it.
- Products change weekly. Each assessment names the version and date it reflects, and older findings may be out of date.
- Open-source projects can pass structural tests while still defaulting to hosted model APIs. We score what the owner can do, not what the default is, and note the default.
- Several tests (P6, V1, A5) can only be settled by hands-on verification, so desk reviews are conservative on them.
Disputes
Anyone can dispute a finding. Include the agent, the test ID (for example P1), and the evidence. We aim to respond within 10 business days, and publish the dispute, our reasoning and any change to the record. Vendors don't get a private channel.
Open a disputeSubmit an agent or format
Any agent format or runtime can be submitted for certification. Submission is free for open-source projects. A format that submits and passes grows the sovereign ecosystem. One that declines tells you something too.
Request certification →