Skip to content
Agents for Humanity
Public draft. All results are preliminary desk assessments against criteria v1.0, updated 27 Sept 2026. No agent has been certified yet. How we assess
Autonomous agenthybridProprietary

Claude Cowork

Anthropic · claude.com/docs/cowork/overview

Cowork is an autonomous task agent built on the Claude Code agent architecture, with readable plugins, skills and hooks, editable Markdown memory, approval modes and, for local sessions, on-device JSON history and an HMAC-chained audit log. Structurally it remains a proprietary Anthropic runtime: sessions now run in Anthropic's cloud by default, only Claude models can be used, the Cowork system prompt is unpublished, connector tokens are held server-side, and the product requires a paid plan.

Strengths
  • Plugins, skills, subagents and hooks are readable, editable files; Git repos can serve as marketplaces
  • Local sessions keep history as JSON on disk with an HMAC-chained audit.jsonl
  • Memory stored as Markdown notes, reviewable and deletable item by item
  • Manual-approval mode and mandatory permission before permanent file deletion
  • Per-app computer-use permissions, app blocklist and sandbox egress allowlist
  • Written ad-free commitment
Gaps
  • No runnable export; runtime is proprietary and needs Anthropic services
  • Cloud sessions (the default) store sessions and files on Anthropic servers
  • Claude models only; no local open-weight model support
  • Cowork system prompt and injected safety context not published
  • Connector OAuth tokens held and used server-side by Anthropic
  • Paid plan required; audit export (OpenTelemetry) limited to Team/Enterprise
Evidence

All 34 findings

Claude Cowork on Pro and Max plans as of 2026-09-27: cloud sessions (default, beta; web, desktop, mobile) and local desktop sessions in a Linux VM on macOS/Windows, including plugins, scheduled tasks, computer use and the built-in browser. Since Sept 2026 Cowork is merged into the Claude app's single conversation surface. Team/Enterprise controls and Claude Desktop on 3P (commercial-terms deployment) noted only where relevant.

Portable · Can you leave, and take the whole agent with you?

8%

Transparent · Can you see everything the agent is, with ordinary tools?

42%

Auditable · Can you reconstruct exactly what the agent did?

20%

Verifiable · Can you prove the agent runs what it claims?

20%
  • V1
    Open, reproducible runtime

    Claude Desktop, the Cowork VM bundle and the cloud sandbox are proprietary; use of the app is governed by Anthropic's terms, and no OSI-licensed, reproducible runtime is published.

    Fail
  • V2
    Active config is inspectable config

    Instructions, memory files and plugins are inspectable, and 3P deployments use inspectable JSON configuration. For consumer accounts, server-side configuration, classifiers and account-synced settings cannot be verified to match what the owner sees.

    Partial
  • V3
    Attributable messages

    Messages go out through the owner's connected accounts and connectors, and recipients cannot verify that they came from this agent under the owner's authority.

    Fail
  • V4
    Independently checkable record

    The local audit log is HMAC-chained with a keychain-protected key, but no verification procedure or open verifier is published, so its integrity cannot be checked independently of Anthropic's app.

    Fail
  • V5
    Comparable state

    Local-session state (transcripts, memory, plugins and the audit log) sits in ordinary files that the owner can snapshot and compare with standard tools. Cloud sessions, the default, and account-synced settings are not fully exposed, so changes there cannot be verified.

    Partial

Modifiable · Can you change anything, without asking?

42%

Controllable · Is your word final?

50%
Something wrong or out of date?

Vendors and the public can dispute any finding with evidence. Disputes and their resolutions are published.

Dispute a finding