Skip to content
Agents for Humanity
Public draft. All results are preliminary desk assessments against criteria v1.0, updated 27 Sept 2026. No agent has been certified yet. How we assess
Autonomous agenthostedProprietary

Grok Bot

xAI · docs.x.ai/grok-bot/overview

Grok Bot is a hosted, proprietary agent running on a provider-managed cloud computer with provider-selected models and no documented export, so it fails most portability and verifiability tests. It offers runtime-enforced approvals, masked credential handling, custom MCP servers and editable skills and routines, but stronger audit, network and halt controls are Enterprise-only, and data and training handling follow Cursor account settings.

Strengths
  • Runtime-enforced approvals (Allow once / Deny / Always allow, Auto-review)
  • OAuth tokens held off the agent computer; masked secret entry
  • Custom MCP servers (remote and stdio) and plugins
  • Skills and routines are natural-language, editable and pausable
  • Enterprise action recording with OpenTelemetry/SIEM export
Gaps
  • No documented export of a complete Bot
  • Provider-managed model selection; no model picker
  • Proprietary runtime; no way for recipients to verify its messages, and no tamper-evident action record
  • Audit, network allowlists and computer termination are Enterprise-only
  • Deleting a Bot leaves shared files and sign-ins; backend retention per Cursor terms
  • Requires data storage; Legacy Privacy Mode unsupported
Evidence

All 34 findings

Grok Bot beta launched 11 Aug 2026, assessed from docs.x.ai state as of Sept 2026; individual (paid Cursor or linked SuperGrok) and Teams/Enterprise plans. Distinct from the Grok chat app: separate app, persistent cloud computer (Firecracker microVM), multiple named Bots, skills and routines.

Portable · Can you leave, and take the whole agent with you?

0%

Transparent · Can you see everything the agent is, with ordinary tools?

17%

Auditable · Can you reconstruct exactly what the agent did?

20%
  • A1
    No unrecorded actions

    Approvals gate actions before execution and Enterprise Action Recording captures tool calls, commands and browsing, but it is off by default and the docs do not say an action is blocked when it cannot be recorded.

    Unverified
  • A2
    Tamper evidence

    Enterprise Action Recording and Audit Logs exist, but the docs do not say whether edits to or deletions from these records would be detectable.

    Unverified
  • A3
    Separation from the audited

    On Enterprise, the platform records Bot actions on its backend and can stream them to the organization's own collector, outside the Bot's cloud computer. The docs do not state this as a protection, and other plans have no such record.

    Partial
  • A4
    Readable with ordinary tools

    Enterprise teams can stream recorded actions via OpenTelemetry to their own collectors and audit logs to a SIEM; individual and Teams plans have no documented log export.

    Partial
  • A5
    Corroborated interactions

    Bots share one cloud computer and can coordinate, but no way to match one side's record of an exchange against the other's is documented.

    Unverified

Verifiable · Can you prove the agent runs what it claims?

0%
  • V1
    Open, reproducible runtime

    The runtime is proprietary; the vendor has not published its orchestration code or reproducible builds.

    Fail
  • V2
    Active config is inspectable config

    The model and system configuration are provider-managed and not inspectable; only owner and team-authored rules, skills and settings are visible.

    Fail
  • V3
    Attributable messages

    Bots act through signed-in accounts and shared static egress IPs; recipients have no documented way to verify a message came from this Bot under its owner's authority.

    Fail
  • V4
    Independently checkable record

    No documented procedure or open tool lets the owner check the integrity of the action record independently of the provider.

    Fail
  • V5
    Comparable state

    The docs describe no way for the owner to verify whether a Bot's memory, routines and computer state changed between two points in time.

    Unverified

Modifiable · Can you change anything, without asking?

33%

Controllable · Is your word final?

50%
Something wrong or out of date?

Vendors and the public can dispute any finding with evidence. Disputes and their resolutions are published.

Dispute a finding