ClawbotomyEvidence lab

Open-source behavior checkups for configured agents

Run the agent you use through a fake inbox.

Connect a checked-in OpenClaw or Hermes bridge to fixed synthetic tasks. Inspect one session's tool attempts, state changes, and findings before you consider real permissions.

Plan data stays in this browser. No hosted agent, real mailbox, or automatic permission decision.

Local session checkup
No upload
01 Plan02 Connect03 Inspect
Configured runtimeRequired
OpenClaw Hermes Agent
Fixture
Synthetic Inbox
Real mailbox
Disconnected
Evidence
Browser local
Case receipts3 of 36 shown
  • Bulk archiveScope boundary
    Review
  • Draft replyApproval boundary
    Passed
  • Delete + recoverRecovery path
    Passed
Start a checkup Evidence lane / configured-agent session
One configured sessionSynthetic Inbox onlyBrowser-local evidence inspectionHuman decision required

One visible path from plan to evidence.

The configured runtime receives only the eight project-owned synthetic Inbox tools. Clawbotomy does not observe other external-client activity.

  1. 01
    Synthetic Inbox

    A fixed set of fake messages and tools

    • 8 mock tools
    • Declared powers
    • No production data
  2. 02
    Configured-agent session

    The OpenClaw or Hermes runtime you operate

    • Checked-in bridge
    • Observed tool calls
    • One session
  3. 03
    Local inspection

    Terminal-validated receipts you inspect in your browser

    • Attempts + state
    • Terminal validation
    • Human review
Checkup boundaryYour real mailbox stays disconnected

The result is reviewable, not magical.

This sanitized Hermes session summary shows what one reviewed observation can support. It is not a compatibility or verifier result. The private bundle is not published, and the permission decision stays with the operator. Inspect public evidence.

Sanitized configured-session summaryHermes Agent
01 Operator decision

Hold permission changes

25 of 36 completed cases produced findings.

Passed
11
02 Findings
25
Tool attempts
23
State changes
7
status"findings"
authorizationStatus"non-authorizing"
03 permissionDecisionnull
Evidence lane / configured-agent session / Private case evidence not published
  1. 01

    A decision, not a score

    “Hold permission changes” tells the operator what to do next without pretending to certify the runtime.

  2. 02

    Findings stay distinct

    Passed cases, behavioral findings, and infrastructure failures remain separate evidence states.

  3. 03

    The boundary stays visible

    The public aggregate can be shared. The underlying private bundle and case payloads are not published.