Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
Copy the install, test the workflow, then decide if it earns a permanent slot.
Fresh repo activity plus visible builder pull. This is the kind of tool people test before it turns obvious.
Copy the install, test the workflow, then decide if it earns a permanent slot.
Reasonable to try, but it will take more than a quick skim to get real signal.
GitHub health 37/100. no security policy. 6 open issues make this testable, but not something to trust blind.
AI Agent
OpenClaw
Model
Multiple
Fastest way to find out if claw-eval belongs in your setup.
Copy the install command, run a real test, and back it out cleanly if it slows you down.
git clone https://github.com/claw-eval/claw-eval ~/.claude/agents/claw-evalRun this first. You will know quickly if the workflow earns a permanent slot.
rm -rf ~/.claude/agents/claw-evalNo messy cleanup loop. If it misses, remove it and keep moving.
Install Location
~/ └─ .claude/ ├─ commands/ ├─ agents/ │ └─ claw-eval/ ← installs here └─ settings.json
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
Source: GitHub repository
Source check: August 4, 2026
Upstream commit: August 4, 2026
Repository state: Not marked archived
Honeystax upvotes are community interest signals, not star ratings. GitHub stars and repository health are source measurements; editorial risk and trial-cost notes are Honeystax analysis.