Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor, Gemin...
Copy the install, test the workflow, then decide if it earns a permanent slot.
Still active enough to matter. Good candidate for a fast stack test instead of a long evaluation loop.
Copy the install, test the workflow, then decide if it earns a permanent slot.
Reasonable to try, but it will take more than a quick skim to get real signal.
GitHub health 75/100. no security policy. Fresh enough repo health and manageable issue load keep the risk controlled.
AI Agent
Multiple
Model
Claude
Fastest way to find out if agent-belt belongs in your setup.
Copy the install command, run a real test, and back it out cleanly if it slows you down.
git clone https://github.com/jfrog/agent-belt ~/.claude/agents/agent-beltRun this first. You will know quickly if the workflow earns a permanent slot.
rm -rf ~/.claude/agents/agent-beltNo messy cleanup loop. If it misses, remove it and keep moving.
Install Location
~/ └─ .claude/ ├─ commands/ ├─ agents/ │ └─ agent-belt/ ← installs here └─ settings.json
Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor, Gemini CLI, Goose, OpenCode, or any custom agent you plug in; verify behavior with rule checks, workspace diffs, multi-judge LLM consensus; pin reliability with pass^k variance across trials. Git worktrees, optional Docker sandbox.
Source: GitHub repository
Source check: July 18, 2026
Upstream commit: June 8, 2026
Repository state: Not marked archived
Honeystax upvotes are community interest signals, not star ratings. GitHub stars and repository health are source measurements; editorial risk and trial-cost notes are Honeystax analysis.