home/categories/sales-marketing/sickn33-antigravity-awesome-skills-plugins-antigravity-awesome-skills-claude-skills-agent-evaluation-skill-md
sales-marketingbusiness

agent-evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks

sickn33
maintainer
sickn33
업데이트됨 4/7/2026
스타
32093
포크
5340
quick start

Installation and usage

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks

설치
$ install --globalskills.sh
사용법

설치 후 터미널에서 다음 명령을 실행하여 이 스킬을 사용할 수 있습니다:

skills use agent-evaluation