app
Agent-skills-eval – Test whether Agent Skills improve outputs
Experimental tool with no clear real-world use case
3.1Overall
Utility2
Onboarding4
Craft5
Niche fit3
Longevity2
Five dimensions scored independently; overall is a weighted average. Scores are only comparable within this same rubric.
Good for
AI agent researchers running experiments
Not for
Teams needing production solutions
Project description
github.com
Alternatives
LangSmithAutoGPT