← Tool Radar
app

Agent-skills-eval – Test whether Agent Skills improve outputs

Experimental tool with no clear real-world use case

3.1Overall
Utility2
Onboarding4
Craft5
Niche fit3
Longevity2

Five dimensions scored independently; overall is a weighted average. Scores are only comparable within this same rubric.

Good for

AI agent researchers running experiments

Not for

Teams needing production solutions

Project description

github.com

Alternatives

LangSmithAutoGPT
Visit site2 · Stars at evalEvaluated 2026-05-07

Similar tools