← Tool Radar
dataset

Deep-XPIA – Prompt injection benchmark for multi-agent AI systems

This is a specialized research benchmark for prompt injection in multi-agent AI systems, offering significant evaluation value for security researchers and system developers.

5.9Overall
Utility6
Onboarding5
Craft6
Niche fit7
Longevity5

Five dimensions scored independently; overall is a weighted average. Scores are only comparable within this same rubric.

Good for

Researchers and security engineers focused on multi-agent AI system security and prompt injection vulnerabilities.

Not for

General AI developers or those not focused on security research.

Project description

freyzo.github.io

Alternatives

None recorded

Visit site1 · Stars at evalEvaluated 2026-06-16

Similar tools