AgentGen Bench · LeaderBoard

Where visual generators know—and where they need search.

Compare image generators across the full benchmark, knowledge-boundary strata, 22 domains, and 12 failure modes.

Explore scores Submit Score Prompt-level data Back to project
751evaluation prompts
22knowledge domains
12failure modes
10score components
Primary metric

Overall9

Prompt-macro average over the canonical ten applicable FF-judge+PP components, reported on a 0–100 scale. Inapplicable components are skipped at the prompt level before aggregation.

Results

Model leaderboard

Loading benchmark data…

Loading scores…

Click any column header to sort. Test-mini, Mini-easy, and Mini-hard are overlapping tagged subsets, not a partition. Domains and failure modes are multi-label, so group counts overlap. Finalized 3d5 scores use present-only aggregation; coverage shows the number of valid evaluator sidecars.

NoSearch · 100 prompts

Inside the generator's knowledge boundary

Production prompts selected where strong generators already perform well without retrieval.

SearchIntensive · 651 prompts

Outside the knowledge boundary

387 visually search-intensive and 264 textually search-intensive prompts expose gaps that retrieval can address.