The short version — teach what you can, search the rest
Modern image generators render gorgeously and lie fluently. Ask for the 2025 Osaka Expo mascot and you get a confident, wrong invention. The failure isn't the pixels — it's the knowledge. On AgentGen Bench, frontier open generators score just 21–28 out of 100 on search-intensive prompts — up to a ~40-point collapse that standard benchmarks never register.
Search is the obvious fix, the way an illustrator consults references. But naive search backfires: it corrupts prompts the generator already handled. The real problem is a knowledge boundary — the line between what a generator can learn and what it must look up. That line is generator-specific, it moves during training, and it can't be hand-drawn. It has to be discovered.
So we discover it, by co-training the generator and the search agent together. Below: how the collapse works, why naive search fails, what the knowledge boundary is, and how a co-trained 8B reasoner on a 4B generator matches a frontier reasoner on the same generator — jump to the finding cards. ↓
They render beautifully. They just make things up.
Ask a frontier image model for the mascot of the 2025 Osaka Expo. You get a polished, confident fabrication. Ask for a historically accurate Spartan phalanx and you get anachronistic armor, rendered in exquisite detail.
The lighting is right. The composition is right. The world is wrong.
This is not a rendering failure. It's a world-knowledge bottleneck. Generators train on fixed corpora with hard knowledge cutoffs; user requests draw on new characters, regional symbols, niche typography, historical artifacts, and events that postdate training.
Worse, generators have no way to flag their own ignorance. They're trained to always output an image — never to say "I don't know what this looks like." So they guess, beautifully, every time.





A 40-point gap that no benchmark shows
Generators that score comparably on standard prompts diverge by nearly 40 points when search-intensive world knowledge is required.
On prompts that need only what a model already learned, open and commercial generators land in the same band (67–75 out of 100). Turn to prompts that need live world knowledge, and the field splits: open generators crater to 21–28, while commercial systems with built-in search barely move. Existing benchmarks test rendering inside known concepts, so they never see this gap at all.
AgentGen Bench contains 751 prompts scored by the finalized 3d5 evaluator. The primary Overall9 metric averages nine applicable components and excludes Text Reference.
Full benchmark breakdown
Scores are on a 0–100 scale; higher is better. Overall9 averages the nine displayed components and excludes Text Reference.
| Generator | Overall | Checklist | Rubric | Prompt | Image quality | Text rendering | AI naturalness | Composition | Physical plausibility | Visual reference |
|---|---|---|---|---|---|---|---|---|---|---|
| NoSearch parametric knowledge is sufficient | ||||||||||
| GPT-Image-2 | 78.9 | 84.8 | 82.7 | 79.3 | 73.4 | 92.3 | 67.6 | 86.5 | 81.2 | 70.9 |
| Grok-Imagine-Image | 74.8 | 81.3 | 78.3 | 72.7 | 72.3 | 83.9 | 65.6 | 79.7 | 77.2 | 66.9 |
| Qwen-Image-3-Max | 73.7 | 79.1 | 76.2 | 69.8 | 70.3 | 86.2 | 63.9 | 82.5 | 78.3 | 64.0 |
| Nano Banana Pro | 73.3 | 78.4 | 76.3 | 70.2 | 70.7 | 85.6 | 65.2 | 80.3 | 77.1 | 63.0 |
| Qwen-Image-2-Pro | 71.2 | 76.8 | 74.1 | 67.3 | 70.5 | 77.8 | 61.8 | 78.3 | 75.3 | 61.2 |
| Qwen-Image-2 | 71.1 | 76.2 | 73.1 | 67.5 | 70.2 | 79.5 | 63.8 | 79.2 | 74.2 | 60.3 |
| SeedDream-4.0 | 70.7 | 76.4 | 73.3 | 65.8 | 69.0 | 76.0 | 61.7 | 80.8 | 75.5 | 58.5 |
| Qwen-Image | 70.0 | 74.9 | 72.0 | 65.5 | 69.2 | 61.1 | 61.7 | 77.0 | 78.6 | 60.1 |
| SeedDream-4.5 | 69.3 | 76.2 | 72.8 | 65.5 | 67.3 | 74.0 | 59.0 | 76.8 | 73.4 | 59.6 |
| Nano-Banana | 66.3 | 70.8 | 68.5 | 61.8 | 67.7 | 49.0 | 59.7 | 77.5 | 70.5 | 56.2 |
| SenseNova-U1 | 62.8 | 68.1 | 65.2 | 58.5 | 64.3 | 58.3 | 56.2 | 72.2 | 65.1 | 51.1 |
| Mage-Flow | 61.2 | 64.2 | 61.8 | 54.5 | 64.7 | 51.2 | 56.3 | 71.8 | 67.4 | 47.1 |
| Flux.2-Klein-9B | 58.5 | 63.7 | 60.7 | 49.5 | 63.2 | 30.0 | 53.5 | 71.0 | 67.0 | 41.1 |
| Flux.2-Klein-4B | 52.3 | 55.8 | 52.6 | 42.3 | 59.7 | 16.1 | 49.8 | 67.7 | 60.1 | 35.3 |
| Bagel | 49.1 | 51.5 | 49.2 | 37.7 | 56.2 | 16.1 | 48.8 | 62.8 | 58.5 | 32.1 |
| OmniGen2 | 47.4 | 49.1 | 46.1 | 33.5 | 56.4 | 7.8 | 47.3 | 61.4 | 60.3 | 30.6 |
| Show-o2 | 34.4 | 32.7 | 30.9 | 21.0 | 47.2 | 2.4 | 38.0 | 49.8 | 40.9 | 19.9 |
| SearchIntensive external knowledge is required | ||||||||||
| GPT-Image-2 | 75.5 | 74.4 | 73.7 | 72.2 | 77.4 | 77.7 | 68.9 | 83.9 | 84.9 | 70.4 |
| Qwen-Image-3-Max | 67.6 | 65.8 | 65.3 | 62.8 | 71.1 | 67.9 | 62.3 | 77.0 | 78.6 | 59.6 |
| Grok-Imagine-Image | 66.7 | 65.9 | 64.6 | 62.8 | 69.5 | 65.7 | 61.7 | 75.6 | 77.8 | 59.7 |
| Nano Banana Pro | 63.6 | 60.7 | 60.1 | 56.4 | 68.1 | 61.3 | 60.0 | 72.7 | 78.1 | 55.4 |
| Qwen-Image-2-Pro | 57.5 | 53.9 | 53.1 | 48.3 | 65.3 | 55.1 | 57.3 | 68.9 | 72.8 | 46.8 |
| Qwen-Image-2 | 54.1 | 49.9 | 49.6 | 44.1 | 62.5 | 50.1 | 54.3 | 66.2 | 71.4 | 42.0 |
| SeedDream-4.5 | 54.1 | 51.7 | 50.8 | 45.2 | 61.8 | 50.4 | 50.8 | 65.9 | 69.3 | 42.1 |
| SeedDream-4.0 | 49.6 | 45.6 | 45.1 | 39.3 | 58.8 | 37.2 | 50.9 | 62.5 | 68.0 | 36.8 |
| Nano-Banana | 47.1 | 43.0 | 42.3 | 36.1 | 56.9 | 31.2 | 48.9 | 60.5 | 67.0 | 35.9 |
| SenseNova-U1 | 31.6 | 29.3 | 27.6 | 21.6 | 39.5 | 15.8 | 32.2 | 42.5 | 48.4 | 21.0 |
| Qwen-Image | 28.6 | 25.0 | 23.8 | 17.4 | 37.9 | 9.2 | 29.9 | 41.0 | 45.5 | 17.7 |
| Mage-Flow | 28.4 | 23.0 | 21.7 | 15.7 | 38.7 | 9.7 | 35.0 | 40.0 | 50.2 | 16.2 |
| Flux.2-Klein-9B | 26.9 | 23.8 | 22.4 | 15.1 | 36.1 | 5.8 | 29.9 | 38.3 | 47.7 | 16.7 |
| Flux.2-Klein-4B | 23.9 | 20.1 | 18.6 | 11.5 | 33.0 | 3.4 | 28.1 | 35.7 | 44.8 | 12.7 |
| Bagel | 22.2 | 18.7 | 17.7 | 12.4 | 30.4 | 2.2 | 25.0 | 32.4 | 38.1 | 14.2 |
| OmniGen2 | 20.4 | 16.4 | 14.9 | 8.8 | 29.2 | 2.5 | 24.3 | 31.4 | 40.8 | 9.5 |
| Show-o2 | 17.5 | 8.1 | 8.0 | 4.2 | 30.2 | 0.7 | 26.5 | 34.3 | 31.7 | 3.4 |
Search should help. Often it hurts.
An illustrator handed an unfamiliar brief looks up references before drawing. Give the generator the same move — a reasoner spots knowledge gaps, search fills them, the results feed generation — and you have agentic visual generation. Natural. And, done naively, harmful.
Naive search actively degrades prompts the generator already handles.
Search everything, blindly, and every generator gets worse on prompts that never needed help. Qwen-Image-2 drops from 70.7 to 60.4 on the no-search stratum — a 14.6% relative loss on prompts it already aced.
| Need Search? | Generator | n | NoSearch | Reasoned | Blind |
|---|---|---|---|---|---|
| NoSearch | Qwen2 | 100 | 70.7 | 76.5 | 60.4 |
| Qwen1 | 99 | 67.5 | 75.0 | 59.5 | |
| Klein-9B | 58 | 57.9 | 66.4 | 52.3 | |
| VisualSearch | Qwen2 | 384 | 37.2 | 49.1 | 45.3 |
| Qwen1 | 384 | 32.8 | 44.4 | 39.7 | |
| Klein-9B | 338 | 29.6 | 39.4 | 36.0 | |
| TextualSearch | Qwen2 | 263 | 22.9 | 34.1 | 32.1 |
| Qwen1 | 263 | 20.5 | 28.5 | 25.3 | |
| Klein-9B | 264 | 18.9 | 26.6 | 24.0 |
Two distinct failures explain it. Concept corruption: search fires on something the model already knew, and the retrieved reference overrides correct internal knowledge — a gating failure, searching when it shouldn't. Copy effect: a reference carries so much detail that the generator copies it wholesale instead of borrowing the one missing fact — a filtering failure, keeping too much.
The line between what a model learns and what it must look up
Some knowledge belongs inside the model. A character's canonical look, a flag's fixed geometry — stable, low-dimensional, learnable once and for all. Fire search for it and you only add noise.
Other knowledge belongs outside. It changes faster than retraining cycles, sits too deep in the long tail to ever learn reliably, or needs per-request reasoning. For this, search isn't a nice-to-have — it's structurally necessary. And the tail is enormous: 93.1% of the 31,537 entities in our data appear in just a single prompt. No feasible training set covers that.
Some knowledge is internalizable and search should not fire for it; other knowledge is contextual and search is structurally necessary.
We call the divide between those two sets the knowledge boundary. Here's the twist that organizes the entire paper: the boundary is generator-specific, and it moves. As a generator learns, concepts migrate from "must search" to "already knows." A search policy tuned for a weak generator is wrong for a strong one.
So you can't hand-specify it. And you don't have to. The boundary is discoverable — it falls out of training the generator and the searcher together.
Build a searcher that resists noise. Then move the boundary.
The method has two halves. First, a searcher that doesn't poison the generator. Second, a training loop that finds — and expands — the boundary.
The noise-resistant reasoner: gate, filter, integrate.
⎇ Gate
Decide whether to search. Only gaps rated critical or important trigger a query; the rest are dropped. At most three queries per prompt, or skip entirely. The defense against concept corruption.
☰ Filter
Decide what to keep. Choose the reference that fills the specific gap with the least extraneous baggage. The defense against the copy effect.
✦ Integrate
Decide how it enters. Route visual references through language, not raw pixels: "following Image 1, render the robe in teal and gold." The generator borrows exactly what's named — nothing else leaks in.
Teach-then-search co-training. The boundary only moves when the generator's weights change, so we act on both sides, in order.
Warm-start
Supervised finetuning gives an 8B reasoner (Qwen3-VL-8B) the gate–filter–integrate protocol. Competent, but generator-agnostic: it searches for anything that might be hard.
Teach the generator
Online iterative Diffusion-DPO feeds it search-augmented examples and reinforces its own best outputs. Two things happen at once: it internalizes stable knowledge (the boundary pushes outward), and it learns to use imperfect references without being dominated by them (noise-robustness).
Recalibrate the searcher
The generator is now stronger, so the search policy is stale. Rejection-sampling finetuning rewards trajectories where search actually helped the new generator and discards the rest. Search fires only for what remains genuinely contextual.
Co-training discovers the boundary — an 8B reasoner matches a frontier oracle
Co-training discovers the knowledge boundary: a calibrated 8B reasoner matches a frontier oracle on the same generator.
On Klein-4B, the full co-training progression climbs monotonically to 31.8 overall — matching the frontier Gemini oracle on the very same generator (31.2). Both phases pull their weight: teaching the generator adds +2.8, recalibrating the searcher adds +2.6.
It's also selective, which is the hard part. On prompts that don't need search, the calibrated reasoner lifts the no-search baseline from 49.9 to 56.9 — it learned when to stay quiet, exactly where naive search did damage. And it's generator-specific: a policy tuned for the strengthened generator scores worse on the base one, confirming the boundary is a joint property of the pair, not a fixed property of the prompt.
The headline is the calibration-per-dollar. An 8B reasoner on a 4B generator reaches the same quality as a frontier commercial reasoner on that generator — and the full recalibration cycle fits in 4×8 GPU-hours.
One honest caveat: this is a comparison of reasoners on a fixed 4B generator, not a claim of frontier image quality. In absolute terms, 31.8 remains below GPT-Image-2's finalized full-benchmark Overall9 score of 76.0.
| Phase | Config | n | NoSearch | Easy | Medium | Hard | Overall |
|---|---|---|---|---|---|---|---|
| Phase 0 | Gen-Agnostic (SFT-8B) + Klein-4B | 602 | 54.6 | 28.9 | 29.2 | 21.2 | 26.4 |
| Phase 1 | Gen-Agnostic (SFT-8B) + Klein-4B-DPO-v2 | 321 | 54.0 | 31.8 | 31.1 | 24.7 | 29.2 |
| Phase 2 | Gen-Adaptive (RFT-8B) + Klein-4B-DPO-v2 | 321 | 56.9 | 34.1 | 33.6 | 27.4 | 31.8 |
| ref | Oracle (frontier API) + Klein-4B-DPO-v2 | 750 | 55.7 | 33.7 | 33.9 | 26.0 | 31.2 |
| ref | No-Search + Klein-4B-DPO-v2 | 751 | 49.9 | 28.2 | 26.3 | 20.6 | 25.0 |
Search-augmented generation you can reproduce without an API key
Reproducing this kind of work usually means paying for a search engine and a fleet of generators — and watching your results drift as those services change underneath you. We froze the whole thing.
We release SearchGen-20K (20,839 world-knowledge-grounded prompts across 12 failure categories and 22 domains, a mean of 5.2 knowledge gaps per prompt), the co-training corpus (90,452 reasoning traces, 281,925 generations), and SearchGen-Corpus-1M — 145,642 archived image and web search sessions, 559,973 unique URLs, and 370,733 cached downloads.
Because every search is pre-executed and frozen, you can replay the entire pipeline offline. No live API keys. No result drift. An expensive research workflow becomes a stable substrate for preference learning, reward modeling, search-policy design, and retrieval studies.
A flywheel — and a principle bigger than search
Our recipe is deliberately minimal: one teaching pass, one recalibration pass. Even so, it improves monotonically — which means it can repeat. Each cycle pushes the generator's boundary further out and tightens the search policy further in, converging toward a system where only genuinely contextual knowledge ever triggers a lookup. That's a recursive self-improvement flywheel for world-knowledge-grounded generation.
The tempting objection is that bigger models will simply learn everything. They won't. Training data is finite; the world is not. No model, at any scale, can hold events after its cutoff, entities too rare for any dataset, or culture that keeps evolving. The boundary shifts outward with scale — it never disappears. Co-training finds where it lies for any generator, at any scale.
And search is only the first tool. The same gate–filter–integrate discipline governs when to invoke any tool — image editing, render-as-code, 3D-asset retrieval, structural control. Each fills a different slice of what a generator can't be taught. The knowledge boundary is a general principle for tool use, and the released harness is built to explore it.
Cite this work
@article{searchgen,
title={Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation},
author={Wang, Haozhe and Feng, Weijia and Yu, Jinpeng and Liu, Che and Nie, Ping and Lin, Fangzhen and Liu, Jiaming and Huang, Ruihua and Lin, Jimmy and Chen, Wenhu and others},
journal={arXiv preprint arXiv:2607.05382},
year={2026}
}