Check Top Image Generators Live Interactive Demo
Agentic Visual Generation

Search Beyond What Can Be Taught Evolving the Knowledge Boundary in Agentic Visual Generation

Haozhe Wang1 · Weijia Feng3 · Jinpeng Yu3 · Che Liu4 · Ping Nie2 · Fangzhen Lin1 · Jiaming Liu3 · Ruihua Huang3 · Jimmy Lin2 · Wenhu Chen2 · Cong Wei2
1 Hong Kong University of Science and Technology · 2 University of Waterloo · 3 Qwen Applications · 4 Imperial College London
✉ Corresponding authors: Jiaming Liu, Cong Wei

Image generators fabricate what they don't know. This one looks it up first — and knows when not to.

📄 arXiv 🌐 Project Page 💻 GitHub
🤗 SearchGen-20K 🤗 SearchGen-Corpus-1M 🤗 AgentGen Bench 🤗 LeaderBoard
AgentGen Bench Live Interactive Demo
Enjoying SearchGen? Help us by upvoting and starring the project.
Two paradigms for knowledge-hungry prompts: prompt rewrite vs SearchGen agentic retrieval.
Two paradigms for knowledge-hungry prompts. Left: prompt-rewriting inflates the text but still generates from stale weights. Right: SearchGen's agent fetches live web + visual context, then conditions the generator — grounding facts a model cannot know. (The animated agent trace is behind the Demo button above.)
TL;DR

The short version — teach what you can, search the rest

Modern image generators render gorgeously and lie fluently. Ask for the 2025 Osaka Expo mascot and you get a confident, wrong invention. The failure isn't the pixels — it's the knowledge. On AgentGen Bench, frontier open generators score just 21–28 out of 100 on search-intensive prompts — up to a ~40-point collapse that standard benchmarks never register.

Search is the obvious fix, the way an illustrator consults references. But naive search backfires: it corrupts prompts the generator already handled. The real problem is a knowledge boundary — the line between what a generator can learn and what it must look up. That line is generator-specific, it moves during training, and it can't be hand-drawn. It has to be discovered.

So we discover it, by co-training the generator and the search agent together. Below: how the collapse works, why naive search fails, what the knowledge boundary is, and how a co-trained 8B reasoner on a 4B generator matches a frontier reasoner on the same generator — jump to the finding cards. ↓

The Hook

They render beautifully. They just make things up.

Ask a frontier image model for the mascot of the 2025 Osaka Expo. You get a polished, confident fabrication. Ask for a historically accurate Spartan phalanx and you get anachronistic armor, rendered in exquisite detail.

The lighting is right. The composition is right. The world is wrong.

This is not a rendering failure. It's a world-knowledge bottleneck. Generators train on fixed corpora with hard knowledge cutoffs; user requests draw on new characters, regional symbols, niche typography, historical artifacts, and events that postdate training.

Worse, generators have no way to flag their own ignorance. They're trained to always output an image — never to say "I don't know what this looks like." So they guess, beautifully, every time.

The lighting is right. The composition is right. The world is wrong.
Miyazaki portrait
Without SearchGenerated portrait without search — a generic, wrong face.
With SearchGenerated portrait with search-augmented grounding.
CRISPR diagram
Without SearchGene-editing diagram without search — misspelled labels.
With SearchGene-editing diagram with search-augmented grounding.
Gallery of failure examples across the 12 failure categories.
Ask for a specific person, a labeled scientific diagram, or live data and today's best generators confidently fabricate. Left of each pair: no search — a generic face, misspelled gene-editing labels. Right: the same generator, this time with search-augmented grounding. Gallery: representative cases spanning the 12 failure categories.
So the obvious move is to let the model look things up. First, let's measure exactly how far it falls without help. ↓
Finding 1

A 40-point gap that no benchmark shows

● Finding 1

Generators that score comparably on standard prompts diverge by nearly 40 points when search-intensive world knowledge is required.

On prompts that need only what a model already learned, open and commercial generators land in the same band (67–75 out of 100). Turn to prompts that need live world knowledge, and the field splits: open generators crater to 21–28, while commercial systems with built-in search barely move. Existing benchmarks test rendering inside known concepts, so they never see this gap at all.

AgentGen Bench contains 751 prompts scored by the finalized 3d5 evaluator. The primary Overall9 metric averages nine applicable components and excludes Text Reference.

Per-generator quality on NoSearch vs Search-Intensive strata across twelve generators.
Generators score well on prompts they can answer from memory (gray). On prompts that require external knowledge (orange), every open generator collapses — up to a 40-point drop — while commercial systems with built-in search hold. The bottleneck is missing knowledge, not rendering skill.
Overall9 AgentGen Bench ranking of seventeen image generators on the full 751-prompt evaluation.
Overall AgentGen Bench ranking. GPT-Image-2 leads, followed by Qwen-Image-3-Max and Grok-Imagine-Image. Scores use Overall9 from the finalized 3d5 evaluator.

Full benchmark breakdown

Scores are on a 0–100 scale; higher is better. Overall9 averages the nine displayed components and excludes Text Reference.

Generator Overall Checklist Rubric Prompt Image quality Text rendering AI naturalness Composition Physical plausibility Visual reference
NoSearch parametric knowledge is sufficient
GPT-Image-278.984.882.779.373.492.367.686.581.270.9
Grok-Imagine-Image74.881.378.372.772.383.965.679.777.266.9
Qwen-Image-3-Max73.779.176.269.870.386.263.982.578.364.0
Nano Banana Pro73.378.476.370.270.785.665.280.377.163.0
Qwen-Image-2-Pro71.276.874.167.370.577.861.878.375.361.2
Qwen-Image-271.176.273.167.570.279.563.879.274.260.3
SeedDream-4.070.776.473.365.869.076.061.780.875.558.5
Qwen-Image70.074.972.065.569.261.161.777.078.660.1
SeedDream-4.569.376.272.865.567.374.059.076.873.459.6
Nano-Banana66.370.868.561.867.749.059.777.570.556.2
SenseNova-U162.868.165.258.564.358.356.272.265.151.1
Mage-Flow61.264.261.854.564.751.256.371.867.447.1
Flux.2-Klein-9B58.563.760.749.563.230.053.571.067.041.1
Flux.2-Klein-4B52.355.852.642.359.716.149.867.760.135.3
Bagel49.151.549.237.756.216.148.862.858.532.1
OmniGen247.449.146.133.556.47.847.361.460.330.6
Show-o234.432.730.921.047.22.438.049.840.919.9
GPT-Image-275.574.473.772.277.477.768.983.984.970.4
Qwen-Image-3-Max67.665.865.362.871.167.962.377.078.659.6
Grok-Imagine-Image66.765.964.662.869.565.761.775.677.859.7
Nano Banana Pro63.660.760.156.468.161.360.072.778.155.4
Qwen-Image-2-Pro57.553.953.148.365.355.157.368.972.846.8
Qwen-Image-254.149.949.644.162.550.154.366.271.442.0
SeedDream-4.554.151.750.845.261.850.450.865.969.342.1
SeedDream-4.049.645.645.139.358.837.250.962.568.036.8
Nano-Banana47.143.042.336.156.931.248.960.567.035.9
SenseNova-U131.629.327.621.639.515.832.242.548.421.0
Qwen-Image28.625.023.817.437.99.229.941.045.517.7
Mage-Flow28.423.021.715.738.79.735.040.050.216.2
Flux.2-Klein-9B26.923.822.415.136.15.829.938.347.716.7
Flux.2-Klein-4B23.920.118.611.533.03.428.135.744.812.7
Bagel22.218.717.712.430.42.225.032.438.114.2
OmniGen220.416.414.98.829.22.524.331.440.89.5
Show-o217.58.18.04.230.20.726.534.331.73.4
Knowledge-sensitive measures test whether requested facts are present. Results use present-only aggregation; coverage is reported in the interactive leaderboard.
If the disease is missing knowledge, the cure is search. Except the cure has a side effect. ↓
The Obvious Fix — and Why It Backfires

Search should help. Often it hurts.

An illustrator handed an unfamiliar brief looks up references before drawing. Give the generator the same move — a reasoner spots knowledge gaps, search fills them, the results feed generation — and you have agentic visual generation. Natural. And, done naively, harmful.

● Finding 2

Naive search actively degrades prompts the generator already handles.

Search everything, blindly, and every generator gets worse on prompts that never needed help. Qwen-Image-2 drops from 70.7 to 60.4 on the no-search stratum — a 14.6% relative loss on prompts it already aced.

Need Search?GeneratornNoSearchReasonedBlind
NoSearchQwen210070.776.560.4
Qwen19967.575.059.5
Klein-9B5857.966.452.3
VisualSearchQwen238437.249.145.3
Qwen138432.844.439.7
Klein-9B33829.639.436.0
TextualSearchQwen226322.934.132.1
Qwen126320.528.525.3
Klein-9B26418.926.624.0

Two distinct failures explain it. Concept corruption: search fires on something the model already knew, and the retrieved reference overrides correct internal knowledge — a gating failure, searching when it shouldn't. Copy effect: a reference carries so much detail that the generator copies it wholesale instead of borrowing the one missing fact — a filtering failure, keeping too much.

The cure and the disease depend entirely on the patient. Search helps only when the generator is actually missing something.
Blind retrieval leaking into outputs: a copied reference boat and a pasted search-result caption.
Search is not free. Fed raw, retrieved content leaks into the image — the model copies a reference boat verbatim (copy effect), or pastes a search-result's caption as the artwork's text (concept corruption). Naive retrieval corrupts the very prompts a model could have answered alone.
So why does the same tool rescue one prompt and wreck another? Because of a line most pipelines never see. ↓
The Key Idea — the hinge

The line between what a model learns and what it must look up

Some knowledge belongs inside the model. A character's canonical look, a flag's fixed geometry — stable, low-dimensional, learnable once and for all. Fire search for it and you only add noise.

Other knowledge belongs outside. It changes faster than retraining cycles, sits too deep in the long tail to ever learn reliably, or needs per-request reasoning. For this, search isn't a nice-to-have — it's structurally necessary. And the tail is enormous: 93.1% of the 31,537 entities in our data appear in just a single prompt. No feasible training set covers that.

✦ Insight

Some knowledge is internalizable and search should not fire for it; other knowledge is contextual and search is structurally necessary.

We call the divide between those two sets the knowledge boundary. Here's the twist that organizes the entire paper: the boundary is generator-specific, and it moves. As a generator learns, concepts migrate from "must search" to "already knows." A search policy tuned for a weak generator is wrong for a strong one.

So you can't hand-specify it. And you don't have to. The boundary is discoverable — it falls out of training the generator and the searcher together.

For a given generator, world knowledge splits in two: what it can absorb into its own weights (internalizable), and what must stay in the prompt as retrieved context (contextual). That split is the knowledge boundary — and it shifts every time the generator improves.
All world knowledge always needs search: fresh facts · rare entities · live data internalizable generate correctly · no search + co-training →
green = internalized · no search orange = fetched · agentic search dashed = knowledge boundary (moves with training)
The boundary moves outward — measured proof in the results ↓
A generator's knowledge has a boundary. Inside it (green) lives what the model can internalize — generate correctly with no search. Outside (orange) lies contextual knowledge that must be fetched. Co-training pushes the boundary outward, converting fetch-only prompts into internalizable ones; whatever remains outside is handled by the agent.
Method

Build a searcher that resists noise. Then move the boundary.

The method has two halves. First, a searcher that doesn't poison the generator. Second, a training loop that finds — and expands — the boundary.

The noise-resistant reasoner: gate, filter, integrate.

Gate

Decide whether to search. Only gaps rated critical or important trigger a query; the rest are dropped. At most three queries per prompt, or skip entirely. The defense against concept corruption.

Filter

Decide what to keep. Choose the reference that fills the specific gap with the least extraneous baggage. The defense against the copy effect.

Integrate

Decide how it enters. Route visual references through language, not raw pixels: "following Image 1, render the robe in teal and gold." The generator borrows exactly what's named — nothing else leaks in.

Teach-then-search co-training. The boundary only moves when the generator's weights change, so we act on both sides, in order.

Phase 0

Warm-start

Supervised finetuning gives an 8B reasoner (Qwen3-VL-8B) the gate–filter–integrate protocol. Competent, but generator-agnostic: it searches for anything that might be hard.

Phase 1

Teach the generator

Online iterative Diffusion-DPO feeds it search-augmented examples and reinforces its own best outputs. Two things happen at once: it internalizes stable knowledge (the boundary pushes outward), and it learns to use imperfect references without being dominated by them (noise-robustness).

Phase 2

Recalibrate the searcher

The generator is now stronger, so the search policy is stale. Rejection-sampling finetuning rewards trajectories where search actually helped the new generator and discards the rest. Search fires only for what remains genuinely contextual.

Phase 1 moves the boundary outward; Phase 2 moves the search policy inward to match. The model is never told where the line is — it learns it from which searches actually helped.
Two coupled training loops: a teach/internalize loop and a search/fetch loop.
Two coupled loops. Bottom: an agent gates each prompt, fetches and filters references only when a knowledge gap warrants it, then integrates them into an enriched prompt. Top: online DPO teaches the generator what it can internalize — expanding the knowledge boundary so the agent has less to fetch over time.
Finding 3 — the boundary is discoverable

Co-training discovers the boundary — an 8B reasoner matches a frontier oracle

● Finding 3

Co-training discovers the knowledge boundary: a calibrated 8B reasoner matches a frontier oracle on the same generator.

On Klein-4B, the full co-training progression climbs monotonically to 31.8 overall — matching the frontier Gemini oracle on the very same generator (31.2). Both phases pull their weight: teaching the generator adds +2.8, recalibrating the searcher adds +2.6.

It's also selective, which is the hard part. On prompts that don't need search, the calibrated reasoner lifts the no-search baseline from 49.9 to 56.9 — it learned when to stay quiet, exactly where naive search did damage. And it's generator-specific: a policy tuned for the strengthened generator scores worse on the base one, confirming the boundary is a joint property of the pair, not a fixed property of the prompt.

The headline is the calibration-per-dollar. An 8B reasoner on a 4B generator reaches the same quality as a frontier commercial reasoner on that generator — and the full recalibration cycle fits in 4×8 GPU-hours.

One honest caveat: this is a comparison of reasoners on a fixed 4B generator, not a claim of frontier image quality. In absolute terms, 31.8 remains below GPT-Image-2's finalized full-benchmark Overall9 score of 76.0.

Frontier-reasoner calibration on the same generator — at a fraction of the reasoner's size and training cost.
Co-training progression chart (panel a) and no-search quality distribution shift (panel b).
(a) Co-training compounds. Each round — reasoner SFT, generator DPO, reasoner RFT — lifts quality across all three difficulty tiers, for both Klein-4B and Bagel-7B. (b) Proof the boundary moved: after co-training, the distribution of per-prompt no-search quality shifts right — more prompts now clear a given quality bar without any search. The shift holds for both Klein-4B and Bagel-7B.
● same generator · ⅟ reasoner cost
PhaseConfignNoSearchEasyMediumHardOverall
Phase 0Gen-Agnostic (SFT-8B) + Klein-4B60254.628.929.221.226.4
Phase 1Gen-Agnostic (SFT-8B) + Klein-4B-DPO-v232154.031.831.124.729.2
Phase 2Gen-Adaptive (RFT-8B) + Klein-4B-DPO-v232156.934.133.627.431.8
refOracle (frontier API) + Klein-4B-DPO-v275055.733.733.926.031.2
refNo-Search + Klein-4B-DPO-v275149.928.226.320.625.0
On a fixed Klein-4B-DPO generator, the co-trained 8B reasoner (31.8) matches the frontier oracle reasoner (31.2) — at a fraction of the reasoner cost. Training moved most of the knowledge inside the boundary.
Naming note (maps to the paper): the paper's condition macros render as No Search → Blind Search (SFT-8B, generator-agnostic) → Generator-Adaptive Search (RFT-8B). We relabel "Blind Search" → "Gen-Agnostic" on the site only, because the paper reuses "blind" for the harmful search-everything policy in the Hook — two different meanings that must not collide on one page. The SFT-8B reasoner is generator-agnostic; it gates like any reasoner but is not calibrated to a specific generator, and is not the naive "search-everything" policy from Finding 2.
None of this is reproducible if you can't replay the searches. So we released them. ↓
The Harness — released, and replayable offline

Search-augmented generation you can reproduce without an API key

Reproducing this kind of work usually means paying for a search engine and a fleet of generators — and watching your results drift as those services change underneath you. We froze the whole thing.

We release SearchGen-20K (20,839 world-knowledge-grounded prompts across 12 failure categories and 22 domains, a mean of 5.2 knowledge gaps per prompt), the co-training corpus (90,452 reasoning traces, 281,925 generations), and SearchGen-Corpus-1M145,642 archived image and web search sessions, 559,973 unique URLs, and 370,733 cached downloads.

Because every search is pre-executed and frozen, you can replay the entire pipeline offline. No live API keys. No result drift. An expensive research workflow becomes a stable substrate for preference learning, reward modeling, search-policy design, and retrieval studies.

20,839
prompts
145,642
search sessions
90,452
agentic traces
281,925
generations
Everything released: the prompts, the searches the agent ran, its full reasoning traces, and every generation used for training and evaluation.
A frozen web. Same query, same result, every run.
Treemap of 22 real-world domains covered by SearchGen-20K.
SearchGen-20K spans 22 real-world domains — from People & Professions and Screen & Performance Media down a long tail through Science, Fashion, and Infrastructure — mirroring how people actually prompt.
Looking Forward

A flywheel — and a principle bigger than search

Our recipe is deliberately minimal: one teaching pass, one recalibration pass. Even so, it improves monotonically — which means it can repeat. Each cycle pushes the generator's boundary further out and tightens the search policy further in, converging toward a system where only genuinely contextual knowledge ever triggers a lookup. That's a recursive self-improvement flywheel for world-knowledge-grounded generation.

The tempting objection is that bigger models will simply learn everything. They won't. Training data is finite; the world is not. No model, at any scale, can hold events after its cutoff, entities too rare for any dataset, or culture that keeps evolving. The boundary shifts outward with scale — it never disappears. Co-training finds where it lies for any generator, at any scale.

And search is only the first tool. The same gate–filter–integrate discipline governs when to invoke any tool — image editing, render-as-code, 3D-asset retrieval, structural control. Each fills a different slice of what a generator can't be taught. The knowledge boundary is a general principle for tool use, and the released harness is built to explore it.

The question isn't how to build a model that knows everything. It's how to build one that knows what it doesn't know.
Citation

Cite this work

BibTeX
@article{searchgen,
  title={Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation},
  author={Wang, Haozhe and Feng, Weijia and Yu, Jinpeng and Liu, Che and Nie, Ping and Lin, Fangzhen and Liu, Jiaming and Huang, Ruihua and Lin, Jimmy and Chen, Wenhu and others},
  journal={arXiv preprint arXiv:2607.05382},
  year={2026}
}
Enjoying SearchGen? Help us by upvoting and starring the project.