Writing an Illustration Brief for an Image Model
10 min read · updated August 4, 2026
Most image prompts fail because they describe a subject and leave every other decision to the model. A brief specifies six things — subject, action, setting, framing, light and finish — and an art director’s vocabulary for each, because those words carry far more information per token than adjectives like “beautiful”.
A brief is not a prompt
A prompt is what you type. A brief is the set of decisions the prompt encodes, and it exists whether or not you wrote it down — if you did not decide the camera height, something decided it for you, and it will be different in the next generation.
The practical consequence is that briefs are reusable and prompts are not. Write the brief once for a series of illustrations and the images share a look, because the same six decisions were made the same way. Write prompts one at a time and you get a set of individually reasonable images that plainly do not belong together, which is the single most common failure in machine-illustrated publications.
Everything below is language, and it works because these models learn from captioned images. The captions were written by people describing pictures, so the descriptive vocabulary that photographers, art directors and museum cataloguers use is over-represented and carries real signal. Vague praise is not in that vocabulary and does almost nothing.
The six slots
| Slot | Description |
|---|---|
| Subject | The one thing the image is of, with the two or three attributes that matter and no more. Long subject descriptions compete with themselves; if everything is specified, the parts get blended rather than all satisfied. |
| Action or state | What the subject is doing, or explicitly at rest. Unstated action produces the posed catalogue photograph, which is the default for almost every subject. |
| Setting | Where, and how much of it is visible. Includes what is behind the subject — a stated background is the cheapest way to stop the model inventing a busy one. |
| Framing | Shot size, camera height, angle, lens character, depth of field. This is the slot people skip and the one that most changes whether an image looks directed. |
| Light | Direction, quality, colour, time of day. Light is what makes a set of images feel like one commission. |
| Medium and finish | Photograph, ink drawing, gouache, screen print, 3D render, and the qualities of that medium. Name a medium and its physical properties rather than an artist's name. |
Art-direction vocabulary that carries information
Each of these terms specifies a decision. That is why they work: they narrow the distribution to images a person would have described that way.
| Term | Description |
|---|---|
| Wide / medium / close / extreme close | Shot size. How much of the frame the subject occupies — the single most reliable compositional control available. |
| Low angle / eye level / high angle / overhead | Camera height relative to the subject. Low reads as dominance, high as vulnerability or as a diagram. |
| Three-quarter view | Subject turned partly away from camera. The default portrait angle, and worth naming because a model left to itself often produces a flat frontal one. |
| Shallow depth of field | Only the subject plane is sharp. Separates subject from background without needing an empty background. |
| Deep focus | Everything sharp front to back. Reads as documentary, technical or observational. |
| Rim light / backlight | Light behind the subject, outlining its edge. Produces separation and drama in one word. |
| Soft / diffused light | Large source, gradual shadow edges. Reads as calm, editorial, expensive. |
| Hard light | Small source, sharp-edged shadows. Reads as harsh, sunlit, mid-century. |
| High key / low key | Overall tonal placement: mostly bright with weak shadows, or mostly dark with small highlights. |
| Negative space | Deliberate empty area. Ask for it explicitly and say where, if text is going over the image. |
| Muted / desaturated / limited palette | Colour restraint, and the strongest single lever for making a series cohere. |
| Flat colour, no gradients | A finish instruction. Far more effective than naming an illustration style, because it describes a property of the surface. |
| Grain / halation / paper texture | Physical artefacts of a medium. These are what make a stated medium read as that medium rather than as a filter. |
Notice what is absent: artist names, style labels, and quality words. Artist names are a blunt instrument that bundles dozens of decisions you have not made, they are the part of this practice most contested on rights grounds, and several platforms restrict them — naming the properties you actually wanted is more precise anyway. Quality words like “masterpiece” and “highly detailed” specify nothing; the words above specify something. More on prompt construction for image models if you want the mechanics underneath.
What you cannot ask for by exclusion
Negatives are weak in a predictable way. Asking for an image without a particular element makes that element more likely in systems that do not have a separate negative-conditioning channel, because the words are in the conditioning text either way. Where a dedicated negative input exists it works better, but the general rule survives: state what you want present, not what you want absent.
- Instead of “no text”, describe a surface that would not carry text: a plain wall, an unbranded object.
- Instead of “no clutter”, name the background: a single flat colour, a swept studio backdrop, fog.
- Instead of “not cartoonish”, name the medium and its physical properties: photographic, 50mm, natural light, visible grain.
- Instead of “no extra people”, say how many people are in frame and what the rest of the frame contains.
Change one slot at a time
The most expensive mistake in image work is changing four things between generations and concluding that the fourth one fixed it. Since output varies between runs anyway, a multi-change comparison tells you nothing you can reuse.
- Fix the seed if the tool exposes one. This is the difference between comparing two prompts and comparing two random draws.
- Generate a baseline with all six slots filled, however roughly.
- Change exactly one slot. Keep both images and label them with the change.
- Decide whether the change did what you expected. Where it did nothing, that phrasing is not a control in this system and you should stop paying tokens for it.
- Keep a file of phrasings that reliably move an output and the ones that never did. After twenty generations that file, not the prompt, is the asset.
The brief template
ILLUSTRATION BRIEF
Purpose: what the image has to do for the reader, in one sentence
Placement: where it appears, at what size, on what background
Text over it: yes/no; if yes, where the negative space must be
SUBJECT one thing, two or three attributes
ACTION what it is doing, or "at rest"
SETTING where, how much is visible, what is behind
FRAMING shot size / camera height / angle / depth of field
LIGHT direction / quality / colour / time of day
FINISH medium and its physical properties
Series constraints (identical across every image in this set):
palette, light quality, finish, camera height
Must not: anything that would make it unusable — recognisable real
people, real brands, depictions of real events
PROMPT (assembled from the slots above)
"Medium close, three-quarter view of a warehouse supervisor checking a
paper manifest, standing at a loading bay door. Deep warehouse behind,
out of focus. Eye level, 50mm, shallow depth of field. Hard morning
light from camera left, long shadow across the floor. Photographic,
muted palette, visible grain. Negative space in the upper right for a
headline."Purpose and placement sit above the slots because they decide several of them. An image that will carry a headline needs stated negative space; an image that will appear 200 pixels wide should not be briefed with fine detail that will not survive the resize.
Before you use the image
- Look at hands, text, reflections, repeated patterns and the edges of the frame. These are where the recognisable artefacts concentrate, and a reader who spots one stops trusting the article.
- Check nothing in frame is a real brand, a real trademark or a recognisable real person. This is a rights and defamation question, not an aesthetic one.
- Check the image does not depict a real event or place in a way that implies documentation. A generated photograph of a real news scene is a different kind of object from a generated illustration.
- Confirm your licence covers this use, at this tier, in this medium — the questions are on the platform terms page.
- Label it in the caption. The image travels without the article, so the article’s disclosure does not follow it; see where disclosure goes.
- Record the prompt, model, version and date alongside the asset. You will need it if anyone asks, and you will need it to make the second image in the series match.