Skip to content

Writing an Illustration Brief for an Image Model

10 min read · updated August 4, 2026

Most image prompts fail because they describe a subject and leave every other decision to the model. A brief specifies six things — subject, action, setting, framing, light and finish — and an art director’s vocabulary for each, because those words carry far more information per token than adjectives like “beautiful”.

A brief is not a prompt

A prompt is what you type. A brief is the set of decisions the prompt encodes, and it exists whether or not you wrote it down — if you did not decide the camera height, something decided it for you, and it will be different in the next generation.

The practical consequence is that briefs are reusable and prompts are not. Write the brief once for a series of illustrations and the images share a look, because the same six decisions were made the same way. Write prompts one at a time and you get a set of individually reasonable images that plainly do not belong together, which is the single most common failure in machine-illustrated publications.

Everything below is language, and it works because these models learn from captioned images. The captions were written by people describing pictures, so the descriptive vocabulary that photographers, art directors and museum cataloguers use is over-represented and carries real signal. Vague praise is not in that vocabulary and does almost nothing.

The six slots

SlotDescription
SubjectThe one thing the image is of, with the two or three attributes that matter and no more. Long subject descriptions compete with themselves; if everything is specified, the parts get blended rather than all satisfied.
Action or stateWhat the subject is doing, or explicitly at rest. Unstated action produces the posed catalogue photograph, which is the default for almost every subject.
SettingWhere, and how much of it is visible. Includes what is behind the subject — a stated background is the cheapest way to stop the model inventing a busy one.
FramingShot size, camera height, angle, lens character, depth of field. This is the slot people skip and the one that most changes whether an image looks directed.
LightDirection, quality, colour, time of day. Light is what makes a set of images feel like one commission.
Medium and finishPhotograph, ink drawing, gouache, screen print, 3D render, and the qualities of that medium. Name a medium and its physical properties rather than an artist's name.

Art-direction vocabulary that carries information

Each of these terms specifies a decision. That is why they work: they narrow the distribution to images a person would have described that way.

TermDescription
Wide / medium / close / extreme closeShot size. How much of the frame the subject occupies — the single most reliable compositional control available.
Low angle / eye level / high angle / overheadCamera height relative to the subject. Low reads as dominance, high as vulnerability or as a diagram.
Three-quarter viewSubject turned partly away from camera. The default portrait angle, and worth naming because a model left to itself often produces a flat frontal one.
Shallow depth of fieldOnly the subject plane is sharp. Separates subject from background without needing an empty background.
Deep focusEverything sharp front to back. Reads as documentary, technical or observational.
Rim light / backlightLight behind the subject, outlining its edge. Produces separation and drama in one word.
Soft / diffused lightLarge source, gradual shadow edges. Reads as calm, editorial, expensive.
Hard lightSmall source, sharp-edged shadows. Reads as harsh, sunlit, mid-century.
High key / low keyOverall tonal placement: mostly bright with weak shadows, or mostly dark with small highlights.
Negative spaceDeliberate empty area. Ask for it explicitly and say where, if text is going over the image.
Muted / desaturated / limited paletteColour restraint, and the strongest single lever for making a series cohere.
Flat colour, no gradientsA finish instruction. Far more effective than naming an illustration style, because it describes a property of the surface.
Grain / halation / paper texturePhysical artefacts of a medium. These are what make a stated medium read as that medium rather than as a filter.

Notice what is absent: artist names, style labels, and quality words. Artist names are a blunt instrument that bundles dozens of decisions you have not made, they are the part of this practice most contested on rights grounds, and several platforms restrict them — naming the properties you actually wanted is more precise anyway. Quality words like “masterpiece” and “highly detailed” specify nothing; the words above specify something. More on prompt construction for image models if you want the mechanics underneath.

What you cannot ask for by exclusion

Negatives are weak in a predictable way. Asking for an image without a particular element makes that element more likely in systems that do not have a separate negative-conditioning channel, because the words are in the conditioning text either way. Where a dedicated negative input exists it works better, but the general rule survives: state what you want present, not what you want absent.

  • Instead of “no text”, describe a surface that would not carry text: a plain wall, an unbranded object.
  • Instead of “no clutter”, name the background: a single flat colour, a swept studio backdrop, fog.
  • Instead of “not cartoonish”, name the medium and its physical properties: photographic, 50mm, natural light, visible grain.
  • Instead of “no extra people”, say how many people are in frame and what the rest of the frame contains.

Change one slot at a time

The most expensive mistake in image work is changing four things between generations and concluding that the fourth one fixed it. Since output varies between runs anyway, a multi-change comparison tells you nothing you can reuse.

  1. Fix the seed if the tool exposes one. This is the difference between comparing two prompts and comparing two random draws.
  2. Generate a baseline with all six slots filled, however roughly.
  3. Change exactly one slot. Keep both images and label them with the change.
  4. Decide whether the change did what you expected. Where it did nothing, that phrasing is not a control in this system and you should stop paying tokens for it.
  5. Keep a file of phrasings that reliably move an output and the ones that never did. After twenty generations that file, not the prompt, is the asset.
Behaviour varies between models and between versions of the same model, and a phrasing that controls one may do nothing in another. That is exactly why the loop above ends in a file you maintain rather than in a list of magic words copied from somewhere else.

The brief template

ILLUSTRATION BRIEF

Purpose:      what the image has to do for the reader, in one sentence
Placement:    where it appears, at what size, on what background
Text over it: yes/no; if yes, where the negative space must be

SUBJECT   one thing, two or three attributes
ACTION    what it is doing, or "at rest"
SETTING   where, how much is visible, what is behind
FRAMING   shot size / camera height / angle / depth of field
LIGHT     direction / quality / colour / time of day
FINISH    medium and its physical properties

Series constraints (identical across every image in this set):
  palette, light quality, finish, camera height

Must not:  anything that would make it unusable — recognisable real
           people, real brands, depictions of real events

PROMPT (assembled from the slots above)
"Medium close, three-quarter view of a warehouse supervisor checking a
paper manifest, standing at a loading bay door. Deep warehouse behind,
out of focus. Eye level, 50mm, shallow depth of field. Hard morning
light from camera left, long shadow across the floor. Photographic,
muted palette, visible grain. Negative space in the upper right for a
headline."

Purpose and placement sit above the slots because they decide several of them. An image that will carry a headline needs stated negative space; an image that will appear 200 pixels wide should not be briefed with fine detail that will not survive the resize.

Before you use the image

  1. Look at hands, text, reflections, repeated patterns and the edges of the frame. These are where the recognisable artefacts concentrate, and a reader who spots one stops trusting the article.
  2. Check nothing in frame is a real brand, a real trademark or a recognisable real person. This is a rights and defamation question, not an aesthetic one.
  3. Check the image does not depict a real event or place in a way that implies documentation. A generated photograph of a real news scene is a different kind of object from a generated illustration.
  4. Confirm your licence covers this use, at this tier, in this medium — the questions are on the platform terms page.
  5. Label it in the caption. The image travels without the article, so the article’s disclosure does not follow it; see where disclosure goes.
  6. Record the prompt, model, version and date alongside the asset. You will need it if anyone asks, and you will need it to make the second image in the series match.