AI Reference Images: How to Control Style, Composition and Subject Consistency
AI reference images separate style, composition and subject controls, helping creative teams direct repeatable visuals without expecting exact copies.

AI reference images are visual inputs that give an image-generation system constraints beyond a written prompt. They do not all perform the same job: a style reference can steer color, texture, lighting, mood, and overall aesthetic, while a composition or structure reference can guide placement, outlines, and depth. In systems that expose those controls separately, a team can tune the look and the layout independently rather than asking one source image to do everything.
That distinction matters in production. A campaign may need the color and material language of an approved key visual, the layout of an existing placement, and the same product or character across a set of new scenes. Treat those as separate constraints whenever the tool allows it. A reference is best understood as weighted guidance, not an instruction to duplicate pixels.
Reference images control three different visual jobs
The first job is style: the visual treatment of a new image. A style image can inform palette, contrast, surface texture, lighting character, mood, and the broader aesthetic without requiring the output to retain the source image’s objects or framing. This is useful when an art director has established a recognizable visual language but needs fresh scenes, subjects, or copy space.
The second job is composition: where major elements sit and how the image is organized in space. Composition references can guide object placement, silhouettes or outlines, and depth relationships. Use one when an approved layout has already solved a practical problem: a product needs to occupy the lower third, a person needs to face into a headline area, or foreground and background need a particular separation.
The third job is subject identity. Here the target is not a general look or a single arrangement but a recurring person, product, animal companion, object, or location. Adobe describes reusable elements for characters, locations, and objects, while Google’s subject-customization documentation frames reference images as few-shot guidance for products, people, and animal companions. Both approaches point to the same production principle: identify what must recur, then provide evidence that defines it rather than hoping a style image will preserve it.
These categories can overlap in one photograph, but they should not be confused. A portrait may have an attractive grade, a useful pose, and a recognizable face. If the request is “new model, same editorial lighting,” it is a style problem. If it is “same product, new setting,” it is a subject problem. If it is “keep the figure, horizon, and negative space in these positions,” it is a composition problem. Naming the job before generation makes review feedback more actionable.
Match the reference type to the constraint
Start with the most consequential constraint, not the most attractive image in the folder. For a stylistic brief, choose a reference with legible color, lighting, texture, and tonal decisions. Avoid treating incidental content in that image as part of the brief unless it actually is. A reference selected only because it contains the right product may bring an unwanted mood or camera treatment into the output.
For layout, choose an image whose geometry is already close to the delivery need. Crop matters: a vertical source can establish a different balance of subject and empty space from a wide banner, even if both show the same scene. Decide the intended output format before evaluating a composition reference, then state the subject, setting, and desired changes in the prompt.
An ordinary composition reference is not always precise enough. Research on ControlNet describes spatial conditioning inputs such as edges, depth maps, segmentation, human pose, normals, and line drawings. These turn selected visual properties into additional signals for a text-to-image diffusion model, addressing the difficulty of specifying exact poses, shapes, and layouts through text alone.
Choose the structural representation that matches the decision being protected. A pose input is suited to limb placement and body orientation. An edge or line drawing is more useful when a silhouette or object boundary matters. A depth map can help preserve near-versus-far relationships, while segmentation is relevant when regions of the image need to remain distinct. The aim is not to add controls by default; it is to supply the least ambiguous condition for the visual requirement.
- Need a campaign look: use a style reference and describe the new scene in text.
- Need an approved placement: use a composition reference, then set an appropriate strength.
- Need a repeatable product or character: use subject references, ideally from more than one view.
- Need exact spatial discipline: use a compatible structural condition, such as pose, edges, depth, or segmentation, where the workflow supports it.
Composition strength and variation
Reference strength is an art-direction decision, not merely a technical setting. In composition-reference workflows, raising strength makes the generated image adhere more closely to the source layout; lowering it gives the model more room to reinterpret the arrangement. Adobe documents this as a fidelity-versus-flexibility trade-off.
High strength is appropriate when the approved composition carries production value. That might be a retail placement with reserved copy space, a social format where the product must read at thumbnail size, or a sequence of frames designed around a fixed focal point. It also narrows the range of useful alternatives. If the source layout is awkward for the newly requested scene, a strong setting can preserve that problem along with the desired structure.
Lower strength is useful for exploration. It lets a team retain a broad spatial cue—such as a close foreground subject against a deep background—while inviting alternative camera relationships and environmental detail. Reviewers should expect more layout drift at this setting. That is the intended exchange, not necessarily a generation failure.
A campaign variation sequence
- Begin with one approved composition that already accommodates the intended crop and text-safe area.
- Use it as the composition reference at a relatively high setting to establish the core layout in a first set of outputs.
- Keep the recurring product or character identified through separate subject references where available, and use a style reference only if the campaign look also needs to carry across.
- For a second set, lower composition strength while retaining the same prompt intent. This produces controlled alternatives rather than a wholly unrelated concept set.
- Compare outputs against the original decision: placement, subject recognition, usable negative space, and the visual treatment. Adjust one control at a time when diagnosing drift.
This sequence makes approval language clearer. Instead of “make it more like the reference,” a reviewer can ask for tighter layout adherence, stronger subject recognition, or a different stylistic treatment. Those requests point to different inputs and settings.

Build reusable subjects from multiple views
A single image rarely defines every feature a later generation may need. The front of a product says little about its side profile; a face-on character portrait gives weak evidence for a three-quarter or rear view. Adobe recommends multiple angles to better define an element, and its Elements workflow is designed around reusable characters, locations, and objects. Google likewise supports multiple subject images for customization.
Build the asset pack around the views needed to distinguish each subject. For products, that may mean front, side, and rear views, along with a detail establishing a distinctive material or marking; for characters, face, profile, full body, and an angle showing a defining garment or accessory. Location references should instead emphasize the durable features the brief requires the model to recognize, not near-duplicates.
The prompt should still identify the referenced subject and state its defining attributes. Google’s documentation specifically recommends identifying the referenced subject in the prompt. Do not make reviewers infer whether a changed result is a deliberate redesign or an unwanted identity shift: name the subject and call out the features that cannot change for the use case.
Keep the brief scoped. “Use the referenced bottle” is different from “use the referenced bottle with its same label placement and shape, photographed from a low three-quarter angle on a wet stone surface.” The first protects identity. The second combines identity, geometry, and a new scene. When several constraints are in play, label them in the production request so the operator knows which reference is carrying which job.
Resolve conflicts through iteration
A reference does not guarantee exact replication. Research on conditioning systems notes that generated details can fail to align with an input condition, particularly where the reference contains complex shapes, low-level detail, or competing visual information. Treat the reference as a weighted constraint from the outset, with a review loop built around crop, prompt, strength, and conditioning changes.
Conflicts commonly arise because the prompt asks for a scene that resists the source geometry, because the crop changes the visual balance, or because style, composition, and subject signals are asking for incompatible outcomes. A rigid composition may not accommodate a new pose. A dense source image may have too much incidental detail to serve as a clean structural cue. More references are not automatically better if each one introduces another unranked instruction.
Use a short diagnostic loop. First, identify the failure in plain visual terms: subject identity drift, incorrect pose, lost negative space, unwanted lighting, or altered proportions. Next, identify the relevant control. Re-crop or replace a composition source for placement issues; adjust strength for too much or too little layout adherence; simplify the prompt when it conflicts with the input; switch to a structural condition when exact geometry is the requirement. Then regenerate and compare against the same approval criteria.
This approach also helps at handoff. Save the reference assets, prompt, crop, selected settings, and the reason each input was chosen alongside the approved result. That record is more useful than a vague instruction to match a previous image, especially when another producer must make format variants or extend a release package later.
Frequently Asked Questions
Can one image control both style and layout?
It can contain cues for both, but separate style and composition controls are preferable when available. They let you preserve a visual treatment while changing the layout, or preserve placement while introducing a different look.
When should I use multiple subject images?
When one image cannot establish a subject’s features or the subject must appear from different angles, provide multiple views. The approach is especially useful for products, characters, and other assets intended to recur across a set.
What does reference strength change?
For composition references, higher strength generally produces closer adherence to the source layout. Lower strength permits more reinterpretation, which can be useful for controlled variation but increases the chance of layout drift.
When is structural conditioning better than an ordinary reference?
Use structural conditioning when a precise pose, silhouette, depth relationship, edge pattern, or segmented region matters more than the overall appearance of the source image. Pose, edges, depth maps, and segmentation each express a different kind of spatial constraint.
Why does the output still drift from the reference?
Complex detail, ambiguous inputs, changed crops, and competing prompt or reference signals can all weaken alignment. Narrow the problem, adjust the relevant control, and judge the next generation against the specific constraint that matters most.



