New to Marqeable? See how it generates leads and wins customers. See the platform

GPT Image 2.5 Reference Images: The Brand Consistency Lever Marketers Have Been Waiting For

For two years the standard complaint about AI images in marketing has been the same: the model has no memory of your brand. Every generation starts from zero, so every generation is a roll of the dice on palette, style, and whether your product still looks like your product. Teams have tried to fix it by writing longer prompts. Longer prompts do not fix it, because the problem is not that the model was not told; it is that it was told in words, and words are a lossy way to describe a picture.

Reference images are the fix that was always going to work, and ChatGPT Images 2.5, released September 8, 2026, made it the headline. OpenAI’s own summary leads with it: the model “is better at preserving the subjects in your reference photos, and follows editing instructions more reliably across multiple turns.” This post is about what that means in practice: what a reference set is, how to build one from what you already have, what it controls and what it still does not, what it costs, and why it finally makes a consistent set of brand images something a team of one can produce.

What actually changed

References were already possible. GPT Image 2 accepted image inputs in the API’s edits endpoint and in ChatGPT, and reviewers found it held a subject better than Google’s Nano Banana 2 in head-to-head tests. The weakness was persistence: the reference held on the first generation and drifted on the edits. Ask for the product on a desk, fine. Ask to warm the lighting, and the label changed. Ask to move it left, and the bottle got taller.

Images 2.5 targets exactly that. The claim is a subject that survives the setting change, the style change, and the third round of “just move the shadow.” Simon Willison’s launch-day test, inserting a new character into an existing chart while the chart’s data stayed intact, is the behavior in miniature: the reference is not inspiration, it is a constraint.

For marketing, the consequence is bigger than “the product looks right.” It means a set is possible. Ten images, one product, ten scenes, and the product is identical in all ten. Or thirty social illustrations in a style you chose once, from one sample, rather than thirty prompts that each describe the style from scratch and land somewhere nearby.

What a reference set is

A reference set is three to five images you attach to every generation, chosen so that together they say “this is what our images look like” better than any paragraph could. Ours, and most we have seen work, break down like this:

ReferenceWhat it anchorsWhere it comes from
The subjectYour product, your mascot, a recurring character, a founderA clean photo or your best existing render; one, not five angles
A photo-style sampleLighting, depth of field, color grade, the “feel” of your photographyYour best published photo, or one generated image you approved
An illustrated-style sampleLine weight, flatness, texture, how abstract ideas get drawnSame: one approved image, not a mood board
A palette swatchYour colors, with rough proportionsA simple image: your primary at 60%, secondary at 30%, accent at 10%
Optional: a layout sampleWhere the subject sits, where copy space livesOne tile from your standard social format

Three notes on choosing them. First, references should be outputs you would be happy to ship, not aspirational images from another brand; the model will reproduce what it sees, including the parts you did not mean. Second, one reference per role. Two product photos from different angles confuse the subject; two style samples that disagree average into mush. Third, the palette swatch matters more than people expect. Palette is the single most visible drift across an AI-generated set, and a swatch anchors it in a way hex codes in the prompt never quite do.

Build one in an hour from what you already have

You do not need a photo shoot. Most brands have the raw material on their own website.

  1. Pull your last twenty published images into one folder and view them as a grid. This is the same 20-image audit we recommend for diagnosing consistency, and it doubles as reference scouting.
  2. Pick the one photo and the one illustration you would be proud to repeat. Not the best-performing post; the one that looks most like the brand you are trying to be.
  3. Isolate the subject. One clean image of the product or person on a simple background. If you do not have one, generate it once, carefully, with Sunburst if you have API access, and approve it. That approved render becomes the subject reference for everything after.
  4. Make the swatch. Any image editor, three rectangles, your colors, roughly in the proportions you use them. Our free image editor does it in a minute.
  5. Write the ten-line instruction block that goes alongside the references: imagery kind for this channel, the two or three subject rules, and the never-list. The references answer “what does it look like”; the block answers “what is it not allowed to be.” Turning Brand Guidelines into an AI Image Brief has the template.

Then test: generate five images for five different briefs with the same set and block attached. Grid them. If they read as siblings, you are done. If one reference is dominating (every image looks like a remix of the photo sample), swap it for a plainer one.

What holds, and what still needs the prompt

Be precise about the boundary, because the failure mode of references is expecting them to do everything.

References hold: the subject’s shape, proportions, label, face; the palette; the lighting mood and texture; the illustration treatment. In 2.5, they are designed to hold through edits as well.

References do not decide: layout and composition (where things sit), text (what the headline says and where), scene content (what is happening), or aspect ratio. Those remain the prompt’s job, and 2.5’s Sketch tool exists precisely because layout is the thing words handle worst. A reference set plus a sketched layout plus a short prompt is the strongest combination available today.

References can over-hold. Attach the same photo sample to fifty generations and you get fifty images that are visibly that photo’s cousins. The fix is a pool larger than the set: six to eight approved images, with three to five rotated in per generation. The set stays consistent; the individual images stop looking cloned.

Why this beats the guidelines PDF. A brand guideline is written for a designer: hex codes, clear space, logo lockups, adjectives. A model reads “clean, confident, modern” through its training data, not through your brand. An image is not ambiguous. Three approved outputs plus a swatch tell the model more about your brand than forty pages, and they tell it in the only language it does not have to translate.

What it costs

In the API, references are billed as image input tokens ($8 per million on both 2.5 models, $2 cached, per OpenAI’s pricing). In our own production runs on GPT Image 2, attaching a reference set of a few images added roughly 40% to the per-image input cost, with no measurable change in generation time. That sounds like a lot until you count what it replaces: the regenerations. A team that needs three tries to get an on-brand image without references and one or two with them is spending less, not more, and getting a consistent set as a by-product. The only cost number worth tracking is cost per accepted image, and references push it down.

In ChatGPT, references are simply uploads; they do not cost money, but they cost the thing ChatGPT is worst at, which is remembering. Every new chat needs the set attached again. Keep them in one folder with the instruction block and treat “paste the set” as step one of every image task.

The one thing references cannot do: show up on their own

This is the boundary that matters most for a small team. References work only when they are present, and in a chat window they are present only when a person remembers. The consistent set survives exactly as long as the discipline does, which in our experience is until the first deadline.

The durable version is a reference set that is attached by default: stored once with the brand, selected per generation, rotated so no single image dominates, and recorded with each output so you can see, a month later, which references produced which image. That is what Marqeable’s brand board is. You approve a handful of images (or let the system propose them from your website), they ride along with every generation inside every email, post, and page, and each generated image carries its provenance: model, prompt, references. Nobody pastes anything. We are in private beta with a small early cohort: get early access if you want to see it on your brand.

Frequently asked questions

Can I use a competitor’s or a stock image as a reference?

Technically yes; practically no. The model reproduces what it sees, a reference you do not own can carry a style or a subject you have no right to reproduce, and the commercial-use rules get murky fast. Use your own approved outputs.

Does the reference need to be high resolution?

No. It needs to be clean and unambiguous. A sharp 1024-pixel image of the subject on a plain background beats a 4K lifestyle shot with six things in it.

Will references make every image look the same?

Only if you use the same references every time. Keep a pool of six to eight and rotate three to five. Consistency is the goal; sameness is the failure mode just past it.

Does this work with Midjourney or Nano Banana too?

Both accept image inputs, and Midjourney has had style and character references for a while. The published tests put GPT ahead on holding a specific subject, and 2.5 is built to widen that. Run the six tests with your own set before you commit.

The bottom line

Brand consistency in AI images was never going to come from better adjectives. It comes from showing the model what you mean, and ChatGPT Images 2.5 is the first mainstream release built around that idea holding through edits. Build the set this week: one subject, one photo sample, one illustrated sample, one swatch, ten lines of rules. Attach it to everything. Then solve the only remaining problem, which is that a set nobody attaches is a folder, and make it show up on its own.


Marqeable runs your campaigns, answers every visitor, text, and email in seconds, and turns them into booked jobs and meetings - even at 9pm on a Saturday. We’re in private beta with a small early cohort. Get early access

Marqeable
© 2026 Marqeable. All rights reserved.