←Back to Journal
02 / Entry· 5 min read

OpenAI GPT Image 2.5:双模型生图与官方最佳 Prompt 全拆解

拆解 OpenAI 最新 gpt-image-2.5 双模型(Sunburst / Flare)的定位、API 参数、官方 prompting 指南与可复用的 Prompt 模板。

🔊 系统朗读

前言

OpenAI 已正式上线 GPT Image 2.5 系列,模型 ID 为 gpt-image-2.5-sunburst 和 gpt-image-2.5-flare。这不再是单模型升级,而是拆成了质量优先和速度优先两条产品线。官方同步发布了完整的 Image Prompting Guide,把生图和修图的 Prompt 写法系统化了。

本文基于 OpenAI 官方文档(2026 年 9 月核验),覆盖模型选型、API 参数、定价、官方最佳实践,以及一套可直接复用的 Prompt 模板。

一、双模型定位

GPT Image 2.5 不是一个模型,而是两个:

模型ID定位适用场景
Sunburstgpt-image-2.5-sunburst质量优先,编辑精度最高复杂修图、精确保留、品牌资产、高质量输出
Flaregpt-image-2.5-flare速度优先,质量对标 GPT Image 2日常生图、快速迭代、延迟敏感场景

两者的默认 snapshot 都是 2026-09-08,输入支持文本+图像,输出为图像。

官方选型建议

官方给了一套很实用的迁移路径:

  • 已有 GPT Image 2 工作流且质量够用 → 先测 Flare,看能否在保住质量的同时降延迟。
  • GPT Image 2 质量不够的复杂场景 → 先测 Sunburst,确认质量达标后再试 Flare 能否替代。
  • 新工作流 → 速度优先选 Flare,质量优先选 Sunburst。质量达标后再考虑降档到 Flare 换延迟。

核心原则:先用 Sunburst 确立质量基线,再用 Flare 测延迟收益。 同一个 quality 标签在两个模型上不代表同样的画质或响应时间。

二、API 参数一览

text
model       : gpt-image-2.5-sunburst | gpt-image-2.5-flare
quality     : auto (默认) | low | medium | high | xhigh | max
size        : auto 或自定义,常见值见下
background  : auto | opaque | transparent
output_format : png | jpeg | webp
n           : 一次生成多张
partial_images : 0-3,流式输出时返回中间帧

常用尺寸

场景尺寸
正方形1024x1024
横版1536x1024
竖版1024x1536
2K 正方形2048x2048
2K 横版2048x1152
4K 横版3840x2160
4K 竖版2160x3840

自定义尺寸约束:每边 ≤ 3840px、两边均为 16 的倍数、长宽比 ≤ 3:1、总像素在 655,360 ~ 8,294,400 之间。超过 2560x1440 的输出属实验性。

Quality 档位策略

官方建议:先用 medium 或 high 建立基线。小字、密集信息、多字体场景建议至少 high。xhigh 和 max 只在有明确质量缺口且延迟预算允许时使用——更高档位不保证每张图都更好。

透明资产:显式传 background="transparent",输出用 PNG/WebP,并检查解码后的 alpha 通道(发丝、玻璃、阴影、物体边缘)。

三、定价

两个模型定价相同(截至核验日):

指标价格
文本输入$5 / 1M tokens
缓存命中文本输入$1.25 / 1M tokens
图像输入$8 / 1M tokens
缓存命中图像输入$2 / 1M tokens
图像输出$30 / 1M tokens

注意:token 费率与 GPT Image 2 一致,但 GPT Image 2 的 token 计算器不适用于 2.5。Rate limits 按 Tier 从 100K TPM / 5 IPM 到 8M TPM / 250 IPM。

四、官方 Prompting 指南核心

这是本文最有价值的部分。OpenAI 把 Prompt 写法归纳为 8 条基本准则:

1. 定义结果

先说清「画什么」和「用来干什么」:产品照、广告、示意图。指定构图、宽高比、关键元素位置。复杂请求用 场景 → 主体 → 细节 → 约束 的分段结构。

2. 选可维护的格式

短句、描述段落、类 JSON 结构、指令、标签都行。选你最容易读和改的格式,不依赖特殊语法。

3. 描述可见细节

写清材质、光线、颜色、视觉媒介。要写实就明确说 photorealistic 或 real photograph。相机参数是外观线索,不是物理仿真保证。宽景、电影感、低光、雨天、霓虹场景要指定尺度、氛围和色彩,别只堆情绪词。

4. 指定人物与动作

写清身体取景、相对比例、视线、与物体的互动。比如「全身可见,含脚」「低头看摊开的书」「双手自然握把」。

5. 精确文字

需要的文字用引号标出,说明位置和排版风格。生僻词或品牌名逐字母拼写。明确要求无额外文字,输出后检查拼写和可读性。小字/密排/多字体至少用 high。

6. 区分变更与约束

编辑时说「只改 X」,列出要保留的细节:身份、几何、布局、光照、标签。排除不要的文字、logo、水印。精确局部编辑还要指定饱和度、对比度、箭头、相机角度和周边物体必须不变。

7. 给参考图分配角色

多图输入时按编号说明用途:主体、风格、服装、背景。解释各元素如何组合、哪些放哪里。

8. 有意识地迭代

把上一轮输出作为下一轮编辑输入,一次只改一个点,重复要保留的细节。「同上风格」可以带上下文,但结果漂移时要重述关键约束。

五、官方 Prompt 模板精选

以下模板直接来自官方 prompting guide,可按需替换主体和约束。

写实人像

text
Create a photorealistic candid photograph of an elderly sailor standing on a small fishing boat.
He has weathered skin with visible wrinkles, pores, and sun texture, and a few faded traditional sailor tattoos on his arms.
He is calmly adjusting a net while his dog sits nearby on the deck.
Shot like a 35mm film photograph, medium close-up at eye level, using a 50mm lens.
Soft coastal daylight, shallow depth of field, subtle film grain, natural color balance.
The image should feel honest and unposed, with real skin texture, worn materials, and everyday detail.
No glamorization, no heavy retouching.

信息图 / 流程图

text
Create a detailed Infographic of the functioning and flow of an automatic coffee machine like a Jura.
From bean basket, to grinding, to scale, water tank, boiler, etc.
I'd like to understand technically and visually the flow.

品牌广告(精确文字)

text
Give me a cool in culture ad / fashion shot for a brand called Thread.
It's a hip young street brand. The ad shows a group of friends hanging out together with the tagline "Yours to Create."
Make it feel like a polished campaign image for a youth streetwear audience: stylish, contemporary, energetic, and tasteful.
Use clean composition, strong color direction, natural poses, and premium fashion photography cues.
Render the tagline exactly once, clearly and legibly, integrated into the ad layout.
No extra text, no watermarks, no unrelated logos.

Logo 设计(透明背景)

text
Create an original, non-infringing logo for a company called Field & Flour, a local bakery.
The logo should feel warm, simple, and timeless. Use clean, vector-like shapes, a strong silhouette, and balanced negative space.
Favor simplicity over detail so it reads clearly at small and large sizes. Flat design, minimal strokes, no gradients unless essential.
Fully transparent background. Deliver a single centered logo with generous padding, clean alpha edges, and no solid backdrop, scenery, checkerboard, or watermark.

UI Mockup

text
Create a realistic mobile app UI mockup for a local farmers market.
Show today's market with a simple header, a short list of vendors with small photos and categories, a small "Today's specials" section, and basic information for location and hours.
Design it to be practical, and easy to use. White background, subtle natural accent colors, clear typography, and minimal decoration.
It should look like a real, well-designed, beautiful app for a small local market.
Place the UI mockup in an iPhone frame.

科学教育图

text
Create a simple biology diagram titled "Cellular Respiration at a Glance" for high school students.
Show how glucose turns into energy inside a cell. Include glycolysis, the Krebs cycle, and the electron transport chain.
Use arrows to connect the steps, and label the main molecules: glucose, pyruvate, ATP, NADH, FADH2, CO2, O2, and H2O.
Make it look like a clean classroom handout or slide, with a white background, simple icons, clear labels, and easy-to-read text.
Avoid tiny text, extra decoration, or anything that makes the diagram hard to understand.

四格漫画

text
Create a short vertical comic-style reel with 4 panels.
Panel 1: The owner leaves through the front door. The pet is framed in the window behind them, small against the glass, eyes wide, paws pressed high, the house suddenly quiet.
Panel 2: The door clicks shut. Silence breaks. The pet slowly turns toward the empty house, posture shifting, eyes sharp with possibility.
Panel 3: The house transformed. The pet sprawls across the couch like it owns the place, crumbs nearby, sunlight cutting across the room like a spotlight.
Panel 4: The door opens. The pet is seated perfectly by the entrance, alert and composed, as if nothing happened.

六、编辑 Prompt 模板

保留身份换服装

text
Edit the image to dress the woman using the provided clothing images.
Do not change her face, facial features, skin tone, body shape, pose, or identity in any way.
Preserve her exact likeness, expression, hairstyle, and proportions.
Replace only the clothing, fitting the garments naturally to her existing pose and body geometry with realistic fabric behavior.
Match lighting, shadows, and color temperature to the original photo so the outfit integrates photorealistically, without looking pasted on.
Do not change the background, camera angle, framing, or image quality, and do not add accessories, text, logos, or watermarks.

合成参考图

text
Place the dog from the second image into the setting of image 1,
right next to the woman, use the same style of lighting, composition and background.
Do not change anything else.

透明抠图

text
Extract the product from the input image and isolate it on a fully transparent background.
Output: centered product, crisp silhouette, no halos/fringing.
Preserve product geometry and label legibility exactly.
Add only light polishing. Do not add a solid backdrop, checkerboard, scenery, or shadow.
Do not restyle the product; remove the background and preserve clean alpha transparency.

线稿转写实

text
Turn this drawing into a photorealistic image.
Preserve the exact layout, proportions, and perspective.
Choose realistic materials and lighting consistent with the sketch intent.
Do not add new elements or text.

家具替换

text
In this room photo, replace ONLY the white chairs with chairs made of wood.
Preserve camera angle, room lighting, floor shadows, and surrounding objects.
Keep all other aspects of the image unchanged.
Photorealistic contact shadows and fabric texture.

七、多轮迭代策略

Responses API 支持多轮图像生成/编辑:

python
from openai import OpenAI
client = OpenAI()

# 第一轮:生成
response = client.responses.create(
    model="gpt-6-astra",
    input="Generate an image of a shampoo billboard on a highway at sunset. Text exactly: 'Fresh and clean'",
    tools=[{"type": "image_generation", "model": "gpt-image-2.5-sunburst"}],
)

# 第二轮:基于上一轮结果继续编辑
follow_up = client.responses.create(
    model="gpt-6-astra",
    previous_response_id=response.id,
    input="Make it look like a winter evening with snowfall.",
    tools=[{"type": "image_generation", "model": "gpt-image-2.5-sunburst"}],
)

action 参数可控制行为:auto(默认,模型自选)、generate(强制新建)、edit(强制编辑)。流式输出可通过 partial_images 参数获取 0-3 张中间帧。

角色一致性技巧

多页绘本/多场景角色保持一致的方法:

  1. 建立角色基准图 — 详细描述外观、比例、服装、性格
  2. 后续场景复用基准图 — 作为编辑输入,重复外观约束
  3. 一次只改一个条件 — 换环境/换动作,不同时改风格
text
Continue the children's book story using the same character.

Scene:
The same young forest hero is gently helping a frightened squirrel
out of a fallen tree after a winter storm.

Character Consistency:
- Same green hooded tunic
- Same facial features, proportions, and color palette
- Same gentle, heroic personality

Style:
Children's book watercolor illustration,
soft lighting, snowy forest environment,
warm and comforting mood.

Constraints:
- Do not redesign the character
- No text
- No watermarks

八、API 调用速查

Image API(单次生成)

python
from openai import OpenAI
import base64

client = OpenAI()
result = client.images.generate(
    model="gpt-image-2.5-sunburst",
    prompt="A photorealistic product shot of a ceramic mug on marble",
    size="1024x1024",
    quality="medium",
)
with open("mug.png", "wb") as f:
    f.write(base64.b64decode(result.data[0].b64_json))

Image API(编辑)

python
result = client.images.edit(
    model="gpt-image-2.5-sunburst",
    image=open("original.png", "rb"),
    prompt="Remove the flower from the man's hand. Do not change anything else.",
    size="1024x1536",
    quality="medium",
)

流式生成

python
stream = client.images.generate(
    model="gpt-image-2.5-sunburst",
    prompt="A river made of white owl feathers in a winter landscape",
    stream=True,
    partial_images=2,
)
for event in stream:
    if event.type == "image_generation.partial_image":
        # 处理中间帧
        save(event.partial_image_index, event.b64_json)

九、关键注意事项

  1. API Organization Verification — 使用 GPT Image 模型前可能需要在开发者控制台完成组织验证。
  2. 透明 ≠ 画格子 — 画出来的棋盘格不是透明背景,必须用 background="transparent" + PNG/WebP。
  3. 文字必须声明 — 不声明文字的 Prompt,模型会自己加英文说明文字。
  4. 精确局部编辑用抠图 — 如果某区域必须像素级不变,把批准的编辑合成回原图,别只靠 Prompt。
  5. quality 标签不跨模型对齐 — 同为 medium,Sunburst 和 Flare 的画质/速度不同。
  6. GPT Image 2 token 计算器不适用于 2.5 — 计费需按实际返回 token 估算。

十、总结

GPT Image 2.5 的核心变化不是「更强了」这么简单,而是:

  • 产品分层:Sunburst 管质量,Flare 管速度,各司其职。
  • 编辑精度提升:官方在身份保留、产品几何、精确文字上都有明确改善。
  • Prompt 方法论系统化:8 条准则 + 按场景的模板库,降低了写出好 Prompt 的门槛。

对于开发者,建议路径是:用 Sunburst + 官方模板建立质量基线 → 换 Flare 测延迟 → 在两者间找到成本/质量/速度的平衡点。


参考文档:

分享
← 返回博客列表
🎁 有邀请福利哦,点击查看
🎁