AI 图像生成正在从碰运气的提示词工程,演变为可组合、可量化的软件工程。本文通过实测 OpenAI gpt-image-1,解析材质替换、Alpha 遮罩控制与多图合成的工程化落地路径。
长期以来,调用 AI 图像生成 API 就像在赌场玩轮盘赌:精心写下提示词,调整修饰语,发出请求,然后默默祈祷。如果生成的画面有 90% 符合预期,唯独有一处细节崩了,开发者通常只能全盘推倒重来,寄希望于下一次抽卡。
OpenAI 开放 gpt-image-1 模型访问后,我们关注的核心不是它能否生成更漂亮的纹理,而是能否基于它构建一套连续迭代的生产流水线:生成基础实体、在保留几何结构的前提下变换材质、对特定区域精准介入,最后将多个独立资产合成为完整画面。
测试的第一步是从基础提示词生成工作对象。以一只置于木桌上的折纸狐狸为例:
img = client.images.generate(
model="gpt-image-1",
prompt=prompt,
background="auto",
n=1,
quality=quality,
size=size,
output_format="png",
moderation="auto",
)
与早期的生图接口相比,新接口直接在返回体中提供了 Token 消耗拆解:
if hasattr(img, "usage") and img.usage:
input_details = getattr(img.usage, "input_tokens_details", None)
if input_details:
print("Prompt tokens:", getattr(input_details, "text_tokens", "N/A"))
print("Input images tokens:", getattr(input_details, "image_tokens", "N/A"))
print("Output image tokens:", getattr(img.usage, "output_tokens", "N/A"))
在生成一张 1024x1024 中等质量图片的测试中,提示词消耗了 23 个文本 Token,模型返回了 1056 个图像 Token(以 Base64 格式返回)。这一机制的价值在于,工程团队无需再依赖外部估算,直接通过代码就能实时计算每次调用的延迟与计算成本。
在视觉生产中,真正的难点从来不是生成第一张图,而是版本迭代中的一致性。
如果需要同一只狐狸在相同视角和相同桌面上,但换成完全不同的材质,旧版模型的常规做法是重新描述场景,但折痕、阴影和姿态往往全变了。gpt-image-1 则支持通过 client.images.edit() 传入原始图片二进制流与修改指令:
with open(img_file_path, "rb") as f:
img = client.images.edit(
model="gpt-image-1",
image=[f],
prompt="Change the origami fox material to glowing iridescent translucent glass with neon blue and purple reflections, keeping the same origami fold shape and desk setting",
n=1,
quality=quality,
size=size,
)
该请求的 Token 消耗如下:
模型将输入的 194 个图像 Token 作为结构约束条件,精确保留了折纸的几何棱角与空间构图,仅将表面重构为半透明蓝色玻璃与折射光效。这让资产跨材质复用具备了极高的稳定性。
全局编辑依然存在局限:模型拥有全图的自由解释权。如果只想在狐狸胸口嵌入一颗宝石,同时确保耳朵、身体折痕和背景桌面像素级不变,单纯依靠自然语言提示词很难约束。
此时可以通过 Mask(遮罩)进行局部干预。为了摆脱对 Photoshop 等外部修图软件的依赖,可以通过 Python Pillow 库动态生成确定性遮罩:
from PIL import Image, ImageDraw
from pathlib import Path
def create_circular_mask(base_image_path: Path, mask_output_path: Path) -> Path:
"""生成白底透明中心的圆形遮罩"""
with Image.open(base_image_path) as im:
width, height = im.size
mask = Image.new("RGBA", (width, height), (255, 255, 255, 255))
draw = ImageDraw.Draw(mask)
cx, cy = width // 2, height // 2
radius = min(width, height) // 4
# 将需要修改的区域设为完全透明 (alpha = 0)
draw.ellipse([cx - radius, cy - radius, cx + radius, cy + radius], fill=(0, 0, 0, 0))
mask.save(mask_output_path, "PNG")
return mask_output_path
API 的运行规则非常明确:完全不透明区域(alpha = 255)保持锁定,透明区域(alpha = 0)开放给模型自由发挥。
with open(img_file_path, "rb") as img_f, open(mask_file_path, "rb") as mask_f:
img = client.images.edit(
model="gpt-image-1",
image=[img_f],
mask=mask_f,
prompt="Place a miniature glowing emerald crystal in the masked area",
quality="medium",
size="1024x1024",
)
模型不仅在透明区域生成了祖母绿宝石,还主动理解了周围纸张的体积感,在邻近折面上投射出绿色环境反光,同时保持了圆圈外部区域的像素级一致。
在 client.images.edit() 中,image 参数支持传入文件列表。这意味着可以把多次独立生成的不同资产,直接送入接口进行多模态合成。
将此前生成的「纸质狐狸」与「玻璃狐狸」作为多图输入:
opened_files = [open(p, "rb") for p in ["example_fox.png", "example_fox_edited.png"]]
try:
img = client.images.edit(
model="gpt-image-1",
image=opened_files,
prompt="Create an artistic gallery exhibition poster showing both the paper origami fox and the glowing glass fox standing side by side on an elegant wooden gallery pedestal, dramatic lighting with typography space",
quality="medium",
size="1536x1024",
)
finally:
for f in opened_files:
f.close()
最终生成的画报同时融合了哑光纸质与发光水晶两种材质特征,在统一的画廊射灯下形成了协调的光影遮蔽,并在顶部生成了边缘整洁、拼写准确的展览排版文字。
AI 生图的范式正在发生改变:它不再是单纯比拼谁更会写 Prompt 的玄学,而是正在转变为标准、可预测且高度可组合的软件工程流水线。
免费获取企业 AI 成熟度诊断报告,发现转型机会
关注公众号

扫码关注,获取最新 AI 资讯
3 步完成企业诊断,获取专属转型建议
已有 200+ 企业完成诊断