OpenAI图像

gpt-image-2

GPT Image 2 通过文本提示提供下一代 AI 图像生成,具备近乎完美的文字渲染、原生 4K 输出和深度世界知识。支持 model、prompt 等参数——非常适合生产工作流程。 图像生成

商业使用

Input

gpt-image-2
Click to upload imageJPEG, PNG or WebP (max 10MB)

Output

Result valid for 24 hours

GPT Image 2 (gpt-image-2)

OpenAI 这款尚未官宣的图像模型曾以 maskingtape-alpha、gaffertape-alpha、packingtape-alpha 三个别名短暂出现在 LMArena 上。早期测试者反馈它的文字渲染近乎完美、原生支持 4K 输出,对真实世界的理解已经更接近摄影而非生成。AiBox 全程跟进发布动态,将在第一天完成接入。

获取 API 密钥
50K+
活跃用户
99.9%
在线率
2x
更快
70%
成本节省

定价详情

透明定价,无隐藏费用。按量付费,用多少付多少。

OpenAIgpt-image-2
规格当前价格官方价格为你节省
1K0.085 Credits/次
~$0.0085
0.10625 Credits/次
~$0.010625
20%
2K0.14 Credits/次
~$0.014
0.175 Credits/次
~$0.0175
20%
4K0.21 Credits/次
~$0.021
0.2625 Credits/次
~$0.02625
20%

* 实际费用以最终输出为准。

常见问题

相关模型

Gemini

gemini-2.5-flash-image-preview

Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.

0.125 Credits
~$0.0125
20%
ByteDance

doubao-seedream-5-0-pro

Seedream 5.0 Pro (doubao-seedream-5-0-pro) is ByteDance's quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.

0.36 Credits/1K tokens
~$0.036
20%
M

midjourney

Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through AiBox's unified API, no Discord required.

0.4504 Credits
~$0.04504
20%
Alibaba

wan2.7-image

wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.

0.216 Credits
~$0.0216
20%