Input
Output
Result valid for 24 hours
Grok Imagine API 核心功能
专业级 AI 图像生成与编辑的强大能力
超写实文生图
使用 grok-imagine-1.5-apimart 从文本提示生成令人惊叹的超写实图像。出色的细节还原、精准的构图控制与真实的色彩表现,专业级输出品质。
自然语言图像编辑
使用 grok-imagine-1.5-edit-apimart,通过文字描述智能修改现有图像。无需蒙版,只需描述你想要的改变,Grok Imagine 自动完成处理。
丰富的风格控制
Grok Imagine 对创意提示词有细腻的理解能力,同一套 API 可生成超写实人像、风格化插画、概念艺术及抽象视觉效果。
极简双参数 API
Grok Imagine API 仅需 model 和 prompt 两个参数,配置极简、灵活性极高,几分钟即可完成集成。
极具竞争力的定价
比直接使用 xAI API 最高节省 70%。AiBox 优化的基础设施以按量计费的方式提供稳定可靠的服务。
企业级可靠性
99.9% 可用率 SLA,冗余基础设施与自动故障转移,AiBox 确保你的图像生成流水线在任何规模下稳定运行。
全球 CDN 加速分发
生成的图像通过 AiBox 全球 CDN 网络加速交付,确保全球各地区低延迟、高可用的图像访问体验。
极速生成响应
Grok Imagine 基于 xAI 高性能推理引擎,秒级返回高质量图像结果,完美适配实时应用与高频交互场景。
定价详情
透明定价,无隐藏费用。按量付费,用多少付多少。
| 规格 | 当前价格 | 官方价格 | 为你节省 |
|---|---|---|---|
| 默认 | 0.1875 Credits/次 ~$0.01875 | 0.1875 Credits/次 ~$0.01875 | 0% |
* 实际费用以最终输出为准。
常见问题
相关模型
gemini-2.5-flash-image-preview
Nano Banana (gemini-2.5-flash-image-preview) is Google DeepMind's fast, conversational image generation and editing model. It delivers natural-language text-to-image, precise multi-turn edits, strong character consistency, and multi-image fusion at low latency.
doubao-seedream-5-0-pro
Seedream 5.0 Pro (doubao-seedream-5-0-pro) is ByteDance's quality-first text-to-image model. It produces cinematic 1K and 2K images with best-in-class text rendering, supports up to 10 reference images for style and character consistency, and unifies generation and editing in a single API call.
midjourney
Midjourney is a leading text-to-image AI model renowned for its painterly aesthetics, strong composition, and cinematic lighting. Generate high-quality images from text prompts or reference images, with fine-grained control over style, aspect ratio, and creative parameters — accessible programmatically through AiBox's unified API, no Discord required.
wan2.7-image
wan‑2‑7‑image is an advanced image generation model in Alibaba’s Wan 2.7 series designed to create high‑quality visuals from text prompts. It excels at generating detailed, realistic images with accurate object representation and rich composition. The model supports multimodal input (e.g., combining text with reference images) to influence style and structure, and it’s well‑suited for creative workflows such as marketing assets, product visuals, social media graphics, and artistic content production. With strong semantic understanding and prompt adherence, wan‑2‑7‑image delivers both fidelity and expressive visual output.