API 更新
新增模型、功能与改进的更新记录。
wan3.0-video-prime
🚀 新品上线
wan3.0-video-prime现已登陆 AiBox,支持通用视频生成、首尾帧视频生成,以及文件/网页转视频,并集成音频生成能力。✨ 核心亮点
- 多模态输入:支持文本、图片、视频、音频、文档及公开网页
- 视频生成模式:文生视频、首帧/尾帧控制、文件/网页生成
- 输出规格:480p / 720p / 1080p,时长 2–30 秒
- 画面比例:支持 16:9、9:16、1:1、4:3、3:4 及自适应
- 参考素材:最多 10 张图片、5 段视频、5 段音频
- 同步音频:支持语音、音效及背景音乐生成
🎯 适用场景
营销视频 · 社交媒体内容 · 产品展示 · 文档/网页内容可视化 · 多媒体创作
📚 [查看 Wan 3.0 API 文档 →](https://aiboxapi.com/en/api-reference/videos/wan3.0-video/generation)
GPT Image 2.5 Ext
🚀 新品上线
gpt-image-2.5-ext现已登陆 AiBox,支持文生图与图生图,提供 Flare、Sunburst 双版本。核心亮点
- 输出分辨率:1K / 2K / 4K
- 批量生成:每次 1–4 张图片
- 参考图片:最多 16 张,无额外输入图费用
- 画面比例:支持自动及 10 种预设比例
- 灵活计费:按版本、分辨率与实际交付张数结算
适用场景
营销素材 · 社交媒体配图 · 图片改创 · 多尺寸视觉创作
[查看 GPT Image 2.5 Ext API 文档 →](https://aiboxapi.com/en/api-reference/images/gpt-image-2.5-ext/generation)
gpt image 2.5
• We are excited to announce the launch of GPT Image 2.5 on AiBox, featuring Flare and Sunburst for high-quality image generation and precise editing, with flexible quality settings and multiple reference images.
## New Models
GPT Image 2.5 Flare
- Model ID: gpt-image-2.5-flare
- Focus: Fast, high-quality image generation for everyday creative workflows
- Use Cases: Social media content, product images, rapid prototyping, and batch generation
GPT Image 2.5 Sunburst
- Model ID: gpt-image-2.5-sunburst
- Focus: Precise image editing and detailed visual refinement
- Use Cases: Advertising creatives, polished product imagery, and edits requiring fine control
## Key Capabilities
- Text-to-Image: Generate images directly from text prompts - Image-to-Image: Edit and transform images using prompts and reference inputs - Multiple References: Combine up to 16 reference images in one request - Resolution Presets: Choose from 1K, 2K, and 4K presets; resolutions above 2560×1440 are experimental - Custom Dimensions: Specify exact pixel dimensions within supported size limits - Quality Settings: Low, Medium, High, XHigh, Max, and Auto - Multiple Aspect Ratios: 1:1, 3:2, 2:3, 4:3, 3:4, 5:4, 4:5, 16:9, 9:16, 2:1, 1:2, 21:9, 9:21, 3:1, and 1:3, plus Auto sizing - Transparent Backgrounds: Generate transparent images in PNG or WebP - Multiple Outputs: Generate 1–4 images per request - Output Formats: PNG, JPEG, and WebP, with adjustable compression for JPEG and WebP
## Documentation
- [GPT Image 2.5 API Documentation](https://aiboxapi.com/en/api-reference/images/gpt-image-2.5/generation)
gemini-omni-1.1-flash
🚀 新品上线
Google 官方视频生成模型 gemini-omni-1.1-flash 现已正式登陆 AiBox。
✨ 核心亮点
生成模式:文生视频、图生视频、首尾帧插值、视频编辑
输出规格:360p / 720p / 1080p / 4K,含音频
画面比例:16:9、9:16
视频时长:约 3–10 秒(由模型决定)
🎯 适用场景
营销素材、社交媒体短视频、创意原型
📚 [查看 Omni-Flash-Ext API 文档](https://aiboxapi.com/en/api-reference/videos/gemini-omni-1.1-flash/generation)
[Notice] Price Adjustment for Seedream-5-0-lite
Due to an upstream pricing adjustment, we will update the price of Seedream-5-0-lite accordingly.
Effective Date: September 3, 2026, at 12:00 PM (UTC+8) / September 3, 2026, at 04:00 AM (UTC+0)
Tasks submitted before this time will still be billed at the original rate.
━━━━━━━━━━━━━━━━━━━━
■ seedream-5-0-lite
Single Image Generation
$0.02275 → $0.028 per image
0.224 Credits → 0.28 Credits
━━━━━━━━━━━━━━━━━━━━
Prices for all other models remain unchanged. Please contact us if you have any questions.
[Notice] Price Adjustment for DeepSeek-V4-Flash
Due to changes in our upstream provider’s billing policy, DeepSeek-V4-Flash will move from a single flat-rate pricing structure to separate Peak and Off-Peak pricing. We have updated the corresponding prices accordingly.
Effective Date: September 3, 2026, at 12:00 PM (UTC+8) / September 3, 2026, at 04:00 AM (UTC+0)
Tasks submitted before this time will still be billed at the original rates.
━━━━━━━━━━━━━━━━━━━━
■ deepseek-v4-flash
Peak
Applicable Period:
UTC+8: Monday–Sunday, 08:00–22:00 UTC+0: Monday–Sunday, 00:00–14:00
Input
$0.3428568 per 1M tokens (3.428568 Credits)
Cached Input
$0.0685712 per 1M tokens (0.685712 Credits)
Output
$1.0285712 per 1M tokens (10.285712 Credits)
Off-Peak
Applicable Period:
UTC+8: Monday–Sunday, 22:00–08:00 the following day UTC+0: Monday–Sunday, 14:00–00:00 the following day
Input
$0.01714284 per 1M tokens (0.1714284 Credits)
Cached Input
$0.00342856 per 1M tokens (0.0342856 Credits)
Output
$0.05142856 per 1M tokens (0.5142856 Credits)
━━━━━━━━━━━━━━━━━━━━
Note on Peak / Off-Peak Pricing:
After the effective time, DeepSeek-V4-Flash will no longer use a single flat rate. Requests will be billed according to the applicable Peak or Off-Peak period.
Prices for all other models remain unchanged. Please contact us if you have any questions.
[Notice] Price Adjustment for Seedream Models
The upstream official vendor for the Seedream image model series recently updated its pricing. Following a recalculation based on the new official rates, we need to adjust the prices for the models listed below accordingly.
Effective Date: August 21, 2026, at 12:00 PM (UTC+8)
Tasks submitted before this time will still be billed at the original rates.
━━━━━━━━━━━━━━━━━━━━
■ seedream-5-0-pro (Tiered by output image pixels)
Single Image Generation
1.5K and below (≤ 2.61M pixels): $0.02928 → $0.02925 (0.2925 Credits)
Above 1.5K (> 2.61M pixels): $0.05856 → $0.0585 (0.585 Credits)
Layer Decomposition (Half price per tier)
1K: $0.01464 → $0.014625 (0.14625 Credits)
2K: $0.02928 → $0.02925 (0.2925 Credits)
Reference Image Input
1st image: Still free
2nd image onward: $0.00195 per image (0.0195 Credits)
■ seedream-5-0-lite
$0.0228 → $0.02275 (0.2275 Credits)
■ seedream-4-5
$0.0228 → $0.026 (0.26 Credits)
■ seedream-4-0
$0.0182 → $0.0195 (0.195 Credits)
━━━━━━━━━━━━━━━━━━━━
Note on Pixel Tiers:
Common dimensions examples:
1536×1536 (~2.36M pixels) falls into the 1.5K and below tier.
2048×2048 (~4.19M pixels) falls into the Above 1.5K tier.
Prices for all other models remain unchanged. Please contact us if you have any questions.
grok-image-2-ext
We are excited to announce the launch of Grok Imagine 2.0 Ext, xAI’s latest high-quality image generation model, now available through AiBox.
## New Model
Grok Imagine 2.0 Ext
- Model ID: grok-imagine-2.0-ext
- Features: High-quality text-to-image generation
- Highlights: Generate up to 12 images in a single request with flexible aspect-ratio support
- Quality: Fixed high-quality generation mode
- Output: URL-based image results valid for 72 hours
## Key Capabilities
- Text-to-Image: Generate detailed images directly from natural-language prompts
- Multiple Image Generation: Generate between 1 and 12 images per request
- Multiple Aspect Ratios: Support for 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, and 16:9
- High-Quality Output: Optimized for detailed, production-ready image generation
- Asynchronous Tasks: Submit generation tasks and retrieve results through task-status polling
- URL-Based Results: Receive downloadable image URLs with a 72-hour validity period
## Documentation
- Grok Imagine 2.0 Ext API Documentation (https://aiboxapi.com/en/api-reference/images/grok-imagine-2.0-ext/generation)
wan3.0
We are excited to announce the launch of Wan 3.0, Alibaba Cloud's latest all-in-one video generation model with multimodal reference support and integrated audio generation capabilities.
## New Model
Wan 3.0
- Model ID: wan3.0-video
- Features: General video generation, frame-to-video generation, and file/webpage-to-video generation
- Highlights: One unified model supporting text prompts, images, videos, audio, documents, and public webpages
- Resolution: 480p / 720p / 1080p
- Duration: 2–30 seconds
## Key Capabilities
- General Generation: Create videos from text prompts alone or combine prompts with image, video, and audio references - Frame-to-Video: Animate a required first frame with an optional last frame for controlled transitions - File and Webpage Generation: Generate videos directly from documents, presentations, PDFs, text files, or public webpages
- Audio Generation: Generate synchronized output audio, including speech, sound effects, and background music
- Multiple Aspect Ratios: 16:9, 9:16, 1:1, 4:3, 3:4, and Adaptive
- Multimodal References: Support for up to 10 images, 5 video clips, and 5 audio clips
## Documentation
- Wan 3.0 API Documentation (https://aiboxapi.com/en/api-reference/videos/wan3.0-video/generation)
doubao-seedance-2-5
We are excited to announce the launch of Doubao Seedance 2.5, an advanced multimodal video generation model supporting longer videos, synchronized audio, and rich reference inputs.
## New Model
Doubao Seedance 2.5
- Model ID: doubao-seedance-2.5
- Features: Text-to-Video, Image-to-Video, video editing, video extension, and multimodal generation
- Highlights: Supports images, videos, and audio as reference inputs with optional synchronized audio generation
- Resolution: 480p / 720p
- Duration: 4–30 seconds or automatic duration
## Key Capabilities
- Text-to-Video: Generate videos directly from text prompts
- Image-to-Video: Animate images using first-frame, last-frame, or reference-image roles
- Video References: Use reference videos for editing, extension, motion, and style guidance
- Audio References: Support audio-guided generation, including pure audio reference workflows
- Audio Generation: Optionally generate synchronized audio alongside the video
- Multimodal Inputs: Combine image, video, and audio references in one request
- Multiple Aspect Ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and Adaptive
- Reference Limits: Up to 30 images, 10 videos, and 10 audio files
- Output Formats: MP4 and MOV
## Documentation
- Doubao Seedance 2.5 API Documentation (https://aiboxapi.com/en/api-reference/videos/doubao-seedance-2-5/generation)
QWEN-IMAGE-3-0
> 🚀 New Model Live > > Image generation models
qwen-image-3.0andqwen-image-3.0-proare now available on AiBox.✨ Key Features
| Capability | Details | | --- | --- | | Generation modes | Text-to-image and image-to-image editing (up to 3 reference images) | | Output resolution | 1K / 2K | | Aspect ratios | 1:1, 4:3, 3:4, 16:9, 9:16, 3:2, 2:3 (custom pixel sizes supported) | | Images per request | 1–6 | | Text & layout | Strong bilingual text rendering; Pro is better for dense layouts such as menus, posters, and storyboards | | Prompt extend | Optional
prompt_extend(direct/agent) |🎯 Best For
- Marketing posters and campaign creatives
- E-commerce menus, product sheets, and text-heavy visuals
- Social media assets and rapid creative prototyping
📚 [View Qwen Image 3.0 API docs](https://aiboxapi.com/en/api-reference/images/qwen-image-3.0/generation)
FLUX-3-VIDEO
> 🚀 New Model Launch > > The
flux-3-videovideo generation model is now available on AiBox.✨ Key Highlights
| Capability | Description | | --- | --- | | Generation Modes | Text-to-video, image-to-video, and video continuation | | Draft Mode | Generate a low-cost draft, then render the final video once satisfied | | Output Quality | HD / FHD | | Video Duration | 5–20 seconds | | Keyframes | Supports 1–10 ordered keyframes | | Aspect Ratios | Auto, 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16 | | Audio Generation | Synchronized audio is enabled by default and can be disabled |
🎯 Use Cases
- Generate videos from text or images
- Control video content with start/end frames or multiple keyframes
- Continue an existing video
- Preview ideas with low-cost drafts before rendering the final video
📚 [View the FLUX 3 Video API Documentation](https://aiboxapi.com/en/api-reference/videos/flux-3-video/generation)
Minimax-h3
We are excited to announce the launch of MiniMax H3, MiniMax’s latest multimodal video generation model with native 2K output and synchronized audio.
New Model
MiniMax H3
• Model ID: MiniMax-H3 • Features: Text-to-Video and Image-to-Video (first/last frame), plus multimodal reference generation (images, video, and audio) • Highlights: Unified omnimodal model for T2V / I2V / reference-to-video with native audio on every result • Resolution: 2K • Duration: 4–15 seconds
Key Capabilities
• Text-to-Video: Generate videos directly from text prompts • Image-to-Video: Animate with first-frame and optional last-frame control • Omni Reference: Guide generation with reference images (≤9), videos (≤3), and audio (≤3; must pair with image or video) • Native Audio: All outputs include synchronized stereo audio • Multiple Aspect Ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and Adaptive (I2V follows input image) • Watermark: Optional AIGC watermark
Documentation
• MiniMax H3 API Documentation (https://aiboxapi.com/en/api-reference/videos/minimax-h3/generation)
Try it on AiBox
• Model page: https://aiboxapi.com/model/minimax-h3
FLUX
We are excited to announce the launch of FLUX.1 Kontext and FLUX.2, Black Forest Labs’ advanced image generation and editing model families, now available through AiBox’s unified image API.
## New Models
### FLUX.1 Kontext
FLUX.1 Kontext Pro
- Model ID:
flux-kontext-pro - Features: Context-aware image editing and text-to-image generation
- Highlights: Faster generation with predictable per-image pricing
- Reference Images: Up to 4 images
FLUX.1 Kontext Max
- Model ID:
flux-kontext-max - Features: Context-aware image editing and text-to-image generation
- Highlights: Higher-quality output for demanding creative workflows
- Reference Images: Up to 4 images
### FLUX.2
FLUX.2 Pro
- Model ID:
flux-2-pro - Features: Image generation and multi-reference image editing
- Highlights: Balanced speed, quality, and cost
- Reference Images: Up to 8 images
FLUX.2 Max
- Model ID:
flux-2-max - Features: High-quality image generation and editing
- Highlights: Maximum output quality for professional workflows
- Reference Images: Up to 8 images
FLUX.2 Flex
- Model ID:
flux-2-flex - Features: Controllable image generation and editing
- Highlights: Adjustable steps and guidance for fine-grained creative control
- Reference Images: Up to 8 images
## Key Capabilities
- Text-to-Image: Generate new images directly from text descriptions
- Image Editing: Modify subjects, objects, backgrounds, colors, compositions, and visual styles
- Multi-Reference Generation: Use multiple reference images to guide identity, composition, and style
- Context-Aware Editing: Preserve visual context while applying natural-language editing instructions
- Flexible Output Sizes: Generate images at 512px, 1K, or 2K resolution with supported aspect ratios
- Prompt Enhancement: Automatically expand prompts with optional
prompt_upsampling - Repeatable Results: Reuse a fixed
seedwith the same inputs to help reproduce an image - Multiple Output Formats: Export generated images as PNG, JPEG, or WebP
- Unified Asynchronous API: Access every FLUX model through the same task submission and polling workflow
## Documentation
- [FLUX.1 Kontext API Documentation](https://aiboxapi.com/en/api-reference/images/flux-kontext/generation)
- [FLUX.2 API Documentation](https://aiboxapi.com/en/api-reference/images/flux-2/generation)
- Model ID:
GPT Image 2: Transparent Token Pricing Now on the Model Page
We have updated the GPT Image 2 model page with transparent token-based pricing details.
What's New
Token Pricing Table
The pricing calculator on the [GPT Image 2 model page](https://aiboxapi.com/model/gpt-image-2) now displays a token price breakdown for
gpt-image-2-official, shown in Credits per 1M tokens:| Type | Price | | --- | --- | | Text input | 50 Credits / 1M tokens | | Cached text input | 12.5 Credits / 1M tokens | | Image input | 80 Credits / 1M tokens | | Cached image input | 20 Credits / 1M tokens | | Image output | 300 Credits / 1M tokens |
Improvements
- Token Price Breakdown: See exactly how each modality (text/image, input/cached input/output) is billed before you generate
- Credits Display: All token prices are converted to Credits, consistent with the rest of the platform
- Live Estimate: The calculator still shows the estimated per-image price with the current token consumption for your selected size, resolution, and quality
Suno
We are excited to announce the launch of the Suno Music API on AiBox, a complete AI music generation and production toolchain with 31 endpoints covering creation, editing, and export.
New Model
Suno
- Model ID:
suno - Features: Text-to-Music generation with a full suite of derivative and post-processing endpoints
- Highlights: Unified API covering generation, extension, covers, stem separation, voice personas, audio editing, and analysis
- Versions: v3.5 / v4 / v4.5 / v4.5+ / v4.5-all / v5 / v5.5
- Output: Two candidate tracks per generation, with streaming audio, cover art, and lyrics
Key Capabilities
- Music Generation: Inspiration mode (describe the song you want) and Custom mode (full control over lyrics, title, style tags, negative tags, vocal gender, and style / weirdness / audio weights)
- Lyrics Generation: Standalone lyrics endpoint with selectable lyrics models (
classic/remi) - Audio Upload: Bring your own audio as the source for extension, covers, and vocal / instrumental additions
- Derivative Works: Extend, Cover, Remaster, Add Vocals, Add Instrumental, Add Stem, Mashup, and Sample-to-Song
- Stem Separation: 4-track separation, full 12-instrument separation, and isolated vocal (Vox) extraction
- Voice Personas: Create reusable voice personas from generated tracks and apply them to new generations
- Audio Editing: Replace Section, Remove Section, Crop, Fade In / Fade Out, Adjust Speed (0.25x–4x with optional pitch preservation), and Full Song synthesis from extended clips
- Analysis & Export: Word-level aligned lyrics timeline, BPM analysis, MIDI generation, Music Video (MP4) rendering, and WAV export
Documentation
- [Suno API Documentation](https://aiboxapi.com/en/api-reference/audios/suno/overview)
- Model ID:
FlowMusic
we are excited to announce that FlowMusic — a full-stack AI music generation suite — is now available on AiBox.
Turn a single prompt or lyric into studio-ready songs, stems, and music videos. Powered by an upgraded Producer engine (default model Lyria 3 Pro), FlowMusic bundles nine music-creation capabilities — generation, extension, replacement, cover, stem separation, import, download, and video rendering — into one unified async task-based API, so you can build an entire song lifecycle end to end.
- Model ID:
flowmusic - Capabilities: Music Generation, Lyrics Generation, Extend, Replace, Cover, Stem Separation, Audio Import, Format Export, Music Video Rendering
- Highlights: 9 endpoints in one model, lossless WAV + streaming M4A, auto-generated cover art, karaoke timing markers, async task API
- Output: M4A (streaming) + WAV (lossless) + optional MP4 music video
- Controls: BPM, length up to 240s, seed for reproducibility
🚀 Key Capabilities
- Text/Lyrics-to-Music: Generate a full song from a sound prompt and/or lyrics, with control over BPM, length (up to 240s), and seed — each request returns one track
- Lyrics Generation: Turn any idea into structured lyrics (≤ 3000 chars), then feed them straight back into music generation
- Music Extend: Continue an existing clip from any timestamp — up to 327 seconds of new audio — guided by a natural-language instruction
- Section Replace: Regenerate a specific time range within a song (e.g. swap a chorus to piano) without touching the rest
- Cover / Restyle: Re-arrange a whole track into a new style with adjustable edit strength (0–1)
- Stem Separation: Split vocals and accompaniment into a downloadable multi-track ZIP
- Audio Import: Bring external audio in via
audio_urlto obtain aclip_idfor downstream editing - Format Export: Download any clip as
wavormp3 - Music Video Rendering: Render a clip into MP4 with
simple/modern/playerpresets - Karaoke Timing: Lyrics timing markers included for word-level highlighting
- Usage-based Billing: Charged on submission with full auto-refund on task failure ($0.06 generate/extend/replace/cover/stems · $0.02 lyrics/download/video · $0.01 import)
🎯 Best For
- Producing complete, lyric-driven songs from a single text idea
- Iterating on tracks — extend, replace a section, or restyle without starting over
- Generating royalty-style background music for video, ads, and games
- Building karaoke experiences with word-level lyric timing
- Extracting stems for remixing and post-production
- Turning existing audio into new covers and arrangements
## 📚 Documentation - [FlowMusic](https://aiboxapi.com/en/api-reference/audios/flow-music/music)
- Model ID:
seedream-5-0-pro
We are excited to announce the launch of Seedream 5.0 Pro, ByteDance's latest quality-first text-to-image model with best-in-class text rendering and unified generation and editing.
New Model
Seedream 5.0 Pro
- Model ID:
doubao-seedream-5-0-pro - Features: Text-to-Image and Image-to-Image generation with unified editing in a single API call
- Highlights: Cinematic, quality-first image generation with industry-leading text rendering
- Resolution: 1K / 2K (default 2K)
- Output: One image per request (PNG / JPEG)
Key Capabilities
- Text-to-Image: Generate cinematic, richly detailed images directly from text descriptions
- Image-to-Image: Edit and refine existing images — remove objects, replace backgrounds, or restyle with simple prompts
- Perfect Text Rendering: Produce accurate, readable text in images for posters, banners, and marketing visuals
- Multi-Reference Fusion: Blend subjects, styles, and outfits from multiple reference images in one request
- Multiple Aspect Ratios: 1:1, 4:3, 3:4, 16:9, 9:16, 3:2, 2:3, 21:9, and Auto
- Reference Images: Support for up to 10 reference images for style and character consistency
- Watermark: Optional watermark on generated images
Notes
- Quality-first model that prioritizes fidelity and detail over raw speed (1K ~90s, 2K ~160s)
- Asynchronous generation: submit to
/v1/images/generations, then poll/v1/tasks/{task_id}until the task completes
Documentation
- [Seedream 5.0 Pro API Documentation](https://aiboxapi.com/en/api-reference/images/seedream-5-0-pro/generation)
- Model ID:
gemini-3.1-flash-lite-image-ext
We've added `gemini-3.1-flash-lite-image-ext` — a fast, low-cost image generation model now available on the Nano Banana page. Select it directly from the model dropdown to get started.
Capabilities
- Text-to-image & image-to-image — generate from a prompt, or guide generation with reference images.
- Up to 14 reference images per request.
- Aspect ratios:
1:1,2:3,3:2,3:4,4:3,4:5,5:4,9:16,16:9,21:9. - Fast & cost-efficient — optimized for low-latency, high-volume generation.
Try it now: [Nano Banana Lite Playground](/model/nano-banana-3-api)
API Docs: https://aiboxapi.com/en/api-reference/images/gemini-3.1-flash/generation-lite
gemini-3.1-flash-lite-image
We're excited to announce Gemini 3.1 Flash Lite Image (Nano Banana Lite) — the fastest and most affordable image model in Google's Gemini 3.1 series. Built for high-volume, real-time, and budget-sensitive workflows, it delivers high-quality images with low latency and low cost.
Model ID:
gemini-3.1-flash-lite-imageHighlights
- Fast & low-cost — optimized for speed and scale, billed by input/output tokens
- Text-to-Image & Image-to-Image — one model for generation and reference-guided editing
- Multi-image reference — up to 14 reference images per request
Key Capabilities
- Text-to-Image — generate images directly from text prompts
- Image-to-Image — edit or restyle from reference images with a natural-language instruction
- Multiple Aspect Ratios — 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9
- Reference Images — supply 1–14 images (e.g. objects + characters) to guide generation
Get Started
See the [Gemini 3.1 Flash Lite Image API Documentation](https://aiboxapi.com/en/api-reference/images/gemini-3.1-flash/generation-lite) for request parameters and examples.
gemini-omni-flash-preview
we are excited to announce that Google's Gemini Omni Flash omni-multimodal video generation model is now available on AiBox.
Generate videos with synced audio from text prompts, reference images, or reference videos — and even mix text + image + video in a single request. Powered by Google's official Gemini Omni Flash, it supports conversational, multi-turn editing so you can iterate on a clip without re-uploading anything, all through one unified async task-based API.
- Model ID:
gemini-omni-flash-preview - Generation Modes: Text-to-Video, Image-to-Video, Video-to-Video
- Highlights: Omni multimodal input (text + image + video mixed), conversational multi-turn editing, audio output included, async task API
- Resolution: 720P
- Sizes: 16:9, 9:16
- Duration: 1 – 24 seconds (content-driven, no duration parameter)
🚀 Key Capabilities
- Text-to-Video: Generate dynamic 720p clips with synced audio directly from natural language prompts, with control over subject, scene, motion, and atmosphere
- Image-to-Video: Provide up to 16 reference images through
image_urlsto guide visual style, subject appearance, or scene composition — multi-subject prompts supported - Video-to-Video (Editing): Supply a reference video via
video_urlsto transform environment, add effects, or restyle footage using text instructions - Omni Multimodal Input: Mix text + image + video in a single request (e.g. one character image + one reference clip + a text instruction)
- Conversational Multi-turn Editing: Use
extend_from_task_idto iterate on your previous result — the model keeps the video context, changes only what you ask, and requires no re-upload - Synced Audio Output: Videos are generated with audio included by default
- Natural-language Timing Control: Since there is no
durationparameter, guide pacing/length directly in the prompt (e.g. "after 3 seconds ...", "[0-3s] ...") - Aspect Ratio Control: Choose
16:9(landscape) or9:16(portrait) to control true output orientation - Async Task API: Submit a job, receive a
task_id, and poll for status and results - Token-based Billing: Pay for real usage on AiBox — settled by actual upstream token consumption, with no charge on task failure
🎯 Best For
- Text-to-video clips with synced audio for social and short-form content
- Iterative, conversation-style video editing without re-uploading source clips
- Turning sketches or reference images into motion drafts
- Restyling and effect experiments on short reference footage
- Multi-subject scene animation from mixed image + text input
- Rapid creative exploration where audio-included previews matter
## 📚 Documentation - [Gemini Omni Flash](https://aiboxapi.com/en/api-reference/videos/gemini-omni-flash-preview/generation)
- Model ID:
Doubao Seedance 2.0 Mini
we are excited to announce that ByteDance's Doubao Seedance 2.0 Mini video generation model is now available on AiBox.
Create fast, cost-efficient AI videos from text prompts, reference images, first/last-frame images, reference videos, or audio guidance — all through one unified task-based API designed for rapid video generation with flexible duration, resolution, and aspect ratio control.
- Model ID:
doubao-seedance-2.0-mini - Generation Modes: Text-to-Video, Image-to-Video, Reference-to-Video, Video-to-Video
- Highlights: Lightweight Seedance 2.0 variant, faster iteration, lower-cost video generation, reference image/video/audio support
- Resolution: 480P / 720P
- Sizes: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, adaptive
- Duration: 4 – 15 seconds
- Web Search Field Notice: Official support for
tools: [{"type":"web_search"}]has not yet been confirmed fordoubao-seedance-2.0-mini. Please avoid using it for now, as requests may fail. We will update this note once official support is confirmed.
🚀 Key Capabilities
- Text-to-Video: Generate dynamic videos directly from natural language prompts with control over subject, scene, motion, camera direction, and atmosphere
- Image-to-Video: Provide reference images through
image_urlsto guide visual style, subject appearance, or scene composition - First / Last Frame Control: Use
image_with_roleswithfirst_frameandlast_frameto guide how the video begins and ends - Reference Video Support: Supply up to 3 reference videos via
video_urlsto guide motion, rhythm, or visual continuity - Reference Audio Support: Add up to 3 audio references via
audio_urlswhen generating videos with stronger audio or timing guidance - Flexible Duration: Choose any integer length from 4 to 15 seconds
- Efficient Resolution Options: Generate lightweight outputs in 480P or 720P for faster previews and lower generation cost
- Adaptive Aspect Ratio: Use standard landscape, portrait, square, ultra-wide, or
adaptivesizing for flexible creative workflows - Optional Audio Generation: Enable
generate_audio=trueto create videos with generated audio when supported - Return Last Frame: Use
return_last_frame=trueto retrieve the final frame for multi-shot continuation workflows - Web Search Tool Support: Enable
tools: [{"type":"web_search"}]for prompt contexts that benefit from fresh external information - Form & JSON Modes: Configure parameters visually in the playground or paste raw JSON for full programmatic control
- Resolution × Duration Tiered Pricing: Access Doubao Seedance 2.0 Mini on AiBox with transparent per-second billing — automatic refunds on task failure
🎯 Best For
- Fast video concept exploration and creative iteration
- Lower-cost AI video previews before upgrading to higher-resolution variants
- Social media clips, ad drafts, and short-form content experiments
- Storyboard motion testing from text or reference images
- First-frame to last-frame transition experiments
- Reference-video based motion studies
- Product, character, or scene animation drafts where speed matters
## 📚 Documentation - [Doubao Seedance 2.0 Mini](https://aiboxapi.com/en/api-reference/videos/doubao-seedance-2-0/generation)
- Model ID:
happyhorse-1.1
we are excited to announce that Alibaba Cloud Bailian's HappyHorse 1.1 video generation model is now available on AiBox.
Create cinematic videos from text prompts, a single first-frame image, or a set of reference images — all through one unified task-based API that automatically routes the generation mode based on the fields you supply, with flexible duration, resolution, and aspect ratio control.
- Model ID:
happyhorse-1.1- Generation Modes: Text-to-Video, Image-to-Video, Reference-to-Video - Highlights: Single-model auto-routing, cinematic motion, first-frame image animation, multi-reference subject/style control - Resolution: 720P / 1080P - Sizes: 16:9, 9:16, 1:1, 4:3, 3:4 - Duration: 3 – 15 seconds## 🚀 Key Capabilities
- Text-to-Video: Generate cinematic videos directly from natural language prompts with full control over subject, scene, camera language, and mood - Image-to-Video: Provide a single
first_frame_image(URL or Base64) to animate a static keyframe into motion — the prompt becomes optional - Reference-to-Video: Supply 1 – 9 reference images viaimage_urlsas subject/style references to generate an entirely new scene - Automatic Mode Routing: Route to T2V / I2V / R2V automatically based on the supplied fields (first_frame_image>image_urls> prompt only) — no explicit mode flag needed - Flexible Duration: Choose any integer length from 3 to 15 seconds - Multi-Resolution Output: Generate in 720P or 1080P - 5 Aspect Ratios: Landscape (16:9, 4:3), portrait (9:16, 3:4), and square (1:1) — image-to-video inherits the ratio of the first frame - Optional Watermark: Toggle the watermark on demand (watermark=true) — omitted by default - Form & JSON Modes: Configure parameters visually in the playground or paste raw JSON for full programmatic control - Resolution × Duration Tiered Pricing: Access HappyHorse 1.1 on AiBox with transparent per-second billing — automatic refunds on task failure## 🎯 Best For
- Cinematic shot exploration and pre-production storyboarding
- Advertising and brand video concept previews
- Product storytelling for landing pages and social media
- Animating a single hero image into living motion
- Multi-reference subject/style narrative experiments
- Rapid iteration when turnaround speed matters more than maximum length
## 📚 Documentation
- [HappyHorse 1.1 Generation API](https://aiboxapi.com/en/api-reference/videos/happyhorse-1.1/generation)
Kling 3.0 Turbo
We are excited to announce the launch of Kling 3.0 Turbo, a fast cinematic AI video generation model, now available on AiBox.
Create cinematic videos from text prompts, a single first-frame image, or multi-shot storyboards — all through one unified task-based API with flexible duration, resolution, and aspect ratio control.
- Model ID:
kling-3.0-turbo - Generation Modes: Text-to-Video, Image-to-Video, Multi-Shot Storyboarding
- Highlights: Turbo-fast generation, cinematic motion, prompt-driven multi-shot control, first-frame image animation
- Resolution: 720p / 1080p
- Sizes: 16:9, 9:16, 1:1
- Duration: 3 – 15 seconds
🚀 Key Capabilities
- Text-to-Video: Generate cinematic videos directly from natural language prompts with full control over subject, scene, camera language, and mood
- Image-to-Video: Provide a single
first_frame_image(URL or Base64) to animate a static keyframe into motion — the prompt becomes optional - Automatic Mode Detection: Route to text-to-video or image-to-video automatically based on whether a first-frame image is supplied — no explicit mode flag needed
- Multi-Shot Storyboarding: Express 1 – 6 sequential shots through a single structured prompt (
Shot n,m,words;), with per-shot durations summing to the total length - Flexible Duration: Choose any integer length from 3 to 15 seconds
- Multi-Resolution Output: Generate in 720p or 1080p
- 3 Aspect Ratios: Landscape (16:9), portrait (9:16), and square (1:1) — image-to-video inherits the ratio of the first frame
- Optional Watermark: Toggle the watermark on demand (
watermark=true) — omitted by default - Form & JSON Modes: Configure parameters visually in the playground or paste raw JSON for full programmatic control
- Resolution + Duration Tiered Pricing: Access Kling 3.0 Turbo on AiBox with transparent per-second billing — automatic refunds on task failure
🎯 Best For
- Cinematic shot exploration and pre-production storyboarding
- Advertising and brand video concept previews
- Product storytelling for landing pages and social media
- Animating a single hero image into living motion
- Multi-shot narrative experiments driven entirely by prompt structure
- Rapid iteration when turnaround speed matters more than maximum length
📚 Documentation
- [Kling 3.0 Turbo Generation API](https://aiboxapi.com/en/api-reference/videos/kling-3.0-turbo/generation)
- Model ID:
Midjourney
We're excited to bring Midjourney to AiBox — the full Midjourney image workflow, now available as a clean, programmable API. No Discord bot, no automation glue: generate, refine, and extend your images entirely over HTTP, then poll a task ID until each job completes.
New Model
Midjourney
- Model ID:
midjourney - Versions: v8.1 / v7 / v6.1 / v5.2 / v5.1 / Niji 7 / Niji 6
- Speed Modes: Relax / Fast / Turbo
Generate
- Imagine: Text-to-image generation, with optional reference image prompts to steer style, composition, or subject
- Blend: Merge 2–4 images into a single new composition
- Edits: Re-generate an image guided by a new prompt
Refine & Extend Results
- Upscale: Pick U1–U4 from a grid to get a single high-resolution image
- Variation / High Variation / Low Variation: Produce new variations from a grid or an upscaled image, at adjustable strength
- Reroll: Re-run the whole grid for a fresh set of results
- Zoom: Zoom out / outpaint to reveal more scene around the image
- Pan: Extend the canvas left, right, up, or down
- Inpaint: Vary Region — repaint a masked area with a new prompt
- Remix: v8 reshape (strong / subtle) for guided re-imagining
More Tools
- Describe: Image-to-text — reverse-engineer a prompt from any image
- Video: Image-to-video (i2v) generation with batch sizes of 1 / 2 / 4
Creative Controls
Fine-tune generation with
stylize,chaos,weird,quality, aspect ratio,seed,tile,--no(negative prompts),raw/draft/hdmodes, and style / character / depth references (sref/cref/dref) with per-reference weights.Workflow
- Async Task Polling: Submit a job and poll the returned task ID until generation completes
- Unified API: One integration, the same request/response pattern as every other model on AiBox
Documentation
- [Midjourney API Documentation](https://aiboxapi.com/en/api-reference/images/midjourney/generation)
- Model ID:
pixverse-v6
We are excited to announce the launch of PixVerse V6, a precise cinematic AI video generation model, now available on AiBox.
Create cinematic videos from text prompts, single reference images, first/last frame transitions, multi-image fusion, or by extending existing clips — all through one unified task-based API with flexible duration, resolution, and aspect ratio control.
- Model ID:
pixverse-v6 - Generation Modes: Text-to-Video, Image-to-Video, First/Last Frame Transition, Multi-Reference Fusion, Video Extend
- Highlights: Precise cinematic control, physics-aware motion, high-fidelity portraits, reference media guidance, video continuation
- Resolution: 360p / 540p / 720p / 1080p
- Sizes: 16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2, 21:9
- Duration: 1 – 15 seconds (Transition mode: 5 or 8 seconds)
🚀 Key Capabilities
- Text-to-Video: Generate cinematic videos directly from natural language prompts with full control over subject, scene, camera language, and mood
- Image-to-Video: Provide a single reference image (via
image_urls) to animate static visuals into motion - First/Last Frame Transition: Supply
first_frame_image+last_frame_imageto generate smooth physics-aware transitions between two keyframes - Multi-Reference Fusion: Combine 1 – 7 reference images (
img_references) to fuse outfits, characters, props, and scenes into one generated video - Video Extend: Continue any previously completed PixVerse V6 task via
extend_from_task_idfor longer narrative shots - Flexible Duration: Choose any integer length from 1 to 15 seconds (5 s or 8 s for Transition)
- Multi-Resolution Output: Generate in 360p, 540p, 720p, or 1080p
- 8 Aspect Ratios: Landscape, portrait, square, and cinematic 21:9 widescreen
- Optional Audio Track: Toggle native audio generation (
audio=true) for sound-aware output - Multi-Clip Mode: Enable
generate_multi_clip_switchfor continuous multi-shot storytelling (Text/Image modes) - Form & JSON Modes: Configure parameters visually in the playground or paste raw JSON for full programmatic control
- Resolution + Audio Tiered Pricing: Access PixVerse V6 on AiBox with transparent per-second billing — automatic refunds on task failure
🎯 Best For
- Cinematic shot exploration and pre-production storyboarding
- Advertising and brand video concept previews
- Product storytelling for landing pages and social media
- High-fidelity portrait and character-driven visuals
- Multi-shot narrative experiments with reference image fusion
- Extending existing AI clips into longer continuous scenes
- Physics-aware motion testing before full production
📚 Documentation
- [PixVerse V6 Generation API](https://aiboxapi.com/en/api-reference/videos/pixverse-v6/generation)
- Model ID:
Upcoming Veo 3.1 Pricing Adjustment
Veo 3.1 pricing will be adjusted due to the recent sharp increase in verification and upstream access costs. ### Pricing Update - veo3.1-lite: expected to increase to $0.15 - veo3.1-fast: expected to increase to $0.18 - veo3.1-quality: unchanged We understand pricing changes can affect your workflow. AiBox will continue to monitor upstream costs and keep pricing as transparent as possible. Thank you for your understanding.
gemini-3.5-flash
We are excited to announce that gemini-3.5-flash is now available on AiBox.
- Model ID:
gemini-3.5-flash- Provider: Google - Type: Chat / Text Generation - Highlights: Fast responses, strong instruction following, coding support, and reliable text generation ## 🚀 Key Capabilities - Fast AI chat and assistant responses - Content writing and rewriting - Code generation and debugging - Summarization and data extraction - Simple integration through AiBox's unified APIOmni-Flash-Ext
We are excited to announce the launch of Omni-Flash-Ext, a fast AI video generation model, now available on AiBox.
Create videos from text prompts or optional reference images with flexible duration presets, multiple resolution options, landscape and portrait sizes, and simple task-based API integration.
- Model ID:
Omni-Flash-Ext - Generation Modes: Text-to-Video, Image-to-Video, Reference Image Fusion
- Highlights: Fast generation, simple API workflow, flexible resolution and duration control
- Resolution: 720p / 1080p / 4K
- Sizes: 16:9, 9:16
- Duration: 4 / 6 / 8 / 10 seconds
🚀 Key Capabilities
- Text-to-Video: Generate videos directly from natural language prompts
- Image-to-Video: Upload 1 reference image to guide video generation
- Reference Image Fusion: Upload 3 reference images to combine visual elements into one generated video
- Flexible Duration: Choose from 4 s, 6 s, 8 s, or 10 s video outputs
- Multi-Resolution Output: Generate in 720p, 1080p, or 4K
- Landscape & Portrait Sizes: Supports 16:9 landscape and 9:16 portrait formats
- Form & JSON Modes: Configure parameters visually or paste raw JSON in the playground
- 20% Off Listed Pricing: Access Omni-Flash-Ext on AiBox with transparent resolution-and-duration billing
🎯 Best For
- Fast social media video drafts
- Marketing concepts and campaign previews
- Product storytelling and landing page visuals
- Educational clips and onboarding content
- Creative prototyping before full production
- Reference-image-guided video experiments
📚 Documentation
- [Omni-Flash-Ext Generation API](https://aiboxapi.com/en/api-reference/videos/omni-flash-ext/generation)
- Model ID:
Sora 2
We are excited to announce the launch of Sora 2, OpenAI's next-generation video model, now available on AiBox.
Create high-quality videos from text prompts or reference images with cinematic motion, realistic physics, flexible duration presets, and transparent per-second billing.
✨ New Model
### Sora 2 - Model ID:
sora-2- Generation Modes: Text-to-Video, Image-to-Video - Highlights: Cinematic motion, strong physical realism, scene consistency - Resolution: 720p - Aspect Ratios: 16:9, 9:16 - Duration: 4 / 8 / 12 / 16 / 20 seconds🚀 Key Capabilities
- Text-to-Video: Generate videos directly from text prompts
- Image-to-Video: Provide a reference image to guide the opening frame; landscape or portrait orientation is detected from the image
- Flexible Duration: Choose from five duration presets: 4 s, 8 s, 12 s, 16 s, or 20 s
- Aspect Ratio Coverage: Supports 16:9 landscape and 9:16 portrait formats
- HD Output:
sora-2provides efficient 720p video generation - Form & JSON Modes: Configure parameters visually or paste raw JSON in the playground
- 20% Off Official Pricing: Access Sora 2 on AiBox at discounted pricing, with transparent per-second pay-as-you-go billing
🎯 Best For
- Social media videos in landscape or vertical format
- Product demos and marketing creatives
- Cinematic concept shots
- Storyboarding and pre-production previews
- Image-guided animation from product, character, or scene references
📚 Documentation
- [Sora 2 Generation API](https://aiboxapi.com/en/api-reference/videos/sora-2/generation)