API 更新

新增模型、功能与改进的更新记录。

  1. Suno

    We are excited to announce the launch of the Suno Music API on AiBox, a complete AI music generation and production toolchain with 31 endpoints covering creation, editing, and export.

    New Model

    Suno

    • Model ID: suno
    • Features: Text-to-Music generation with a full suite of derivative and post-processing endpoints
    • Highlights: Unified API covering generation, extension, covers, stem separation, voice personas, audio editing, and analysis
    • Versions: v3.5 / v4 / v4.5 / v4.5+ / v4.5-all / v5 / v5.5
    • Output: Two candidate tracks per generation, with streaming audio, cover art, and lyrics

    Key Capabilities

    • Music Generation: Inspiration mode (describe the song you want) and Custom mode (full control over lyrics, title, style tags, negative tags, vocal gender, and style / weirdness / audio weights)
    • Lyrics Generation: Standalone lyrics endpoint with selectable lyrics models (classic / remi)
    • Audio Upload: Bring your own audio as the source for extension, covers, and vocal / instrumental additions
    • Derivative Works: Extend, Cover, Remaster, Add Vocals, Add Instrumental, Add Stem, Mashup, and Sample-to-Song
    • Stem Separation: 4-track separation, full 12-instrument separation, and isolated vocal (Vox) extraction
    • Voice Personas: Create reusable voice personas from generated tracks and apply them to new generations
    • Audio Editing: Replace Section, Remove Section, Crop, Fade In / Fade Out, Adjust Speed (0.25x–4x with optional pitch preservation), and Full Song synthesis from extended clips
    • Analysis & Export: Word-level aligned lyrics timeline, BPM analysis, MIDI generation, Music Video (MP4) rendering, and WAV export

    Documentation

    • [Suno API Documentation](https://aiboxapi.com/en/api-reference/audios/suno/overview)
  2. FlowMusic

    we are excited to announce that FlowMusic — a full-stack AI music generation suite — is now available on AiBox.

    Turn a single prompt or lyric into studio-ready songs, stems, and music videos. Powered by an upgraded Producer engine (default model Lyria 3 Pro), FlowMusic bundles nine music-creation capabilities — generation, extension, replacement, cover, stem separation, import, download, and video rendering — into one unified async task-based API, so you can build an entire song lifecycle end to end.

    • Model ID: flowmusic
    • Capabilities: Music Generation, Lyrics Generation, Extend, Replace, Cover, Stem Separation, Audio Import, Format Export, Music Video Rendering
    • Highlights: 9 endpoints in one model, lossless WAV + streaming M4A, auto-generated cover art, karaoke timing markers, async task API
    • Output: M4A (streaming) + WAV (lossless) + optional MP4 music video
    • Controls: BPM, length up to 240s, seed for reproducibility

    🚀 Key Capabilities

    • Text/Lyrics-to-Music: Generate a full song from a sound prompt and/or lyrics, with control over BPM, length (up to 240s), and seed — each request returns one track
    • Lyrics Generation: Turn any idea into structured lyrics (≤ 3000 chars), then feed them straight back into music generation
    • Music Extend: Continue an existing clip from any timestamp — up to 327 seconds of new audio — guided by a natural-language instruction
    • Section Replace: Regenerate a specific time range within a song (e.g. swap a chorus to piano) without touching the rest
    • Cover / Restyle: Re-arrange a whole track into a new style with adjustable edit strength (0–1)
    • Stem Separation: Split vocals and accompaniment into a downloadable multi-track ZIP
    • Audio Import: Bring external audio in via audio_url to obtain a clip_id for downstream editing
    • Format Export: Download any clip as wav or mp3
    • Music Video Rendering: Render a clip into MP4 with simple / modern / player presets
    • Karaoke Timing: Lyrics timing markers included for word-level highlighting
    • Usage-based Billing: Charged on submission with full auto-refund on task failure ($0.06 generate/extend/replace/cover/stems · $0.02 lyrics/download/video · $0.01 import)

    🎯 Best For

    • Producing complete, lyric-driven songs from a single text idea
    • Iterating on tracks — extend, replace a section, or restyle without starting over
    • Generating royalty-style background music for video, ads, and games
    • Building karaoke experiences with word-level lyric timing
    • Extracting stems for remixing and post-production
    • Turning existing audio into new covers and arrangements

    ## 📚 Documentation - [FlowMusic](https://aiboxapi.com/en/api-reference/audios/flow-music/music)

  3. seedream-5-0-pro

    We are excited to announce the launch of Seedream 5.0 Pro, ByteDance's latest quality-first text-to-image model with best-in-class text rendering and unified generation and editing.

    New Model

    Seedream 5.0 Pro

    • Model ID: doubao-seedream-5-0-pro
    • Features: Text-to-Image and Image-to-Image generation with unified editing in a single API call
    • Highlights: Cinematic, quality-first image generation with industry-leading text rendering
    • Resolution: 1K / 2K (default 2K)
    • Output: One image per request (PNG / JPEG)

    Key Capabilities

    • Text-to-Image: Generate cinematic, richly detailed images directly from text descriptions
    • Image-to-Image: Edit and refine existing images — remove objects, replace backgrounds, or restyle with simple prompts
    • Perfect Text Rendering: Produce accurate, readable text in images for posters, banners, and marketing visuals
    • Multi-Reference Fusion: Blend subjects, styles, and outfits from multiple reference images in one request
    • Multiple Aspect Ratios: 1:1, 4:3, 3:4, 16:9, 9:16, 3:2, 2:3, 21:9, and Auto
    • Reference Images: Support for up to 10 reference images for style and character consistency
    • Watermark: Optional watermark on generated images

    Notes

    • Quality-first model that prioritizes fidelity and detail over raw speed (1K ~90s, 2K ~160s)
    • Asynchronous generation: submit to /v1/images/generations, then poll /v1/tasks/{task_id} until the task completes

    Documentation

    • [Seedream 5.0 Pro API Documentation](https://aiboxapi.com/en/api-reference/images/seedream-5-0-pro/generation)
  4. gemini-3.1-flash-lite-image-ext

    We've added `gemini-3.1-flash-lite-image-ext` — a fast, low-cost image generation model now available on the Nano Banana page. Select it directly from the model dropdown to get started.

    Capabilities

    • Text-to-image & image-to-image — generate from a prompt, or guide generation with reference images.
    • Up to 14 reference images per request.
    • Aspect ratios: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9.
    • Fast & cost-efficient — optimized for low-latency, high-volume generation.

    Try it now: [Nano Banana Lite Playground](/model/nano-banana-3-api)

    API Docs: https://aiboxapi.com/en/api-reference/images/gemini-3.1-flash/generation-lite

  5. gemini-3.1-flash-lite-image

    We're excited to announce Gemini 3.1 Flash Lite Image (Nano Banana Lite) — the fastest and most affordable image model in Google's Gemini 3.1 series. Built for high-volume, real-time, and budget-sensitive workflows, it delivers high-quality images with low latency and low cost.

    Model ID: gemini-3.1-flash-lite-image

    Highlights

    • Fast & low-cost — optimized for speed and scale, billed by input/output tokens
    • Text-to-Image & Image-to-Image — one model for generation and reference-guided editing
    • Multi-image reference — up to 14 reference images per request

    Key Capabilities

    • Text-to-Image — generate images directly from text prompts
    • Image-to-Image — edit or restyle from reference images with a natural-language instruction
    • Multiple Aspect Ratios — 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9
    • Reference Images — supply 1–14 images (e.g. objects + characters) to guide generation

    Get Started

    See the [Gemini 3.1 Flash Lite Image API Documentation](https://aiboxapi.com/en/api-reference/images/gemini-3.1-flash/generation-lite) for request parameters and examples.

  6. gemini-omni-flash-preview

    we are excited to announce that Google's Gemini Omni Flash omni-multimodal video generation model is now available on AiBox.

    Generate videos with synced audio from text prompts, reference images, or reference videos — and even mix text + image + video in a single request. Powered by Google's official Gemini Omni Flash, it supports conversational, multi-turn editing so you can iterate on a clip without re-uploading anything, all through one unified async task-based API.

    • Model ID: gemini-omni-flash-preview
    • Generation Modes: Text-to-Video, Image-to-Video, Video-to-Video
    • Highlights: Omni multimodal input (text + image + video mixed), conversational multi-turn editing, audio output included, async task API
    • Resolution: 720P
    • Sizes: 16:9, 9:16
    • Duration: 1 – 24 seconds (content-driven, no duration parameter)

    🚀 Key Capabilities

    • Text-to-Video: Generate dynamic 720p clips with synced audio directly from natural language prompts, with control over subject, scene, motion, and atmosphere
    • Image-to-Video: Provide up to 16 reference images through image_urls to guide visual style, subject appearance, or scene composition — multi-subject prompts supported
    • Video-to-Video (Editing): Supply a reference video via video_urls to transform environment, add effects, or restyle footage using text instructions
    • Omni Multimodal Input: Mix text + image + video in a single request (e.g. one character image + one reference clip + a text instruction)
    • Conversational Multi-turn Editing: Use extend_from_task_id to iterate on your previous result — the model keeps the video context, changes only what you ask, and requires no re-upload
    • Synced Audio Output: Videos are generated with audio included by default
    • Natural-language Timing Control: Since there is no duration parameter, guide pacing/length directly in the prompt (e.g. "after 3 seconds ...", "[0-3s] ...")
    • Aspect Ratio Control: Choose 16:9 (landscape) or 9:16 (portrait) to control true output orientation
    • Async Task API: Submit a job, receive a task_id, and poll for status and results
    • Token-based Billing: Pay for real usage on AiBox — settled by actual upstream token consumption, with no charge on task failure

    🎯 Best For

    • Text-to-video clips with synced audio for social and short-form content
    • Iterative, conversation-style video editing without re-uploading source clips
    • Turning sketches or reference images into motion drafts
    • Restyling and effect experiments on short reference footage
    • Multi-subject scene animation from mixed image + text input
    • Rapid creative exploration where audio-included previews matter

    ## 📚 Documentation - [Gemini Omni Flash](https://aiboxapi.com/en/api-reference/videos/gemini-omni-flash-preview/generation)

  7. Doubao Seedance 2.0 Mini

    we are excited to announce that ByteDance's Doubao Seedance 2.0 Mini video generation model is now available on AiBox.

    Create fast, cost-efficient AI videos from text prompts, reference images, first/last-frame images, reference videos, or audio guidance — all through one unified task-based API designed for rapid video generation with flexible duration, resolution, and aspect ratio control.

    • Model ID: doubao-seedance-2.0-mini
    • Generation Modes: Text-to-Video, Image-to-Video, Reference-to-Video, Video-to-Video
    • Highlights: Lightweight Seedance 2.0 variant, faster iteration, lower-cost video generation, reference image/video/audio support
    • Resolution: 480P / 720P
    • Sizes: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, adaptive
    • Duration: 4 – 15 seconds
    • Web Search Field Notice: Official support for tools: [{"type":"web_search"}] has not yet been confirmed for doubao-seedance-2.0-mini. Please avoid using it for now, as requests may fail. We will update this note once official support is confirmed.

    🚀 Key Capabilities

    • Text-to-Video: Generate dynamic videos directly from natural language prompts with control over subject, scene, motion, camera direction, and atmosphere
    • Image-to-Video: Provide reference images through image_urls to guide visual style, subject appearance, or scene composition
    • First / Last Frame Control: Use image_with_roles with first_frame and last_frame to guide how the video begins and ends
    • Reference Video Support: Supply up to 3 reference videos via video_urls to guide motion, rhythm, or visual continuity
    • Reference Audio Support: Add up to 3 audio references via audio_urls when generating videos with stronger audio or timing guidance
    • Flexible Duration: Choose any integer length from 4 to 15 seconds
    • Efficient Resolution Options: Generate lightweight outputs in 480P or 720P for faster previews and lower generation cost
    • Adaptive Aspect Ratio: Use standard landscape, portrait, square, ultra-wide, or adaptive sizing for flexible creative workflows
    • Optional Audio Generation: Enable generate_audio=true to create videos with generated audio when supported
    • Return Last Frame: Use return_last_frame=true to retrieve the final frame for multi-shot continuation workflows
    • Web Search Tool Support: Enable tools: [{"type":"web_search"}] for prompt contexts that benefit from fresh external information
    • Form & JSON Modes: Configure parameters visually in the playground or paste raw JSON for full programmatic control
    • Resolution × Duration Tiered Pricing: Access Doubao Seedance 2.0 Mini on AiBox with transparent per-second billing — automatic refunds on task failure

    🎯 Best For

    • Fast video concept exploration and creative iteration
    • Lower-cost AI video previews before upgrading to higher-resolution variants
    • Social media clips, ad drafts, and short-form content experiments
    • Storyboard motion testing from text or reference images
    • First-frame to last-frame transition experiments
    • Reference-video based motion studies
    • Product, character, or scene animation drafts where speed matters

    ## 📚 Documentation - [Doubao Seedance 2.0 Mini](https://aiboxapi.com/en/api-reference/videos/doubao-seedance-2-0/generation)

  8. happyhorse-1.1

    we are excited to announce that Alibaba Cloud Bailian's HappyHorse 1.1 video generation model is now available on AiBox.

    Create cinematic videos from text prompts, a single first-frame image, or a set of reference images — all through one unified task-based API that automatically routes the generation mode based on the fields you supply, with flexible duration, resolution, and aspect ratio control.

    - Model ID: happyhorse-1.1 - Generation Modes: Text-to-Video, Image-to-Video, Reference-to-Video - Highlights: Single-model auto-routing, cinematic motion, first-frame image animation, multi-reference subject/style control - Resolution: 720P / 1080P - Sizes: 16:9, 9:16, 1:1, 4:3, 3:4 - Duration: 3 – 15 seconds

    ## 🚀 Key Capabilities

    - Text-to-Video: Generate cinematic videos directly from natural language prompts with full control over subject, scene, camera language, and mood - Image-to-Video: Provide a single first_frame_image (URL or Base64) to animate a static keyframe into motion — the prompt becomes optional - Reference-to-Video: Supply 1 – 9 reference images via image_urls as subject/style references to generate an entirely new scene - Automatic Mode Routing: Route to T2V / I2V / R2V automatically based on the supplied fields (first_frame_image > image_urls > prompt only) — no explicit mode flag needed - Flexible Duration: Choose any integer length from 3 to 15 seconds - Multi-Resolution Output: Generate in 720P or 1080P - 5 Aspect Ratios: Landscape (16:9, 4:3), portrait (9:16, 3:4), and square (1:1) — image-to-video inherits the ratio of the first frame - Optional Watermark: Toggle the watermark on demand (watermark=true) — omitted by default - Form & JSON Modes: Configure parameters visually in the playground or paste raw JSON for full programmatic control - Resolution × Duration Tiered Pricing: Access HappyHorse 1.1 on AiBox with transparent per-second billing — automatic refunds on task failure

    ## 🎯 Best For

    • Cinematic shot exploration and pre-production storyboarding
    • Advertising and brand video concept previews
    • Product storytelling for landing pages and social media
    • Animating a single hero image into living motion
    • Multi-reference subject/style narrative experiments
    • Rapid iteration when turnaround speed matters more than maximum length

    ## 📚 Documentation

    - [HappyHorse 1.1 Generation API](https://aiboxapi.com/en/api-reference/videos/happyhorse-1.1/generation)

  9. Kling 3.0 Turbo

    We are excited to announce the launch of Kling 3.0 Turbo, a fast cinematic AI video generation model, now available on AiBox.

    Create cinematic videos from text prompts, a single first-frame image, or multi-shot storyboards — all through one unified task-based API with flexible duration, resolution, and aspect ratio control.

    • Model ID: kling-3.0-turbo
    • Generation Modes: Text-to-Video, Image-to-Video, Multi-Shot Storyboarding
    • Highlights: Turbo-fast generation, cinematic motion, prompt-driven multi-shot control, first-frame image animation
    • Resolution: 720p / 1080p
    • Sizes: 16:9, 9:16, 1:1
    • Duration: 3 – 15 seconds

    🚀 Key Capabilities

    • Text-to-Video: Generate cinematic videos directly from natural language prompts with full control over subject, scene, camera language, and mood
    • Image-to-Video: Provide a single first_frame_image (URL or Base64) to animate a static keyframe into motion — the prompt becomes optional
    • Automatic Mode Detection: Route to text-to-video or image-to-video automatically based on whether a first-frame image is supplied — no explicit mode flag needed
    • Multi-Shot Storyboarding: Express 1 – 6 sequential shots through a single structured prompt (Shot n,m,words;), with per-shot durations summing to the total length
    • Flexible Duration: Choose any integer length from 3 to 15 seconds
    • Multi-Resolution Output: Generate in 720p or 1080p
    • 3 Aspect Ratios: Landscape (16:9), portrait (9:16), and square (1:1) — image-to-video inherits the ratio of the first frame
    • Optional Watermark: Toggle the watermark on demand (watermark=true) — omitted by default
    • Form & JSON Modes: Configure parameters visually in the playground or paste raw JSON for full programmatic control
    • Resolution + Duration Tiered Pricing: Access Kling 3.0 Turbo on AiBox with transparent per-second billing — automatic refunds on task failure

    🎯 Best For

    • Cinematic shot exploration and pre-production storyboarding
    • Advertising and brand video concept previews
    • Product storytelling for landing pages and social media
    • Animating a single hero image into living motion
    • Multi-shot narrative experiments driven entirely by prompt structure
    • Rapid iteration when turnaround speed matters more than maximum length

    📚 Documentation

    • [Kling 3.0 Turbo Generation API](https://aiboxapi.com/en/api-reference/videos/kling-3.0-turbo/generation)
  10. Midjourney

    We're excited to bring Midjourney to AiBox — the full Midjourney image workflow, now available as a clean, programmable API. No Discord bot, no automation glue: generate, refine, and extend your images entirely over HTTP, then poll a task ID until each job completes.

    New Model

    Midjourney

    • Model ID: midjourney
    • Versions: v8.1 / v7 / v6.1 / v5.2 / v5.1 / Niji 7 / Niji 6
    • Speed Modes: Relax / Fast / Turbo

    Generate

    • Imagine: Text-to-image generation, with optional reference image prompts to steer style, composition, or subject
    • Blend: Merge 2–4 images into a single new composition
    • Edits: Re-generate an image guided by a new prompt

    Refine & Extend Results

    • Upscale: Pick U1–U4 from a grid to get a single high-resolution image
    • Variation / High Variation / Low Variation: Produce new variations from a grid or an upscaled image, at adjustable strength
    • Reroll: Re-run the whole grid for a fresh set of results
    • Zoom: Zoom out / outpaint to reveal more scene around the image
    • Pan: Extend the canvas left, right, up, or down
    • Inpaint: Vary Region — repaint a masked area with a new prompt
    • Remix: v8 reshape (strong / subtle) for guided re-imagining

    More Tools

    • Describe: Image-to-text — reverse-engineer a prompt from any image
    • Video: Image-to-video (i2v) generation with batch sizes of 1 / 2 / 4

    Creative Controls

    Fine-tune generation with stylize, chaos, weird, quality, aspect ratio, seed, tile, --no (negative prompts), raw / draft / hd modes, and style / character / depth references (sref / cref / dref) with per-reference weights.

    Workflow

    • Async Task Polling: Submit a job and poll the returned task ID until generation completes
    • Unified API: One integration, the same request/response pattern as every other model on AiBox

    Documentation

    • [Midjourney API Documentation](https://aiboxapi.com/en/api-reference/images/midjourney/generation)
  11. pixverse-v6

    We are excited to announce the launch of PixVerse V6, a precise cinematic AI video generation model, now available on AiBox.

    Create cinematic videos from text prompts, single reference images, first/last frame transitions, multi-image fusion, or by extending existing clips — all through one unified task-based API with flexible duration, resolution, and aspect ratio control.

    • Model ID: pixverse-v6
    • Generation Modes: Text-to-Video, Image-to-Video, First/Last Frame Transition, Multi-Reference Fusion, Video Extend
    • Highlights: Precise cinematic control, physics-aware motion, high-fidelity portraits, reference media guidance, video continuation
    • Resolution: 360p / 540p / 720p / 1080p
    • Sizes: 16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2, 21:9
    • Duration: 1 – 15 seconds (Transition mode: 5 or 8 seconds)

    🚀 Key Capabilities

    • Text-to-Video: Generate cinematic videos directly from natural language prompts with full control over subject, scene, camera language, and mood
    • Image-to-Video: Provide a single reference image (via image_urls) to animate static visuals into motion
    • First/Last Frame Transition: Supply first_frame_image + last_frame_image to generate smooth physics-aware transitions between two keyframes
    • Multi-Reference Fusion: Combine 1 – 7 reference images (img_references) to fuse outfits, characters, props, and scenes into one generated video
    • Video Extend: Continue any previously completed PixVerse V6 task via extend_from_task_id for longer narrative shots
    • Flexible Duration: Choose any integer length from 1 to 15 seconds (5 s or 8 s for Transition)
    • Multi-Resolution Output: Generate in 360p, 540p, 720p, or 1080p
    • 8 Aspect Ratios: Landscape, portrait, square, and cinematic 21:9 widescreen
    • Optional Audio Track: Toggle native audio generation (audio=true) for sound-aware output
    • Multi-Clip Mode: Enable generate_multi_clip_switch for continuous multi-shot storytelling (Text/Image modes)
    • Form & JSON Modes: Configure parameters visually in the playground or paste raw JSON for full programmatic control
    • Resolution + Audio Tiered Pricing: Access PixVerse V6 on AiBox with transparent per-second billing — automatic refunds on task failure

    🎯 Best For

    • Cinematic shot exploration and pre-production storyboarding
    • Advertising and brand video concept previews
    • Product storytelling for landing pages and social media
    • High-fidelity portrait and character-driven visuals
    • Multi-shot narrative experiments with reference image fusion
    • Extending existing AI clips into longer continuous scenes
    • Physics-aware motion testing before full production

    📚 Documentation

    • [PixVerse V6 Generation API](https://aiboxapi.com/en/api-reference/videos/pixverse-v6/generation)
  12. Upcoming Veo 3.1 Pricing Adjustment

    Veo 3.1 pricing will be adjusted due to the recent sharp increase in verification and upstream access costs. ### Pricing Update - veo3.1-lite: expected to increase to $0.15 - veo3.1-fast: expected to increase to $0.18 - veo3.1-quality: unchanged We understand pricing changes can affect your workflow. AiBox will continue to monitor upstream costs and keep pricing as transparent as possible. Thank you for your understanding.

  13. gemini-3.5-flash

    We are excited to announce that gemini-3.5-flash is now available on AiBox.

    - Model ID: gemini-3.5-flash - Provider: Google - Type: Chat / Text Generation - Highlights: Fast responses, strong instruction following, coding support, and reliable text generation ## 🚀 Key Capabilities - Fast AI chat and assistant responses - Content writing and rewriting - Code generation and debugging - Summarization and data extraction - Simple integration through AiBox's unified API

  14. Omni-Flash-Ext

    We are excited to announce the launch of Omni-Flash-Ext, a fast AI video generation model, now available on AiBox.

    Create videos from text prompts or optional reference images with flexible duration presets, multiple resolution options, landscape and portrait sizes, and simple task-based API integration.

    • Model ID: Omni-Flash-Ext
    • Generation Modes: Text-to-Video, Image-to-Video, Reference Image Fusion
    • Highlights: Fast generation, simple API workflow, flexible resolution and duration control
    • Resolution: 720p / 1080p / 4K
    • Sizes: 16:9, 9:16
    • Duration: 4 / 6 / 8 / 10 seconds

    🚀 Key Capabilities

    • Text-to-Video: Generate videos directly from natural language prompts
    • Image-to-Video: Upload 1 reference image to guide video generation
    • Reference Image Fusion: Upload 3 reference images to combine visual elements into one generated video
    • Flexible Duration: Choose from 4 s, 6 s, 8 s, or 10 s video outputs
    • Multi-Resolution Output: Generate in 720p, 1080p, or 4K
    • Landscape & Portrait Sizes: Supports 16:9 landscape and 9:16 portrait formats
    • Form & JSON Modes: Configure parameters visually or paste raw JSON in the playground
    • 20% Off Listed Pricing: Access Omni-Flash-Ext on AiBox with transparent resolution-and-duration billing

    🎯 Best For

    • Fast social media video drafts
    • Marketing concepts and campaign previews
    • Product storytelling and landing page visuals
    • Educational clips and onboarding content
    • Creative prototyping before full production
    • Reference-image-guided video experiments

    📚 Documentation

    • [Omni-Flash-Ext Generation API](https://aiboxapi.com/en/api-reference/videos/omni-flash-ext/generation)
  15. Sora 2

    We are excited to announce the launch of Sora 2, OpenAI's next-generation video model, now available on AiBox.

    Create high-quality videos from text prompts or reference images with cinematic motion, realistic physics, flexible duration presets, and transparent per-second billing.

    ✨ New Model

    ### Sora 2 - Model ID: sora-2 - Generation Modes: Text-to-Video, Image-to-Video - Highlights: Cinematic motion, strong physical realism, scene consistency - Resolution: 720p - Aspect Ratios: 16:9, 9:16 - Duration: 4 / 8 / 12 / 16 / 20 seconds

    🚀 Key Capabilities

    • Text-to-Video: Generate videos directly from text prompts
    • Image-to-Video: Provide a reference image to guide the opening frame; landscape or portrait orientation is detected from the image
    • Flexible Duration: Choose from five duration presets: 4 s, 8 s, 12 s, 16 s, or 20 s
    • Aspect Ratio Coverage: Supports 16:9 landscape and 9:16 portrait formats
    • HD Output: sora-2 provides efficient 720p video generation
    • Form & JSON Modes: Configure parameters visually or paste raw JSON in the playground
    • 20% Off Official Pricing: Access Sora 2 on AiBox at discounted pricing, with transparent per-second pay-as-you-go billing

    🎯 Best For

    • Social media videos in landscape or vertical format
    • Product demos and marketing creatives
    • Cinematic concept shots
    • Storyboarding and pre-production previews
    • Image-guided animation from product, character, or scene references

    📚 Documentation

    • [Sora 2 Generation API](https://aiboxapi.com/en/api-reference/videos/sora-2/generation)
  16. Doubao Seedance 2.0 - Avatar Material Review

    We are excited to announce the launch of Doubao Seedance 2.0 Avatar Review, Doubao's unified avatar asset review pipeline for both virtual subjects and verified real persons, now available on AiBox. Submit images, videos, or audio for content review and turn them into reusable assets you can plug straight into Seedance 2.0 video generation, with auto group lifecycle, real-person H5 verification, and transparent token-based access. ### Doubao Seedance 2.0 Avatar Review - Endpoints: private-avatar, real-avatar - Generation Workflow: Group Creation + Batch Asset Review + Reusable Asset Activation - Highlights: Auto group lifecycle, 3-step real-person flow, batch up to 20 assets per request - Models: doubao-seedance-2.0, doubao-seedance-2.0-fast - Asset Types: image, video, audio - Core Inputs: asset_type, assets (required), group_id or group_name (optional) - Real-Person Inputs: callback_url, byted_token, group_id - Pipelines: private-avatar, real-avatar - Task Polling: Unified task status query ## 🚀 Key Capabilities - Two Independent Pipelines: Use private-avatar for direct batch review, or real-avatar for full H5 identity verification - Auto Group Lifecycle: Submit without group_id to auto-create an AIGC group, or reuse an existing one to keep reviewed assets organized - 3-Step Real-Person Flow: Create H5 session, query verification result, then batch submit verified assets in one stateful UI - Batch Asset Submission: Up to 20 images, videos, or audio files per request, each with optional name - Reusable Approved Assets: Verified assets surface as ready-to-use references for Seedance 2.0 video generation - Form & JSON Modes: Configure parameters visually or paste raw JSON in the playground - Multilingual UI: Full labels in en, ja, ko, ru, zh on day one - Token-Based Access, No Pre-Billing: Material review endpoints run on token auth and rate limit, with no credit pre-deduction ## 🎯 Best For - Brand-safe AI video generation backed by pre-reviewed subject material - Real-person endorsement and spokesperson videos requiring identity verification - Studios building reusable avatar libraries with clear approved versus raw separation - Compliance-driven workflows that need an audit trail of group and asset IDs - Teams integrating Seedance 2.0 with existing object-storage pipelines - Multilingual product surfaces shipping to en, ja, ko, ru, zh markets ## 📚 Documentation - [Submit Private Virtual Avatar](https://aiboxapi.com/en/api-reference/videos/doubao-seedance-2-0/private-avatar) - [Submit Real-Person Avatar (H5 Verification)](https://aiboxapi.com/en/api-reference/videos/doubao-seedance-2-0/real-avatar)

  17. Kling Motion Control

    We are excited to announce the launch of Kling Motion Control, Kling's reference-driven motion control video model, now available on AiBox. Create high-quality videos from a reference image and reference video with controllable motion transfer, realistic physics, flexible orientation controls, and transparent per-second billing. ### Kling Motion Control - Model IDs: kling-v3-motion-control, kling-v2-6-motion-control - Generation Workflow: Reference Image + Reference Video Motion Control - Highlights: Stable subject consistency, controllable motion transfer, realistic movement physics - Core Inputs: image_url, video_url (required), prompt (optional) - Character Orientation: image, video - Mode: std, pro - Audio & Watermark: keep_original_sound, watermark_info - Mode: std, pro ## 🚀 Key Capabilities - Reference-Guided Motion Control: Generate videos by combining a subject reference image with a motion reference video - Subject Consistency: Keep character/object identity stable while transferring motion patterns - Controllable Orientation: Choose character_orientation (image or video) to match composition intent - Flexible Quality Mode: Select std or pro mode based on speed and quality requirements - Prompt-Enhanced Generation: Add optional prompt text to refine scene style and generation intent - Form & JSON Modes: Configure parameters visually or paste raw JSON in the playground - Transparent Per-Second Billing: Access Kling Motion Control on AiBox with clear pay-as-you-go pricing ## 🎯 Best For - Reference-driven social media video creation - Product and character motion showcase videos - Storyboarding and pre-production previs - Motion-controlled marketing creatives - Teams requiring API-based controllable video generation - Workflows that need stable identity + predictable motion transfer ## 📚 Documentation - [Kling Motion Control Generation API](https://aiboxapi.com/cn/api-reference/videos/kling-v2-6/kling-v2-6-motion-control-generation)

  18. Sora 2 Pro Now Available on AiBox

    We are excited to announce the launch of Sora 2 Pro, OpenAI's next-generation professional video model, now available on AiBox.

    Create high-quality videos from text prompts or reference images with cinematic motion, realistic physics, flexible duration presets, full HD output, and transparent per-second billing.

    Sora 2 Pro

    • Model ID: sora-2-pro
    • Generation Modes: Text-to-Video, Image-to-Video
    • Highlights: Higher fidelity, sharper detail, cinematic motion, strong physical realism
    • Resolution: 720p / 1024p / 1080p
    • Aspect Ratios: 16:9, 9:16
    • Duration: 4 / 8 / 12 / 16 / 20 seconds

    🚀 Key Capabilities

    • Text-to-Video: Generate videos directly from text prompts
    • Image-to-Video: Provide a reference image to guide the opening frame; landscape or portrait orientation is detected from the image
    • Flexible Duration: Choose from five duration presets: 4 s, 8 s, 12 s, 16 s, or 20 s
    • Aspect Ratio Coverage: Supports 16:9 landscape and 9:16 portrait formats
    • Full HD Output: Generate high-fidelity videos with 720p, 1024p, and 1080p resolution tiers
    • Form & JSON Modes: Configure parameters visually or paste raw JSON in the playground
    • 20% Off Official Pricing: Access Sora 2 Pro on AiBox at discounted pricing, with transparent per-second pay-as-you-go billing

    🎯 Best For

    • Social media videos in landscape or vertical format
    • Product demos and marketing creatives
    • Cinematic concept shots
    • Storyboarding and pre-production previews
    • Image-guided animation from product, character, or scene references
    • High-fidelity video generation requiring sharper detail and full HD output

    📚 Documentation

    • [Sora 2 Pro Generation API](https://aiboxapi.com/cn/api-reference/videos/sora-2/generation)
  19. Claude Opus 4.7 & DeepSeek V4

    We are excited to announce the launch of three powerful new text generation models on AiBox: Claude Opus 4.7, DeepSeek V4 Flash, and DeepSeek V4 Pro, all available today.

    New Models

    Claude Opus 4.7

    • Model ID: claude-opus-4-7
    • Provider: Anthropic
    • Best For: Complex reasoning, advanced code generation, long-document understanding and analysis
    • Integration: Fully compatible with OpenAI Chat Completions format

    DeepSeek V4 Flash

    • Model ID: deepseek-v4-flash
    • Provider: DeepSeek
    • Best For: High-throughput, low-latency, and real-time interactive applications
    • Highlights: Lightweight and ultra-fast — significantly lower response time compared to the Pro variant

    DeepSeek V4 Pro

    • Model ID: deepseek-v4-pro
    • Provider: DeepSeek
    • Best For: Deep reasoning, complex multi-turn conversations, and long-context tasks
    • Highlights: Flagship reasoning capability, optimized for accuracy-critical use cases

  20. 4K Mode for Kling V3 & Kling V3 Omni

    We are excited to announce that Kling V3 and Kling V3 Omni now support 4K resolution, bringing ultra-high-definition video generation to the Kling model family.

    What's New

    4K Mode for Kling V3 & Kling V3 Omni

    • Models: kling-v3, kling-v3-omni
    • New Parameter Value: mode: "4k" (in addition to existing std 720P and pro 1080P)
    • Highlights: Ultra-high-definition output for cinematic-grade video generation
    • Audio Support: 4K mode is fully compatible with synchronized audio generation

    Key Capabilities

    • 4K Resolution: Generate ultra-HD videos with mode: "4k" for both Text-to-Video and Image-to-Video
    • Synchronized Audio: Combine mode: "4k" with audio generation for fully immersive output
    • Aspect Ratios: 16:9, 9:16, 1:1
    • Duration: 5 / 10 seconds
    • Note: For kling-v3-omni, the 4K tier does not support reference video input — reference images remain supported

    Documentation

    • https://aiboxapi.com/en/api-reference/videos/kling-v3/generation
    • https://aiboxapi.com/en/api-reference/videos/kling-v3-omni/generation
  21. GPT Image 2

    We are excited to announce the launch of GPT Image 2, OpenAI's next-generation image model — available on AiBox through both the primary channel and an official-channel fallback. New Model

    GPT Image 2

    • Model ID: gpt-image-2 (primary) / gpt-image-2-official (official channel)
    • Features: Text-to-Image and Image-to-Image generation with up to 16 reference images
    • Highlights: LMArena preview model — near-perfect text rendering, native 4K output, and deep world knowledge
    • Resolution: 1K / 2K / 4K
    • Aspect Ratios: 13 ratios + Auto mode

    Key Capabilities

    • Text-to-Image: Generate images directly from text prompts
    • Image-to-Image: Guide generation with up to 16 reference images for style, subject, or composition control
    • Native 4K Output: True 4K rendering on widescreen ratios (16:9, 9:16, 2:1, 1:2, 21:9, 9:21); 1K and 2K available across all 13 ratios
    • Aspect Ratio Coverage: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 5:4, 4:5, 2:1, 1:2, 21:9, 9:21, plus Auto
    • Dual-Channel Reliability: Enable official_fallback on the primary channel to automatically retry on gpt-image-2-official if the primary route fails
    • Form & JSON Modes: Build requests via the visual form or paste a raw JSON config in the playground

    Documentation

    • https://aiboxapi.com/en/api-reference/images/gpt-image-2/generation
    • https://aiboxapi.com/en/api-reference/images/gpt-image-2/official
  22. HappyHorse 1.0

    We are excited to announce the launch of HappyHorse 1.0, Alibaba's latest video generation model with native audio synthesis and multilingual lip-sync. New Model HappyHorse 1.0

    • Model ID: happyhorse-1.0
    • Features: Text-to-Video, Image-to-Video, Reference-to-Video, and Video Editing in a single unified model
    • Highlights: Native AI-generated audio with 7-language lip-sync support
    • Resolution: 720P / 1080P
    • Duration: 3-15 seconds

    Key Capabilities

    • Text-to-Video (T2V): Generate videos directly from text prompts (up to 2,500 characters)
    • Image-to-Video (I2V): Animate a single first-frame image into a full video clip
    • Reference-to-Video (R2V): Guide generation with 1-9 reference images for consistent style and subjects
    • Video Editing (EDIT): Re-edit an existing video (MP4/MOV, 3-60s, ≤100MB) with optional 0-5 style reference images
    • Native Audio Generation: Synthesized voice, sound effects, and background music; in EDIT mode choose auto (regenerate) or origin (keep source audio)
    • Multiple Aspect Ratios: 16:9, 9:16, 1:1, 4:3, 3:4
    • Reproducibility: Optional seed (0 - 2,147,483,647) and watermark toggle

    Documentation

    • https://aiboxapi.com/en/api-reference/videos/happyhorse-1.0/generation
  23. Add 10 new GPT-5 series text models

    We have launched the OpenAI GPT-5 series with 10 brand-new text models, all designed for conversation and code generation. These models cover a wide range of scenarios including general chat, complex reasoning, and code completion. All models are accessible through the unified Chat Completions API — no additional integration required.

    New Models 🚀 Flagship & General Conversation

    • gpt-5.5 — Next-generation flagship text model with the strongest overall capabilities
    • gpt-5.4-pro — Built for complex reasoning and long-context text tasks
    • gpt-5.4 — A balanced, mainstream conversation model offering strong performance at lower cost
    • gpt-5.2 — General-purpose chat and text generation

    ⚡ Lightweight & Cost-Effective

    • gpt-5.4-mini — Mid-sized model with faster response times and lower cost
    • gpt-5.4-nano — Ultra-lightweight, ideal for high-concurrency, low-latency text scenarios

    💻 Codex Code-Specialized

    • gpt-5.3-codex — Code generation and code understanding
    • gpt-5.2-codex — A balanced version for code-related tasks
    • gpt-5.1-codex-max — Large-context code tasks, ideal for repository-level code analysis
    • gpt-5.1-codex-mini — Lightweight code completion with low latency

    For the complete model list, please visit /model.

  24. New Model | Seedance 1.5 Pro Video Generation with Audio

    We are excited to announce the launch of Seedance 1.5 Pro, BytePlus's latest video generation model with integrated audio generation capabilities.

    New Model

    Seedance 1.5 Pro

    • Model ID: doubao-seedance-1.5-pro
    • Features: Text-to-Video and Image-to-Video generation with optional synchronized audio
    • Highlights: Unified model supporting with AI-generated audio
    • Resolution: 480p / 720p
    • Duration: 4-12 seconds

    Key Capabilities

    • Text-to-Video: Generate videos directly from text descriptions
    • Image-to-Video: Animate static images with motion
    • Audio Generation: Optional synchronized audio including voice, sound effects, and background music
    • Multiple Aspect Ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9
    • Reference Images: Support for 1-2 reference images for guided generation

    Documentation

    • [Seedance 1.5 Pro API Documentation](https://aiboxapi.com/en/api-reference/videos/doubao-seedance-1-5-pro/generation)