Frontier models from Runway and other labs
The best video, image, audio, and real-time models, available through one API.
runwayGen-4.5Runway's state-of-the-art text to video and image to video model.
runwayAleph 2.0Runway's flagship in-context video editing model with keyframe-guided control.
bytedanceSeedance 2.5Cinematic video up to 30 seconds at 480p, 720p, and 1080p, with a large reference budget for images, videos, and audio.
runwayRubyConvert any video to true HDR (BT.2020 with PQ or HLG), delivered as HEVC, ProRes, or EXR sequence.
googleGemini Omni Flash 1.1Google’s multimodal video model with native audio, real-world grounding, and cinematic control.
openaiGPT Image 2.5 SunburstOpenAI's highest-quality image generation and editing model with up to 4K resolution.Access new models
the day they release
Frontier models land on the API the day they release. Integrated endpoints can update automatically, based on your preference.
Available models by capability

google/Gemini Omni Flash 1.1
Google’s multimodal video model with native audio, real-world grounding, and cinematic control.

bytedance/Seedance 2.5
Cinematic video up to 30 seconds at 480p, 720p, and 1080p, with a large reference budget for images, videos, and audio.

runwayml/Aleph 2.0
Runway's flagship in-context video editing model with keyframe-guided control.

bytedance/Seedance 2.0
Cinematic video generation with fine-grained control over duration, ratio, and audio.

bytedance/Seedance 2.0 Fast
Faster cinematic video at 480p and 720p with duration, aspect ratio, references, and optional audio.

bytedance/Seedance 2.0 Mini
The most cost-efficient Seedance 2.0 tier: cinematic video at 480p and 720p with duration, aspect ratio, references, and optional audio.

minimax/MiniMax H3
Multimodal video generation with text, image, video, and audio references.

minimax/MiniMax H3 Max
Text-to-video and image-to-video with first/last frames, seed, prompt expansion, and always-on native audio at 480p and 768p.

alibaba/WAN 3.0 Prime
High-speed Alibaba video generation with native audio, 2–30s duration, and image, video, and audio references at 480p, 720p, and 1080p.

alibaba/WAN 3.0
Alibaba video generation with native audio, 2–30s duration, and image, video, and audio references at 480p, 720p, and 1080p.

xai/Grok Imagine Video 1.5
Fast video generation with native audio, prompt-addressable image references, and audio-driven performance.

alibaba/HappyHorse 1.0
Alibaba video generation with native audio, flexible duration (3–15s), and ten output ratios for text- and image-to-video.

google/Gemini Omni Flash
Google video generation and editing from text, an image, or a source video.

google/Veo 3.1
High quality video generation with audio and speech.

runwayml/Act Two
Runway's next-generation motion capture model.

runwayml/Gen-4 Turbo
Runway's fastest Image to Video generation model.

runwayml/Gen-4.5
Runway's state-of-the-art text to video and image to video model.

openai/GPT Image 2.5 Flare
OpenAI's faster image generation model with up to 4K resolution.

openai/GPT Image 2.5 Sunburst
OpenAI's highest-quality image generation and editing model with up to 4K resolution.

bytedance/Seedream 5.0 Pro
Reasoning image editing with layer control and multi-reference fusion.

bytedance/Seedream 5.0 Lite
Image generation with multi-image fusion and reference editing.

openai/GPT Image 2
OpenAI's latest image generation model with up to 4K resolution.

xai/Grok Imagine Image 2
Image generation and editing from up to three references, across a wide range of aspect ratios.

meta/Muse Image
Image generation and editing from up to 10 references, across eight aspect ratios.

google/Nano Banana 2
Google's fast image generation model with flexible resolution tiers up to 4K.

google/Nano Banana Pro
Google's most capable image generation model with 4K resolution support.

google/Nano Banana
State-of-the-art image generation and editing model.

runwayml/Gen-4 Image Turbo
Runway's fastest and most cost efficient image generation model.

runwayml/Gen-4 Image
Runway's best in class image generation model.
runwayml/Character Video
Generate a video of a character speaking from a text script or an audio file, powered by GWM-1.

magnific/Precision Upscaler V2
Image upscaling up to 16x with sharpness and grain controls.

magnific/Video Upscaler
Video upscaling to 4K with creative detail and FPS boost.

runwayml/Ruby
Convert any video to true HDR (BT.2020 with PQ or HLG), delivered as HEVC, ProRes, or EXR sequence.

bytedance/Seed Audio 1.0
Generate speech, sound effects, and rich audio scenes from text with cloned voices.

elevenlabs/Eleven v3
Generate expressive speech with audio tags like [laughs] and [whispers] in the script.

elevenlabs/ElevenLabs Text to Speech
Generate lifelike speech with nuanced intonation and emotion.

elevenlabs/Voice Isolation
Remove background noise from audio.

elevenlabs/ElevenLabs Sound Effect
Turn text into sound effects for your videos, voice-overs or video games.

elevenlabs/Voice Dubbing
Translate audio to up to 29 other languages.

elevenlabs/Speech to Speech
Change voice while preserving emotion and tone.
runwayml/gwm-avatars
Real-time, interactive avatars powered by GWM-1.
Frequently
asked questions
Which models are available?
Video, image, audio and real-time models from Runway and leading labs. The full catalog is above.
Do you support models from other labs?
Yes. Use models from Runway, Google, OpenAI, ByteDance, ElevenLabs and more through the same API.
When can I use a newly released model?
We aim to make frontier models available through the API on the day they are released.
How much does a generation cost?
Pricing varies by model and generation settings. Each model page shows its current credit cost.
How do I switch models?
Choose a different model identifier on the same endpoint, or use Model Router to select one automatically.
Can I try a model before integrating it?
Yes. Sign in to the developer portal and use the model playground before adding it to your application.
Which model should I use?
Filter the catalog by capability, or use Model Router to optimize for quality, latency, or cost.