Frontier models from Runway and other labs

The best video, image, audio, and real-time models, available through one API.

Access new models
the day they release

Frontier models land on the API the day they release. Integrated endpoints can update automatically, based on your preference.

Available models by capability

41 models

google/Gemini Omni Flash 1.1

Google’s multimodal video model with native audio, real-world grounding, and cinematic control.

bytedance/Seedance 2.5

Cinematic video up to 30 seconds at 480p, 720p, and 1080p, with a large reference budget for images, videos, and audio.

runwayml/Aleph 2.0

Runway's flagship in-context video editing model with keyframe-guided control.

bytedance/Seedance 2.0

Cinematic video generation with fine-grained control over duration, ratio, and audio.

bytedance/Seedance 2.0 Fast

Faster cinematic video at 480p and 720p with duration, aspect ratio, references, and optional audio.

bytedance/Seedance 2.0 Mini

The most cost-efficient Seedance 2.0 tier: cinematic video at 480p and 720p with duration, aspect ratio, references, and optional audio.

minimax/MiniMax H3

Multimodal video generation with text, image, video, and audio references.

minimax/MiniMax H3 Max

Text-to-video and image-to-video with first/last frames, seed, prompt expansion, and always-on native audio at 480p and 768p.

alibaba/WAN 3.0 Prime

High-speed Alibaba video generation with native audio, 2–30s duration, and image, video, and audio references at 480p, 720p, and 1080p.

alibaba/WAN 3.0

Alibaba video generation with native audio, 2–30s duration, and image, video, and audio references at 480p, 720p, and 1080p.

xai/Grok Imagine Video 1.5

Fast video generation with native audio, prompt-addressable image references, and audio-driven performance.

alibaba/HappyHorse 1.0

Alibaba video generation with native audio, flexible duration (3–15s), and ten output ratios for text- and image-to-video.

google/Gemini Omni Flash

Google video generation and editing from text, an image, or a source video.

google/Veo 3.1

High quality video generation with audio and speech.

runwayml/Act Two

Runway's next-generation motion capture model.

runwayml/Gen-4 Turbo

Runway's fastest Image to Video generation model.

runwayml/Gen-4.5

Runway's state-of-the-art text to video and image to video model.

openai/GPT Image 2.5 Flare

OpenAI's faster image generation model with up to 4K resolution.

openai/GPT Image 2.5 Sunburst

OpenAI's highest-quality image generation and editing model with up to 4K resolution.

bytedance/Seedream 5.0 Pro

Reasoning image editing with layer control and multi-reference fusion.

bytedance/Seedream 5.0 Lite

Image generation with multi-image fusion and reference editing.

openai/GPT Image 2

OpenAI's latest image generation model with up to 4K resolution.

xai/Grok Imagine Image 2

Image generation and editing from up to three references, across a wide range of aspect ratios.

meta/Muse Image

Image generation and editing from up to 10 references, across eight aspect ratios.

google/Nano Banana 2

Google's fast image generation model with flexible resolution tiers up to 4K.

google/Nano Banana Pro

Google's most capable image generation model with 4K resolution support.

google/Nano Banana

State-of-the-art image generation and editing model.

runwayml/Gen-4 Image Turbo

Runway's fastest and most cost efficient image generation model.

runwayml/Gen-4 Image

Runway's best in class image generation model.

runwayml/Character Video

Generate a video of a character speaking from a text script or an audio file, powered by GWM-1.

magnific/Precision Upscaler V2

Image upscaling up to 16x with sharpness and grain controls.

magnific/Video Upscaler

Video upscaling to 4K with creative detail and FPS boost.

runwayml/Ruby

Convert any video to true HDR (BT.2020 with PQ or HLG), delivered as HEVC, ProRes, or EXR sequence.

bytedance/Seed Audio 1.0

Generate speech, sound effects, and rich audio scenes from text with cloned voices.

elevenlabs/Eleven v3

Generate expressive speech with audio tags like [laughs] and [whispers] in the script.

elevenlabs/ElevenLabs Text to Speech

Generate lifelike speech with nuanced intonation and emotion.

elevenlabs/Voice Isolation

Remove background noise from audio.

elevenlabs/ElevenLabs Sound Effect

Turn text into sound effects for your videos, voice-overs or video games.

elevenlabs/Voice Dubbing

Translate audio to up to 29 other languages.

elevenlabs/Speech to Speech

Change voice while preserving emotion and tone.

runwayml/gwm-avatars

Real-time, interactive avatars powered by GWM-1.

Frequently
asked questions

Which models are available?

Video, image, audio and real-time models from Runway and leading labs. The full catalog is above.

Do you support models from other labs?

Yes. Use models from Runway, Google, OpenAI, ByteDance, ElevenLabs and more through the same API.

When can I use a newly released model?

We aim to make frontier models available through the API on the day they are released.

How much does a generation cost?

Pricing varies by model and generation settings. Each model page shows its current credit cost.

How do I switch models?

Choose a different model identifier on the same endpoint, or use Model Router to select one automatically.

Can I try a model before integrating it?

Yes. Sign in to the developer portal and use the model playground before adding it to your application.

Which model should I use?

Filter the catalog by capability, or use Model Router to optimize for quality, latency, or cost.