AI Models — Video, Image, Audio & Avatar Generation
Generate with 18 best-in-class AI models in one platform — Seedance, Veo 3.1, Kling, Nano Banana, Seedream, ElevenLabs and more. Compare what each model does and try any of them free.
Start Creating FreeVideo 9 models
Seedance 2.5 generates up to 30 seconds of video in a single job, with up to 50 reference assets, controllable video editing, high-fidelity temporal extension, any aspect ratio from 0.4 to 2.5, and native prompting in 10+ languages. Live on Popcraft.
Generate 2 to 30-second AI videos in one pass with Wan 3.0, Alibaba's video model: 1080p at 30 fps, dialogue and music included, up to 4 images, 5 clips and 5 voices as references. From 11 credits a second on Popcraft.
Gemini Omni 1.1 Flash is Google's video model that creates and edits video from any input. Up to 40 seconds by scene extension, 1080p and 4K output, first-and-last-frame interpolation, up to 10 reference images, and speech, music and sound effects generated with the picture. Live on Popcraft.
MiniMax H3 generates 2K video at 24fps with a synchronised stereo soundtrack — effects, ambience and dialogue in one job. Multi-shot from a single prompt, 5 to 15 seconds, first and last frame control. Live on Popcraft.
Generate cinematic 4K video from text, images, or clips with Seedance 2.0 on Popcraft. Native synced audio, multi-shot sequences, and razor-sharp detail at up to 4K (2160p).

Seedance 2.0 Mini is ByteDance's leanest video tier — up to 50% cheaper than standard Seedance 2.0 and ~2× faster than Fast, with text, image, video, and audio prompting. Live on Popcraft.

Turn images and prompts into cinematic video with Seedance 2.0 on Popcraft. Reference-to-video, first/last frame, and multi-aspect outputs at up to 4K.

Generate cinematic 1080p video with built-in audio using Google's Veo 3.1 on Popcraft. Reference-to-video, first/last frame, and synchronized sound from a single prompt.

Generate multi-shot AI video with Kling 3.0 Omni on Popcraft. Fuse up to 7 reference images into connected scenes with synchronized native audio at 1080p.
Image 4 models
OpenAI GPT Image 2.5, Fast and standard builds: sketch to render, reference fidelity, in-context edits, 16 references, 4K. From 10 credits an image on Popcraft.

Generate images with Google Nano Banana 2 on Popcraft. Character consistency, 14 aspect ratios, 4K resolution, web-grounded knowledge — free to try.

Generate images with ByteDance Seedream 5 Lite on Popcraft. Live web search, visual reasoning, 2K output in seconds — free to try.

Generate premium images with Google Nano Banana Pro on Popcraft. Complex multi-element compositions, 4K output, pro-grade typography and style transfer — free to try.
Audio 3 models
Eleven v4 and Eleven v4 Turbo on Popcraft: ElevenLabs' most emotive text-to-speech yet. Inline audio tags, 90+ languages, 10,000-character takes, and a Turbo cut built for real-time voice agents.
Score videos in seconds with Music V2.5 on Popcraft. Generate AI background music, instrumental BGM, and mood-matched tracks from 3 to 200 seconds long.
Generate custom sound effects from text with ElevenLabs SFX on Popcraft. 0.5–22s clips, tunable prompt influence, and 48 kHz studio-ready audio.
Character 2 models
Turn a single portrait and audio into a lifelike talking video with OmniHuman 1.5 on Popcraft. 1080p lip-sync, 9:16/16:9/1:1, up to 30 seconds.
Generate long-form talking avatar videos with Kling Avatar on Popcraft. 9:16, 16:9, and 1:1 outputs up to 60 seconds from a single image and audio.
How to choose an AI model
Different models excel at different jobs. For the longest unbroken take, Seedance 2.5 holds thirty seconds in a single generation; for maximum resolution, Seedance 2.0 renders native 4K; for dialogue-driven scenes with built-in audio, Veo 3.1 or MiniMax H3; for multi-shot sequences from several references, Kling 3.0 Omni. Image work splits between Nano Banana 2's speed and Seedream 5's fidelity, while avatars come from OmniHuman 1.5 and Kling Avatar. Use the table below to match a model to your use case.
Model comparison
| Model | Type | Best for | Max output | Provider |
|---|---|---|---|---|
| Seedance 2.0 4K | Video | Cinematic, high-res | 4K | ByteDance |
| Veo 3.1 | Video | Dialogue + audio | 1080p | |
| Kling 3.0 Omni | Video | Multi-shot | 1080p | Kuaishou |
| Nano Banana 2 | Image | Fast iterations | 4K | |
| Seedream 5 Lite | Image | High-fidelity stills | — | ByteDance |
| ElevenLabs TTS | Audio | Voiceover | — | ElevenLabs |
| OmniHuman 1.5 | Avatar | Talking avatars | — | ByteDance |
Frequently asked questions
It depends on the job — Seedance 2.5 for a full thirty seconds in one take, Seedance 2.0 for native 4K, Veo 3.1 or MiniMax H3 for synced audio, Kling 3.0 Omni for multi-shot. The comparison above maps each to its strength.
Your starter credits work on a selection of models straight away, no card required. The full catalogue — including Seedance 2.5, Veo 3.1, Kling and GPT Image 2 — unlocks on any paid plan, from $4.99 a month.
Seedance 2.5 generates up to thirty seconds in a single job and accepts up to fifty reference assets, with editing and extension afterwards. Seedance 2.0 tops out at fifteen seconds but renders in native 4K. Open either model page to try them.