MODEL DIRECTORY · UPDATED AUGUST 2026
AI image and video model library
Compare current text-to-image, image editing, text-to-video, image-to-video, reference-to-video and native-audio models through the production lens of MindVideo AI 2.0 Beta.
Create with Polox AI ↗Seedance 2.5
Cinematic text-to-video, image-to-video and video editing with native synchronized audio.
Capabilities, inputs and workflow →MiniMax · AI video generatorMiniMax H3
Open-weights audiovisual generation with text, first/last frames and multimodal references.
Capabilities, inputs and workflow →Alibaba · AI video generatorWan 3.0
Longer multimodal video generation with flexible duration, references, audio and continuity controls.
Capabilities, inputs and workflow →Black Forest Labs · Multimodal generatorFLUX 3
Early-access multimodal model direction spanning video, image, synchronized audio and action prediction.
Capabilities, inputs and workflow →Krea · AI image generatorKrea 2
Aesthetic image foundation model designed for expressive generation and creative control.
Capabilities, inputs and workflow →Alibaba Qwen · AI image generatorQwen Image 3
Current Qwen image family for prompt-led generation, editing and visually precise iteration.
Capabilities, inputs and workflow →ByteDance · AI image generatorSeedream 5
High-quality text-to-image and multi-reference editing with flexible output sizing.
Capabilities, inputs and workflow →ByteDance · AI image generatorSeedream 4.5
Production image generation and editing model for detailed prompts and controlled visual revisions.
Capabilities, inputs and workflow →ByteDance · AI image generatorSeedream 2.5
Earlier Seedream image model useful for comparing prompt interpretation, style and editing progress across releases.
Capabilities, inputs and workflow →OpenAI · AI image generatorGPT Image 2
Natural-language image generation and multi-reference editing for precise creative revision.
Capabilities, inputs and workflow →Google · AI image generatorNano Banana 2
Fast Gemini image generation and editing with high-resolution output and localization features.
Capabilities, inputs and workflow →Google · AI image generatorNano Banana Pro
Higher-detail Gemini image route focused on 4K-capable generation and editing.
Capabilities, inputs and workflow →Google DeepMind · AI video generatorVeo 3.1
Cinematic video generation family emphasizing prompt adherence, image animation and native audio.
Capabilities, inputs and workflow →Kuaishou · AI video generatorKling 3.0
Professional video family covering text, image, motion control and multimodal shot direction.
Capabilities, inputs and workflow →Kuaishou · Multimodal video generatorKling Omni O3
Omni-style Kling workflow for more flexible multimodal generation and editing.
Capabilities, inputs and workflow →OpenAI · AI video generatorSora 2
OpenAI video generation route for text-led scenes, image animation and coherent visual storytelling.
Capabilities, inputs and workflow →MiniMax · AI video generatorHailuo 2.3
Image and text driven video model for expressive character motion and cinematic short clips.
Capabilities, inputs and workflow →xAI · AI video generatorGrok Imagine Video 1.5
Image and prompt-led video generation for rapid visual concepts and motion experiments.
Capabilities, inputs and workflow →ShengShu Technology · AI video generatorVidu Q3
Vidu generation family for reference-aware motion, character consistency and short narrative clips.
Capabilities, inputs and workflow →Black Forest Labs · AI image generatorFLUX 2 Klein
Compact FLUX image route for efficient generation, editing and controlled creative iteration.
Capabilities, inputs and workflow →