Updated
Black Forest Labs
FLUX 3
Black Forest Labs' FLUX 3: video clips up to 20 seconds with native audio, and keyframes at chosen moments.
What is FLUX 3?
FLUX 3 is Black Forest Labs' multimodal model, trained on video, images and audio together.
The company opened early access on July 23, 2026, and on August 4 made FLUX 3 Video generally available for text-to-video, image-to-video and keyframes, with clips up to 20 seconds. On Flashloop it runs from 5 to 20 seconds at 720p or 1080p.
Keyframes take 1 to 10 images, each pinned to a time you choose, so a scene can pass through a product close-up at 2 seconds and a pack shot at 8 on its way from the first frame to the last. Black Forest Labs also documents lip-synced speech in more than a dozen languages, with sound effects and ambience generated together with the picture.
In Black Forest Labs' internal evaluation, human raters preferred FLUX 3 for text-to-video, and for image-to-video it tied Seedance 2.0 and beat the other models tested. The three clips at the top of this page are FLUX 3 outputs made on Flashloop, 10 seconds each at 720p, with the audio from the same generation. Their prompts are below.
- Provider
- Black Forest Labs
- Length
- 20 seconds
- Resolution
- 720p, 1080p
- Aspect ratios
- auto, 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16
- Audio
- Native audio
What you can do with FLUX 3
Up to 20 seconds per generation
Black Forest Labs documents clips of 5 to 20 seconds, and Flashloop offers every whole second in that range
Native audio
dialogue, sound effects and ambience come out of the same generation as the picture, with lip-synced speech in English, Chinese, Spanish, French, German, Japanese, Portuguese, Russian and the other languages Black Forest Labs lists
Keyframes
pin up to 10 images to chosen moments of the clip, and the model generates the motion that connects them
Start and end frames
animate a single image, or set both ends and let the model fill in the motion between them
Multi-shot and typography
Black Forest Labs says one generation can cut between scenes and camera angles and render text inside the scene
720p or 1080p
720p comes straight off the generation, and Black Forest Labs makes 1080p by upscaling the same shot
Prompts to try with FLUX 3
“A young woman with curly hair sits by a sunlit window in a cozy cafe and talks straight to the camera, smiling and gesturing with one hand. She says: "I tried the new spring menu and honestly, the lavender latte is my new favorite." Soft cafe chatter and cups clinking in the background. Handheld close-up, warm natural light, one continuous shot.”
Where FLUX 3 fits best
Scenes that pass through chosen images
Pin a product shot, a close-up and a final pack shot at set seconds, and the model fills in the motion between them.
Multilingual dialogue
A character speaks to camera in one of the documented languages, with lip-sync and room sound from the same generation.
Long single clips
Up to 20 seconds per generation, enough for a scene with a setup and a payoff.
Animating a still
Start from a portrait, a render or a product photo and add motion and sound.
Transitions between two images
Set a start and an end frame and generate the move from one to the other.
How FLUX 3 compares
| Model | Length | Top resolution | Aspect ratios | Native audio |
|---|---|---|---|---|
| FLUX 3 | 20 seconds | 1080p | auto, 21:9, 2:1 +5 | Yes |
| Veo 3.1 | 8 seconds | 1080p | 16:9 | Yes |
| Kling 3.0 | 15 seconds | Pro (1080p) | 16:9, 9:16, 1:1 | No |
| Kling 3.0 Turbo | 15 seconds | 1080p | 16:9, 9:16, 1:1 | No |
| Kling 4.0Coming soon | 30 seconds | To be confirmed | To be confirmed | Yes |
Good to know before you generate
- Black Forest Labs also documents video continuation, draft mode, 2K and 4K output and a 9:21 ratio, which Flashloop does not offer
- No reference images and no @characters on Flashloop, so start from an image when a face has to stay the same
- Keyframes and start or end frames cannot be combined: each generation uses one or the other
- Black Forest Labs' comparisons with other models come from its own internal evaluation, so test it on your own prompts
Start creating with FLUX 3
Create with FLUX 3Select FLUX 3
Open Flashloop and choose FLUX 3 from the model selector in the video creator.
Enter your prompt
Describe what you want to see in your video.
Generate and download
Hit generate and your video will be ready in seconds. Download or share directly.
Questions about FLUX 3
What can FLUX 3 do on Flashloop?
Text-to-video, and image-to-video from a start frame, a start and end frame pair, or up to 10 keyframes, all with native audio. Clips run from 5 to 20 seconds at 720p or 1080p.
When did FLUX 3 come out?
Black Forest Labs opened FLUX 3 in early access on July 23, 2026 and made FLUX 3 Video generally available on August 4, 2026. Its docs also describe FLUX 3 Image, a separate endpoint for images, and list open weights for FLUX 3 Dev as coming soon.
How long can FLUX 3 videos be?
Black Forest Labs documents clips up to 20 seconds in one generation. Flashloop offers whole-second lengths from 5 to 20 seconds.
What are keyframes?
Keyframes are images pinned to moments in the clip. FLUX 3 takes from 1 to 10 of them, each at a time you choose, and generates the motion between them.
Does FLUX 3 generate audio?
Yes. Black Forest Labs documents audio generated with the video: multilingual speech with lip-sync, sound effects and ambience.
Who makes FLUX 3?
Black Forest Labs, the company behind the FLUX image models. Flashloop hosts models from several makers and does not build FLUX 3 itself.
More video models on Flashloop
Veo 3.1Google DeepMind's Veo 3.1: 8-second videos with dialogue, sound effects and ambience generated in one pass.
Kling 3.0Kuaishou's Kling 3.0: multi-shot video up to 15 seconds, with consistent characters across cuts.
- Kling 3.0 Turbo
Kuaishou's faster, cheaper Kling 3.0: text and image to video at 720p or 1080p, up to 15 seconds.
Kling 4.0Kuaishou's Kling 4.0, announced September 28, 2026: native 30-second clips, up to 10 keyframes, full release in October.