TLDR
FLUX 3 is Black Forest Labs’ July 23, 2026 multimodal foundation model. Official positioning: one architecture jointly learning from images, video, and audio, with an action-prediction path for robotics partners (BFL blog; BFL models). Public video highlights include clips up to about 20 seconds, multi-shot sequences from a single prompt, and native audio (sound effects, ambience, dialogue) (The Decoder summary of BFL materials; early access notes from Justine Moore / a16z). On FLUX 3 (this site), open the video generator or image generator and start free—no card required to try.
Key Takeaways
- Released: July 23, 2026 for the public FLUX 3 launch (@bfl_ai; bfl.ai/blog/flux-3).
- Scope: Image + Video + Audio + Action-Prediction under one multimodal design (BFL models; wire release).
- Video headline features: up to ~20s generation, multi-shot from one prompt, native audio (The Decoder; venturetwins early notes).
- Series context: FLUX.1 (2024 image era) → FLUX.2 image/editing (Nov 2025) → FLUX 3 multimodal real-world models (BFL FLUX.2 post; BFL FLUX 3 post).
- Try path here: FLUX 3 video (default for short clips) and FLUX 3 image in the same product workspace.
What Actually Shipped
1. A unified multimodal pitch—not “video bolted onto image”
Black Forest Labs describes FLUX 3 as jointly learning from images, video, and audio inside a unified architecture, because no single modality captures the world alone (BFL blog; The Decoder). The company frames this as progress toward real-world visual intelligence—models that perceive, predict, and (via partners) act (BFL announcement thread).

Everyday-to-cinematic still language · showcase style frame on FLUX 3
2. Video: length, multi-shot, native audio
Public materials and launch-week reporting emphasize:
| Capability | Public claim | Primary sources | | --- | --- | --- | | Duration | Video clips up to about 20 seconds | The Decoder; CryptoBriefing | | Multi-shot | Multiple shots / angles from one prompt | Justine Moore (a16z) | | Native audio | Sound aligned with the clip (SFX, ambient, dialogue paths) | BFL launch framing; The Decoder | | Interfaces | Text-to-video, image-to-video, and related video workflows in the public narrative | The Decoder; product path on this site’s video generator |
Early preference tables published in launch coverage (for example, short 720p comparisons vs other video systems) are vendor-reported early signals, not independent film-industry certifications (The Decoder chart summary). Use them as context, then judge your own prompts.

Prompt-first motion language · showcase keyframe on FLUX 3
3. Image path in the same family
Official posts show FLUX 3 Image samples and describe synthesis/editing across styles, aspect ratios, and resolutions as part of the multimodal stack (BFL thread PS samples; BFL blog). On this product, stills live on the image generator with FLUX 3 as the house image default.
4. Action / robotics (context, not the main creator CTA)
BFL also announced FLUX-mimic work with Mimic Robotics, including industrial testing narratives involving Audi (BFL action post in launch thread; wire release). That is important for the “world model” story. For social ads, product demos, and short films, the practical surface is still video + image generation.
5. Series timeline (FLUX → FLUX.2 → FLUX 3)
| Generation | Public theme | Anchor | | --- | --- | --- | | FLUX.1 era | Open/image generation wave that made FLUX a creator default | Series history on BFL / Wikipedia overview | | FLUX.2 | Production image generation & editing (multi-reference, higher fidelity) | FLUX.2 blog, Nov 25, 2025 | | FLUX 3 | Multimodal video + audio + image backbone | FLUX 3 blog, July 23, 2026 |
Deep dive: who FLUX 3 is for
Social and UGC marketers
Short clips that hold faces, motion, and everyday lighting—without hiring a full production day—are the jobs multi-shot + native audio enable. Start on text to video.
Product and growth teams
Image-to-video turns a packaging still or hero product frame into a first-pass motion ad. Use a clean still, then animate on the same video workspace.
AI filmmakers and storyboarders
One-prompt multi-shot is the differentiator launch-week creators keep posting about (venturetwins). Treat early outputs as editable first passes, then refine prompts shot-by-shot.
Still designers who already live in FLUX
If you know FLUX for stills, FLUX 3’s public story is “same family, now time + sound.” Keep stills on image and motion on video.
What We Know vs. What We Don’t
| We know (sourced) | We don’t know / won’t claim | | --- | --- | | Public launch July 23, 2026 with multimodal positioning (@bfl_ai; BFL blog) | That every host, resolution, or duration toggle is identical across all products worldwide | | Video narrative includes ~20s, multi-shot, native audio (The Decoder) | Guaranteed first-take perfection for every face, hand, and full-body turn | | Early preference scores exist in launch coverage and are preliminary (The Decoder) | Independent, peer-reviewed leaderboards that crown a permanent “#1 video model” | | FLUX.2 established a strong image chapter (BFL FLUX.2) | That stills and video always share the exact same aesthetic fingerprint on every seed | | This site exposes browser video and image generators for FLUX 3 workflows | Cross-vendor dollar-per-second price tables (not compared here) |
Why this matters for builders
- Video + audio in one generation reduces tool-hopping for first-pass ads and explainers.
- Multi-shot prompting changes how you write briefs—write cuts, not only vibes (multi-shot early notes).
- FLUX brand continuity helps teams that already prompt in FLUX stills language migrate motion work without a new mental model.
- Browser try paths beat waiting for a perfect local install when you only need to know if your product shot survives motion.
How to evaluate FLUX 3 yourself on this site
- Open the FLUX 3 video generator.
- Paste Prompt A below (multi-shot).
- Watch once at full screen: faces, cut clarity, audio fit (if present), and crop survival for 9:16.
- Save keepers; optional second pass with image to video from a still you already like on image.
Prompt A — multi-shot backyard joust
Multi-shot short film, 12 seconds, playful daylight backyard.
Shot 1 wide: two kids jousting with bright pool noodles, green grass, summer sun.
Shot 2 medium: slow motion near-miss as noodles clash, fabric and hair motion clear.
Shot 3 close-up: both kids laughing, natural faces, soft bokeh fence in background.
Natural ambient audio feel, no music bed, no text overlays, no watermark.
Prompt B — product hero spin
Product commercial, 8 seconds, clean studio.
Shot 1: matte black wireless earbuds case on white seamless, soft key light.
Shot 2: lid opens smoothly, earbuds catch a specular highlight.
Shot 3: earbud floats closer to camera, macro detail on mesh.
Subtle whoosh SFX energy, no logos other than the product shape, no watermark.
Prompt C — rainy neon street (mood)
Cinematic street clip, 10 seconds, night rain, neon reflections.
Single continuous push-in toward a woman laughing under a translucent umbrella,
wet asphalt, cyan and magenta signs, handheld micro-motion, natural skin texture.
City ambience and light rain SFX, no dialogue, no on-screen text, no watermark.
What to watch next
- Creator multi-shot recipes (shot lists as prompts)—see our FLUX 3 Video: Multi-Shot, ~20s Clips & Native Audio.
- Image workflows for stills that feed image-to-video (image generator).
- Official BFL research notes on Self-Flow / multimodal backbone as they publish more (BFL research hub via The Decoder’s architecture summary).
FAQ
What is FLUX 3?
FLUX 3 is Black Forest Labs’ July 23, 2026 multimodal foundation model for image, video, audio, and action-prediction, trained jointly rather than as separate single-purpose models (BFL blog; launch thread).
Is FLUX 3 only a video model?
No. Public materials describe video, image, audio, and action paths under one multimodal design (BFL models). On this site you can start with video or image.
How long can FLUX 3 videos be?
Launch reporting cites clips up to about 20 seconds in BFL’s public narrative (The Decoder). Always confirm the duration control available in the generator UI for your run.
Does FLUX 3 support multi-shot in one prompt?
Yes—that is a launch-week creator highlight: multiple shots from a single prompt (venturetwins). Write explicit Shot 1 / Shot 2 / Shot 3 lines for best control.
What is native audio in FLUX 3?
Native audio means the model can produce sound with the video—ambience, SFX, and dialogue-oriented demos appear in public coverage (The Decoder; BFL launch).
How is FLUX 3 different from FLUX.2?
FLUX.2 (Nov 2025) is the production image generation & editing chapter (BFL FLUX.2). FLUX 3 expands the family into a multimodal stack with first-class video + audio storytelling (BFL FLUX 3).
Can I try FLUX 3 free online?
Yes. Open the FLUX 3 video generator on flux3.video to start free—no card required to begin. Stills: image generator.
Text to video or image to video—which first?
Use text to video when inventing a scene from scratch. Use image to video when you already have a product photo or character still you want to animate—both live under video.
Are the published win-rate comparisons final?
No. Early preference rates in launch articles are preliminary vendor-reported signals under specific test conditions (The Decoder). Run your own prompts for decisions that matter.
Where should I start right now?
Paste Prompt A into the video generator, watch full screen once, then iterate shot descriptions—not just adjectives.
About Avery Lin
Avery Lin is an AI model release analyst who tracks FLUX-family launches and turns public model facts into clear try paths for creators on FLUX 3.
