Four, Six or Eight Seconds
Selectable generation length per clip.
Google DeepMind's Veo 3.1 generates short clips with synchronised 48kHz audio and extends them into longer sequences by carrying the closing frames forward.
Abre esta página en un navegador de escritorio para empezar a crear.
Veo 3.1 is Google DeepMind's video model, and it takes a different approach to duration than most of the field.
Rather than stretching a single generation, it produces clips of four, six, or eight seconds and then extends them — the closing frames of one clip become the opening of the next, so a sequence grows shot by shot with continuity carried across each join. That structure gives you a decision point between every segment instead of one long roll of the dice, which is a meaningfully different way to build a piece of video.
Use Veo 3.1 when the finished thing is longer than a single shot. Assemble narrative sequences, campaign films, walkthroughs, and any piece where the second half depends on how the first half turned out.
Veo 3.1 is an AI video generation model developed by Google DeepMind, offered in Virse alongside a lower-cost variant called Veo 3.1 Fast. Both accept a first frame, a last frame, or both, together with a written brief.
The model is built around three major strengths:
Generations run four, six, or eight seconds, in 16:9 and vertical 9:16. The two entries differ in cost and rendering quality rather than in what they accept, so a brief written for one runs unchanged on the other.
Selectable generation length per clip.
Closing frames of one clip become the opening of the next.
True 3840×2160 output, with 1080p and 720p below it.
Dialogue, sound effects, and ambient noise generated with the picture.
Landscape and vertical from the same written brief.
Veo 3.1 and Veo 3.1 Fast, taking identical inputs.
Scene extension turns duration into a series of decisions rather than one commitment. If segment three goes wrong, you regenerate segment three — not the whole piece.
Because the extension starts from the previous clip's closing frames, the subject, lighting, and setting persist across the cut rather than being re-derived from a description.
At 48kHz stereo, the sound is at a specification that survives a real edit rather than needing replacement before anything goes to air.
Four, six, or eight seconds is a genuine choice, and it matters because pacing is set by duration more than by anything in the brief.
3840×2160 rendered directly, for work that will be projected, cropped into, or delivered at full size.
Veo 3.1 Fast accepts identical inputs, so the exploratory pass and the delivery pass are the same prompt on different entries.
A chained sequence is only manageable if you can see the chain. Segment four is built on segment three, and if segment three gets replaced, everything downstream needs rechecking. Virse keeps the whole run laid out in order on one canvas. Each extension sits after the clip it grew from, so the structure of the sequence stays visible rather than living in a filename convention.
Every segment sits beside the one it extends, which makes a broken join obvious at a glance.
Replace one link in the chain and rebuild forward from there rather than starting the whole piece again.
Move between Veo 3.1 and 30+ other image and video models without leaving the canvas or rewriting the brief.
Produce the establishing frame with an image model in the same workspace and start the chain from it.
Multi-segment pieces assembled shot by shot with continuity across the joins.
Branded video built to a length that a single generation cannot reach.
Extended demonstrations that carry the product consistently from segment to segment.
Location footage where ambience carries as much as the picture.
9:16 output for feeds, from the same brief as the landscape version.
Storyboards turned into moving reference, extended into full sequences.
Fast for exploration, the standard entry for delivery. Pick four, six, or eight seconds by the pacing you want.
Upload the composition the sequence starts on, and the closing one if the shot has to land somewhere specific.
State the movement, the camera's behaviour, and the mix — and be specific about how the segment ends.
Use the closing frames to generate the following shot, and review each join before continuing.
A useful Veo 3.1 prompt usually includes five elements:
En lugar de escribir
A dramatic shot of waves hitting rocks, cinematic, with epic music.
Escribe
A wide shot of grey waves breaking against black volcanic rock under overcast light. Each swell rises and collapses at an unhurried pace, spray carrying to the right on the wind. Camera locked off on a tripod, slightly above the waterline. Audio is water and wind only, no music. End on a wave receding and the rock left glistening.
Open on the composition in Image 1: a rural petrol station at dusk, single canopy light on, no cars and no people. Insects move through the beam of the canopy light. Everything else is still. Camera tracks slowly right at a constant speed, keeping the canopy in frame throughout. Cicadas, the electrical buzz of the light, one distant vehicle passing without appearing. No music. End with the canopy at the left edge and open road filling the right.
Open on a pair of over-ear headphones resting closed on a dark walnut surface, lit from a single source above and behind. The earcups rotate open slowly and evenly over the first two seconds while the camera pushes in a short distance and stops. Low room tone with one soft mechanical click as the cups reach their open position. No music. End with the headphones fully open, filling the lower half of the frame.
A man in a canvas jacket stands in a hardware shop aisle, holding a tin of paint and reading the label. He looks up toward someone off camera and says, "This one's the matt, right?" Camera static in a medium shot from the end of the aisle, no movement. Fluorescent hum, distant shelving noise, his voice at conversational level, no music. End on him looking back down at the tin.
| Dimensión | Veo 3.1 Fast | Veo 3.1 |
|---|---|---|
| Positioned for | Exploration and timing | Delivery |
| Inputs | Identical to standard | Identical to Fast |
| Clip lengths | Four, six, or eight seconds | Four, six, or eight seconds |
| Scene extension | Yes | Yes |
| Audio | Optional, 48kHz stereo | Optional, 48kHz stereo |
| Output range | 720p to 4K | 720p to 4K |
The closing state becomes the next segment's opening. A loose ending propagates through everything you build after it.
Check the transition between segments before generating the next one. Errors compound down a chain.
"No music, room tone and footsteps only" is directable. "Epic score" is a feeling the model has to guess at.
Four to eight seconds holds one thing happening. Stacking events is what segmenting exists to avoid.
Build the sequence in pieces, check each join, and keep the decision points that a single long generation would take away from you.