AI モデルカタログ

モデルライブラリ

Virse で使えるすべてのモデルを、つくれるものごとにまとめました。モデルを選んで詳細ページを開くか、そのままキャンバスで制作を始めましょう。

動画モデル 7 件

Google DeepMind's Veo 3.1 generates short clips with synchronised 48kHz audio and extends them into longer sequences by carrying the closing frames forward.

Google

ByteDance's Seedance 2.0 produces picture and sound in one pass, from a first and last frame or from reference images, at sizes from 480P to native 4K.

ByteDance

MiniMax's Hailuo 03 model produces silent 2K video, driven either by a first and last frame or by a set of reference images.

MiniMax

Kuaishou's Kling 3.0 speaks — lip-synced dialogue in five languages, a different voice per character, and a storyboard mode that lays out a whole sequence.

Kuaishou

The Happy Horse 1.0 AI video model needs one still image and a sentence, producing motion from the frame you already have.

Alibaba

Google's natively multimodal video model, taking text, images, video, and audio as input, offered in Virse as frame-driven and reference-driven entries.

Google

Generate cinematic 30-second videos with Seedance 2.5, with synchronized audio, multimodal references, and precise creative control.

ByteDance

画像モデル 13 件

A 6-billion-parameter model compressed to an eight-step inference pipeline, built to return a usable picture before your train of thought moves on.

Alibaba

ByteDance's newest Seedream generation, available as Seedream 5.0 Pro and Seedream V5 Lite, both taking the same written brief.

ByteDance

ByteDance's Seedream 4.0 and Seedream 4.5 generate from written descriptions and from images you supply, turning one source picture into as many directions as a project needs.

ByteDance

Reve AI's second-generation model builds an addressable layout of positioned elements before it renders, so a specified arrangement comes back arranged.

Reve

Alibaba's Qwen Image 3.0 Pro renders type as language across a dozen scripts, holding legibility down to ten pixels in dense layouts.

Alibaba

Generate studio-grade 4K images with Nano Banana Pro, blending multiple references while holding characters, products, and typography consistent.

Google

Google's Nano Banana 2 pairs flagship-level image quality with Flash-class speed, across four output sizes from rapid 0.5K prototypes to finished 4K.

Google

A text-to-image foundation model trained from scratch around layout, with bounding-box placement, hex colour conditioning, and native 2K output.

Ideogram

Google's Gemini 2.5 Flash Image — the original Nano Banana — generates and edits pictures at conversation speed, so a correction takes a sentence instead of a session.

Google

OpenAI's GPT Image 2.0 reads a hundred-word brief and attempts all of it, with independent quality and resolution controls from 1K to 4K.

OpenAI

Black Forest Labs' editing models, FLUX Kontext Pro and Kontext Max, built to apply one described change while the rest of the frame survives untouched.

Black Forest Labs

Black Forest Labs' production model reads briefs up to 32,000 tokens, so brand rules, colour values, and compositional constraints can all go in at once.

Black Forest Labs

The generation of Black Forest Labs' FLUX line that came before FLUX 2, kept available as two entries — the standard model and Ultra.

Black Forest Labs