How to Replace a Character in Video with MiniMax H3 Without Identity Drift

Vincent12 min de lectura ·

How to Replace a Character in Video with MiniMax H3 Without Identity Drift

To replace a character in video with MiniMax H3 without identity drift, use the reference image as the identity source and the source video as the performance source. A simple rule is: image = identity, video = performance, prompt = constraints. This helps H3 preserve the new character’s face and appearance while keeping the original motion, timing, pose, and camera behavior.

The main challenge is identity drift. A character may look correct at first and then shift back toward the original performer, while fast motion, occlusion, long clips, or conflicting references can make consistency worse. Our review of documented H3 workflows includes one six-hour study with more than 400 generations, alongside multi-character, high-motion, framing, and local ComfyUI tests. Across these cases, clear reference roles, short test clips, strong identity anchors, and explicit preservation rules consistently emerge as the most useful workflow principles.

For creative teams managing reference-heavy AI workflows like this, Virse provides a more structured way to organize the process. Instead of forcing every decision into a chat box, Virse works on an infinite canvas where references, outputs, and creative relationships stay visible, while multiple Agents can work across the same project and share context. With long-term memory for team preferences, brand standards, and project knowledge, Virse helps professional designers iterate and scale AI-assisted production without giving up creative control.

virse workforce

How Does MiniMax H3 Character Replacement Work?

H3 Character Replacement Is More Than Face Swap

A traditional face swap mainly changes facial appearance. MiniMax H3 can support broader replacement in which the face, hairstyle, clothing, colors, silhouette, and body identity change while the original movement and cinematography remain.

The most useful model is simple:

Image = identity. Video = performance. Prompt = routing and constraints.

The image answers what the new character should look like. The video answers how that character should move, perform, and interact with the scene.

This distinction explains several common failures. In documented workflows, some outputs changed hair or clothing but kept too much of the original face. Others preserved the source video so aggressively that the requested character replacement barely happened.

Image, Video, and Audio References Need Clear Roles

Reference quality matters more than reference count.

Input

Best Role

Character image

Face, hair, outfit, proportions, identity

Source video

Motion, pose, expression, position, timing

Source video

Camera, framing, lighting, scene, occlusion

Audio reference

Voice, tone, or audio context when needed

Prompt

Defines replacement, inheritance, and preservation

The goal is not to fill every available reference slot. It is to make each reference responsible for one clear type of information.

How to Prepare the Best MiniMax H3 Character Reference

Use One Strong Portrait Before Adding More Images

For close-ups and medium shots, one clear portrait is often the best starting point.

Our review found workflows where reducing the number of reference images improved identity accuracy. Additional images can introduce small contradictions in hairstyle, facial structure, costume, lighting, or proportions.

A strong first reference should provide:

  • a clearly visible face
  • recognizable hairstyle
  • consistent clothing
  • minimal obstruction
  • enough facial detail to anchor identity

Add another view only when it solves a specific missing-information problem.

Use a Character Board for Full-Body Motion and Turns

A single portrait becomes less useful when the performer rotates, appears in profile, or turns away from the camera.

For these clips, a character board combining front view, rear view, and close-up face can provide better coverage. This is especially useful for distinctive hairstyles, jackets, footwear, or costumes that must remain coherent during turns.

From a design workflow perspective, this works like a traditional character turnaround sheet: the model receives the visual information before that view appears in motion.

Match the Reference Framing to the Target Shot

Framing is an overlooked variable.

One documented workflow improved consistency after adapting the reference to approximately 864 × 480, closer to the target composition. In another case, an 848 × 1264 full-body PNG was used against a 1280 × 720, 30 fps, 14.4-second source clip, yet noticeable identity drift remained.

These are not controlled benchmarks, but they support a practical rule: make the identity-defining features large and clear enough for the scale at which the character appears in the final shot.

Scatter plot of three documented H3 media dimensions: an adapted 864 by 480 reference, an 848 by 1264 full-body reference, and a 1280 by 720 source video.

How to Replace a Character in Video with MiniMax H3 Step by Step

Step 1: Make the Source Video the Performance Master

Let the source video control movement, pose, position, body rotation, facial performance, speed, timing, and choreography.

Do not rewrite every movement in text if the source video already contains the performance you want. Re-describing the action can create another interpretation layer rather than improve fidelity.

Step 2: Assign the Reference Image as the Identity Source

State clearly that the original performer should be replaced by the character shown in the reference.

Define the visual attributes that must remain stable, including face, hairstyle, outfit, colors, proportions, and distinctive details.

In one documented study involving more than 400 generations over roughly six hours, weak subject definitions were associated with identity loss in around half of some test runs. The lesson is not that every workflow will behave the same way, but that vague identity labels are risky when consistency matters.

Step 3: Transfer the Original Performance

The replacement character should inherit the original performer’s movement, pose, position, expression, rotation, speed, and timing.

This instruction is critical because it tells H3 that the source performer provides performance, not identity.

Step 4: Preserve Camera, Lighting, and Scene Structure

Protect everything outside the intended character edit.

Preserve the background, props, camera path, framing, focus, motion blur, lighting, shadows, composition, and edit timing.

The fewer unrelated elements H3 must redesign, the easier it is to focus model capacity on the replacement itself.

Step 5: Preserve Occlusion

Occlusion matters whenever hands, hair, props, weapons, clothing, or other characters pass in front of the subject.

If a hand crosses the face in the source, it should cross the replacement face at the same moment. The same applies to bows, equipment, foreground objects, and character overlap.

This is especially important in dance, running, boxing, archery, fast turns, and action scenes.

What Is the Best MiniMax H3 Character Swap Prompt?

Use Replace, Match, Inherit, Preserve, and Prevent

The strongest general prompt framework is:

Replace: Replace the original performer with the reference character.

Match: Match the character’s face, hair, outfit, colors, proportions, and recognizable identity details.

Inherit: Preserve the original performer’s movement, pose, expression, position, speed, rotation, and timing.

Preserve: Keep the original camera, background, lighting, shadows, props, framing, focus, and edit.

Prevent: Do not introduce new actions, objects, cuts, camera movement, text, or accessories.

A concise production prompt can therefore read:

In Video 1, replace the original performer with the character shown in Image 1. Match the reference character’s face, hairstyle, outfit, colors, body proportions, and visible identity details throughout the clip. Preserve the original performer’s movement, pose, position, rotation, expression, speed, and timing. Keep the background, props, camera path, framing, focus, motion blur, lighting, shadows, composition, and editing unchanged. Maintain the original occlusion relationships when hands, hair, clothing, or objects pass in front of the character. Do not add new actions, objects, camera motion, cuts, text, or accessories.

Simple Prompts Can Beat Over-Structured Prompts

Longer prompts are not automatically more controllable.

In one documented head-replacement workflow, a simplified instruction produced strong results in roughly 70% of attempts, while a more elaborate version sometimes preserved so much of the source that little replacement occurred.

Structured prompting becomes more useful when it solves a specific problem such as multi-character mapping, persistent identity drift, or conflicting references.

How to Fix MiniMax H3 Identity Drift

Start with Short Clips Before Increasing Duration

Identity drift often appears gradually. A character may begin correctly, move toward the source performer, and then partially recover.

A 14.4-second, 1280 × 720, 30 fps test using an 848 × 1264 full-body reference still showed substantial drift on a 16GB RTX 5070 Ti workflow.

A more efficient debugging order is:

  1. Test approximately five seconds.
  2. Strengthen the face reference.
  3. Match the framing more closely.
  4. Remove unnecessary references.
  5. Clarify identity and performance roles.
  6. Add a rear view for turning shots.
  7. Test another seed.
  8. Increase quality settings for difficult motion.

High Motion May Need More Conservative Settings

In the 400-plus-generation workflow, five-second clips were commonly used for testing, with additional 15-second experiments.

Four-step Turbo settings were useful for iteration, but difficult motion sometimes lost identity consistency. For harder sequences, that workflow returned to approximately 20 steps.

Dumbbell chart comparing 4-step Turbo iteration with about 20 steps used for more difficult high-motion sequences in one documented H3 workflow.

This is not a universal preset. It illustrates a broader tradeoff: faster generation can become less reliable when the model must solve identity, pose, motion, occlusion, and temporal continuity at once.

Radar-style profile summarizing one documented H3 study with more than 400 generations over roughly six hours, common 5-second tests, 15-second extended tests, 4-step Turbo iteration, and about 20 steps for high-motion sequences.

Keep H3 Character Identity Consistent Across Cuts and Longer Videos

Longer sequences and hard cuts introduce additional opportunities for the source identity to reappear or for references to compete.

For production work, validate the replacement in short, independently reviewed segments before extending the sequence. Keep the same identity description and reference mapping across clips, and check continuity at every cut.

Do not assume that a character established correctly in the first shot will automatically remain dominant in later shots. Consistency across clips should be treated as an explicit production task, not a one-time prompt decision.

How to Replace Multiple Characters with MiniMax H3

Map Every Character Explicitly

For two-person replacement, each identity needs a stable mapping.

A practical structure is:

  • Image 1 defines Character A
  • Image 2 defines Character B
  • Video 1 provides both performances
  • Character A replaces the left performer
  • Character B replaces the right performer

One documented two-character workflow reached approximately 90% reported success after improving subject mapping and retention, although results remained seed-dependent.

Horizontal bar chart showing about 70% strong results for a simplified H3 head-replacement workflow and about 90% reported success for a separate two-character workflow after improved mapping and retention.

Prevent Identity Mixing During Crossovers

The hardest moments occur when performers cross positions, overlap, or block one another.

Clear Character A and Character B assignments are safer than vague labels such as first person and second person.

For multi-character work, reference role clarity is more valuable than maximum reference count.

What Do MiniMax H3 ComfyUI Performance Tests Show?

12GB VRAM Can Work with Strong Tradeoffs

One RTX 3060 12GB workflow reported approximately:

  • 5 seconds at around 2MP
  • 8 seconds at around 1.34MP
  • about 25 minutes for the 1.34MP generation

Another 12GB test used nine HD PNG references at 0.4MP. A normal 30-second image-to-video generation took around 14 minutes, while the multi-reference Ref2V workflow took around 20 minutes.

These results are workflow examples rather than universal benchmarks, but they show the tradeoff between duration, resolution, reference count, and generation cost.

Line chart for an RTX 3060 12GB workflow showing about 8 seconds at 1.34MP and about 5 seconds at 2MP, with about 25 minutes generation time reported for the 1.34MP run.

More VRAM Does Not Automatically Fix Identity Drift

A 16GB workflow still experienced drift on a longer replacement clip.

A 24GB RTX 3090 Ti test using 56 frames at 0.5MP described video editing as roughly twice as demanding as simpler image-reference generation.

A higher-end 5090 optimization workflow reported more than 10GB of VRAM savings and roughly 14GB less shared GPU-memory use after switching to more memory-efficient components, while generation speed remained approximately similar.

The key point is that GPU capacity affects what settings you can run, but reference quality, duration, motion complexity, and prompt structure still determine identity stability.

Comparison chart of documented local H3 workflows, including RTX 3060 12GB duration and resolution data, a separate 12GB nine-reference workflow, a 16GB RTX 5070 Ti identity-drift case, and a 24GB RTX 3090 Ti 56-frame editing case.

What Are the Most Common MiniMax H3 Character Swap Mistakes?

Adding More References Before Diagnosing the Failure

More references can create more contradictions.

Start with one strong identity source. Add another view only when the target clip exposes information the first image cannot provide.

Testing Long Videos Too Early

A failed 15-second action clip can have many possible causes.

A five-second test makes it easier to isolate identity, framing, motion, seed, resolution, or occlusion problems before spending more compute.

Changing Too Many Elements at Once

If the task is character replacement, avoid simultaneously changing the background, lighting, camera, animation, and edit style.

Professional AI video workflows become more controllable when each generation has a narrow transformation boundary.

FAQ

Why Does H3 Keep the Original Face?

The most common causes are weak identity references, ambiguous reference roles, mismatched framing, excessive references, or preservation instructions that overpower the edit. Start with one clear portrait, assign the image to identity and the video to performance, and validate the replacement on a short clip before increasing complexity.

Can H3 Replace Two Characters at the Same Time?

Yes. Multi-character replacement can work when each reference is mapped explicitly to one performer. The main challenge is identity mixing when characters overlap, cross positions, or occlude one another. Stable Character A and Character B mappings are more reliable than vague positional descriptions.

How Many Reference Images Should I Use with H3?

Start with one strong image. Add another view only when the source video reveals information the first image cannot provide, such as a rear hairstyle or costume view. In several documented workflows, fewer references improved consistency because the model had fewer conflicting visual signals to reconcile.

How Much VRAM Does H3 Need?

There is no single fixed requirement. Documented local workflows include workable cases on 12GB, 16GB, and 24GB GPUs, but resolution, duration, steps, quantization, model configuration, and reference count all affect memory use. For constrained hardware, reduce clip duration and resolution first.

Conclusion

The best MiniMax H3 character replacement workflow is not built around one “ultimate prompt.” It comes from making every information source responsible for the right thing: the reference image defines identity, the source video defines performance and cinematography, and the prompt defines what must change, transfer, and remain untouched. Across hundreds of documented generations, multi-character cases, high-motion tests, framing experiments, and local GPU workflows, the strongest pattern is consistent: start with a short clip, use the minimum number of strong references, separate identity from performance, preserve camera and occlusion explicitly, and add complexity only when a specific failure requires it.

Más del blog de Virse