MiniMax H3 Pose and Depth Control: What It Does, How It Works, and When to Use It

Vincent10 min de lecture ·

MiniMax H3 Pose and Depth Control: What It Does, How It Works, and When to Use It

MiniMax H3 Pose and Depth Control gives creators more precise control over motion and spatial structure through Fun ControlNet. Pose constrains body posture and movement, while Depth constrains scene geometry, subject position, and camera-relative structure. This separates identity, motion, geometry, and creative direction instead of forcing H3 to infer everything from one reference.

The problem is that understanding a reference is not the same as reproducing it precisely. In character swaps, dance transfer, and motion retargeting, small errors in timing, body position, scale, or camera distance can cause drift and repeated generations.

Pose and Depth make H3 more directable by giving each input a clearer job: Reference controls appearance and identity, Pose controls movement, Depth controls spatial geometry, and the prompt controls creative direction. For teams that want to manage these references visually, Virse brings AI Agents, connected assets, and project context into an infinite canvas, making it easier to organize and iterate on controlled H3 workflows without treating every generation as an isolated prompt. MiniMax H3, Seedance 2.5 and Seedance 2.0, are now available on Virse, so teams can explore and compare multiple leading video models in the same creative workflow.

virse workforce

What Is MiniMax H3 Pose and Depth Control?

MiniMax H3 Pose Control constrains body posture and motion, while Depth Control constrains spatial geometry and camera-relative structure. These capabilities come from MiniMax-H3-Fun-Controlnet-Union rather than the original H3 launch feature set.

The Fun ControlNet checkpoint is approximately 6.8 GB and adds a control branch to 5 of H3’s 50 transformer blocks. A single union checkpoint can support several structural conditions.

Control

Best For

Main Constraint

Pose

Dance, action, character swaps

Body motion

Depth

Scene composition, camera distance

3D geometry

Canny

Sketch-driven animation

Edges and layout

HED

Softer structural guidance

Object boundaries

MLSD

Architecture, products

Straight lines

The important design shift is that creators can decide what must remain structurally fixed and what H3 can reinterpret creatively.

MiniMax H3 Reference Video vs Pose Control: What Is the Difference?

The clearest distinction is:

Reference Video provides semantic guidance. Pose and Depth provide structural constraints.

H3’s reference system can work with multiple images, videos, and audio inputs, making it strong at understanding character appearance, actions, scenes, expressions, and camera behavior.

When H3 Reference Video Is Enough

Reference Video is useful when you want to transfer the overall performance or visual intent while leaving some freedom for H3 to reinterpret the result.

Our review of documented H3 workflows shows a recurring limitation: understanding an action is not the same as reproducing its timing and body position precisely.

That is often acceptable for creative exploration. It becomes less acceptable when choreography or subject placement is a hard requirement.

When H3 Pose Control Is Better

Pose is more useful when the movement itself must remain consistent, such as:

  • dance transfer;
  • actor replacement;
  • walking or running cycles;
  • animation blocking;
  • motion retargeting.

A useful mental model is:

Reference controls who the subject is. Pose controls what the subject does. Depth controls where the subject exists in space.

MiniMax H3 Pose vs Depth: Which Should You Use?

Pose and Depth solve different problems, and subject size can affect how useful each control becomes.

Why Pose Worked Better for a Small Character in One Test

One documented workflow generated at 1280 × 704, producing a latent spatial grid of roughly 40 × 22. The character occupied only around 3 of the 22 vertical latent positions.

In that case, Depth still allowed the figure to drift toward the camera and grow in scale, while Pose maintained subject position more reliably.

This should not be generalized into “Pose is always better than Depth.” It shows that a small subject may provide relatively weak depth information after downsampling.

When Depth Is the Better Choice

Depth is more appropriate for:

  • foreground-background relationships;
  • movement toward or away from camera;
  • rooms and environments;
  • large products;
  • non-human objects;
  • scene geometry.

The correct question is therefore not “Pose or Depth?” but “Do I need to lock articulated movement or spatial geometry?”

How Pose Control Changes H3 Character Swap and Motion Retargeting

Character replacement is one of the strongest use cases because Fun ControlNet helps separate appearance transfer from motion transfer.

Our review of recurring user questions identified three common problems:

  1. The replacement character does not follow the source performance closely enough.
  2. Identity becomes less stable during complex motion.
  3. Visual traits from the original performer leak into the new character.

Pose does not solve identity by itself. Instead, it allows each input to perform a clearer role:

Target character → Reference
Source performance → Pose
Scene geometry → Depth when needed
Creative styling → Prompt

This is closer to a professional design system: if movement is wrong, adjust movement. If appearance is wrong, adjust the reference. You no longer need to change every variable at once.

How to Tune H3 Pose and Depth Without Over-Control

More conditioning does not automatically mean better control.

Treat Pose and Depth as a Shared Control Budget

In one multi-control workflow, Pose strength 1.0 plus Depth strength 1.0 produced visible saturation. Increasing a single control to around 1.6 also failed to correct the structural problem while reducing visual quality.

A safer workflow is to choose the most important control first, then introduce the secondary control gradually.

Control Timing Matters as Much as Strength

Structural control can also become counterproductive if it remains too strong late in diffusion.

One documented test found that ending control at roughly 60% of the sampling schedule preserved important structure while giving later steps more freedom to recover texture and surface detail.

The practical principle is simple:

Use early steps to establish structure. Give later steps enough freedom to finish the image.

There is not yet enough standardized testing to recommend one universal strength or end percentage.

Two-panel MiniMax H3 ControlNet tuning graphic showing Pose 1.0 plus Depth 1.0 causing saturation, a single control around 1.6 degrading quality, and one test ending control at roughly 60% of sampling.


MiniMax H3 VRAM Requirements and Reference Video Memory

H3 can run on consumer GPUs, but minimum VRAM and comfortable production VRAM are not the same thing.

GPU

Workload

Observed Time

RTX 3060 12GB

480p, 5 sec, 20 steps

under 9 min

RTX 4070 Ti 12GB

0.4MP, 15 sec

about 18 min

RTX 4070 Ti 12GB

0.2MP, 15 sec

under 7 min

RTX 3060 12GB

about 2MP, 5 sec I2V

about 25 min

Our research also found working 8GB and 6GB workflows, but they rely more heavily on quantization, offloading, staged loading, and system RAM.

Horizontal bar chart comparing four documented MiniMax H3 workloads on 12GB GPUs, ranging from under 7 minutes to about 25 minutes depending on resolution, duration, and workflow.

Reference Length Can Be a Hidden VRAM Bottleneck

A particularly useful 12GB case showed:

  • 20-second reference → roughly 5-second output at 0.4–0.5MP
  • 5-second reference → roughly 15-second output at 0.4MP

The longer reference repeatedly pushed the workflow toward OOM.

The lesson is important: conditioning media is part of your memory budget. If H3 is running out of VRAM, trim the reference before automatically sacrificing final output quality.

Comparison chart showing a 20-second H3 reference producing roughly 5 seconds of output, while a 5-second reference allowed roughly 15 seconds of output in a documented 12GB workflow.

MiniMax H3 GPU Performance and RunPod Cost

Comparable workflow data also shows that more VRAM does not translate linearly into faster H3 generation.

GPU

VRAM

Approx. 5-sec Time

RTX 3090 local

24GB

14m 36s

RTX 3090 cloud

24GB

16m 05s

RTX 4090 cloud

24GB

6m 47s

RTX 5090 cloud

32GB

4m 59s

RTX PRO 6000

96GB

3m 36s

The 4090 was more than twice as fast as the 3090 in this case despite having the same VRAM capacity.

One optimized RTX 4090 workflow using INT8, SageAttention, Turbo LoRA, eight steps, approximately 0.9MP, and 10-second output took about 9–10 minutes, with an estimated compute cost of roughly $0.12 per generation under that specific cloud configuration.

That is not a universal H3 price. The more useful metric is cost per usable clip.

Memory optimization and speed optimization should also be separated. One documented 5090 workflow reduced shared GPU memory by about 14GB with optimized model components but saw little change in raw generation speed. Another 4090 optimization stack involving quantization, CUDA/Triton changes, and attention optimization reported roughly 2.6× overall speed improvement.

Radar chart comparing documented approximate five-second MiniMax H3 generation times across RTX 3090 local and cloud, RTX 4090, RTX 5090, and RTX PRO 6000 GPUs.

What Is the Best Low-Cost H3 Workflow?

The most reliable cost optimization is often previewing before rendering the final shot.

  1. Generate a 1–2 second low-resolution preview.
  2. Check framing and subject scale.
  3. Confirm Pose actually follows the intended movement.
  4. Evaluate whether Depth helps or over-constrains the shot.
  5. Check identity and general visual direction.
  6. Increase duration and resolution only after structure is stable.

In one 12GB workflow, short previews could complete in around one minute or less, versus many minutes for a full sequence.

This follows a basic professional design principle: validate structure before investing in polish.

What Are the Most Common H3 ControlNet Problems?

Fun ControlNet adds precision, but it also introduces new failure modes.

Common issues in our workflow review include:

  • dimension mismatches;
  • incorrect frame counts;
  • shape mismatches;
  • checkpoint incompatibility;
  • incorrect control batches;
  • attention-backend conflicts;
  • control signals that silently stop working.

A particularly useful attention test compared the same 90-frame, 1280 × 704 workflow:

  • standard attention: 332 seconds
  • Sol-Attn without Morton reordering: 220 seconds
  • Sol-Attn with Morton reordering: 240 seconds

The final configuration still generated plausible video, but structural control effectively disappeared because token ordering no longer aligned correctly.

This is why every new ControlNet setup should include a simple control-on versus control-off A/B test using the same seed.

Line chart showing a 90-frame H3 ControlNet test taking 332 seconds with standard attention, 220 seconds with Sol-Attn and Morton off, and 240 seconds with Morton on, where structural control disappeared.

What Pose and Depth Control Still Cannot Solve

Pose and Depth improve structural control, but they do not automatically guarantee:

  • perfect facial identity;
  • clothing consistency;
  • long-video color stability;
  • perfect hands;
  • lower VRAM use;
  • faster generation;
  • zero rerolls.

Long-form continuity remains especially difficult. One documented workflow generated approximately 10-second segments, then reused the final 2 seconds, or about 48 frames at 24 FPS, as the reference for the next segment.

The approach was extended to roughly two minutes, but gradual color drift remained visible.

This distinction is important: motion consistency and visual consistency are related, but they are not the same problem.

Timeline illustrating an H3 continuation workflow using roughly 10-second segments with the final 2 seconds, about 48 frames at 24 FPS, reused as reference and extended to roughly two minutes with color drift remaining.

Conclusion

MiniMax H3 Pose and Depth Control matters because Fun ControlNet turns H3 into a more modular video workflow: Reference can carry identity and appearance, Pose can constrain articulated motion, Depth can constrain spatial geometry, and prompts can define creative direction. The documented cases also show why the system still requires careful production thinking: multi-control can saturate, small subjects may respond differently to Depth, long references can consume significant VRAM, attention optimizations can silently break control, and hardware performance varies dramatically across configurations. The real improvement is therefore not simply “more control.” It is the ability to debug, direct, and iterate on motion, geometry, and appearance as separate creative variables.

FAQs

Does H3 support Pose and Depth natively?

No. Pose and Depth are not part of the original native H3 launch feature set. They are added through the later Fun ControlNet extension, which also supports structural controls such as Canny, HED, and MLSD.

What is the difference between H3 Reference Video and Pose Control?

Reference Video gives H3 rich semantic guidance about a performance, while Pose provides a more explicit body-motion constraint. Use Reference when creative interpretation is acceptable and Pose when choreography or subject movement needs to follow a clearer structure.

Can H3 Pose keep character identity consistent?

Not by itself. Pose controls body structure and movement, not character identity. For character replacement, use Reference for appearance and identity, Pose for movement, and Depth when the shot also needs stronger spatial control.

Can H3 run on 12GB VRAM?

Yes. Documented RTX 3060 and RTX 4070 Ti workflows show that 12GB VRAM can run H3, but resolution, duration, reference length, quantization, and offloading significantly affect performance. Lower-VRAM workflows are practical for experimentation, but they usually require more compromises than higher-VRAM production setups.

Plus d’articles du blog Virse