Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
77 commits
Select commit Hold shift + click to select a range
027c8b3
feat(hunyuan3d): add native 2.1 PBR paint UNet port and loader
jethac Jul 21, 2026
13918bf
feat(hunyuan3d): add 2.1 paint multiview nodes
jethac Jul 21, 2026
0cad4ab
test(hunyuan3d): unit tests for 2.1 paint model, detection and sampler
jethac Jul 21, 2026
f9bf88e
feat(hunyuan3d): torch-native mesh rasterizer, geometry renderer and …
jethac Jul 21, 2026
3538a37
feat(save3d): support a metallic-roughness texture in the textured GL…
jethac Jul 21, 2026
14761fc
feat(hunyuan3d): wire mesh rendering + multiview bake into the 2.1 pa…
jethac Jul 21, 2026
a74dd5c
test(hunyuan3d): renderer, UV unwrap, baker and textured-GLB write-back
jethac Jul 21, 2026
7575d49
fix(hunyuan3d): make the multiview baker device-robust; add end-to-en…
jethac Jul 21, 2026
a6188d3
fix(hunyuan3d): match Tencent's rasterizer pixel convention (y-flip)
jethac Jul 21, 2026
10fa23f
Clarify camera-matrix provenance comment in paint renderer
jethac Jul 22, 2026
988ec99
Upgrade paint bake hole fill to gutter dilation plus push-pull inpaint
jethac Jul 22, 2026
fe8352e
Test bake view weighting, cosine falloff and the two-stage hole fill
jethac Jul 22, 2026
d8f4abe
Document external-UV bake targeting in Hunyuan3DBakeMultiView
jethac Jul 22, 2026
09265a2
Clarify paint module provenance and restate PoseRoPE helpers
jethac Jul 22, 2026
eec1f2a
Assert paint DDIM schedule parity with core model sampling
jethac Jul 22, 2026
397356d
Add chirality-cube convention armor with mutation self-test
jethac Jul 22, 2026
f8b5ffb
Document renderer coordinate conventions
jethac Jul 22, 2026
9e34909
Improve paint loader errors for wrong-family and truncated checkpoints
jethac Jul 22, 2026
4a576e3
Validate renderer inputs and test GLB texture colorspace round-trip
jethac Jul 22, 2026
7db47da
Add paint parity harness with tiny CI goldens
jethac Jul 22, 2026
bc1eb87
Make reference capture runnable in a slim CPU venv
jethac Jul 22, 2026
e09cffc
Fix reference attention value packing to match trained weights
jethac Jul 22, 2026
99be586
Route paint attention through comfy optimized_attention
jethac Jul 22, 2026
7a175c3
Drive the paint model through core sampling and model management
jethac Jul 22, 2026
953618d
Address review: drop einops, no_grad wrappers; harden edge cases
jethac Jul 22, 2026
d7984c8
Fix CONVENTIONS.md scale wording (diameter, not radius); hoist test i…
jethac Jul 22, 2026
64c1b9a
Assert nonzero bake coverage before MR channel checks
jethac Jul 22, 2026
6f56d77
Append texture_mr to MESH rather than inserting it mid-signature
jethac Aug 5, 2026
536a1ee
Skip the metallic-roughness encode when there is no base texture
jethac Aug 5, 2026
4d76ec0
Fail loudly on a mispacked paint latent instead of dropping per-view …
jethac Aug 5, 2026
6e9e296
Make the paint test weight init reproducible; tidy the parity harness
jethac Aug 5, 2026
090a266
Allocate the learned material embeddings with torch.empty
jethac Aug 5, 2026
8320964
Find rasterizer chunk boundaries once instead of rescanning every face
jethac Aug 5, 2026
fc132f8
Bake albedo and metallic-roughness in a single pass
jethac Aug 5, 2026
d6147e5
feat(hunyuan3d): add native 2.1 PBR paint UNet port and loader
jethac Jul 21, 2026
5e1969d
feat(hunyuan3d): add 2.1 paint multiview nodes
jethac Jul 21, 2026
7ebf720
test(hunyuan3d): unit tests for 2.1 paint model, detection and sampler
jethac Jul 21, 2026
8e5a027
feat(hunyuan3d): torch-native mesh rasterizer, geometry renderer and …
jethac Jul 21, 2026
d8d0e66
feat(save3d): support a metallic-roughness texture in the textured GL…
jethac Jul 21, 2026
74f2b27
feat(hunyuan3d): wire mesh rendering + multiview bake into the 2.1 pa…
jethac Jul 21, 2026
3d6e0f8
test(hunyuan3d): renderer, UV unwrap, baker and textured-GLB write-back
jethac Jul 21, 2026
9ef5570
fix(hunyuan3d): make the multiview baker device-robust; add end-to-en…
jethac Jul 21, 2026
d36d876
fix(hunyuan3d): match Tencent's rasterizer pixel convention (y-flip)
jethac Jul 21, 2026
5eed053
Clarify camera-matrix provenance comment in paint renderer
jethac Jul 22, 2026
c648387
Upgrade paint bake hole fill to gutter dilation plus push-pull inpaint
jethac Jul 22, 2026
8081ee5
Test bake view weighting, cosine falloff and the two-stage hole fill
jethac Jul 22, 2026
93a98a5
Document external-UV bake targeting in Hunyuan3DBakeMultiView
jethac Jul 22, 2026
db9a0bf
Clarify paint module provenance and restate PoseRoPE helpers
jethac Jul 22, 2026
8120377
Assert paint DDIM schedule parity with core model sampling
jethac Jul 22, 2026
671d1d3
Add chirality-cube convention armor with mutation self-test
jethac Jul 22, 2026
660f218
Document renderer coordinate conventions
jethac Jul 22, 2026
cae7f58
Improve paint loader errors for wrong-family and truncated checkpoints
jethac Jul 22, 2026
1ccef2a
Validate renderer inputs and test GLB texture colorspace round-trip
jethac Jul 22, 2026
a14e2c4
Add paint parity harness with tiny CI goldens
jethac Jul 22, 2026
0e78a52
Make reference capture runnable in a slim CPU venv
jethac Jul 22, 2026
3f21076
Fix reference attention value packing to match trained weights
jethac Jul 22, 2026
c3a71b6
Route paint attention through comfy optimized_attention
jethac Jul 22, 2026
a00b05e
Drive the paint model through core sampling and model management
jethac Jul 22, 2026
f248425
Address review: drop einops, no_grad wrappers; harden edge cases
jethac Jul 22, 2026
0d8081d
Fix CONVENTIONS.md scale wording (diameter, not radius); hoist test i…
jethac Jul 22, 2026
1144201
Assert nonzero bake coverage before MR channel checks
jethac Jul 22, 2026
8eaafcf
Append texture_mr to MESH rather than inserting it mid-signature
jethac Aug 5, 2026
b6342b8
Skip the metallic-roughness encode when there is no base texture
jethac Aug 5, 2026
2da13bc
Fail loudly on a mispacked paint latent instead of dropping per-view …
jethac Aug 5, 2026
4a6c9f5
Make the paint test weight init reproducible; tidy the parity harness
jethac Aug 5, 2026
57d5241
Allocate the learned material embeddings with torch.empty
jethac Aug 5, 2026
f54e62e
Find rasterizer chunk boundaries once instead of rescanning every face
jethac Aug 5, 2026
048f5e2
Bake albedo and metallic-roughness in a single pass
jethac Aug 5, 2026
e3a3869
Stop test_weight_function_compatibility polluting the shared weight_f…
jethac Aug 30, 2026
ab9c923
Clone tensors before saving parity bundles (safetensors rejects share…
jethac Aug 30, 2026
398cb00
Clean-room rewrite of Hunyuan3D 2.1 paint tainted surfaces
jethac Aug 30, 2026
33d1f4d
Fix three Tier-2 parity bugs found against the real paint UNet
jethac Aug 30, 2026
1377283
Merge pull request #4 from jethac/devin/1788073351-paint-cleanroom
devin-ai-integration[bot] Aug 30, 2026
a47f4cb
Merge current upstream master into Luna paint audit
jethac Sep 6, 2026
c588e42
Fix paint mesh save integration
jethac Sep 6, 2026
09ff1d7
Merge superseded PR implementation ancestry
jethac Sep 6, 2026
2d1c662
Make paint sampler fixture match device loading
jethac Sep 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions comfy/latent_formats.py
Original file line number Diff line number Diff line change
Expand Up @@ -1009,6 +1009,17 @@ class Hunyuan3Dv2mini(LatentFormat):
latent_dimensions = 1
scale_factor = 1.0188137142395404

class Hunyuan3Dv2_1Paint(SD15):
"""SD-2.1 image latents with the paint model's albedo/MR views packed on a
non-batch axis: (B, 4, n_pbr * views, H, W)."""
latent_dimensions = 3

def __init__(self):
super().__init__()
# The SD15 TAESD decoder assumes (B, C, H, W); previews of the packed
# layout go through latent2rgb, which handles the 5D case (first view).
self.taesd_decoder_name = None

class ACEAudio(LatentFormat):
latent_channels = 8
latent_dimensions = 2
Expand Down
103 changes: 103 additions & 0 deletions comfy/ldm/hunyuan3d/paint/CONVENTIONS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
# Coordinate & image conventions — Hunyuan3D 2.1 paint renderer/baker

Every frame this module touches, in pipeline order, and where each (non-)flip
lives. The reference-parity basis is Tencent's `hy3dpaint` renderer
(`DifferentiableRenderer/{MeshRender,camera_utils}.py`, `custom_rasterizer`),
which produced the conditioning images the released
`hunyuan3d-paintpbr-v2-1` UNet was trained on. Matching those weights — not
"looking natural" in any single buffer — is the design goal. Self-consistency
and chirality are enforced by
`tests-unit/comfy_test/test_hunyuan3d_paint_conventions.py` (chirality-cube
armor with a mutation self-test).

## 1. Input mesh frame

glTF/GLB convention: **Y-up, front = +Z**, right-handed, outward triangle
winding. This is what core `MESH` tensors carry.

## 2. World (renderer) frame — `normalize_mesh()`

`(x, y, z)_glb -> (-x, z, -y)_world`, then center on the bbox midpoint and
scale so the bounding-sphere *diameter* is `MESH_SCALE_FACTOR = 1.15`, i.e.
max vertex radius 0.575 (matches `MeshRender.set_mesh(auto_center=True)`;
this is what makes `0.5 - p / 1.15` land in `[0, 1]` for the position maps).

Two properties to be aware of:

- The map has **det = -1** (it is a mirror). Face normals computed from the
*unchanged* winding therefore flip: a glTF-outward-wound front (+Z) face has
world-frame flat normal `(0, -1, 0)`, i.e. it points *away* from the front
camera. This is the reference behaviour; the rasterizer and baker are
winding-agnostic by construction (z-buffer visibility, `|cos|` view
weighting), and the normal-map colors the UNet was trained on encode exactly
this convention (front face ≈ `[0.5, 0.0, 0.5]` after the `(n+1)/2` map).
- glTF "up" (+Y) maps to world **-Z**, which the camera model compensates
(below). Neither half of this pair may be changed independently.

## 3. Camera — `view_matrix()` (matches `camera_utils.get_mv_matrix`)

Z-up look-at orbit with `elev := -elev` and `azim := azim + 90` applied
internally, orthographic projection (`ortho_scale = 1.2`, near 0.1, far 100,
camera distance 1.45). Net mapping of the standard views to glTF axis faces
(enforced by the armor tests):

| view | elev | azim | sees glTF face |
|--------|------|------|----------------|
| front | 0 | 0 | +Z |
| right | 0 | 90 | +X |
| back | 0 | 180 | -Z |
| left | 0 | 270 | -X |
| top | 90 | 0 | +Y |
| bottom | -90 | 180 | -Y |

In the front view, screen-x tracks glTF **+X** (a character's left hand
appears on the viewer's right, as when facing someone) and camera-up is world
+Z = glTF **-Y**.

## 4. Raster buffer — `rasterize()`

Pixel row grows **with** +NDC y (`sy = (ndc_y * 0.5 + 0.5) * H`, no
top/bottom flip), matching Tencent's `custom_rasterizer`
(`barycentricFromImgcoordCPU`). Because camera-up is glTF -Y (section 3), the
composition leaves the raw buffer **upright**: glTF-up lands at row 0.

There is **no explicit image y-flip anywhere in this pipeline** — not in the
rasterizer, the geometry maps, the baker, or the GLB writer. The upright
appearance is the composition of the two deliberate sign pairs above
(mesh `-y` swap x elevation negation). If you add a flip in one place you must
remove its partner, and the trained UNet will disagree with you; the armor
tests will fail first.

## 5. Conditioning maps — `render_geometry_maps()`

- Normal maps: **object(world)-space flat normals**, `(n + 1) / 2`, white
background (mask complement). See section 2 for the winding/mirror caveat.
- Position maps: `0.5 - p_world / 1.15`, white background ("no geometry" =
white). These feed both the VAE conditioning latents and PoseRoPE
(`multires_voxel_indices` treats all-ones pixels as background).

## 6. UV / texture space — `bake_multiview()`, `pack_per_triangle_uv()`

`(u, v) -> texel (row = v * (T-1), col = u * (T-1))`; v = 0 is texture row 0.
The baked tensor's row 0 is written as the PNG top row by the save path, and
glTF's `TEXCOORD_0` origin is top-left with v growing downward — so UVs are
written to the GLB **verbatim** (`save_glb` performs no V flip) and the
round-trip is consistent end-to-end. Baking targets a mesh's existing
per-vertex UVs when present; UV-less meshes get the per-triangle atlas.

## 7. Bake weighting

Per-view weight = `view_weight * |cos(view angle)| ** bake_exp` with a 75°
grazing cutoff, z-buffer visibility per view, weighted average across views
(reference `bake_exp = 4`, per-view weights `[1.0, 0.1, 0.5, 0.1, 0.05,
0.05]`). `|cos|` (not signed cos) keeps the baker winding-agnostic under the
det = -1 world map.

## Reference-parity basis

Conventions were transcribed from `hy3dpaint` sources (read-only reference:
`MeshRender`, `camera_utils`, `custom_rasterizer`) and validated two ways:
end-to-end chest renders with the released weights (correct texture placement
at ss4 supersampling), and the procedural chirality-cube armor in this repo,
whose mutation self-test proves each convention class (V flip, axis negation,
winding reversal, camera-sign swap) is actually caught.
4 changes: 4 additions & 0 deletions comfy/ldm/hunyuan3d/paint/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
from .unet import UNet2p5DConditionModel
from .loader import detect_paint_config, load_paint_unet

__all__ = ["UNet2p5DConditionModel", "detect_paint_config", "load_paint_unet"]
162 changes: 162 additions & 0 deletions comfy/ldm/hunyuan3d/paint/attention.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,162 @@
# Specialized attention for the Hunyuan3D 2.1 paint UNet: per-material projection
# heads, wide-head reference-injection attention, and the 3D rotary embedding
# (voxelized canonical-coordinate maps) used by cross-view attention.

import torch
import torch.nn as nn

from comfy.ldm.modules.attention import CrossAttention, attention_pytorch, optimized_attention


class MaterialAttention(CrossAttention):
"""CrossAttention with additional per-material projections.

The base projections serve the albedo material; every other material gets its
own set under ``processor`` (checkpoint layout: ``attn.processor.to_q_<mat>``,
``to_out_<mat>.0`` etc.). ``full_qkv=False`` builds only the value/output pair
used by reference-injection attention.
"""

def __init__(self, query_dim, context_dim=None, heads=8, dim_head=64, extra_materials=(),
full_qkv=True, dtype=None, device=None, operations=None):
super().__init__(query_dim, context_dim=context_dim, heads=heads, dim_head=dim_head,
dtype=dtype, device=device, operations=operations)
inner_dim = heads * dim_head
kv_dim = query_dim if context_dim is None else context_dim
extra = {}
for mat in extra_materials:
if full_qkv:
extra[f"to_q_{mat}"] = operations.Linear(query_dim, inner_dim, bias=False, dtype=dtype, device=device)
extra[f"to_k_{mat}"] = operations.Linear(kv_dim, inner_dim, bias=False, dtype=dtype, device=device)
extra[f"to_v_{mat}"] = operations.Linear(kv_dim, inner_dim, bias=False, dtype=dtype, device=device)
extra[f"to_out_{mat}"] = nn.Sequential(operations.Linear(inner_dim, query_dim, dtype=dtype, device=device))
self.processor = nn.ModuleDict(extra)

def forward_per_material(self, tokens, materials, transformer_options={}):
"""Self-attention within each frame, with projections selected by material.

tokens: ``(B, M, V, L, C)`` grouped per material; returns the same shape.
"""
b, m, v, l, c = tokens.shape
outs = []
for i, mat in enumerate(materials):
x = tokens[:, i].reshape(b * v, l, c)
if i == 0:
q, k, val = self.to_q(x), self.to_k(x), self.to_v(x)
out = optimized_attention(q, k, val, self.heads, transformer_options=transformer_options)
outs.append(self.to_out(out))
else:
q = self.processor[f"to_q_{mat}"](x)
k = self.processor[f"to_k_{mat}"](x)
val = self.processor[f"to_v_{mat}"](x)
out = optimized_attention(q, k, val, self.heads, transformer_options=transformer_options)
outs.append(self.processor[f"to_out_{mat}"](out))
return torch.stack(outs, dim=1).reshape(b, v, m, l, c).transpose(1, 2)

def forward_reference(self, albedo_tokens, bank_tokens, materials, transformer_options={}):
"""Reference-injection attention with wide-head value packing.

Queries come from the albedo tokens only (``(B, V*L, C)``); keys from the
bank tokens; the per-material value projections are channel-concatenated
and split head-first into ``M*dim_head``-wide slices, so each head attends
over a contiguous chunk of the concatenated channels (mixing materials),
then each head's output is split back per material and routed through that
material's output projection. Returns ``(M, B, V*L, C)``.
"""
b, tq = albedo_tokens.shape[:2]
n_mat = len(materials)
q = self.to_q(albedo_tokens)
k = self.to_k(bank_tokens)
vals = [self.to_v(bank_tokens)]
for mat in materials[1:]:
vals.append(self.processor[f"to_v_{mat}"](bank_tokens))
v = torch.cat(vals, dim=-1)

q = q.view(b, tq, self.heads, self.dim_head).transpose(1, 2)
k = k.view(b, -1, self.heads, self.dim_head).transpose(1, 2)
v = v.view(b, -1, self.heads, n_mat * self.dim_head).transpose(1, 2)
if n_mat == 1:
out = optimized_attention(q, k, v, self.heads, skip_reshape=True,
transformer_options=transformer_options)
return self.to_out(out).unsqueeze(0)
# value head width differs from the query/key head width; fused backends
# derive it from the query, so this call needs the plain SDPA path
out = attention_pytorch(q, k, v, self.heads, skip_reshape=True, skip_output_reshape=True)
out = out.view(b, self.heads, tq, n_mat, self.dim_head)
out = out.permute(3, 0, 2, 1, 4).reshape(n_mat, b, tq, self.heads * self.dim_head)
outs = [self.to_out(out[0])]
for i, mat in enumerate(materials[1:]):
outs.append(self.processor[f"to_out_{mat}"](out[i + 1]))
return torch.stack(outs, dim=0)


def cross_view_attention(attn, tokens, rope=None, transformer_options={}):
"""Joint self-attention over all views of one (batch, material) group.

tokens: ``(B*M, V*L, C)``; ``rope`` is an optional ``(cos, sin)`` pair
(fp32, ``(B*M, 1, V*L, dim_head // 2)``) applied to Q and K.
"""
b = tokens.shape[0]
q = attn.to_q(tokens)
k = attn.to_k(tokens)
v = attn.to_v(tokens)
q = q.view(b, -1, attn.heads, attn.dim_head).transpose(1, 2)
k = k.view(b, -1, attn.heads, attn.dim_head).transpose(1, 2)
v = v.view(b, -1, attn.heads, attn.dim_head).transpose(1, 2)
if rope is not None:
cos, sin = rope
q = apply_rotary(q, cos, sin)
k = apply_rotary(k, cos, sin)
out = optimized_attention(q, k, v, attn.heads, skip_reshape=True,
transformer_options=transformer_options)
return attn.to_out(out)


def apply_rotary(x, cos, sin):
"""Pairwise rotary rotation of ``(..., T, D)`` by half-length cos/sin tables."""
x1, x2 = x.float().unflatten(-1, (-1, 2)).unbind(-1)
out = torch.stack([x1 * cos - x2 * sin, x1 * sin + x2 * cos], dim=-1)
return out.flatten(-2).to(x.dtype)


def voxelize_position_maps(position_maps, grid_h, grid_w, voxel_resolution):
"""Average canonical-coordinate maps into a token grid of integer voxel indices.

position_maps: ``(B, V, 3, Hp, Wp)`` in ``[0, 1]`` with exact-white background.
A pixel is background if any channel is exactly 1.0; cells with fewer valid
pixels than ``cell_area // 16`` quantize to the origin voxel. The masked
average runs in fp16 (trained rounding behavior). Returns ``(B, V, gh*gw, 3)``
int64 voxel coordinates.
"""
b, v, _, hp, wp = position_maps.shape
ch, cw = hp // grid_h, wp // grid_w
cells = position_maps.reshape(b, v, 3, grid_h, ch, grid_w, cw).half()
valid = (cells != 1.0).all(dim=2, keepdim=True).half()
count = valid.sum(dim=(4, 6))
mean = (cells * valid).sum(dim=(4, 6)) / count.clamp(min=1.0)
enough = count >= float((ch * cw) // 16)
coords = torch.where(enough, mean, torch.zeros_like(mean))
coords = coords.clamp(0.0, 1.0) * (voxel_resolution - 1)
coords = coords.round().long()
return coords.permute(0, 1, 3, 4, 2).reshape(b, v, grid_h * grid_w, 3)


def rotary_tables(voxels, dim_head, voxel_resolution):
"""Per-token cos/sin tables for the 3-axis rotary embedding.

``dim_head`` splits 3:3:2 eighths across x/y/z (position-map channel order);
each axis uses a standard 1-D rotary table with base theta 10000 over the
voxel range. voxels: ``(B, V, T, 3)`` int64. Returns fp32 ``(cos, sin)`` of
shape ``(B, V*T, dim_head // 2)`` (half length, pairwise application).
"""
axis_dims = (3 * dim_head // 8, 3 * dim_head // 8, dim_head // 4)
b, v, t, _ = voxels.shape
pos = torch.arange(voxel_resolution, device=voxels.device, dtype=torch.float32)
cos_parts, sin_parts = [], []
for axis, dim in enumerate(axis_dims):
freqs = 10000.0 ** (-torch.arange(0, dim, 2, device=voxels.device, dtype=torch.float32) / dim)
angles = pos[:, None] * freqs[None]
per_token = angles[voxels[..., axis].reshape(b, v * t)]
cos_parts.append(per_token.cos())
sin_parts.append(per_token.sin())
return torch.cat(cos_parts, dim=-1), torch.cat(sin_parts, dim=-1)
Loading
Loading