ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

Mingda Lin1,† Weijie Wang1,†* Zeyu Zhang1 Bowen Cui1 Yefei He1 Haoyu Zhao1 Yuanyu He1 Donny Y. Chen2 Feng Chen3,* Bohan Zhuang1 Equal contribution * Corresponding authors

1 Zhejiang University 2 Monash University 3 University of Adelaide

1token on ShapeNet
4tokens on TRELLIS
5shared refinement passes

Spend fewer tokens. Let shared computation recover the detail.

Fixed-budget tokenizers ask every token to share the load. ZipTok3D instead learns an ordered family of prefixes: early tokens preserve object-wide geometry, while later tokens add residual detail. A recurrent decoder then turns a short code into a complete triplane without generative completion.

32x shorter than COD-VAE-32 at the ShapeNet headline point, with matched rounded CD and F1.
Comparison of Ground Truth, VecSet with 512 tokens, COD-VAE with 32 tokens, and ZipTok3D with one token and five refinement steps.

Prefix length × refinement depth explorer

Refinement