Mesh-controllable multi-level-of-detail text-to-3D generation

被引：1

作者：

Huang, Dongjin ^{[1
]}

Wang, Nan ^{[1
]}

Huang, Xinghan ^{[2
]}

Qu, Jiantao ^{[1
]}

Zhang, Shiyu ^{[1
]}

机构：

[1] Shanghai Univ, Shanghai Film Acad, Shanghai 200072, Peoples R China

[2] Newcastle Univ, Sch Comp, Newcastle Upon Tyne NE4 5TG, England

来源：

COMPUTERS & GRAPHICS-UK | 2024年 / 123卷

关键词：

Level of detail; Multi-face Janus problem; 3D Gaussian splatting;

D O I：

10.1016/j.cag.2024.104039

中图分类号：

TP31 [计算机软件];

学科分类号：

081202 ; 0835 ;

摘要：

Text-to-3D generation is a challenging but significant task and has gained widespread attention. Its capability to rapidly generate 3D digital assets holds huge potential application value infields such as film, video games, and virtual reality. However, current methods often face several drawbacks, including long generation times, difficulties with the multi-face Janus problem, and issues like chaotic topology and redundant structures during mesh extraction. Additionally, the lack of control over the generated results limits their utility in downstream applications. To address these problems, we propose a novel text-to-3D framework capable of generating meshes with high fidelity and controllability. Our approach can efficiently produce meshes and textures that match the text description and the desired level of detail (LOD) by specifying input text and LOD preferences. This framework consists of two stages. In the coarse stage, 3D Gaussians are employed to accelerate generation speed, and weighted positive and negative prompts from various observation perspectives are used to address the multi-face Janus problem in the generated results. In the refinement stage, mesh vertices and faces are iteratively refined to enhance surface quality and output meshes and textures that meet specified LOD requirements. Compared to the state-of-the-art text-to-3D methods, extensive experiments demonstrate that the proposed method performs better in solving the multi-face Janus problem, enabling the rapid generation of 3D meshes with enhanced prompt adherence. Furthermore, the proposed framework can generate meshes with enhanced topology, offering controllable vertices and faces with textures featuring UV adaptation to achieve multi-level-of-detail(LODs) outputs. Specifically, the proposed method can preserve the output's relevance to input texts during simplification, making it better suited for mesh editing and rendering efficiency. User studies also indicate that our framework receives higher evaluations compared to other methods.

引用

页数：12

共 49 条

[1] Armandpour M, 2023, Arxiv, DOI arXiv:2304.04968
[2] Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
Barron, Jonathan T.
Mildenhall, Ben
Verbin, Dor
Srinivasan, Pratul P.
Hedman, Peter
[J]. 2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 5460 - 5469
[3] Chang AX., 2015, ARXIV151203012, V1512, P03012
[4] Text2Shape: Generating Shapes from Natural Language by Learning Joint Embeddings
Chen, Kevin
Choy, Christopher B.
Savva, Manolis
Chang, Angel X.
Funkhouser, Thomas
Savarese, Silvio
[J]. COMPUTER VISION - ACCV 2018, PT III, 2019, 11363 : 100 - 116
[5] Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation
Chen, Rui
Chen, Yongwei
Jiao, Ningxin
Jia, Kui
[J]. 2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, : 22189 - 22199
[6] Chen YW, 2023, Arxiv, DOI arXiv:2308.11473
[7] MeshCNN: A Network with an Edge
Hanocka, Rana
Hertz, Amir
Fish, Noa
Giryes, Raja
Fleishman, Shachar
Cohen-Or, Daniel
[J]. ACM TRANSACTIONS ON GRAPHICS, 2019, 38 (04):
[8] Haque A, 2023, Arxiv, DOI arXiv:2303.12789
[9] Ho J., 2020, Advances in Neural Information Processing Systems, V33, P6840
[10] Hong S, 2023, Arxiv, DOI arXiv:2303.15413

← 1 2 3 4 5 →