Pixel Codec Avatars

被引:65
作者
Ma, Shugao [1 ]
Simon, Tomas [1 ]
Saragih, Jason [1 ]
Wang, Dawei [1 ]
Li, Yuecheng [1 ]
De La Torre, Fernando [1 ]
Sheikh, Yaser [1 ]
机构
[1] Facebook Real Labs Res, Menlo Pk, CA 94025 USA
来源
2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021 | 2021年
关键词
D O I
10.1109/CVPR46437.2021.00013
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Telecommunication with photorealistic avatars in virtual or augmented reality is a promising path for achieving authentic face-to-face communication in 3D over remote physical distances. In this work, we present the Pixel Codec Avatars (PiCA): a deep generative model of 3D human faces that achieves state of the art reconstruction performance while being computationally efficient and adaptive to the rendering conditions during execution. Our model combines two core ideas: (1) a fully convolutional architecture for decoding spatially varying features, and (2) a rendering-adaptive per-pixel decoder. Both techniques are integrated via a dense surface representation that is learned in a weakly-supervised manner from low-topology mesh tracking over training images. We demonstrate that PiCA improves reconstruction over existing techniques across testing expressions and views on persons of different gender and skin tone. Importantly, we show that the PiCA model is much smaller than the state-of-art baseline model, and makes multi-person telecommunicaiton possible: on a single Oculus Quest 2 mobile VR headset, 5 avatars are rendered in realtime in the same scene.
引用
收藏
页码:64 / 73
页数:10
相关论文
共 28 条
  • [1] A Decoupled 3D Facial Shape Model by Adversarial Training
    Abrevaya, Victoria Fernandez
    Boukhayma, Adnane
    Wuhrer, Stefanie
    Boyer, Edmond
    [J]. 2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 9418 - 9427
  • [2] Aliev K.-A., 2019, ARXIV190608240
  • [3] [Anonymous], 2017, The IEEE International Conference on Computer Vision (ICCV)
  • [4] Modeling Facial Geometry using Compositional VAEs
    Bagautdinov, Timur
    Wu, Chenglei
    Saragih, Jason
    Fua, Pascal
    Sheikh, Yaser
    [J]. 2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 3877 - 3886
  • [5] A morphable model for the synthesis of 3D faces
    Blanz, V
    Vetter, T
    [J]. SIGGRAPH 99 CONFERENCE PROCEEDINGS, 1999, : 187 - 194
  • [6] Cheng S, 2019, MeshGAN: Non-linear 3D Morphable Models of FacesJ
  • [7] Chu Hang, 2020, EXPRESSIVE TELEPRESE
  • [8] Goodfellow I., 2020, ADV NEUR IN, V63, P139, DOI DOI 10.1145/3422622
  • [9] Kingma DP, 2014, ADV NEUR IN, V27
  • [10] Lewis J. P., 2014, EUROGRAPHICS