PhD CIFRE, PhD – Next-generation realistic avatar representation for videoconferencing
Description
About InterDigital
InterDigital is a global research and development company focused primarily on wireless, video, artificial intelligence (“AI”), and related technologies. We design and develop foundational technologies that enable connected, immersive experiences in a broad range of communications and entertainment products and services. We license our innovations worldwide to companies providing such products and services, including makers of wireless communications devices, consumer electronics, IoT devices, cars and other motor vehicles, and providers of cloud-based services such as video streaming. As a leader in wireless technology, our engineers have designed and developed a wide range of innovations that are used in wireless products and networks, from the earliest digital cellular systems to 5G and today’s most advanced Wi-Fi technologies. We are also a leader in video processing and video encoding/decoding technology, with a significant AI research effort that intersects with both wireless and video technologies. Founded in 1972, InterDigital is listed on Nasdaq.
InterDigital is a registered trademark of InterDigital, Inc.
For more information, visit: www.interdigital.com .
Scientific & Industrial Challenges
With the development of Virtual Reality applications, avatars have become a major feature for improving the user experience, impacting both users’ performance [Rybarczyk et al. 2014] and their appreciation of these experiences [Yee and Bailenson 2007]. However, several factors typically affect how users accept their avatars as their virtual representation in the virtual experience, which is often evaluated through the sense of embodiment [Kilteni et al. 2012]. Among these factors, several elements have already been identified as particularly important for eliciting a strong sense of embodiment, in particular the degree of realism of appearance and animation control [Argelaguet et al. 2016, Fribourg et al. 2020, Gorisse et al. 2017].
While these factors are becoming increasingly well understood in the context of single-user applications, the recent enthusiasm for large-scale multi-user applications such as the metaverse raises novel challenges in terms of creating and animating these life-like user representations. The sheer number of users expected to interact in such virtual locations raises questions related to the transmission of each user’s representation, both in terms of visual appearance and animations. What is the minimum information that needs to be available to represent a user in a shared application? Should some features be prioritized over others, e.g., facial features vs. body features? What novel representations should be proposed to account for such a context? How can such representations provide an appropriate trade-off between realism and the volume of data required to be transferred to display and animate these avatars? What is the effect of displaying different levels of realism on different parts of the avatar (e.g., realistic appearance vs. low-quality animations, or realistic facial animations vs. static hair or body)?
Answering these questions is important for the development of the next generation of videoconference and metaverse applications. In particular, in the context of its standard video and immersive activities, InterDigital is aiming to provide semantic-based data solutions for such applications. The current challenge is to stream relevant data that enables the editability, controllability, and interactivity of the content, while keeping data throughput low enough to enable the use of existing and future networks. So far, the core of InterDigital’s technology has focused on the human face and already enables the extraction of facial parameters from an input video stream (head pose and facial expressions). These parameters are then encoded and streamed to a video AI decoder capable of reconstructing a full and complete image on the decoder side. The expertise of the Inria teams revolves around body-based animation and interaction of avatars, e.g., investigating new paradigms of animation for multi-user virtual reality experiences and evaluating the impact that the resulting animation quality can have on users’ perception and behaviour.
To advance future videoconference and metaverse applications, the main goal of this Ph.D. is to explore novel approaches including both full-body and facial elements, by extending the current state of the art to enable full-body and facial encoding and decoding for multi-user immersive experiences and the evaluation of quality of experience.
Compatibility with Real-Time Avatar Communication
The surveyed methods can be grouped according to their runtime architecture. Fully explicit methods, such as SplattingAvatar, GaussianBlendshape, and GaussianAvatars, do not require runtime neural inference and can be driven directly by standard skeletal joints and blendshape weights. Hybrid methods, such as FlashAvatar and HUGS, use small neural networks to add residual motion while retaining a conventional animation interface. Fully neural methods require more complex training and runtime pipelines and are less straightforward to standardise as baseline interoperable tools.
For interoperability, the decisive issue is not only whether a neural network is present, but whether the animation interface remains unchanged. If an avatar representation can be driven solely by joints and blendshape weights, then an existing animation stream remains sufficient, and the decoder primarily needs to select an appropriate renderer. If additional per-frame latent codes, feature maps, or learned temporal states are required, the transport, synchronization, and conformance requirements become more complex. Determinism is therefore central in real-time communication: explicit methods are naturally more predictable, while hybrid methods require constrained operator sets, model distribution rules, and numerical tolerances to preserve portability across devices.
Overall, 3DGS avatar methods provide a promising basis for the representation and coding of realistic digital characters in immersive telepresence. The most relevant approaches for this Ph.D. are those that combine high visual quality with compact, semantic, and standard-compatible animation control. In particular, mesh-embedded Gaussians and Gaussian blendshapes align well with the objective of extending avatar coding beyond the face toward full-body and facial representations while maintaining editability, controllability, interactivity, and low transmission requirements.
Proposed Research Topic
To tackle these challenges, this Ph.D. will investigate Gaussian-avatar representations for multi-user immersive telepresence. The objective is to move from face-centric avatar coding toward a unified representation in which head, face, body, and visually important secondary components can be encoded, animated, rendered, and evaluated within a compact semantic framework.
Geometry-aware Gaussian representations for head and face avatars. A first direction will be to study how 3D Gaussian Splatting can be combined with geometric priors for view-dependent rendering of realistic 3D head models. Recent work such as Geometry-aware 3D Gaussian Splatting for View-Dependent Rendering of 3D Head Models shows that Gaussians can be guided by head geometry to improve novel-view rendering of non-Lambertian human faces (Boyer et al., 2025). In this Ph.D., this direction will be explored as a basis for avatar representation, with particular attention to the compatibility between high-quality view-dependent rendering and controllable animation parameters such as facial expressions, head pose, and mesh-based priors.
Semantic and standard-compatible control of Gaussian avatars. A second direction will focus on compact semantic representations that allow Gaussians to be animated through the same signals used by conventional avatars, notably skeletal joints a