PhD Research Intern
Date: Aug 10, 2026
Location: Beijing, Chaoyang District, CN
Company: Dolby Laboratories, Inc.
Join the leader in entertainment innovation and help us design the future. At Dolby, science meets art, and high tech means more than computer code. As a member of the Dolby team, you’ll see and hear the results of your work everywhere, from movie theaters to smartphones. We continue to revolutionize how people create, deliver, and enjoy entertainment worldwide. To do that, we need the absolute best talent. We’re big enough to give you all the resources you need, and small enough so you can make a real difference and earn recognition for your work. We offer a collegial culture, challenging projects, and excellent compensation and benefits, not to mention a Flex Work approach that is truly flexible to support where, when, and how you do your best work.
Dolby Overview:
Join the leader in entertainment innovation and help us design the future. At Dolby, science meets art, and high tech means more than computer code. As a member of the Dolby team, you’ll see and hear the results of your work everywhere, from movie theaters to smartphones. We continue to revolutionize how people create, deliver, and enjoy entertainment worldwide. To do that, we need the absolute best talent. We’re big enough to give you all the resources you need, and small enough so you can make a real difference and earn recognition for your work. We offer a collegial culture, challenging projects, and excellent compensation and benefits.
Advanced Technology Group (ATG) is the research and technology arm of Dolby Labs. It has multiple competencies that innovate technologies in audio, video, AR/VR, gaming, music, and movies. Many areas of expertise related to computer science and electrical engineering, such as AI/ML, computer vision, image processing, algorithms, digital signal processing, audio engineering, data science & analytics, distributed systems, cloud, edge & mobile computing, natural language processing, knowledge engineering and management, social network analysis, computer graphics, image & signal compression, computer networking, IoT are highly relevant to our research.
Current Dolby ATG Beijing team is looking for a talented, self-motivated Research Intern who dedicates to research deep learning algorithms for speech and audio processing, you will be involved into investigating various models and transfer the learned knowledge to the in-house deep learning models.
This position will be in the Dolby Beijing office.
Essential Job Functions
· Work with researchers to design and implement a controllable multi-modal emotional speech generation system.
· Investigate and apply cutting-edge generative models for speech, including diffusion models, flow matching, and related approaches, to enable high-quality emotion-aware dialogue editing.
· Build and align multi-modal emotion conditioning modules from text (e.g., natural language descriptions), audio (e.g., reference emotional speech), and image (e.g., facial expression) inputs into a shared emotion embedding space.
· Work with researchers on the full cycle of research with the goal of pushing state-of-the-art.
· Extensive literature reading, creative thinking, hands-on development, experiment design, result analysis and patent/academic paper writing.
Education, Skills, Abilities, and Experience Required
Desired Qualifications:
· Candidates working towards a PhD degree in the field of deep learning for speech and audio processing. PhD candidates are strongly preferred.
· Strong hands-on experience in one or more research areas in deep learning for emotional/expressive speech generation, such as text-to-speech (TTS), speech editing, or expressive speech synthesis.
· Solid understanding of and practical experience with generative models for speech/audio, such as Diffusion Models, Flow Matching, GANs, or VAEs.
· Experience with deep learning models such as spoken language models, audio/multi-modal language models, Transformer, etc.
· Familiarity with speech disentanglement techniques for separating speaker identity, linguistic content, and speaking style/emotion is a plus.
· Experience with multi-modal learning involving text, audio, and/or image modalities is a plus.
· Strong coding capability with deep learning tools, such as PyTorch.
· Demonstrable experience programming in Python.
· Excellent analytical skills and ability to communicate complex information rapidly and efficiently. Good written communication skills in English.
Nice to have:
· Publications in top conferences/journals (such as ICASSP, INTERSPEECH, NeurIPS, ICLR, ICML, ACL, etc.) in areas related to speech generation, emotional voice generation, expressive TTS, or multi-modal speech generation is a big plus.
· Prior project or research experience directly related to speech emotion transfer, emotional speech synthesis, or dialogue generation is strongly preferred.
#LI-JZ1