INTERNSHIP DETAILS

PhD Research Intern - Emotional Speech Generation

CompanyDolby Laboratories, Inc.
LocationNot specified
Work ModeOn Site
PostedAugust 10, 2026
Internship Information
Core Responsibilities
The intern will design and implement a controllable multi-modal emotional speech generation system using cutting-edge generative models. They will also conduct literature reviews, perform hands-on development, and contribute to academic paper or patent writing.
Internship Type
full time
Company Size
1993
Visa Sponsorship
No
Language
English
Working Hours
40 hours
Apply Now →

You'll be redirected to
the company's application page

About The Company
We're the rain on the roof in a movie. The music flowing through your earbuds when you're at the gym. The footsteps lurking behind you in a video game. The voice of a colleague on a call who seems to be right next to you. The sight of a breathtakingly bright and vivid sunset on your TV. Making experiences come alive through technology is what we do. It's been our mission since day one. It began with our founder, Ray Dolby, a visionary scientist and inventor. As a young engineer and music lover, he was driven to improve the listening experience. And with that simple motivation, plus countless hours of experimentation, he created a solution—a solution that was elegant and practical, highly sophisticated, and wholly devoted to the artist's vision. Even as we've become a global company, Dolby Laboratories continues to reflect Ray Dolby's values. Here, science meets art. And high tech goes far beyond computer code. Founded in 1965 and headquartered in San Francisco, Dolby has grown into a leading global innovator and developer of audio, imaging and voice technologies for cinema, home theaters, PCs, mobile phones, and games. Our products include Dolby Digital Plus, TrueHD, Dolby Voice, Dolby Atmos and Dolby Vision. Today, over 2,000 individuals around the globe share their talents and energy to enable the most immersive experiences that technology can deliver.
About the Role

Join the leader in entertainment innovation and help us design the future. At Dolby, science meets art, and high tech means more than computer code. As a member of the Dolby team, you’ll see and hear the results of your work everywhere, from movie theaters to smartphones. We continue to revolutionize how people create, deliver, and enjoy entertainment worldwide. To do that, we need the absolute best talent. We’re big enough to give you all the resources you need, and small enough so you can make a real difference and earn recognition for your work. We offer a collegial culture, challenging projects, and excellent compensation and benefits, not to mention a Flex Work approach that is truly flexible to support where, when, and how you do your best work.

 

 

Dolby Overview:

Join the leader in entertainment innovation and help us design the future. At Dolby, science meets art, and high tech means more than computer code. As a member of the Dolby team, you’ll see and hear the results of your work everywhere, from movie theaters to smartphones. We continue to revolutionize how people create, deliver, and enjoy entertainment worldwide. To do that, we need the absolute best talent. We’re big enough to give you all the resources you need, and small enough so you can make a real difference and earn recognition for your work. We offer a collegial culture, challenging projects, and excellent compensation and benefits.

 

Advanced Technology Group (ATG) is the research and technology arm of Dolby Labs. It has multiple competencies that innovate technologies in audio, video, AR/VR, gaming, music, and movies. Many areas of expertise related to computer science and electrical engineering, such as AI/ML, computer vision, image processing, algorithms, digital signal processing, audio engineering, data science & analytics, distributed systems, cloud, edge & mobile computing, natural language processing, knowledge engineering and management, social network analysis, computer graphics, image & signal compression, computer networking, IoT are highly relevant to our research.

 

Current Dolby ATG Beijing team is looking for a talented, self-motivated Research Intern who dedicates to research deep learning algorithms for speech and audio processing, you will be involved into investigating various models and transfer the learned knowledge to the in-house deep learning models.

This position will be in the Dolby Beijing office.

 

Essential Job Functions

· Work with researchers to design and implement a controllable multi-modal emotional speech generation system.

· Investigate and apply cutting-edge generative models for speech, including diffusion models, flow matching, and related approaches, to enable high-quality emotion-aware dialogue editing.

· Build and align multi-modal emotion conditioning modules from text (e.g., natural language descriptions), audio (e.g., reference emotional speech), and image (e.g., facial expression) inputs into a shared emotion embedding space.

· Work with researchers on the full cycle of research with the goal of pushing state-of-the-art.

· Extensive literature reading, creative thinking, hands-on development, experiment design, result analysis and patent/academic paper writing.

 

Education, Skills, Abilities, and Experience Required

 

Desired Qualifications:

 

· Candidates working towards a PhD degree in the field of deep learning for speech and audio processing. PhD candidates are strongly preferred.

· Strong hands-on experience in one or more research areas in deep learning for emotional/expressive speech generation, such as text-to-speech (TTS), speech editing, or expressive speech synthesis.

· Solid understanding of and practical experience with generative models for speech/audio, such as Diffusion Models, Flow Matching, GANs, or VAEs.

· Experience with deep learning models such as spoken language models, audio/multi-modal language models, Transformer, etc.

· Familiarity with speech disentanglement techniques for separating speaker identity, linguistic content, and speaking style/emotion is a plus.

· Experience with multi-modal learning involving text, audio, and/or image modalities is a plus.

· Strong coding capability with deep learning tools, such as PyTorch.

· Demonstrable experience programming in Python.

· Excellent analytical skills and ability to communicate complex information rapidly and efficiently. Good written communication skills in English.

 

Nice to have:

· Publications in top conferences/journals (such as ICASSP, INTERSPEECH, NeurIPS, ICLR, ICML, ACL, etc.) in areas related to speech generation, emotional voice generation, expressive TTS, or multi-modal speech generation is a big plus.

· Prior project or research experience directly related to speech emotion transfer, emotional speech synthesis, or dialogue generation is strongly preferred.

 

#LI-JZ1

Key Skills
Deep learningSpeech processingAudio processingGenerative modelsDiffusion modelsFlow matchingPyTorchPythonTransformerMulti-modal learningSpeech synthesisText-to-speechData scienceAlgorithmsDigital signal processing
Categories
TechnologyScience & ResearchSoftwareData & AnalyticsEngineering
Benefits
Excellent compensationFlex work approach