Jobs and Careers
TE
Senior Researcher: Artificial General Intelligence (Audio, Speech and Multimodal Processing)
TencentUnited Statesfull_timeVerifiedPosted 25 Nov 2024
💰 $219,600/yr($129,600/yr – $219,600/yr)
About the role
Business Unit
Technology Engineering Group (TEG) is responsible for supporting the company and its business groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers, TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.What the Role Entails
About the jobTencent is seeking researchers in artificial general intelligence (AGI) with a focus in audio, speech and multimodal processing at the senior and principal levels to join our AI Lab in Seattle, Beijing, and Shenzhen. We are looking for recognized experts and thought leaders specializing in speech, audio and multimodal processing to tackle a variety of tasks, including (but not limited to) speech enhancement, speech recognition, audio/speech synthesis, speech codec, music processing, and spatial audio in unified multi-modal foundation models. The ideal candidates are those who are self-motivated and passionate about advancing the state of the art of AGI by developing novel model architectures and algorithms and solving real-world problems. The job level will be determined based on the experience and accomplishments of the candidate.
- Work with other researchers to identify new and upcoming research areas, long-term ambitious research goals, and intermediate milestones by interacting with potential external and internal collaborators. Own long-term research strategy and plans to expand the impact of Tencent AI Lab.
- Identify undefined problems in existing technology and develop theoretically sound novel models and algorithms to address them.
- Design experiments, write reusable code, run evaluations, and analyze results.
- Collaborate with other researchers and engineers across functional groups to push forward the state-of-the art of AGI.
- Prioritize research that can be applied to Tencent's products. Deploy promising ideas quickly and broadly.
- Author research papers to share and generate the impact of research results across organizations and in the research community.
- Share research trends and best practices in the community by reviewing academic papers, serving on program committees and grant panels, speaking at Tencent events or research conferences, or organizing research conferences and visioning activities.
Who We Look For
- Currently has or is in the process of obtaining a PhD degree in AI, computer science, electrical engineering, math, physics, or related technical fields.
- Proven record of influential publications in AI or speech, music and audio-specific conferences/journals (e.g., NeuIPS, ICML, IEEE Trans. ASLP, ICASSP, Interspeech, ISMIR, AES.)
- Expertise in speech, music and audio processing from both a signal processing standpoint and machine learning standpoint and ability to integrate traditional signal processing techniques with deep learning models to advance current speech, music and audio systems.
- Proficient in building and optimizing models for speech recognition, synthesis, enhancement, or other audio-related tasks.
- Hands-on experience with deep learning frameworks such as PyTorch. Has proven ability to design, train, and deploy deep learning models for speech, music and audio processing tasks with ability to write efficient, reusable code for processing large volumes of high-dimensional audio data.
- Strong communication skills for articulating research ideas, results, and the impact of innovations both within the organization and in the broader research community.
- Work authorization in the country of employment at the time of hire and maintains ongoing work authorization during employment.
Qualifications (Preferred):
- Familiarity with state-of-the-art (SOTA) approaches in speech, music and audio processing, such as transformer-based models, self-supervised learning (SSL) for speech, or end-to-end speech recognition and text-to-speech systems.
- Understanding of related fields such as acoustics, auditory perception, computer vision, natural language processing, or neuroscience as they apply to speech, music and audio processing. Ability to incorporate insights from these fields into the development of novel speech, music and audio technologies.
- Experience working with large-scale speech, music, audio and video datasets and developing big models that scale across multiple GPUs or cloud-based systems.
- Experience in multi-modal foundation models.
- Experience in model optimizat
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s