Staff Research Engineer, Speech Machine Learning (TTS)
Samsung Research AmericaAbout the role
Lab Summary:
Bixby is an intelligent personal assistant which is only available as a built-in application on Samsung flagship devices and wearables. This application uses Natural Language Understanding to perform tasks on these devices using voice/ text, including but not limited to making phone calls, sending text messages, setting up meetings, opening apps, setting alarms and timers, getting directions, answering general questions, providing information about restaurants and other businesses, etc.
Position Summary:
For this position we are expanding our Advanced Intelligence Labs (AILs) voice technology and features to include advanced research and projects in Wake word detection, Automatic Speech Recognition (ASR), that includes Acoustic and Language Modeling, and personalization. We also work on language and gender detection using speech signals, Speaker identification, verification and diarization techniques. At AIL we perform state-of-the-art research in multi-lingual/accents research and bringing those research ideas to production. We are looking for candidates with extensive expertise in Digital Signal/Speech Processing with Speech recognition specialization, demonstrated research expertise by publishing papers in reputed journals/conferences, excellent knowledge of Deep/Machine Learning with 5+ years of industry experience. Candidates are expected to work in a fast paced environments.
Position Overview:
For this position we are expanding our Advanced Intelligence Labs (AILs) voice technology and features to include advanced research and projects in Text to speech synthesis, Automatic Speech Recognition (ASR), that includes Acoustic and Language Modeling, and personalization. We also work on audio generation, speech cloning technologies. At AIL we perform state-of-the-art research in multi-lingual/accents research and bringing those research ideas to production. We are looking for candidates with extensive expertise in Digital Signal/Speech Processing with Text to Speech Specialization, demonstrated research expertise by publishing papers in reputed journals/conferences, excellent knowledge of Deep/Machine Learning with 5+ years of industry experience. Candidates are expected to work in a fast paced environments.
Position Responsibilities:
- Architect and design end to end Automatic Speech Recognition products, applications and solutions for specific business needs and provide implementation guidance during delivery
- Leverage, customize and implement TTS models, algorithms, and methodologies to improve the overall quality TTS in various applications and systems
- Analyze and evaluate the performance TTS systems and provide design recommendations
- Analyze and make right technological choices for generative ai solutions
- Design and prototype reusable components for LLM based solutions for TTS
- Architect components of an TTS solution to address Responsible AI & Security
- Collaborate seamlessly with diverse, cross-functional teams to accurately identify and prioritize requirements, ensuring that the language model meets the needs and expectations of various stakeholders
- Create and maintain comprehensive technical documentation that comprehensibly captures the intricate details of the language model, facilitating seamless understanding, efficient troubleshooting, and future development
- Harness the power of transformer architecture, a cutting-edge deep learning model widely employed in natural language processing and computer vision, to optimize the language model's performance and efficiency
- Exploiting the transformative capabilities of transformer architectures to seamlessly process and reshape vast volumes of data, empowering the language model to achieve unprecedented levels of accuracy and versatility
- Ensure ethical AI development practices, prioritizing fairness, transparency, and privacy
Required Skills:
- MS or Ph.D. in Computer Science or Digital Signal Processing or equivalent combination of education, training, and experience
- 5+ years of relevant professional experience in Machine Learning or relevant field
- Experience with Tensorflow or Pytorch or similar frameworks
- Worked on advance architectures such as Tacotron, WavNet, Fastspeech, Vall-e and other advanced models for TTS systems
- Experience working in voice cloning, neural style transfer and machine synthesis of speech from speakers
- Experience in Prosody modeling for more natural generation of speech
- Working experience on TTS in large scale production systems
- Working on various vocoder techniques for production
- Experience in modeling ML algorithms on GPUs at scale
- Experience with multi-lingual TTS, low resour
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s