Published Date
3 weeks ago
Work Arrangement
Hybrid • Redmond, WA
Open Positions
2 openings
Experience Level
Senior
About the opportunity
Develop state-of-the-art neural speech synthesis, multi-speaker diarization, and streaming voice translation models.
What you will do
- check_circle Train modern autoregressive and diffusion models for high-fidelity zero-shot expressive voice synthesis.
- check_circle Build streaming end-to-end automatic speech recognition (ASR) pipelines with sub-200ms latency.
- check_circle Implement noise-robust audio feature extractors and self-supervised speech representations (wav2vec, Whisper variants).
- check_circle Deploy optimized ONNX and TensorRT speech inference engines to Azure cloud endpoints.
What we are looking for
- arrow_circle_right MS or PhD in Computer Science, Electrical Engineering, or equivalent practical experience.
- arrow_circle_right 4+ years training deep learning models specifically for audio processing, speech recognition, or acoustics.
- arrow_circle_right Strong proficiency with PyTorch, ONNX, C++, and Python audio DSP libraries (torchaudio, librosa).
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.