Published Date
3 weeks ago
Work Arrangement
Hybrid • Seattle, WA
Open Positions
3 openings
Experience Level
Senior
About the opportunity
Engineer distributed training fabrics and petabyte-scale feature stores supporting global ranking algorithms.
What you will do
- check_circle Develop fault-tolerant training primitives in PyTorch supporting automatic checkpoint save/resume.
- check_circle Optimize networking pipelines utilizing RoCEv2 and InfiniBand for ultra-low latency all-reduce communication.
- check_circle Profile CPU-GPU data ingestion pipelines to maximize GPU compute utilization percentages.
- check_circle Design self-healing cluster orchestration tools for long-running pretraining workflows.
What we are looking for
- arrow_circle_right 5+ years experience in systems programming (C++, Python, Rust) and distributed systems.
- arrow_circle_right Strong familiarity with PyTorch internals, FSDP, DeepSpeed, and high-performance networking.
- arrow_circle_right Experience troubleshooting large-scale compute infrastructure in hyperscale cloud environments.
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.