Published Date
4 weeks ago
Work Arrangement
Hybrid • Menlo Park, CA
Open Positions
2 openings
Experience Level
Senior
About the opportunity
Optimize distributed training, speculative decoding, and model quantization for open-weights Llama foundation models.
What you will do
- check_circle Implement speculative decoding and KV cache optimizations reducing LLM time-to-first-token by 40%.
- check_circle Quantize foundation models using AWQ, GPTQ, and FP8 formats for fast consumer GPU deployment.
- check_circle Manage massive petabyte synthetic data generation and filtering pipelines.
- check_circle Collaborate with the open-source community to maintain Llama GitHub reference implementations.
What we are looking for
- arrow_circle_right 4+ years experience in deep learning systems and large model training.
- arrow_circle_right Proficiency in PyTorch, C++, CUDA, Triton, and vLLM or TensorRT-LLM.
- arrow_circle_right Strong understanding of LLM architecture modifications (RoPE, GQA, MoE).
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.