Published Date
2 weeks ago
Work Arrangement
Hybrid • Santa Clara, CA
Open Positions
4 openings
Experience Level
Senior
About the opportunity
Accelerate transformer training and multi-node inference pipelines using CUDA, TensorRT, and Megatron-LM.
What you will do
- check_circle Optimize deep neural network execution kernels using CUDA, Triton, and TensorRT.
- check_circle Implement pipeline, tensor, and data parallelism strategies for trillion-token pretraining runs.
- check_circle Benchmark and reduce inference latency for generative video and vision models on edge and data center GPUs.
- check_circle Engage with open-source AI frameworks (PyTorch, vLLM) to upstream acceleration features.
What we are looking for
- arrow_circle_right Bachelor’s or Master’s in Computer Science, Electrical Engineering, or related discipline.
- arrow_circle_right 5+ years hands-on experience training and profiling deep learning models on NVIDIA GPU clusters.
- arrow_circle_right Mastery of PyTorch internals, CUDA C/C++, and distributed communication libraries (NCCL).
- arrow_circle_right Deep understanding of transformer attention optimizations (FlashAttention, KV cache compression).
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.