Published Date
1 week ago
Work Arrangement
Remote • Remote — USA
Open Positions
2 openings
Experience Level
Associate / Mid
About the opportunity
Optimize deep learning models for high-throughput GPU inference and test accelerated AI pipelines.
Key Focus Areas:
• Deep learning model quantization (FP16, INT8) and post-training pruning.
• High-throughput batch inference optimization on NVIDIA TensorRT.
• Deploying models on Triton Inference Server in Kubernetes.
• Benchmarking throughput, memory bandwidth, and GPU compute efficiency.
What you will do
- check_circle Benchmark model inference speed and memory footprint across modern GPU architectures.
- check_circle Apply quantization (FP16, INT8) and pruning techniques to vision and language models.
- check_circle Containerize deployment pipelines using Docker, Triton Inference Server, and Kubernetes.
- check_circle Collaborate with application engineers to debug pipeline bottlenecks.
What we are looking for
- arrow_circle_right Solid programming skills in Python and foundational knowledge of C++ or CUDA concepts.
- arrow_circle_right Hands-on experience with PyTorch or TensorFlow model development and export (ONNX).
- arrow_circle_right Understanding of GPU computing, memory bandwidth, and neural network latency trade-offs.
- arrow_circle_right Passion for hardware-software co-optimization.
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.