Skip to main content
Machine Learning Engineer - Model Optimization & Inference — NVIDIA
NVIDIA logo
Machine Learning · Verified Opening #58

Machine Learning Engineer - Model Optimization & Inference

business NVIDIA location_on Remote — USA (Remote) home_work Remote Full Time

Compensation

$98K – 125K/yr

bolt Apply for this role
calendar_month

Published Date

1 week ago

home_work

Work Arrangement

Remote • Remote — USA

group_add

Open Positions

2 openings

workspace_premium

Experience Level

Associate / Mid

overview Role Overview

About the opportunity

Optimize deep learning models for high-throughput GPU inference and test accelerated AI pipelines.

NVIDIA is hiring a Machine Learning Engineer to work on deep learning deployment and runtime performance. You will profile, quantize, and optimize neural network models using TensorRT and ONNX Runtime, ensuring peak efficiency and minimal latency across cloud and edge hardware.

Key Focus Areas:
• Deep learning model quantization (FP16, INT8) and post-training pruning.
• High-throughput batch inference optimization on NVIDIA TensorRT.
• Deploying models on Triton Inference Server in Kubernetes.
• Benchmarking throughput, memory bandwidth, and GPU compute efficiency.
task Core Responsibilities

What you will do

  • check_circle Benchmark model inference speed and memory footprint across modern GPU architectures.
  • check_circle Apply quantization (FP16, INT8) and pruning techniques to vision and language models.
  • check_circle Containerize deployment pipelines using Docker, Triton Inference Server, and Kubernetes.
  • check_circle Collaborate with application engineers to debug pipeline bottlenecks.
verified_user Candidate Profile

What we are looking for

  • arrow_circle_right Solid programming skills in Python and foundational knowledge of C++ or CUDA concepts.
  • arrow_circle_right Hands-on experience with PyTorch or TensorFlow model development and export (ONNX).
  • arrow_circle_right Understanding of GPU computing, memory bandwidth, and neural network latency trade-offs.
  • arrow_circle_right Passion for hardware-software co-optimization.
code_blocks Technologies & Competencies

Skills & Tech Stack

PyTorch TensorRT Python ONNX Docker GPU Acceleration
handshake Direct Referral Network

Why candidate applications stand out

verified

Verified Technical Credentials

Applications include direct proof-of-work repositories and instructor verification endorsements.

person_check

Fast-Track Hiring Visibility

Direct internal referral channels through enterprise partners bypass automated resume discard filters.

work Planning your next career move? Get professional course guidance from TutorAI.

cookie

We value your privacy

We use cookies to run the site and, with your consent, to understand how it's used so we can improve it. See our Cookie Policy and Privacy Policy.