Published Date
3 weeks ago
Work Arrangement
Hybrid • Sunnyvale, CA
Open Positions
2 openings
Experience Level
Lead
About the opportunity
Pioneer SRE principles, error budgets, and chaos engineering practices across Google Cloud infrastructure.
What you will do
- check_circle Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budget policies.
- check_circle Architect automated self-healing infrastructure reducing Mean Time to Resolution (MTTR) by over 50%.
- check_circle Conduct chaos engineering drills and disaster simulations to validate multi-region failover automation.
- check_circle Lead rigorous, blameless post-mortem investigations on complex production outages.
What we are looking for
- arrow_circle_right 6+ years experience in Site Reliability Engineering, distributed systems, or high-scale DevOps.
- arrow_circle_right Fluency in Go, Python, or C++ with deep knowledge of Linux kernel internals and networking.
- arrow_circle_right Strong background managing containerized services on Kubernetes and Borg.
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.