Senior Site Reliability Engineer, DGX Cloud

NVIDIA

10 - 15 years

20 - 25 Lacs

kolkata mumbai new delhi hyderabad pune chennai bengaluru

Posted:2 months ago| Platform:

Apply

Skills Required

computer science capacity management networking linux gcp consulting support services gaming operations python

Work Mode

Work from Office

Job Type

Full Time

Job Description

Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real time monitoring, logging and alerting
Define SLOs/SLIs, monitor error budgets, and streamline reporting
Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews
Maintain services once they are live by measuring and monitoring availability, latency and overall system health
Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds
Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity
Lead triage and root-cause analysis of high-severity incidents
Practice balanced incident response and blameless postmortems
Participate in on-call rotation to support production services

What we need to see:

BS in Computer Science or related technical field, or equivalent experience
10+ years of experience operating production services
Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture
Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet)
Proficiency in at least one high-level programming language (e.g., Python, Go)
In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards
Proficient knowledge of SRE principles, encompassing SLOs, SLIs, error budgets, and incident handling
Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.

Ways to stand out from the crowd:

Operating GPU-accelerated clusters with KubeVirt in production
Applying generative-AI techniques to reduce operational toil
Automating incidents with Shoreline or StackStorm

More Jobs at NVIDIA

Senior System Software Engineer – Simulation and Virtualization

Mumbai Metropolitan Region

5 - 5 yrs

Salary: Not disclosed

Senior System Software Engineer – Simulation and Virtualization

Gurugram, Haryana, India

5 - 5 yrs

Salary: Not disclosed

Senior System Software Engineer

Pune, Maharashtra, India

Experience: Not specified

Salary: Not disclosed

Senior Site Reliability Engineer

Pune, Maharashtra, India

Experience: Not specified

Salary: Not disclosed

Senior System Software Engineer, GPU Firmware

Pune, Maharashtra, India

Experience: Not specified

Salary: Not disclosed

Mock Interview

Practice Video Interview with JobPe AI

Start Python Interview

Start Your Job Search Today

Browse through a variety of job opportunities tailored to your skills and preferences. Filter by location, experience, salary, and more to find your perfect fit.