Senior Software Engineer, DGX Cloud Production Engineering

Company: Nvidia

Location: US, CA, Santa Clara (Remote)

Salary: $152k - $241.5k per year

Type: Full-time

Remote: Yes

Posted: 2026-08-18

About this role

NVIDIA is seeking a Senior Software engineer to build the next generation of our Kubernetes platform. Our teams build foundational capabilities for self-service GPU infrastructure, managed Kubernetes control planes, management of cluster operations, and automation for large-scale AI and improved computational environments.

In this role, you will work at the intersection of platform architecture and production-grade Kubernetes lifecycle systems. You will help with technical strategy and execution for systems that make clusters easier to provision, upgrade, operate, and scale across cloud and on-premises environments.

What you will be doing:

  • Collaborate with architecture and development of core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations.
  • Design and build highly reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
  • Define technical requirements, validation criteria, production-readiness practices, and the direction for declarative workflows and automation across the Kubernetes stack.
  • Collaborate across engineering teams to create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
  • Collaborate to diagnosis and find a resolution of complex platform issues spanning infrastructure, runtime, networking, hardware, and operations, improving the scalability, resilience, and operability of systems supporting large-scale AI deployments.**​**

What we need to see:

  • BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 5+ years of relevant software engineering experience, including experience building and operating large-scale production systems.
  • Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
  • A strong background in distribu...

Create Your Job Alert

Other Senior Jobs

Other Jobs in US