Senior Software Engineer, DGX Cloud Production Engineering
Company: Nvidia
Location: US, CA, Santa Clara (Remote)
Salary: $152k - $241.5k per year
Type: Full-time
Remote: Yes
Posted: 2026-08-18
About this role
NVIDIA is seeking a Senior Software engineer to build the next generation of our Kubernetes platform. Our teams build foundational capabilities for self-service GPU infrastructure, managed Kubernetes control planes, management of cluster operations, and automation for large-scale AI and improved computational environments.
In this role, you will work at the intersection of platform architecture and production-grade Kubernetes lifecycle systems. You will help with technical strategy and execution for systems that make clusters easier to provision, upgrade, operate, and scale across cloud and on-premises environments.
What you will be doing:
- Collaborate with architecture and development of core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations.
- Design and build highly reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
- Define technical requirements, validation criteria, production-readiness practices, and the direction for declarative workflows and automation across the Kubernetes stack.
- Collaborate across engineering teams to create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
- Collaborate to diagnosis and find a resolution of complex platform issues spanning infrastructure, runtime, networking, hardware, and operations, improving the scalability, resilience, and operability of systems supporting large-scale AI deployments.****
What we need to see:
- BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 5+ years of relevant software engineering experience, including experience building and operating large-scale production systems.
- Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
- A strong background in distribu...