Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure
Company: NVIDIA
Location: Austin, TX 78717
Salary: $152,000 - $287,500 a year
Type: Full-time
Posted: 2026-09-01
About this role
Sr Software Engineer - Distributed Systems Engineer, EDA Infrastructure
NVIDIA is hiring engineers to build and scale the infrastructure that supports our Electronic Design Automation (EDA) workloads. We are looking for engineers with strong programming skills, a deep understanding of distributed systems, experience operating large-scale production infrastructure, and excellent communication and planning abilities. You will help design reliable automation and platform services that manage large fleets of GPU-based and CPU-based compute systems used by engineering teams across NVIDIA.
The ideal candidate is comfortable working across software, operating systems, cluster schedulers, networking, storage, and physical hardware. You should enjoy solving complex operational problems, eliminating repetitive work through automation, and building systems that remain reliable as infrastructure grows. If you are creative, pragmatic, and motivated to improve how critical engineering workloads are delivered, we would like to hear from you.
What You Will Be Doing:
- Design and build platforms that automate the provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute infrastructure.
- Develop monitoring, health-management, and remediation systems that improve the reliability, availability, and utilization of EDA compute environments.
- Automate hardware deployment, operating-system configuration, firmware and software updates, cluster enrollment, and recovery workflows.
- Build reliable services and workflows that integrate with workload schedulers, infrastructure management systems, and observability platforms.
- Use hardware diagnostics, operating-system signals, scheduler data, and network and storage telemetry to identify failures and return unhealthy systems to service.
- Work with EDA, infrastructure, networking, storage, and hardware engineering teams to deliver scalable solutions for critical chip-design workloads.
- Parti...