Staff Software Engineer (Compute Architecture)
Company: CoreWeave
Location: New York, NY / Sunnyvale, CA
Salary: $188k - $275k per year
Type: Full-time
Posted: 2026-08-14
About this role
- As a Staff Software Engineer within our Compute Architecture Org, you’ll shape the backbone of our GPU-driven data centers—powering some of the most advanced workloads in AI and large-scale computing
- This isn’t just about keeping the lights on; it’s about architecting the next generation of reliable, secure, and massively scalable infrastructure
- The METALDEV team builds and operates a suite of Go-based services that power large-scale datacenter deployments
- These platforms automate complex workflows while providing deep observability and monitoring for tens of thousands of GPU servers and diverse infrastructure components—including CDUs, PDUs, and NVLink switches
- Our tooling is designed for next-generation rack systems like NVIDIA GB200 and GB300, as well as a broad range of GPU server platforms
- Provide technical leadership in designing, architecting, and operating large-scale infrastructure services for GPU servers, with a focus on security, reliability, and scalability
- Build and enhance infrastructure services and automation, including inventory management systems and lifecycle management solutions using open source technologies
- Drive strategic direction for infrastructure automation, lifecycle management, and service orchestration, making MetalDev core services more scalable and resilient
- Define best practices for API development (REST/gRPC), distributed databases, and Kubernetes orchestration—while mentoring engineers to follow your lead
- Partner with hardware, software, and operations teams to align infrastructure with business impact
- Contribute to open source communities (e.g., Go, Redfish) through collaboration and technical thought leadership
- Lead and improve CI/CD pipelines for hardware compliance, firmware management, and data systems
- Champion reliability and operational excellence by driving observability (Prometheus/Grafana), production incident response, and continuous service improvement- Skilled in applying a data-driven approa...