Senior System Software Engineer, Agentic Inference - Dynamo

Company: Nvidia

Location: US, CA, Santa Clara (Remote)

Salary: $224k - $356.5k per year

Type: Full-time

Remote: Yes

Posted: 2026-07-27

About this role

We are now looking for a Senior System Software Engineer to work on [Dynamo](https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdeveloper.nvidia.com%2Fdynamo&data=05%7C02%7Cpvelasquez%40nvidia.com%7Ca51b8a633427445f569308dee423bb6a%7C43083d15727340c1b7db39efd9ccc17a%7C0%7C0%7C639199039317198740%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=VbtoerO2IbnuUy9XfH%2BuvJp%2Fb1MYVN%2BPk62cvQlKTK0%3D&reserved=0). NVIDIA is hiring software engineers for its GPU-accelerated deep learning software team. Academic and commercial groups around the world are using GPUs to power a revolution in AI, enabling breakthroughs in problems from image classification to speech recognition to natural language processing. We are a fast-paced team building Generative AI inference platform to make design and deployment of new AI models easier and accessible to all users.

What you'll be doing:

  • In this role, you will develop open source software to serve inference of trained AI models running on GPUs.
  • Contribute to the development of disaggregated serving for Dynamo-supported inference engines (vLLM, SGLang, TRT-LLM) and expand these capabilities to support agentic inference workloads, including long-horizon reasoning, tool calling, and stateful, multi-turn execution.
  • Innovate in inference-state management for long-running agents, including KV- and prefix-cache reuse and transfer across heterogeneous memory and storage hierarchies with NIXL, to reduce repeated prompt processing, improve latency and token throughput, maximize GPU utilization, and lower per-token and per-task costs for self-hosted LLMs.
  • Build and evolve Dynamo’s distributed inference frontend across vLLM, SGLang, and TensorRT-LLM, delivering day-0 support for new models, model-specific request parameters, upstream API compatibility, and stateful Responses API semantics.
  • Balance a variety of objectives: build robus...

Create Your Job Alert

Other Senior Jobs

Other Jobs in US