Senior DevOps / MLOps Engineer
Company: SimpliGov
Location: Baltimore, MD (Remote)
Salary: From $150,000 a year
Type: Full-time
Remote: Yes
Posted: 2026-08-11
About this role
Role Overview:
You will own the Azure platform behind SimpliGov’s AI-native delivery model—infrastructure, Kubernetes, networking, observability, AI serving, and cost discipline. This is hands-on production engineering within a FedRAMP-conscious environment, where security, auditability, and reliability are core responsibilities.
Responsibilities:
- Deploy and operate our Azure platform: AKS, networking, identity, storage, and environments from development through production
- Own infrastructure as code end to end: environments are reproducible, drift is detected, and nothing reaches an environment without platform visibility
- Operate the AI infrastructure layer: self-hosted observability and evaluation tooling (Langfuse), product telemetry, model gateway and per-workload routing, and compliant GovCloud inference paths
- Own cloud and AI cost: metering, budgets, unit economics, MACC drawdown strategy, and active remediation; cost is an engineering metric here, not a finance afterthought
- Harden production access and controls: least privilege, secrets management, audit evidence, and a FedRAMP-conscious security posture
- Partner with AI Operations on the deploy-and-release path: Octopus Deploy, environment promotion, progressive rollout, and rollback
- Build platform reliability: monitoring, alerting, incident response, and capacity planning
- Give the microservices decomposition the platform primitives it needs: service infrastructure, scaling patterns, and clean environment boundaries
Qualifications:
- 5+ years in DevOps, platform engineering, or site reliability engineering in SaaS environments
- Deep Azure experience: AKS, networking, identity (Entra), and monitoring; you have run production Kubernetes
- Infrastructure as code as your default (Terraform, Bicep, or similar), plus strong scripting; you automate before you document
- MLOps experience: deploying and operating LLM or ML systems in production, including model gateways, ...