Cloud AI Engineer
Company: Peraton
Location: US (Remote)
Salary: $104k - $166k per year
Type: Full-time
Remote: Yes
Posted: 2026-08-26
About this role
## Responsibilities
Peraton is seeking a Mid-Level Cloud AI Engineer to support the development, deployment, and operation of artificial intelligence and machine learning solutions across a multi-cloud government environment serving 70+ customer tenants and growing. The environment spans AWS, Microsoft Azure, Google Cloud Platform (GCP), and Oracle Cloud Infrastructure (OCI).
Location: Remote, but must reside and perform all work within the United States
Work Hours: This position requires working online from 8:00 AM Eastern to 5:00 PM Eastern
## Day to Day Roles and Responsibilities:
### AI/ML Development and Deployment
- Build, train, and deploy machine learning models using managed AI/ML services across AWS (SageMaker, Bedrock), Azure (Azure ML, Azure OpenAI Service), GCP (Vertex AI), and OCI (OCI Data Science, OCI Generative AI)
- Develop and maintain ML pipelines for data ingestion, feature engineering, model training, evaluation, and deployment
- Implement model serving infrastructure including real-time inference endpoints, batch prediction workflows, and API integration patterns
- Support the integration of large language models and generative AI capabilities into government applications with appropriate guardrails and compliance controls
### Data Engineering and Processing
- Design and implement data processing workflows using cloud-native services for ETL, data lake management, and feature stores
- Work with structured and unstructured data sources to prepare training datasets, ensuring data quality, lineage, and governance requirements are met
- Optimize data pipelines for performance, cost, and reliability across cloud platforms
### Monitoring, Operations, and Optimization
- Monitor deployed models for performance degradation, data drift, and bias using platform-native and third-party monitoring tools
- Troubleshoot and resolve issues across AI/ML workloads, including training failures, inference latency, and resource utilization pro...