Basic Qualifications:
- Bachelors degree and 5 years of experience or an Associates degree and 7 years of experience or a High School diploma/equivalentand 9 years of experience.
- Must be a U.S. Citizen with the ability to obtain/maintain a DHS Public Trust.
- 3 to 5 years of experience in AI/ML engineering, data engineering, or applied machine learning using cloud based technologies.
- Hands on experience with managed AI/ML services on at least two of the following cloud platforms: AWS, Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).
- Proficiency in Python and experience with machine learning frameworks and libraries such as TensorFlow, PyTorch, scikit learn, or equivalent technologies.
- Experience designing, building, deploying, and maintaining ML pipelines and model serving infrastructure in production cloud environments.
- Experience with cloud based AI services, including generative AI, large language models, machine learning platforms, or related AI capabilities.
- Familiarity with responsible AI practices, model governance, data governance, and compliance requirements associated with deploying AI solutions in federal government environments.
- Strong communication, analytical, problem solving, and technical documentation skills.
Preferred Qualifications:
- DHS Public Trust or higher clearance
- Relevant cloud or AI/ML certification, such as AWS Machine Learning Specialty, Azure AI Engineer Associate, Google Professional Machine Learning Engineer, OCI AI Foundations Associate, or an equivalent certification.
- Experience working with large language models, Retrieval Augmented Generation (RAG), and the integration of generative AI services.
- Familiarity with MLOps practices, processes, and tools such as MLflow, Kubeflow, SageMaker Pipelines, Azure ML Pipelines, or equivalent technologies.
- Experience using containerization technologies, including Docker and Kubernetes, to support AI/ML workloads.
- Knowledge of data governance frameworks, policies, and tools applicable to federal data environments and data handling requirements.
- Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, CloudFormation, or equivalent technologies.
- Additional cloud certifications across multiple cloud service providers.
- Relevant Agile certification or demonstrated experience working in Agile development environments.
Peraton is seeking a Mid-Level Cloud AI Engineer to support the development, deployment, and operation of artificial intelligence and machine learning solutions across a multi-cloud government environment serving 70+ customer tenants and growing. The environment spans AWS, Microsoft Azure, Google Cloud Platform (GCP), and Oracle Cloud Infrastructure (OCI).
Location: Remote, but must reside and perform all work within the United States
Work Hours: This position requires working online from 8:00 AM Eastern to 5:00 PM Eastern
Day to Day Roles and Responsibilities:
AI/ML Development and Deployment
- Build, train, and deploy machine learning models using managed AI/ML services across AWS (SageMaker, Bedrock), Azure (Azure ML, Azure OpenAI Service), GCP (Vertex AI), and OCI (OCI Data Science, OCI Generative AI)
- Develop and maintain ML pipelines for data ingestion, feature engineering, model training, evaluation, and deployment
- Implement model serving infrastructure including real-time inference endpoints, batch prediction workflows, and API integration patterns
- Support the integration of large language models and generative AI capabilities into government applications with appropriate guardrails and compliance controls
Data Engineering and Processing
- Design and implement data processing workflows using cloud-native services for ETL, data lake management, and feature stores
- Work with structured and unstructured data sources to prepare training datasets, ensuring data quality, lineage, and governance requirements are met
- Optimize data pipelines for performance, cost, and reliability across cloud platforms
Monitoring, Operations, and Optimization
- Monitor deployed models for performance degradation, data drift, and bias using platform-native and third-party monitoring tools
- Troubleshoot and resolve issues across AI/ML workloads, including training failures, inference latency, and resource utilization problems
- Optimize cloud resource usage and costs for AI/ML workloads including GPU/accelerator allocation and spot/preemptible instance strategies
Collaboration and Knowledge Sharing
- Collaborate with data scientists, application developers, and infrastructure engineers to operationalize AI/ML solutions
- Document AI/ML architecture decisions, deployment procedures, and operational runbooks
- Support service delivery metrics and reporting in coordination with the Service Delivery Manager and ISR Product Owner
- Adhere to Change Management procedures for all production AI/ML deployments