A leading tech solutions firm in Virginia is seeking an AI/ML Engineer with strong skills in Python and AWS. In this full-time role, you will design and implement advanced AI/ML solutions, develop cloud-native applications, and build efficient microservices. The position requires expertise in various AWS services and the ability to create robust MLOps workflows. This is an onsite role requiring 5 days in the office, offering a dynamic environment for AI-driven projects.
Responsibilities
Design and implement end-to-end AI/ML and Generative AI solutions using Python.
Build and maintain cloud native applications on AWS.
Develop high performance Python microservices enabling scalable data pipelines.
Architect and operationalize RAG pipelines and LLM powered automation.
Implement CI/CD pipelines and infrastructure as code.
Build robust MLOps workflows including model versioning and monitoring.
Skills
AI/ML
Python
AWS
Generative AI
Fast API
Flask
Terraform
CloudFormation
CI/CD
MLOps
Tools
AWS Lambda
ECS
S3
DynamoDB
RDS/Aurora
SageMaker
GitHub
GitLab
CodePipeline
Job description
Overview
Role: AI/ML Engineer - Python & AWS
Duration: Fulltime/Contract
Onsite: 5 Days Onsite
Responsibilities
Core Responsibilities (AI/ML, Python, AWS, GenAI)
Design and implement end-to-end AI/ML and Generative AI solutions using Python, including model training, evaluation, optimization, and deployment.
Build and maintain cloud native applications on AWS using services such as Lambda, ECS/Fargate, S3, API Gateway, DynamoDB, RDS/Aurora, SageMaker, and Bedrock.
Develop high performance Python microservices (Fast API/Flask) enabling scalable data pipelines, model inference, and real time analytics.
Architect and operationalize RAG pipelines, embeddings, vector databases, and LLM powered automation (chatbots, summarization, semantic search, anomaly detection).
Implement CI/CD pipelines (GitHub/GitLab/CodePipeline) and infrastructure as code (Terraform/CloudFormation) for reliable, automated deployments.
Build robust MLOps workflows, including model versioning, containerized training/inference, automated retraining, monitoring, and performance tuning.