Our client, a technology organization focused on advanced AI and cloud infrastructure is seeking an AWS AI Infrastructure Engineer to join their team. As an AWS AI Infrastructure Engineer, you will be part of the AI Infrastructure Engineering organization supporting large-scale AI/ML platform delivery and distributed model training. The ideal candidate will have hands-on AWS accelerator experience, strong distributed training and LLM understanding, and the ability to optimize and benchmark AI workloads for performance and scalability.
Job Title
AWS AI Infrastructure Engineer
Location
Remote (USA)
Pay Range
$45/hr - $50/hr
What's the Job?
- Design, build, and support scalable AI/ML infrastructure on AWS for distributed model training and inference.
- Implement and operate AWS AI accelerator workflows using AWS Trainium/Inferentia, EC2 Trn instances, and the AWS Neuron SDK.
- Deploy and manage AI workloads using SageMaker and Kubernetes (EKS), ensuring reliable, cloud-native operations.
- Optimize AI performance through benchmarking, profiling, and tuning, including migration from GPU-based environments to AWS AI accelerators.
- Collaborate on architecture and engineering practices for LLMs, generative AI systems, and scalable distributed training pipelines.
What's Needed?
- 5+ years of experience in Cloud, Data, or AI Infrastructure, with hands-on support of large-scale AI/ML training workloads in AWS environments.
- Strong experience with AWS Trainium, Inferentia, EC2 Trn instances, and AWS Neuron SDK for AI model training and inference.
- Solid understanding of LLMs, generative AI, distributed training concepts, and AI/ML infrastructure fundamentals.
- Hands-on experience with Kubernetes (EKS), Docker, Python, PyTorch, and cloud-native architectures.
- Exposure to Trainium or Inferentia is highly preferred, along with experience migrating or optimizing workloads for accelerator-based execution.
What's in it for me?
- Work on mission-critical AI infrastructure that enables large-scale distributed training and inference.
- Build real-world expertise with AWS AI accelerators (Trainium/Inferentia) and the Neuron SDK.
- Contribute to performance optimization through benchmarking, tuning, and migration from GPU-based setups.
- Collaborate with teams focused on scalable, cloud-native AI platform engineering.
- Opportunity to support immediate-start initiatives with a confidential submission process.
Upon completion of waiting period consultants are eligible for:
- Medical and Prescription Drug Plans
- Dental Plan
- Vision Plan
- Health Savings Account
- Health Flexible Spending Account
- Dependent Care Flexible Spending Account
- Supplemental Life Insurance
- Short Term and Long Term Disability Insurance
- Business Travel Insurance
- 401(k), Plus Match
- Weekly Pay