AI Infrastructure Engineer
Position Overview
We are seeking anAI Infrastructure Engineer to design, build, and scale the infrastructure that powers our artificial intelligence and machine learning workloads. This role sits at the intersection ofAI/ML, cloud infrastructure, distributed systems, and DevOps/MLOps.
The ideal candidate has experience building highly available, scalable infrastructure for training, deploying, and operating machine learning and generative AI applications. You will partner closely with Machine Learning Engineers, Data Scientists, Software Engineers, and Platform Engineering teams to ensure AI workloads can run reliably, securely, and efficiently at scale.
Key Responsibilities
- Design, build, and maintain scalable infrastructure forAI, machine learning, and Generative AI workloads
- Build and manage cloud infrastructure acrossAWS, Azure, and/or Google Cloud Platform
- Deploy and operate GPU-based compute environments for model training and inference
- Design infrastructure supportingLLMs, model training, fine-tuning, inference, and AI applications
- Build and manage containerized workloads usingDocker and Kubernetes
- Develop infrastructure-as-code using tools such asTerraform, CloudFormation, or Pulumi
- Build CI/CD and MLOps pipelines supporting model development and deployment
- Optimize GPU/CPU utilization, infrastructure performance, scalability, and cloud costs
- Implement monitoring, logging, observability, and alerting for AI infrastructure and services
- Support distributed training and high-performance computing environments
- Build secure, highly available systems capable of supporting production AI workloads
- Partner with ML Engineers and Data Scientists to move models from experimentation into production
- Troubleshoot infrastructure, networking, performance, and deployment issues
- Evaluate emerging AI infrastructure technologies and recommend improvements to the platform
Required Qualifications
- 3+ years of experience inCloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure
- Strong experience with at least one major cloud platform:AWS, Azure, or GCP
- Experience withKubernetes and Docker
- Experience with Infrastructure-as-Code tools such asTerraform
- Strong scripting/programming skills inPython, Bash, Go, or similar languages
- Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or similar
- Knowledge of networking, Linux systems, distributed computing, and cloud architecture
- Experience implementing monitoring and observability solutions
- Understanding of machine learning development and deployment workflows
Preferred Qualifications
- Experience managingGPU infrastructure, including NVIDIA GPUs and CUDA environments
- Experience with AI/ML frameworks such asPyTorch, TensorFlow, JAX, or Hugging Face
- Experience supportingLLM training, fine-tuning, RAG, or inference workloads
- Experience with MLOps platforms such asMLflow, Kubeflow, SageMaker, Vertex AI, or Azure Machine Learning
- Experience with distributed training technologies such asRay, DeepSpeed, PyTorch Distributed, or Horovod
- Familiarity with AI inference technologies such asvLLM, NVIDIA Triton, or TensorRT
- Experience managing Kubernetes-based GPU clusters
- Understanding of model serving, vector databases, and modern Generative AI architecture
- Experience optimizing infrastructure for performance and cloud/GPU cost efficiency
What Success Looks Like
In this role, you will help create the infrastructure foundation that allows AI teams toexperiment faster, train models efficiently, deploy AI applications reliably, and scale them into production. You will reduce friction between AI development and production while improving reliability, performance, security, and infrastructure cost.