Position Summary
We are seeking an exceptionally skilled and experienced MLOps / Platform Engineer to join our enterprise infrastructure and machine learning operations team on a full-time onsite basis in Charlotte, NC, via Racedog Technologies. In this critical cloud platform and MLOps role, you will be responsible for designing secure cloud networking, managing IAM governance, configuring load balancers and API gateways, and streamlining ML model deployment pipelines. The ideal candidate will bring 12+ years of professional platform engineering and DevOps experience, deep mastery in AWS/Azure infrastructure, and robust MLOps tooling expertise. If you possess a distinguished background in cloud platform architecture and MLOps, we invite you to apply.
Detailed Job Description
As an MLOps / Platform Engineer in Charlotte, NC via Racedog Technologies, you will serve as a principal technical authority responsible for building and scaling enterprise MLOps platforms and secure cloud infrastructure. Your daily responsibilities will encompass configuring VPC/VNets, managing Identity and Access Management (IAM), setting up enterprise load balancers and API gateways, and automating ML model training/deployment pipelines. You will collaborate closely with data scientists, machine learning engineers, and security teams. Candidates are required to submit their updated resume via email to Raju Varma at raju.varma@racedogtechnologies.com.
Key Responsibilities
- Architect, build, and maintain scalable MLOps platforms and cloud infrastructure supporting machine learning model lifecycles.
- Configure and manage secure cloud networking including VPC/VNets, subnets, routing tables, and peering connections.
- Implement robust Identity and Access Management (IAM) policies, role-based access control, and credential rotation frameworks.
- Deploy and manage enterprise load balancers, reverse proxies, and API gateways for high-availability traffic routing.
- Automate ML model CI/CD pipelines, containerization (Docker/Kubernetes), model registry, and automated model validation.
- Collaborate proactively with data scientists and ML engineers to streamline model training, tuning, and production inference.
- Monitor platform performance, resource utilization, GPU clusters, and logging observability using Prometheus, Grafana, or Datadog.
- Enforce strict enterprise security baselines, compliance standards, and vulnerability scanning across cloud assets.
- Mentor junior DevOps and platform engineers on infrastructure-as-code (Terraform) and MLOps best practices.
- Author comprehensive infrastructure architecture diagrams, MLOps runbooks, and disaster recovery procedures.
Required Qualifications & Skills
- Bachelor’s degree in Computer Science, Information Technology, Software Engineering, or a related discipline.
- Minimum of 12+ years of professional IT experience, with at least 6+ years in DevOps, cloud platform engineering, and MLOps.
- Deep technical proficiency in cloud networking (VPC/VNet), IAM governance, load balancers, and API gateways.
- Strong practical experience implementing MLOps pipelines using MLflow, Kubeflow, SageMaker, or Vertex AI.
- Extensive hands-on expertise with Infrastructure-as-Code (Terraform, CloudFormation) and container orchestration (Kubernetes/Docker).
- Outstanding analytical problem-solving, incident management, and cross-functional communication capabilities.
- Mandatory residency/work authorization: Ability to operate fully onsite in Charlotte, NC.
Nice-to-Have Skills
- Professional certifications such as AWS Certified DevOps Engineer Professional, Certified Kubernetes Administrator (CKA), or Azure DevOps Engineer Expert.
- Prior professional experience building enterprise MLOps platforms within banking or financial services institutions in Charlotte.
- Familiarity with GPU resource management and distributed training clusters.
- Immediate availability for onsite mobilization in Charlotte.
Application Information
- Recruiter: Racedog Technologies / Technical Staffing Division
- Contact Name: Raju Varma (Account Manager / Sr. Recruiter)
- Email: raju.varma@racedogtechnologies.com
- Phone:
- Application URL:
- Salary/Rate: Competitive market rate
- Deadline: Immediate
- Notice Period: Immediate to 30 Days
- Contract Duration: Full Time / Permanent
Recruitment Pro Tip
When applying for this MLOps / Platform Engineer position with Racedog Technologies, ensure your resume explicitly highlights your VPC/IAM/API Gateway configuration expertise, your MLOps pipeline automation portfolio, and your Terraform/Kubernetes proficiency.