Senior AI Systems Engineer

ARA

Raleigh (NC)

Hybrid

USD 150,000 - 210,000

Full time

16 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

ARA is seeking a senior AI infrastructure engineer to lead deployment, integration, and operations of enterprise AI platforms across on-prem and cloud environments. You will design scalable AI infrastructure with server, cloud, and platform teams and operationalize ML workflows from development to production.

You will build robust CI/CD and MLOps pipelines, implement IaC automation, and provide ongoing support, logging, metrics, and incident response.

Qualifications

  • Bachelor’s degree in computer science, engineering or related STEM field with 8-10 years of engineering experience.
  • 2+ years supporting AI/ML platforms, MLOps workflows, model deployment, or AI-enabled infrastructure.
  • Strong coding and automation skills in Python, Bash, or similar scripting languages.

Responsibilities

  • Lead the deployment, integration, and operational support of AI platforms, tools, and services, ensuring compatibility with existing systems and enterprise processes.
  • Design, implement, monitor, and optimize AI infrastructure, working with server, cloud, and platform engineering teams.
  • Operationalize machine learning workflows and support AI-enabled applications from development through production deployment and sustainment.
  • Build and maintain CI/CD and MLOps pipelines for model packaging, testing, deployment, rollback, and lifecycle management.
  • Implement infrastructure automation using scripting, Infrastructure as Code, and configuration management practices.
  • Provide ongoing technical support, troubleshooting, root cause analysis, and documentation for AI platforms and user-facing AI services.
  • Maintain observability across AI systems through logging, metrics, performance monitoring, alerting, and incident response practices.
  • Ensure security, compliance, and governance requirements are met, including participation in audits, vulnerability management, and secure architecture reviews.
  • Assess and implement system enhancements to improve performance, scalability, reliability, and cost efficiency.
  • Collaborate across divisions to support diverse AI initiatives and align technical implementations with mission and business objectives.
  • Evaluate emerging AI tools, frameworks, and infrastructure approaches for operational fit, supportability, and long-term value.
  • Develop and maintain technical documentation, runbooks, architecture diagrams, and operational procedures.

Skills

AI platforms
Python
Bash
Kubernetes
CI/CD
MLOps
Security/compliance (NIST/CMMC)
Communication

Education

Bachelor’s degree in CS/Engineering or related field

Tools

PyTorch
Hugging Face
MLflow
Kubeflow

Job description

  • Lead the deployment, integration, and operational support of AI platforms, tools, and services, ensuring compatibility with existing systems and enterprise processes.
  • Design, implement, monitor, and optimize AI infrastructure, working with server, cloud, and platform engineering teams.
  • Operationalize machine learning workflows and support AI-enabled applications from development through production deployment and sustainment.
  • Build and maintain CI/CD and MLOps pipelines for model packaging, testing, deployment, rollback, and lifecycle management.
  • Implement infrastructure automation using scripting, Infrastructure as Code, and configuration management practices.
  • Provide ongoing technical support, troubleshooting, root cause analysis, and documentation for AI platforms and user-facing AI services.
  • Maintain observability across AI systems through logging, metrics, performance monitoring, alerting, and incident response practices.
  • Ensure security, compliance, and governance requirements are met, including participation in audits, vulnerability management, and secure architecture reviews.
  • Assess and implement system enhancements to improve performance, scalability, reliability, and cost efficiency.
  • Collaborate across divisions to support diverse AI initiatives and align technical implementations with mission and business objectives.
  • Evaluate emerging AI tools, frameworks, and infrastructure approaches for operational fit, supportability, and long-term value.
  • Develop and maintain technical documentation, runbooks, architecture diagrams, and operational procedures.
Essential Functions
  • Lead the deployment, integration, and operational support of AI platforms, tools, and services, ensuring compatibility with existing systems and enterprise processes.
  • Design, implement, monitor, and optimize AI infrastructure, working with server, cloud, and platform engineering teams.
  • Operationalize machine learning workflows and support AI-enabled applications from development through production deployment and sustainment.
  • Build and maintain CI/CD and MLOps pipelines for model packaging, testing, deployment, rollback, and lifecycle management.
  • Implement infrastructure automation using scripting, Infrastructure as Code, and configuration management practices.
  • Provide ongoing technical support, troubleshooting, root cause analysis, and documentation for AI platforms and user-facing AI services.
  • Maintain observability across AI systems through logging, metrics, performance monitoring, alerting, and incident response practices.
  • Ensure security, compliance, and governance requirements are met, including participation in audits, vulnerability management, and secure architecture reviews.
  • Assess and implement system enhancements to improve performance, scalability, reliability, and cost efficiency.
  • Collaborate across divisions to support diverse AI initiatives and align technical implementations with mission and business objectives.
  • Evaluate emerging AI tools, frameworks, and infrastructure approaches for operational fit, supportability, and long-term value.
  • Develop and maintain technical documentation, runbooks, architecture diagrams, and operational procedures.
Experience And Skills Required
  • Bachelor’s degree in computer science, Engineering, Information Technology, or a related STEM field with 8-10 years of engineering experience.
  • 2+ years of experience supporting AI/ML platforms, MLOps workflows, model deployment, or AI-enabled infrastructure.
  • Strong coding and automation skills in Python, Bash, or similar scripting languages.
  • Experience with AI/ML frameworks and tooling such as PyTorch, Hugging Face, or similar ecosystems.
  • Proficiency with DevOps and MLOps practices, including CI/CD pipelines, Git-based workflows, containerization, and Kubernetes.
  • Experience deploying AI/ML models or AI services into operational environments, including containerized, cloud, or high-performance computing environments.
  • Familiarity with security frameworks and compliance standards such as NIST and CMMC.
  • Familiarity with AI security functionality in enterprise environments including OAuth
  • Strong communication skills and the ability to collaborate effectively across technical and non-technical teams.
Preferred
  • Advanced degree or certifications related to AI or machine learning.
  • Experience integrating AI models into scientific workflows.
  • Familiarity with large language model (LLM) APIs and orchestration frameworks such as OpenAI, Hugging Face, LangGraph, or LangChain.
  • Experience with model serving, inference optimization, or AI platform tools such as MLflow, Kubeflow, vLLM, or similar.
  • Experience with simulations for scientific or engineering projects, particularly physical systems simulations.
  • Experience with GPU-based systems or running AI models in HPC environments.
  • Experience writing and deploying MCP Servers on Kubernetes
  • DoD experience
  • Secret Security Clearance – Active or Inactive
Education
  • Bachelor’s degree in CS, Software Engineering or other IT-related field or equivalent experience

REMOTE WORK NOTICE: This position may be performed fully remote, hybrid, or onsite at an ARA office. Preference will be given to candidates located onsite in the Albuquerque, NM and Raleigh, NC area.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer
Senior AI Systems Engineer

ARA Brand • Raleigh (NC), Albuquerque (NM)

Hybrid
USD 120,000 - 170,000
Senior AI Systems Engineer
Senior AI Systems Engineer

ARA • Albuquerque (NM)

Hybrid
USD 105,000 - 165,000
Senior AI Software Engineer & Developer
Senior AI Software Engineer & Developer

Powerhouse Institute Inc • Washington

Remote
USD 155,000 - 189,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Dairy Farmers of America • Kansas City (KS)

On-site
USD 130,000 - 170,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Kansas Ag Connection • Kansas City (KS)

On-site
USD 140,000 - 190,000
Gen AI Architect
Gen AI Architect

GlobalPoint • Charlotte (NC)

On-site
USD 180,000 - 240,000
AI Infrastructure & Platform Engineer
AI Infrastructure & Platform Engineer

International Materials, LLC • Delray Beach (FL), Northern (KY)

On-site
USD 120,000 - 180,000
Senior Software Engineer/Developer - AI
Senior Software Engineer/Developer - AI

Pyramid Systems, Inc. • Washington, Northern (KY)

On-site
USD 180,000 - 280,000
AI Software Developer
AI Software Developer

Puris Corporation, Llc • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Senior AI Engineer
Senior AI Engineer

Veritas Search Group • Tustin (CA)

Hybrid
USD 150,000 - 190,000