AI Platform Engineer & DevOps Observability Lead

Ernst & Young Oman

McLean (VA)

Hybrid

USD 126,000 - 230,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model
Medical and dental coverage
Paid time off

Job summary

EY seeks an AI Systems Engineer to own delivery, model-serving, routing, and observability for an AI-native platform across cloud, on-prem, edge, and air-gapped environments. You will manage CI/CD/CV pipelines, governance, and cost visibility while ensuring secure, repeatable AI workloads.

This role sits at the intersection of DevOps, MLOps, FinOps and observability, requiring deep GPU inference experience and strong cost-management skills for multi-tenant environments.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or related technical field.
  • 8+ years in DevOps, MLOps, platform, or observability engineering, with hands-on production ownership of AI or high-throughput services.
  • Hands-on DevOps experience, including CI/CD/CV pipelines and GitOps tooling for automated build, test, release, and rollback.
  • Hands-on expertise operating inference/model-serving frameworks on GPU infrastructure.
  • Strong observability experience with metrics, logs, traces, and OpenTelemetry.
  • Experience with API gateways and request routing, including streaming responses.
  • Experience with cost management / FinOps tooling and quota enforcement.
  • Familiarity with model/artifact registries and governance tooling.
  • Proven track record delivering AI or service infrastructure under compliance and security constraints.

Responsibilities

  • Own DevOps and delivery for AI workloads: CI/CD/CV pipelines, automated build, test, verification, release and rollback.
  • Own governance and discovery for AI assets with registries and metadata.
  • Own resource and cost management to keep AI execution economically bounded per tenant.
  • Own the full observability stack: metrics, logs, traces, dashboards, and debugging tools.
  • Own the OpenTelemetry collection layer and multi-tenant telemetry routing.
  • Automate GitOps-based delivery and verification with policy gates.
  • Close the loop between delivery and observability using telemetry and cost signals to guide deployments.

Skills

DevOps
MLOps
Observability
Cost governance
Communication

Education

Bachelor’s or Master’s in Computer Science

Tools

ArgoCD
Helm
GitHub Actions
GitLab CI
Ray Serve
vLLM/Triton/NIM
Prometheus
Grafana
OpenTelemetry

Job description

EY seeks an AI Systems Engineer to own delivery, model-serving, routing, and observability for an AI-native platform across cloud, on-prem, edge, and air-gapped environments. You will manage CI/CD/CV pipelines, governance, and cost visibility while ensuring secure, repeatable AI workloads.

This role sits at the intersection of DevOps, MLOps, FinOps and observability, requiring deep GPU inference experience and strong cost-management skills for multi-tenant environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer — DevOps, Observability
Senior AI Systems Engineer — DevOps, Observability

Ernst & Young Advisory Services Sdn Bhd • Atlanta (GA)

Hybrid
USD 107,000 - 177,000
AI Systems Engineer & DevOps Observability Lead
AI Systems Engineer & DevOps Observability Lead

Ernst & Young Advisory Services Sdn Bhd • Atlanta (GA)

Hybrid
USD 126,000 - 230,000
AI Systems Engineer — Delivery, Observability & MLOps
AI Systems Engineer — Delivery, Observability & MLOps

Ernst & Young Oman • Toledo (OH)

Hybrid
USD 112,000 - 177,000
AI Platform Delivery & Observability Lead
AI Platform Delivery & Observability Lead

Ernst & Young Oman • Tucson (AZ)

Hybrid
USD 126,000 - 230,000
AI Systems Engineer: DevOps, Observability & FinOps
AI Systems Engineer: DevOps, Observability & FinOps

Ernst & Young Oman • McLean (VA)

Hybrid
USD 107,000 - 177,000
Hybrid work model
Medical and dental coverage
Paid time off
Scalable AI Platform Engineer - Kubernetes & GPU Infra
Scalable AI Platform Engineer - Kubernetes & GPU Infra

Nyu Langone Hospitals • New York (NY)

Hybrid
USD 151,000 - 262,000
Hybrid work model
Medical and dental coverage
401(k) / Pension
Senior AI Data & State Infrastructure Engineer
Senior AI Data & State Infrastructure Engineer

Ernst & Young Advisory Services Sdn Bhd • Atlanta (GA)

Hybrid
USD 107,000 - 177,000
Hybrid work model
Medical and dental coverage
Paid time off
Senior AI Data & State Platform Engineer
Senior AI Data & State Platform Engineer

Ernst & Young Oman • Rogers (AR)

Hybrid
USD 107,000 - 177,000
Hybrid work model
Medical and dental coverage
Pension and 401(k) plans
+1
Senior AI Data & State Infrastructure Engineer
Senior AI Data & State Infrastructure Engineer

Ernst & Young Oman • Toledo (OH)

On-site
USD 107,000 - 177,000
Senior Platform Engineer – AI/ML, NLP & Cloud
Senior Platform Engineer – AI/ML, NLP & Cloud

EY • Hoboken (NJ)

Hybrid
USD 134,000 - 259,000
Hybrid work model
Total Rewards package