Remote AI Platform Engineer: Delivery & Observability

EY

Richmond (VA)

Hybrid

USD 126,000 - 230,000

Full time

1 hour ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model
Total Rewards package
Paid time off

Job summary

EY is seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform across cloud, on‑prem, edge, and air‑gapped environments. You will automate pipelines, secure model execution, and govern AI assets with strong FinOps and compliance awareness.

This role sits at the intersection of DevOps, MLOps, FinOps, and observability, requiring deep expertise in delivering and operating AI infrastructure.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or related technical field.
  • 8+ years in DevOps, MLOps, platform, or observability engineering, with hands‑on production ownership of AI or high‑throughput services.
  • Hands‑on expertise operating inference/model‑serving frameworks (Ray Serve, vLLM, Triton, or NIM) on GPU infrastructure.
  • Strong hands‑on DevOps experience, including CI/CD/CV pipelines and GitOps tooling (ArgoCD, Helm, GitHub Actions/GitLab CI, or equivalents).
  • Experience with observability stacks (Prometheus, Grafana, Loki, Tempo/Jaeger) and OpenTelemetry.
  • Experience with API gateways and request routing (Envoy or equivalent), including streaming responses.
  • Experience with cost management / FinOps tooling (OpenCost, Kubecost, or equivalent) and quota/rate‑limit enforcement.
  • Familiarity with model/artifact registries and supply‑chain scanning (Harbor, MLflow, Trivy/SBOM).
  • Proven track record operating AI or service infrastructure under compliance, security, or regulatory constraints.
  • Ability to define clean ownership boundaries and consumption contracts with platform, trust, and data teams.

Responsibilities

  • Own DevOps and delivery for AI workloads: build and operate the CI/CD/CV pipelines that ship AI services, agents, and runtime components, including automated build, test, continuous verification, release, and rollback, so AI workloads are delivered repeatably and safely into every environment.
  • Own governance and discovery for AI assets, including service catalog/registry (Artifactory/Nexus, Harbor), experiment tracking and model metadata (MLflow), upstream registries/mirrors (HuggingFace/NGC), CVE/SBOM scanning (Trivy), lineage contracts (OpenLineage), and license management.
  • Own resource and cost management, including quotas and rate limits, cost attribution and utilization (Apptio/OpenCost/Kubecost), so AI execution stays economically bounded and controllable per tenant and engagement.
  • Own the full observability stack, including metrics (Prometheus/Mimir), logs (Loki), traces (Tempo/Jaeger), dashboards (Grafana), LLM debugging and evaluation (LangSmith/Langfuse), and SLA/alert notifications.
  • Own the OpenTelemetry collection layer, including multi‑tenant receiver, exporters and queues (Kafka sink), DCGM exporter for GPU telemetry, processor batching, and dynamic filtering, so every signal is captured and routed reliably.
  • Automate GitOps‑based delivery and continuous verification; embedding quality, integrity, and cost gates into pipelines so releases are policy‑compliant by default rather than by manual review.
  • Close the loop between delivery and observability by using telemetry, evaluation, and cost signals to drive deployment decisions, progressive rollout, and automated rollback of AI workloads.
  • Ensure cost and telemetry are identity‑stamped and per‑tenant, so consumption and behavior are attributable end‑to‑end, keeping FinOps and observability tied to the workloads that generate the load.

Skills

DevOps ownership
AI platform experience
CI/CD pipelines
GitOps tooling
Model serving frameworks
Observability stacks
OpenTelemetry
Cost governance
Artifact registries
Regulatory compliance

Education

Bachelor's/Master's in CS or related field

Tools

Ray Serve
vLLM/Triton/NIM
Prometheus
Grafana
Loki
Tempo/Jaeger
Envoy
Harbor/MLflow
Trivy
OpenCost/Kubecost

Job description

EY is seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform across cloud, on‑prem, edge, and air‑gapped environments. You will automate pipelines, secure model execution, and govern AI assets with strong FinOps and compliance awareness.

This role sits at the intersection of DevOps, MLOps, FinOps, and observability, requiring deep expertise in delivering and operating AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platforms Engineer: Delivery, Governance & Observability
AI Platforms Engineer: Delivery, Governance & Observability

EY • Chantilly (VA)

Hybrid
USD 126,000 - 230,000
Medical and dental coverage
Pension and 401(k)
Paid time off
AI Platform Delivery & Observability Lead
AI Platform Delivery & Observability Lead

EY • Fort Worth (TX)

Hybrid
USD 126,000 - 230,000
Hybrid work model
Total rewards package: medical/dental,
401(k) plans
AI Platform Delivery & Observability Manager
AI Platform Delivery & Observability Manager

Ernst & Young Oman • Honolulu (HI)

Hybrid
USD 126,000 - 230,000
Hybrid work model
Flexible vacation policy
AI Platform Delivery & Observability Lead
AI Platform Delivery & Observability Lead

EY • San Antonio (TX)

Hybrid
USD 126,000 - 230,000
Hybrid work model
Medical and dental coverage
Pension and 401(k)
+1
AI Platform DevOps & Observability Lead
AI Platform DevOps & Observability Lead

EY • Portland (OR)

Hybrid
USD 126,000 - 230,000
Medical & dental coverage
Pension & 401(k)
Paid time off
AI Platform Engineer: DevOps & Observability Lead
AI Platform Engineer: DevOps & Observability Lead

Ernst & Young Oman • City of Albany (NY)

On-site
USD 151,000 - 262,000
AI Platform Delivery and Observability Lead
AI Platform Delivery and Observability Lead

EY • Tallahassee (FL)

Hybrid
USD 126,000 - 230,000
Hybrid model
Medical and dental coverage
Pension and 401(k)
Remote AI Systems Engineer & DevOps Observability Lead
Remote AI Systems Engineer & DevOps Observability Lead

EY • Seattle (WA)

Hybrid
USD 126,000 - 262,000
AI Systems Engineer: Observability, Delivery & FinOps
AI Systems Engineer: Observability, Delivery & FinOps

EY • Baton Rouge (LA)

Hybrid
USD 107,000 - 177,000
Hybrid work model
Medical and dental coverage
Pension and 401(k) plans
+1
AI Platform Engineer: DevOps, Observability & Governance
AI Platform Engineer: DevOps, Observability & Governance

Ernst & Young Oman • City of Albany (NY)

Hybrid
USD 107,000 - 177,000
Hybrid work model
Competitive compensation
Total Rewards package
+1