AI Platform Engineer — DevOps, Observability & FinOps

EY

Tampa (FL)

Hybrid

USD 107,000 - 177,000

Full time

15 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical and dental coverage
Pension and 401(k) plans
Paid time off

Job summary

EY is seeking an AI Systems Engineer to own delivery, model-serving, routing, and observability for its AI-native platform. The role spans GPU inference, automated CI/CD pipelines, governance, cost attribution, and end-to-end telemetry across cloud, on‑prem, edge, and air‑gapped environments.

The ideal candidate blends DevOps, MLOps, and FinOps expertise with strong communication skills to guide engineers, architects, and leadership through capacity planning, cost control, and secure, repeatable

Qualifications

  • 8+ years in DevOps/MLOps/platform or observability engineering with production ownership.
  • Hands-on CI/CD/CV pipelines and GitOps tooling (ArgoCD, Helm, GitHub Actions/GitLab CI).
  • Hands-on expertise operating inference/model-serving frameworks (Ray Serve, vLLM, Triton, or NIM) on GPU infrastructure.
  • Strong observability stack experience (Prometheus, Grafana, Loki, Tempo/Jaeger) and OpenTelemetry.
  • Experience with API gateways and request routing (Envoy or equivalent), including streaming responses.
  • Experience with cost management / FinOps tooling (OpenCost, Kubecost, or equivalent) and quota/rate‑limit enforcement.
  • Familiarity with model/artifact registries and supply-chain scanning (Harbor, MLflow, Trivy).
  • Proven track record operating AI or service infrastructure under compliance, security, or regulatory constraints.
  • Ability to define clean ownership boundaries and consumption contracts with platform, trust, and data teams.

Responsibilities

  • Supports DevOps and delivery for AI workloads: build and operate the CI/CD/CV pipelines that ship AI services, agents, and runtime components, including automated build, test, continuous verification, release, and rollback, so AI workloads are delivered repeatably and safely into every environment.
  • Own governance and discovery for AI assets, including service catalog/registry (Artifactory/Nexus, Harbor), experiment tracking and model metadata (MLflow), upstream registries/mirrors (HuggingFace/NGC), CVE/SBOM scanning (Trivy), lineage contracts (OpenLineage), and license management.
  • Own resource and cost management, including quotas and rate limits, cost attribution and utilization (Apptio/OpenCost/Kubecost), so AI execution stays economically bounded and controllable per tenant and engagement.
  • Own the full observability stack, including metrics (Prometheus/Mimir), logs (Loki), traces (Tempo/Jaeger), dashboards (Grafana), LLM debugging and evaluation (LangSmith/Langfuse), and SLA/alert notifications.
  • Own the OpenTelemetry collection layer, including multi‑tenant receiver, exporters and queues (Kafka sink), DCGM exporter for GPU telemetry, processor batching, and dynamic filtering, so every signal is captured and routed reliably.
  • Automate GitOps-based delivery and continuous verification; embedding quality, integrity, and cost gates into pipelines so releases are policy‑compliant by default rather than by manual review.
  • Close the loop between delivery and observability by using telemetry, evaluation, and cost signals to drive deployment decisions, progressive rollout, and automated rollback of AI workloads.
  • Ensure cost and telemetry are identity‑stamped and per‑tenant, so consumption and behavior are attributable end‑to‑end, keeping FinOps and observability tied to the workloads that generate the load.

Skills

DevOps ownership
MLOps experience
Observability
Cost attribution
Regulatory compliance
Cross-team communication

Education

Bachelor's or Master's degree in CS or related field

Tools

Ray Serve
vLLM
Triton
NIM
Prometheus
Grafana
Loki
Tempo/Jaeger
Envoy
Harbor
MLflow
Trivy
Kubecost/OpenCost
ArgoCD
Helm
GitHub Actions
GitLab CI

Job description

EY is seeking an AI Systems Engineer to own delivery, model-serving, routing, and observability for its AI-native platform. The role spans GPU inference, automated CI/CD pipelines, governance, cost attribution, and end-to-end telemetry across cloud, on‑prem, edge, and air‑gapped environments.

The ideal candidate blends DevOps, MLOps, and FinOps expertise with strong communication skills to guide engineers, architects, and leadership through capacity planning, cost control, and secure, repeatable

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Engineer: DevOps, Observability & FinOps
AI Platform Engineer: DevOps, Observability & FinOps

EY • Wichita (KS)

Hybrid
USD 126,000 - 230,000
Hybrid work model
Medical and dental coverage
Pension and 401(k)
+1
AI Platform Engineer: DevOps, Observability & Governance
AI Platform Engineer: DevOps, Observability & Governance

EY • Seattle (WA)

On-site
USD 107,000 - 177,000
Hybrid work model (implied)
Medical & dental coverage
401(k)
AI Platform Engineer: DevOps, Observability & FinOps
AI Platform Engineer: DevOps, Observability & FinOps

EY • Memphis (TN)

Hybrid
USD 126,000 - 230,000
Hybrid work model
Medical and dental coverage
Pension and 401(k) plans
+1
AI Platform DevOps & Observability Engineer
AI Platform DevOps & Observability Engineer

EY • City of Albany (NY)

Hybrid
USD 107,000 - 177,000
Hybrid work model
Total Rewards package
Medical & dental coverage
AI Platform Engineer: DevOps, Observability & FinOps
AI Platform Engineer: DevOps, Observability & FinOps

EY • Columbus (OH)

On-site
USD 107,000 - 177,000
AI Platform Engineer - DevOps, Observability & FinOps
AI Platform Engineer - DevOps, Observability & FinOps

EY • St. Louis (MO)

On-site
USD 107,000 - 177,000
Hybrid work model
Total Rewards package
Paid time off
+2
AI Platform Engineer: Observability & DevOps Lead
AI Platform Engineer: Observability & DevOps Lead

EY • McLean (VA)

Hybrid
USD 126,000 - 230,000
Hybrid work model
Total Rewards package
Paid time off
+1
AI Platform DevOps & Observability Engineer
AI Platform DevOps & Observability Engineer

EY • Atlanta (GA)

Hybrid
USD 126,000 - 230,000
Hybrid work model
Medical and dental coverage
Pension and 401(k)
+2
AI Platform DevOps & Observability Lead
AI Platform DevOps & Observability Lead

EY • Nashville (TN)

Hybrid
USD 126,000 - 230,000
AI Platform DevOps & Observability Lead
AI Platform DevOps & Observability Lead

EY • Washington

Hybrid
USD 150,000 - 230,000
Hybrid work model
Medical and dental coverage
401(k) plans