Data & AI Reliability Engineering Consultant/Architect

EPAM Systems

United States

On-site

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

EPAM Systems Inc. is hiring for Data & AI Reliability Engineering Consultants and Architects to lead client-facing discovery, presales and advisory work on the reliability of data platforms and AI systems.

You will own client conversations, shape solutions and proposals, and guide the engineering team that builds them. The role covers observability, data platform reliability and AI/LLM system reliability, including defining SLOs/SLIs, tool selection, and cost optimization.

Qualifications

  • 7+ years in engineering with 2+ years in a client-facing role: consultant, solution architect, presales engineer or technical lead with direct client ownership.
  • Proven presales record: led discovery or assessment workshops, produced estimates and proposals, presented to senior stakeholders.
  • Able to structure an ambiguous client problem into scope, options, trade-offs and a recommendation, in writing and live English B2+ with confident spoken delivery; can run a workshop and handle objections without support.
  • Hands-on background with at least one enterprise observability platform: New Relic, Datadog, Splunk, Dynatrace, Grafana stack or Elastic OpenTelemetry, distributed tracing, metrics and log pipelines; alert design, event correlation and noise reduction
  • SRE practice in production: SLO/SLI, error budgets, incident management, postmortems
  • Cloud (AWS, Azure or GCP), Kubernetes, Terraform or other IaC; Python or similar for automation
  • Understands how modern data platforms work and fail: Databricks, Snowflake or a cloud-native equivalent; orchestration (Airflow or similar); batch and streaming
  • Data reliability practice: data quality checks, freshness and volume monitoring, lineage, pipeline SLAs, cost observability
  • Working understanding of LLM application architecture (RAG, agents, model gateways) and what has to be measured: quality evaluation, latency, token cost, drift, guardrails
  • Experience instrumenting or operating at least one AI/ML workload in production, or designing such a solution for a client
  • Self-driven and reliable on commitments: owns deadlines for proposals and client deliverables without supervision
  • Comfortable switching between several presales and one delivery engagement
  • Uses AI assistants in daily engineering and documentation work
  • Nice to have: Vendor certifications: New Relic, Datadog, Splunk, Dynatrace; Databricks or Snowflake; cloud architect level (AWS, Azure, GCP)
  • Data observability tooling: Monte Carlo, Soda, Great Expectations, Databricks Lakehouse Monitoring, Unity Catalog
  • LLM observability and evaluation tooling: Langfuse, LangSmith, Arize, MLflow, OpenTelemetry GenAI conventions
  • AIOps and ITSM integration: ServiceNow, PagerDuty, event correlation engines
  • FinOps for observability and data platforms; licence and ingestion cost optimisation
  • AI security and governance basics: guardrails, red teaming, data masking
  • Domain experience in retail, finance or manufacturing
  • Public profile: conference talks, articles, community leadership

Responsibilities

  • Lead the technical side of presales: qualify the request, run client workshops, define scope, assumptions and estimates
  • Run discovery and maturity assessments of a client's observability, data reliability and AI operations; deliver findings and a prioritised roadmap
  • Write the solution part of proposals and RFP responses; present and defend it to client technical and business stakeholders
  • Design target architectures for observability and reliability of data platforms and AI/LLM workloads, including tool selection and migration paths
  • Define SLOs, SLIs and error budgets for data products and AI services; translate them into alerting, incident and governance processes
  • Build the business case: cost of incidents, tooling cost optimisation, expected effect of the change
  • Act as the trusted advisor for client engineering leads, SDMs and directors during the engagement
  • Lead the first phase of delivery after a won deal, then hand over to the engineering team while staying accountable for the solution
  • Review the work of engineers, set technical standards (alert-as-code, dashboards-as-code, IaC), unblock decisions
  • Turn project experience into reusable assets: offerings, accelerators, assessment frameworks, reference architectures
  • Mentor engineers growing towards consulting; take part in technical interviews
  • Represent Data & AI externally: talks, articles, vendor partnerships

Skills

Client-facing
Presales
Observability
SRE
Cloud platforms
IaC
Python automation
LLM reliability

Tools

New Relic
Datadog
Splunk
Grafana
OpenTelemetry
Terraform
Kubernetes

Job description

We are hiring Data & AI Reliability Engineering Consultants and Architects to lead client-facing discovery, presales and advisory work on the reliability of data platforms and AI systems.

The role is a consultant: the person owns the conversation with the client, shapes the solution and the proposal, and then guides the engineering team that builds it.

The consultant covers three connected areas: observability and SRE practice, data platform reliability (pipelines, quality, lineage, cost), and AI/LLM system reliability (evaluation, telemetry, guardrails).

Responsibilities
  • Lead the technical side of presales: qualify the request, run client workshops, define scope, assumptions and estimates
  • Run discovery and maturity assessments of a client's observability, data reliability and AI operations; deliver findings and a prioritised roadmap
  • Write the solution part of proposals and RFP responses; present and defend it to client technical and business stakeholders
  • Design target architectures for observability and reliability of data platforms and AI/LLM workloads, including tool selection and migration paths
  • Define SLOs, SLIs and error budgets for data products and AI services; translate them into alerting, incident and governance processes
  • Build the business case: cost of incidents, tooling cost optimisation, expected effect of the change
  • Act as the trusted advisor for client engineering leads, SDMs and directors during the engagement
  • Lead the first phase of delivery after a won deal, then hand over to the engineering team while staying accountable for the solution
  • Review the work of engineers, set technical standards (alert-as-code, dashboards-as-code, IaC), unblock decisions
  • Turn project experience into reusable assets: offerings, accelerators, assessment frameworks, reference architectures
  • Mentor engineers growing towards consulting; take part in technical interviews
  • Represent Data & AI externally: talks, articles, vendor partnerships
Requirements
  • 7+ years in engineering, of which 2+ years in a client-facing role: consultant, solution architect, presales engineer or technical lead with direct client ownership
  • Proven presales record: led discovery or assessment workshops, produced estimates and proposals, presented to senior stakeholders
  • Able to structure an ambiguous client problem into scope, options, trade-offs and a recommendation, in writing and live English B2+ with confident spoken delivery; can run a workshop and handle objections without support
  • Hands-on background with at least one enterprise observability platform: New Relic, Datadog, Splunk, Dynatrace, Grafana stack or Elastic OpenTelemetry, distributed tracing, metrics and log pipelines; alert design, event correlation and noise reduction
  • SRE practice in production: SLO/SLI, error budgets, incident management, postmortems
  • Cloud (AWS, Azure or GCP), Kubernetes, Terraform or other IaC; Python or similar for automation
  • Understands how modern data platforms work and fail: Databricks, Snowflake or a cloud-native equivalent; orchestration (Airflow or similar); batch and streaming
  • Data reliability practice: data quality checks, freshness and volume monitoring, lineage, pipeline SLAs, cost observability
  • Working understanding of LLM application architecture (RAG, agents, model gateways) and what has to be measured: quality evaluation, latency, token cost, drift, guardrails
  • Experience instrumenting or operating at least one AI/ML workload in production, or designing such a solution for a client
  • Self-driven and reliable on commitments: owns deadlines for proposals and client deliverables without supervision
  • Comfortable switching between several presales and one delivery engagement
  • Uses AI assistants in daily engineering and documentation work
Nice to have
  • Vendor certifications: New Relic, Datadog, Splunk, Dynatrace; Databricks or Snowflake; cloud architect level (AWS, Azure, GCP)
  • Data observability tooling: Monte Carlo, Soda, Great Expectations, Databricks Lakehouse Monitoring, Unity Catalog
  • LLM observability and evaluation tooling: Langfuse, LangSmith, Arize, MLflow, OpenTelemetry GenAI conventions
  • AIOps and ITSM integration: ServiceNow, PagerDuty, event correlation engines
  • FinOps for observability and data platforms; licence and ingestion cost optimisation
  • AI security and governance basics: guardrails, red teaming, data masking
  • Domain experience in retail, finance or manufacturing
  • Public profile: conference talks, articles, community leadership
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Engineer - Data Engg & AI
Lead Engineer - Data Engg & AI

Anblicks • Dallas (TX)

On-site
USD 150,000 - 190,000
Senior SRE
Senior SRE

Accelerant • United States

Remote
USD 140,000 - 210,000
Applied AI - Technical Architect
Applied AI - Technical Architect

Insight Global • Georgia

Hybrid
USD 140,000 - 190,000
Data Operations Lead
Data Operations Lead

Zohorecruit • Irvine (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Data & AI Reliability Architect (Client-Facing Lead)
Data & AI Reliability Architect (Client-Facing Lead)

EPAM Systems • United States

Remote
USD 120,000 - 180,000
Lead AI Engineer
Lead AI Engineer

RedStream Technology • Lewisville (TX)

On-site
USD 180,000 - 240,000
Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Cephas Consultancy Services Private Limited • Cary (NC)

Hybrid
USD 150,000 - 210,000
Senior AI Agentic Lead - On-Site & Automation Architect
Senior AI Agentic Lead - On-Site & Automation Architect

Vidorra Consulting Group • Mountain View (CA)

On-site
USD 120,000 - 150,000
Principal Data Architect
Principal Data Architect

Altimetrik • New Jersey

On-site
USD 180,000 - 280,000
REMOTE Agentic AI Technical Architect
REMOTE Agentic AI Technical Architect

Insight Global • Dallas (TX)

On-site
USD 140,000 - 210,000