Observability and Evaluation Engineer

NTT DATA, Inc.

Charlotte (NC)

On-site

USD 140,000 - 190,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NTT DATA, Inc. seeks an Overwatch Observability & Evaluation Engineer to build telemetry, dashboards, and evaluation suites for Tachyon agent releases. This role ensures production AI systems can be monitored, evaluated, improved, and supported with clear operational visibility.

You will design monitoring requirements, automate evidence collection, and create runbooks to enable rapid release readiness and governance in Agile teams.

Qualifications

  • 7+ years of engineering experience with observability, monitoring, test automation, platform operations, or AI/ML systems.
  • Experience with dashboards, metrics, alerts, traces, logs, SLOs, and production monitoring.
  • Understanding of LLM evaluation, prompt evaluation, RAG evaluation, or AI quality assessment approaches.
  • Experience working in Agile engineering teams and production support environments.

Responsibilities

  • Implement observability and telemetry for LLM-powered applications, agents, tools, and platform services.
  • Build evaluation suites for agent behavior, prompt quality, response quality, retrieval performance, latency, reliability, and safety signals.
  • Develop dashboards, alerts, traces, metrics, service objectives, and reporting for production readiness.
  • Work with platform engineers and Product Owners to define monitoring requirements and evaluation metrics.
  • Automate evidence collection for release readiness, operational reviews, and governance checkpoints.
  • Create runbooks and support documentation for priority agent releases.
  • Analyze production behavior and recommend improvements to reliability, performance, and quality.

Skills

Observability
Python
Telemetry
Dashboards
SLOs
Traces
Alerts
Monitoring
Agile
Production Ops

Tools

OpenTelemetry

Job description

Position Summary

The Overwatch Observability & Evaluation Engineer will build telemetry, tracing, dashboards, evaluation suites, alerts, service objectives, runbooks, and readiness evidence for Tachyon agent releases. This role ensures production AI systems can be monitored, evaluated, improved, and supported with clear operational visibility.

Key Responsibilities
  • Implement observability and telemetry for LLM-powered applications, agents, tools, and platform services.
  • Build evaluation suites for agent behavior, prompt quality, response quality, retrieval performance, latency, reliability, and safety signals.
  • Develop dashboards, alerts, traces, metrics, service objectives, and reporting for production readiness.
  • Work with platform engineers and Product Owners to define monitoring requirements and evaluation metrics.
  • Automate evidence collection for release readiness, operational reviews, and governance checkpoints.
  • Create runbooks and support documentation for priority agent releases.
  • Analyze production behavior and recommend improvements to reliability, performance, and quality.
Required Qualifications
  • 7+ years of engineering experience with observability, monitoring, test automation, platform operations, or AI/ML systems.
  • Experience with dashboards, metrics, alerts, traces, logs, SLOs, and production monitoring.
  • Understanding of LLM evaluation, prompt evaluation, RAG evaluation, or AI quality assessment approaches.
  • Experience working in Agile engineering teams and production support environments.
Required Skills / Knowledge
  • Python, telemetry, tracing, monitoring, dashboards, alerting, SLOs, evaluation frameworks, test automation, and production operations.
  • Understanding of LLMs, agents, RAG, prompt performance, retrieval quality, latency, and reliability metrics.
  • Experience with observability tools and open telemetry concepts.
Preferred Qualifications
  • Experience with GenAI observability, AI evaluation tools, ML monitoring, or platform reliability engineering.
  • Experience in regulated environments with evidence and readiness documentation.
  • Kubernetes, cloud platforms, and CI/CD experience.
Expected Outcomes
  • Operational dashboards and evaluation suites for priority agent releases.
  • Clear readiness evidence, alerts, SLOs, and runbooks.
  • Improved quality, reliability, and trust in production Agentic AI systems.
About NTT DATA

NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com

NTT DATA endeavors to make https://us.nttdata.com accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at https://us.nttdata.com/en/contact-us. This contact information is for accommodation requests only and cannot be used to inquire about the status of applications.

NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Observability and Evaluation Engineer
Observability and Evaluation Engineer

JobDiva, Inc. • Charlotte (NC)

On-site
USD 120,000 - 190,000
Observability & Evaluation Engineer
Observability & Evaluation Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 130,000 - 185,000
Observability & Evaluation Engineer
Observability & Evaluation Engineer

NTT DATA Americas, Inc. • Charlotte (NC)

On-site
USD 140,000 - 190,000
Observability & Evaluation Engineer
Observability & Evaluation Engineer

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 97,000 - 145,000
Medical, dental, and vision insurance
401k with company match
Paid time off
AI Observability & Evaluation Engineer
AI Observability & Evaluation Engineer

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 120,000 - 190,000
Senior AI Systems Engineer (Agentic & LLM Production)
Senior AI Systems Engineer (Agentic & LLM Production)

NTT DATA North America • New York (NY)

On-site
USD 82,656,000 - 165,312,000
Senior AI-Driven DevOps Engineer
Senior AI-Driven DevOps Engineer

JobDiva, Inc. • Atlanta (GA)

Hybrid
USD 90,000 - 98,000
AI Observability & Evaluation Engineer
AI Observability & Evaluation Engineer

NTT DATA, Inc. • Charlotte (NC)

On-site
USD 140,000 - 190,000
AI Architect
AI Architect

Talentify • Dallas (TX)

On-site
USD 180,000 - 250,000
Senior AI Ops / DevOps Engineer
Senior AI Ops / DevOps Engineer

JobDiva, Inc. • Atlanta (GA)

On-site
USD 90,000 - 98,000