DevOps Engineer

Rivago Infotech Inc

Charlotte (NC)

On-site

USD 120,000 - 180,000

Full time

7 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Rivago Infotech Inc. seeks an AI DevOps/Observability Engineer to ensure reliability, performance, and readiness of Generative AI and LLM agent releases. You will bridge AI development and production operations with telemetry, evaluation pipelines, and comprehensive monitoring.

Responsibilities include deep tracing, dashboarding, alerting, SLOs, and automated evaluation suites to optimize cost, speed, and accuracy in production environments.

Qualifications

  • Experience with framework-based evaluation tools (Ragas, DeepEval, TruLens) to measure hallucination, faithfulness, and relevancy.
  • Proficiency with LLM tracing tools (LangSmith, LangFuse, Phoenix, Arize) and OpenTelemetry.
  • Hands-on experience building production dashboards in Datadog, Prometheus, Grafana, or New Relic.
  • Ability to define meaningful SLIs/SLOs and minimize alert fatigue.
  • Strong background in integrating automated test frameworks into CI/CD for AI features.
  • Analytical mindset to benchmark prompts, track regressions, and profile latency.

Responsibilities

  • Observability & Telemetry: implement deep tracing and telemetry for LLM architectures and multi-agent workflows.
  • Monitoring & Insights: design dashboards tracking system health, infra metrics, and AI performance indicators.
  • Operational Readiness: establish alerting systems, define SLOs, curate runbooks, provide production readiness evidence.
  • Evaluation & Testing Pipelines: develop automated continuous evaluation suites for LLM behavior and safety.
  • Performance Analysis: analyze prompt efficiency, token usage, latency, and model performance.
  • Production Operations: support deployment pipelines and incident management for live AI services.

Skills

LLM evaluation
Telemetry
Dashboards
SLOs
Test automation
Prompt analysis
Python
Cloud infra
DevOps

Tools

Datadog
Prometheus
Grafana
LangSmith
LangFuse
Phoenix
Arize
OpenTelemetry
LangChain
LlamaIndex

Job description

We are seeking a highly skilled AI DevOps/Observability Engineer to join our production operations team. In this role, you will be responsible for the reliability, performance, and operational readiness of our priority Generative AI and LLM agent releases. You will bridge the gap between AI development and production operations by implementing robust telemetry, automated evaluation pipelines, and comprehensive monitoring systems. Your work will ensure our intelligent agents are stable, accurate, efficient, and scale seamlessly in production environments.

Key Responsibilities
  • Observability & Telemetry: Implement and maintain deep tracing, logging, and telemetry solutions specifically tailored for LLM application architectures and multi-agent workflows.
  • Monitoring & Insights: Design, build, and maintain production dashboards that track systemic health, infrastructure metrics, and specialized AI performance indicators.
  • Operational Readiness: Establish actionable alerting systems, define Service Level Objectives (SLOs), curate runtime runbooks, and provide concrete engineering evidence for production readiness.
  • Evaluation & Testing Pipelines: Develop automated continuous evaluation suites to assess LLM agent behavior, safety, and output quality prior to and during deployment.
  • Performance Analysis: Analyze prompt efficiency, token usage, latency, and overall model performance to optimize cost, speed, and accuracy.
  • Production Operations: Support the deployment pipeline, participate in incident management, and continually improve the resilience of our live AI services.
Required Skills and Qualifications
Core Technical Skills
  • LLM & Agent Evaluation: Experience with framework-based evaluation tools (e.g., Ragas, DeepEval, TruLens) to measure hallucination, faithfulness, and relevancy.
  • Tracing & Telemetry: Proficiency with LLM-specific tracing tools (e.g., LangSmith, LangFuse, Phoenix, Arize) and open standards like OpenTelemetry.
  • Metrics & Dashboards: Hands-on experience building monitoring views in platforms like Datadog, Prometheus/Grafana, New Relic, or cloud-native suites.
  • Production Alerting & SLOs: Proven ability to define meaningful Service Level Indicators (SLIs) and SLOs, minimizing alert fatigue while maximizing system reliability.
  • Test Automation: Strong background in integrating automated test frameworks into CI/CD pipelines for continuous integration of AI features.
  • Prompt & Model Analysis: Analytical mindset to benchmark prompt variants, track regression in model behavior, and profile latency across model providers.
Programming & Operations
  • Python Mastery: Advanced Python programming skills, including experience with async execution, API integration, and AI frameworks (e.g., LangChain, LlamaIndex).
  • Production Operations: Solid understanding of cloud infrastructure (AWS/GCP/Azure), containerization (Docker, Kubernetes), and DevOps best practices.
Preferred Qualifications
  • 3+ years of experience operationalizing LLMs or generative AI applications in production.
  • Experience managing vector databases (e.g., Pinecone, Milvus, Chroma) and tracking RAG pipeline performance.
  • Background in Site Reliability Engineering (SRE) or specialized MLOps roles.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM / GenAI Engineer
LLM / GenAI Engineer

Evlo AI • New York (NY)

On-site
USD 140,000 - 200,000
Senior LLMOps Engineer
Senior LLMOps Engineer

UNAVAILABLE • McLean (VA)

On-site
USD 180,000 - 240,000
LLMOps Engineer
LLMOps Engineer

UNAVAILABLE • McLean (VA)

On-site
USD 150,000 - 210,000
AI Engineer
AI Engineer

Teserac, Inc. • Santa Clara (CA)

On-site
USD 140,000 - 200,000
Health Care Plan
Paid Time Off
Free Food & Snacks
+2
Senior AI Engineer: Scalable LLMs & MLOps Leader
Senior AI Engineer: Scalable LLMs & MLOps Leader

Compunnel, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Forward Deployed Engineers
Senior Forward Deployed Engineers

Conquer AI • United States

On-site
USD 140,000 - 220,000
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
Application Engineer – LLM
Application Engineer – LLM

Salt Digital Recruitment • United States

On-site
USD 120,000 - 180,000
Agentic AI Engineer
Agentic AI Engineer

Compunnel, Inc. • Dallas (TX)

On-site
USD 120,000 - 150,000
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000