Lead Observability Engineer – Sumo Logic

E-Solutions

United States

Remote

USD 120,000 - 150,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A leading technology solutions provider is seeking a Lead Observability Engineer to implement Sumo Logic for a client transitioning from Dynatrace. The ideal candidate will possess deep expertise in Sumo Logic and SRE practices, driving enhancements in observability through designing scalable dashboards and alerting frameworks. This remote role offers a unique opportunity to leverage advanced skills in AWS and Kubernetes to ensure service-level reliability and operational excellence.

Qualifications

  • Expert-level experience with Sumo Logic, including dashboarding and alerting.
  • Background in SRE, with a focus on SLIs/SLOs and error budgets.
  • Proficiency in AWS, specifically CloudWatch and EKS.
  • Experience with OpenTelemetry for distributed tracing.
  • Strong understanding of Kubernetes metrics and pod health.

Responsibilities

  • Lead the implementation of the Sumo Logic observability platform.
  • Migrate assets from Dynatrace to Sumo Logic.
  • Define SLIs/SLOs and reliability metrics for containerized services.
  • Deploy and configure Sumo Logic collectors.
  • Design dashboards for service health and performance.
  • Implement alerting workflows and ensure best practices.

Skills

Expert-level experience with Sumo Logic
Strong background in Site Reliability Engineering (SRE)
Proficiency in AWS services
Hands-on experience with OpenTelemetry
Strong understanding of Kubernetes metrics
Strong scripting experience with Terraform and Helm
Excellent communication skills

Job description

Overview

Role : Lead Observability Engineer – Sumo Logic & SRE

Location : Remote

We are seeking a highly skilled Lead Observability Engineer to lead a critical implementation of Sumo Logic for a client migrating from Dynatrace. This role requires deep expertise in Sumo Logic, Site Reliability Engineering (SRE) practices, and Kubernetes (EKS) observability. The ideal candidate will design and implement scalable dashboards, alerts, and tracing strategies, drive service-level reliability, and enable a steady-state SRE operations model.

Responsibilities
  • Lead the end-to-end implementation of Sumo Logic observability platform for AWS and EKS environments.
  • Migrate monitoring and alerting assets from Dynatrace to Sumo Logic.
  • Define and implement SLIs/SLOs, error budgets, and reliability metrics for containerized services.
  • Deploy and configure Sumo Logic collectors across AWS and Kubernetes workloads (EKS).
  • Configure log, metric, and trace ingestion pipelines using OpenTelemetry and Sumo Logic apps.
  • Design and maintain dashboards for service health, performance, and reliability insights.
  • Implement intelligent alerting and notification workflows, using thresholds, baselines, and anomaly detection.
  • Collaborate with DevOps, SRE, and development teams to ensure complete tracing coverage across services.
  • Ensure best practices for alert noise reduction, escalation policies, and incident response are in place.
  • Contribute to observability runbooks, operational handover, and training for the client SRE team.
Focus Areas on Sumo Logic
  • Strong knowledge of the new UI navigation.
  • Proven expertise in building and optimizing queries.
  • Advanced troubleshooting skills.
  • The ability to go beyond task execution and provide proactive recommendations to improve our setup and overall efficiency.
Required Skills & Qualifications
  • Expert-level experience with Sumo Logic, including dashboarding, alerting, collector deployment, and ML features.
  • Strong background in Site Reliability Engineering (SRE), including SLIs/SLOs, error budgets, MTTR/MTTD metrics.
  • Proficiency in AWS services (especially CloudWatch, CloudTrail, Lambda, RDS) and EKS (Amazon Kubernetes Service).
  • Hands-on experience with OpenTelemetry for distributed tracing and service maps.
  • Strong understanding of Kubernetes metrics, pod health, container resource usage, and cluster monitoring.
  • Proven ability to define alert thresholds, configure notification routing (e.g. Slack, PagerDuty, ServiceNow), and manage alert fatigue.
  • Strong scripting experience with tools like Terraform, Helm, YAML, and GitOps workflows.
  • Experience with incident triage, RCA documentation, and building operational maturity in observability teams.
  • Excellent communication and stakeholder engagement skills.
Preferred Qualifications
  • Sumo Logic certifications (Admin, Advanced Analytics) are a plus.
  • Experience with Dynatrace (for migration purposes).
  • Familiarity with integrating observability into CI/CD pipelines.
  • Exposure to service mesh (Istio/Linkerd) and monitoring microservices in that context.
Deliverables This Role Will Drive
  • Sumo Logic observability reference architecture
  • EKS and AWS observability configuration
  • SLI/SLO documentation and tracking
  • Alerting and tracing setup across services
  • Production-ready dashboards and runbooks
  • Knowledge transfer and enablement sessions for SRE/DevOps teams
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Observability Lead - Sumo Logic & SRE
Remote Observability Lead - Sumo Logic & SRE

E-Solutions • United States

Remote
USD 120,000 - 150,000
Director Platform Engineering - SRE / Observability
Director Platform Engineering - SRE / Observability

Request Technology, LLC • Chicago (IL)

On-site
USD 180,000 - 240,000
Senior SRE (Site Reliability Engineer)
Senior SRE (Site Reliability Engineer)

Vytwo • Dallas (TX)

On-site
USD 130,000 - 160,000
Flexible work from home options
Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Ontrac Solutions • New York (NY)

On-site
USD 120,000 - 190,000
Senior Observability Engineer (Splunk / SignalFx) - US Citizen or Green Card only
Senior Observability Engineer (Splunk / SignalFx) - US Citizen or Green Card only

Zohorecruit • Fairfax (VA)

On-site
USD 140,000 - 190,000
Senior Splunk & Observability Engineer
Senior Splunk & Observability Engineer

System One • Lafayette (LA)

On-site
USD 120,000 - 180,000
Senior Splunk & Observability Engineer
Senior Splunk & Observability Engineer

System One • Knoxville (TN)

On-site
USD 120,000 - 180,000
ZR_2776_JOB
ZR_2776_JOB

Zohorecruit • Orlando (FL)

On-site
USD 140,000 - 190,000
Observability Operations Engineer
Observability Operations Engineer

Tata Consultancy Services • Phoenix (AZ)

On-site
USD 100,000 - 120,000
Senior Observability Engineer
Senior Observability Engineer

Calance • Orlando (FL)

On-site
USD 80,000 - 120,000