Observability Engineer

Evolutyz Corp

Brea (CA)

On-site

USD 150,000 - 190,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evolutyz Corp is seeking a Senior Observability Engineer to design, implement, and optimize modern observability across cloud-native environments. You will work on IaC with Terraform, backend APIs in Python/FastAPI, and observability via OpenTelemetry and Datadog.

The role requires deep experience in cloud infrastructure, distributed systems, and CI/CD, with a focus on reliability and proactive monitoring for critical applications and services.

Qualifications

  • Degree in computer science, engineering, or related field or equivalent practical experience.
  • 10 years of experience in cloud infrastructure, software engineering, or observability engineering.
  • Strong hands-on experience with Terraform and Infrastructure as Code practices.
  • Proficiency in Python and backend API development using FastAPI.
  • Experience implementing observability frameworks using OpenTelemetry.
  • Strong experience with Datadog, including dashboards, monitors, alerting, and SLO management.
  • Solid understanding of distributed systems, cloud-native architectures, and microservices.
  • Experience with CI/CD tools and deployment automation.
  • Strong troubleshooting, analytical, and problem-solving skills.

Responsibilities

  • IaC: design, build, and manage scalable cloud infrastructure using Terraform.
  • Develop reusable infrastructure modules and provisioning standards.
  • Implement and maintain CI/CD pipelines to automate deployments.
  • Improve deployment reliability through automation and best practices.
  • Backend: design, develop, and maintain RESTful APIs using Python and FastAPI.
  • Build services that support observability workflows, telemetry processing, and integrations.
  • Optimize services for performance, scalability, reliability, and maintainability.
  • Observability: design and implement solutions using OpenTelemetry for tracing, metrics, logging.
  • Configure and maintain Datadog dashboards, monitors, alerts, and SLOs.
  • Develop monitoring strategies to improve visibility and reduce incidents.
  • Analyze telemetry to identify performance bottlenecks and reliability issues.
  • Establish alerting thresholds and monitoring standards for production ops.

Skills

Terraform
Python
FastAPI
OpenTelemetry
Datadog
CI/CD
Distributed Systems

Education

Bachelor's degree in CS

Tools

Kubernetes
Prometheus
Grafana
Splunk

Job description

Summary

We are seeking a highly skilled Senior Observability Engineer to design, implement, and optimize modern observability solutions across cloud-native environments. The ideal candidate will have strong expertise in infrastructure automation, backend development, and monitoring platforms, with hands-on experience in Terraform, Python, FastAPI, OpenTelemetry, and Datadog. This role will focus on building scalable observability frameworks, improving system reliability, and enabling proactive monitoring for critical applications and infrastructure.

Responsibilities
  • Infrastructure as Code (IaC):
    • Design, build, and manage scalable cloud infrastructure using Terraform.
    • Develop reusable infrastructure modules and maintain infrastructure provisioning standards.
    • Implement and maintain CI/CD pipelines to automate infrastructure and application deployments.
    • Improve deployment reliability through automation and infrastructure best practices.
  • Backend Development:
    • Design, develop, and maintain RESTful APIs using Python and FastAPI.
    • Build backend services that support observability workflows, telemetry processing, and integrations.
    • Optimize services for performance, scalability, reliability, and maintainability.
    • Troubleshoot application and API performance issues.
  • Observability & Monitoring:
    • Design and implement observability solutions using OpenTelemetry for distributed tracing, metrics, and logging.
    • Configure and maintain Datadog dashboards, monitors, alerts, and Service Level Objectives (SLOs).
    • Develop monitoring strategies to improve system visibility and reduce incident response times.
    • Analyze telemetry data to identify performance bottlenecks and reliability issues.
    • Establish alerting thresholds and monitoring standards to support production operations.
Requirements
  • Degree in Computer Science, Engineering, or related field (or equivalent practical experience).
  • 10 years of experience in cloud infrastructure, software engineering, or observability engineering.
  • Strong hands-on experience with Terraform and Infrastructure as Code practices.
  • Proficiency in Python and backend API development using FastAPI.
  • Experience implementing observability frameworks using OpenTelemetry.
  • Strong experience with Datadog, including dashboards, monitors, alerting, and SLO management.
  • Solid understanding of distributed systems, cloud-native architectures, and microservices.
  • Experience with CI/CD tools and deployment automation.
  • Strong troubleshooting, analytical, and problem-solving skills.
Preferred Skills
  • Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud.
  • Knowledge of container orchestration platforms such as Kubernetes.
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, or Splunk.
  • Experience supporting high-availability production environments.
Benefits

Details about benefits are not provided in the draft. Please specify benefits if applicable.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Architect
Observability Architect

TechDigital Group • Atlanta (GA)

On-site
USD 120,000 - 150,000
Senior Observability Engineer: Build Scalable Telemetry
Senior Observability Engineer: Build Scalable Telemetry

Evolutyz Corp • Brea (CA)

On-site
USD 150,000 - 190,000
Observability Engineer: OpenTelemetry & Open-Source
Observability Engineer: OpenTelemetry & Open-Source

Softility Tech Pvt. Ltd. • United States

Hybrid
Lead DevOps Engineer
Lead DevOps Engineer

Trekrecruit • Richmond (VA)

On-site
USD 90,000 - 120,000
Infrastructure Engineer, Observability
Infrastructure Engineer, Observability

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000
Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000
Sr Observability Engineer
Sr Observability Engineer

IT Associates • Irvine (CA)

Hybrid
USD 150,000 - 210,000
Observability Engineer
Observability Engineer

ManpowerGroup Global, Inc. • Denver (CO), Town of Norway (WI)

On-site
USD 117,000 - 143,000
Medical and Prescription Drug Plans
Dental Plan
Vision Plan
+3
Senior Observability Engineer – NS2JP00000386
Senior Observability Engineer – NS2JP00000386

Prestige Staffing • Oak Hill (WV)

Remote
USD 100,000 - 130,000
Software Engineer (Observability & Monitoring) - West Des Moines, IA
Software Engineer (Observability & Monitoring) - West Des Moines, IA

AHU Technologies Inc • Washington

On-site
USD 120,000 - 150,000