Staff SRE - Observability

Focused

Chicago (IL)

Hybrid

USD 160,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Focused in Chicago is seeking a knowledgeable professional for a role specializing in observability and site reliability engineering. The position entails designing and implementing OpenTelemetry solutions, establishing observability infrastructure, and engaging with CI/CD systems.

The ideal candidate will have 3-7 years of experience in monitoring and distributed systems, strong knowledge of DevOps practices, and proficiency in tools like Terraform and Kubernetes. The position requires three days a week in the office and offers a base salary range of $160,000 - $200,000.

Qualifications

  • 3-7 years of experience in observability, monitoring, and distributed systems.
  • Deep hands‑on experience with OpenTelemetry ecosystem.
  • 5+ years of Platform Engineering or DevOps experience.

Responsibilities

  • Design and implement end-to-end OpenTelemetry solutions across technology stacks.
  • Establish Service Level Indicators, Objectives, Agreements.
  • Deploy and manage observability infrastructure across cloud platforms.

Skills

OpenTelemetry ecosystem experience
Platform Engineering or DevOps
CI/CD platforms proficiency
Cloud providers observability services
Container technologies knowledge
Multiple programming languages proficiency

Education

Bachelor's degree in a related field

Tools

Terraform
Kubernetes
AWS
GitHub Actions

Job description

At Focused, we move quickly to deliver quality software that achieves client outcomes and meets their customer’s needs. We strategically partner with our clients to leverage our expertise in design and software, while our clients bring their own domain expertise. We work with a variety of clients from different industries, collaborating as we get new products to market, modernizing legacy systems, or helping teams learn the skills they need to be successful.

Our Values
  • Listen first: We are experts in product practices but lifelong learners in the domain of our customers. We research, collaborate, and understand.
  • Learn why: We ask questions and talk to users to understand problem spaces, objectives, and goals, which allows us to deeply invest and drive towards the outcomes of our clients.
  • Love your craft: We love diving into a variety of domains and solving problems. We take pride in delivering value, in communicating progress, and guiding our clients to success.
Key Responsibilities
OpenTelemetry & Observability
  • Design and implement end-to-end OpenTelemetry solutions across diverse technology stacks
  • Configure and deploy OpenTelemetry Collectors for efficient data collection, processing, sampling, and routing
  • Establish telemetry pipelines for metrics, traces, and logs across microservices architectures
  • Optimize collector configurations for performance, reliability, and cost-effectiveness
  • Augment existing infrastructure with integrated observability solutions
  • Implement Infrastructure as Code (IaC) solutions using Terraform, Pulumi, CloudFormation, etc.
  • Architect and manage Kubernetes clusters with comprehensive monitoring and logging
  • Build CI/CD pipelines with embedded observability and automated testing
Site Reliability Engineering (SRE)
  • Establish and maintain Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs)
  • Implement error budgets, toil reduction strategies, and capacity planning
  • Support incident response procedures and post‑mortem processes
  • Deploy and manage observability infrastructure across AWS, GCP, and Azure
  • Establish security, compliance, and governance frameworks for telemetry data
  • Experience automating Agent Evaluations in CI/CD pipelines and observability backends
Required Qualifications
Core Observability & OpenTelemetry
  • 3-7 years of experience in observability, monitoring, and distributed systems
  • Deep hands‑on experience with OpenTelemetry ecosystem, including SDKs, APIs, and specifications
  • Proficiency with OpenTelemetry Collector configuration, processors, exporters, and receivers
  • Strong understanding of telemetry data models, semantic conventions, and instrumentation best practices
  • 5+ years of Platform Engineering or DevOps experience with focus on site reliability, observability, and incident response
  • Proficiency with Infrastructure as Code tools (Terraform, Pulumi, CloudFormation, CDK)
  • Strong experience with CI/CD platforms (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
  • Hands‑on experience with major cloud providers (AWS, GCP, Azure) and their observability services
  • Experience with container technologies (Docker, Podman) and container registries
  • Knowledge of networking, security, load balancing, and distributed systems concepts
Site Reliability Engineering
  • Experience implementing SRE practices including error budgets and toil metrics
  • Proficiency in incident management, on‑call procedures, and post‑mortem culture
  • Experience with capacity planning, performance optimization, and scalability design
Programming & Automation
  • Proficiency in multiple programming languages preferred (Go, Python, Java, Node.js, Rust)
  • Strong scripting and automation skills (Bash, Python, PowerShell)
  • Understanding of software engineering best practices and testing methodologies
Preferred Qualifications (Exceptional Candidates)
AI & Agentic Frameworks
  • Understanding of Large Language Models (LLMs) and their application in DevOps
  • Knowledge of vector databases, embeddings, and retrieval‑augmented generation (RAG)
  • Experience with AI/ML model deployment and monitoring in production environments
Leadership & Communication
  • Strong technical writing and documentation skills
  • Ability to present complex technical concepts to diverse stakeholders
  • A passion for knowledge sharing
  • Systems thinking and ability to design holistic observability solutions
  • Strong analytical and troubleshooting skills for complex distributed systems
  • Curiosity about emerging technologies, particularly AI applications in operations
  • Adaptability to rapidly evolving cloud‑native and observability technologies
  • Collaborative mindset with focus on enabling developer productivity and system reliability
What Sets Exceptional Candidates Apart
  • Experience with Honeycomb
  • Contributions to open‑source observability or AI framework projects
  • Track record of implementing platform engineering solutions that significantly improved developer experience
  • Experience scaling observability infrastructure to handle high event volume
What to Know Before You Apply
  • This role will require being in the Chicago office three days per week and up to 20% travel within the United States.
  • Focused is unable to sponsor or take over sponsorship of the employment Visa process at this time.
  • The Chicago base salary range for this role is $160,000 - $200,000.

As set forth in Focused’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Observability
Staff Software Engineer, Observability

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Senior Observability & SRE Engineer – OpenTelemetry
Senior Observability & SRE Engineer – OpenTelemetry

Focused • Chicago (IL)

Hybrid
USD 160,000 - 200,000
Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000
SRE/Observability Engineer
SRE/Observability Engineer

BlueSky Resource Solutions • United States

Remote
USD 100,000 - 130,000
Site Reliability Engineer - Observability
Site Reliability Engineer - Observability

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 180,000
Hybrid work model
Senior Site Reliability Engineer, Observability New York, NY, United States
Senior Site Reliability Engineer, Observability New York, NY, United States

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Observability Engineer Site Reliability Engineer
Observability Engineer Site Reliability Engineer

Ontrac Solutions • Arizona

Hybrid
USD 140,000 - 190,000
Verification cost reimbursement
Observability Engineer Site Reliability Engineer
Observability Engineer Site Reliability Engineer

Ontrac Solutions • United States

On-site
USD 120,000 - 180,000
Software Engineer, Observability
Software Engineer, Observability

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 175,000
Equity
Healthcare
Mentorship & events
+2
Observability Engineer
Observability Engineer

BCforward • Phoenix (AZ)

Hybrid
USD 120,000 - 140,000