Site Reliability Engineer

Optomi

Town of Florida (NY)

On-site

USD 145,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A leading recruiting company in New York is seeking a Senior Site Reliability Engineer to join their data platform team. This role involves building scalable infrastructure, automating processes, and ensuring system reliability using AWS. The ideal candidate has over 6 years of experience in software engineering, strong AWS knowledge, and expertise in observability practices. This full-time position offers competitive compensation within the specified range.

Qualifications

  • 6+ years of professional software engineering experience focusing on reliability, infrastructure, or platform engineering.
  • Strong programming skills in Python and at least one statically typed language.
  • Deep hands-on experience with AWS services.

Responsibilities

  • Build, deploy, and maintain scalable infrastructure for data pipelines.
  • Automate operational processes and reduce toil.
  • Monitor and improve system performance using observability tools.
  • Collaborate with data engineering and product teams.

Skills

Software engineering experience
Programming skills in Python
AWS services expertise
Operating and scaling distributed systems
Observability and telemetry design
CI/CD automation
SQL/NoSQL understanding
Agile development workflows
Incident response leadership
Communication skills

Education

Bachelor of Science

Tools

Terraform
AWS CDK
DataDog
CloudWatch

Job description

Overview

This range is provided by Optomi. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.

Base pay range

$145,000.00/yr - $160,000.00/yr

Cloud & Infrastructure Technical Recruiter @ Optomi | Bachelor of Science

Site Reliability Engineer

Optomi, in partnership with a leading global media organization are seeking a Senior Site Reliability Engineer to join their data platform team. This position operates at the intersection of DevOps, data engineering, and platform reliability, working closely with cross-functional teams to ensure the scalability, observability, and reliability of high-throughput data systems.

Requirements and skills
  • 6+ years of professional software engineering experience, with a focus on reliability, infrastructure, or platform engineering
  • Strong programming skills in Python and at least one statically typed language (e.g., Java, TypeScript, Go)
  • Deep hands-on experience with AWS services (Lambda, ECS/EKS, S3, IAM, API Gateway, SNS/SQS, Kinesis)
  • Proven experience operating and scaling distributed systems in production environments
  • Expertise in observability and telemetry design: tracing, metrics, logging
  • Proficiency in CI/CD automation, infrastructure-as-code (e.g., Terraform, AWS CDK), and DevOps best practices
  • Solid understanding of SQL/NoSQL data stores and architectural trade-offs
  • Familiarity with agile development workflows, code reviews, and collaborative SDLC processes
  • Experience leading incident response, root cause analysis, and driving continuous improvement
  • Ability to design and maintain SLAs, SLOs, and SLIs in production systems
  • Strong communication and cross-functional collaboration skills
Key responsibilities
  • Build, deploy, and maintain highly available and scalable infrastructure for data pipelines and platform services using AWS and infrastructure-as-code tools like Terraform or AWS CDK
  • Automate operational processes and reduce toil through scripting (Python, Go, etc.), CI/CD pipelines, and workflow automation
  • Monitor, analyze, and improve system performance, latency, and reliability using tools like CloudWatch, DataDog, and custom telemetry
  • Manage observability for services—design and implement SLIs, SLOs, and SLAs; maintain dashboards and alerts for distributed systems
  • Lead incident response, root cause analysis, and post-mortem reviews; continuously improve incident detection and remediation processes
  • Collaborate with data engineering and product teams to ensure reliable integration of new services and features into the data platform
  • Optimize cost and performance of cloud infrastructure, including autoscaling, provisioning strategies, and storage lifecycle policies
  • Participate in on-call rotation, ensuring timely resolution of production issues and follow-up improvements
  • Maintain compliance and security best practices within data infrastructure, including IAM, auditing, and resource governance
  • Review code, architecture, and infrastructure changes to ensure adherence to reliability and scalability standards
Seniority level
  • Mid-Senior level
Employment type
  • Full-time
Job function
  • Information Technology
Industries
  • Entertainment Providers
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Optomi • Irving (TX)

Hybrid
USD 140,000 - 150,000
Medical insurance
Vision insurance
401(k)
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Optomi • Fort Worth (TX), Arlington (TX)

Hybrid
Medical insurance
Vision insurance
Senior Director, Site Reliability and Platform Engineering
Senior Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

Hybrid
USD 150,000 - 200,000
Medical insurance
Vision insurance
System Engineer (OS, DevOps, QA)
System Engineer (OS, DevOps, QA)

Optomi • United States

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optomi • Dallas (TX)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • United States

On-site
USD 120,000 - 180,000
Sr. Cloud DevOps Engineer
Sr. Cloud DevOps Engineer

Optomi • Orlando (FL)

Hybrid
USD 110,208 - 119,851
Medical insurance
Vision insurance
Site Reliability Engineer
Site Reliability Engineer

Jobot • Akron (OH)

Remote
USD 100,000 - 150,000
Comprehensive health insurance
Vision insurance
Dental insurance
+3
Site Reliability Engineer
Site Reliability Engineer

Optomi • Charlotte (NC)

Hybrid
USD 140,000 - 170,000
Sr. Director, Site Reliability and Platform Engineering
Sr. Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

On-site
USD 150,000 - 200,000