Sr Site Reliability Engineer

Realtor.com

Austin (TX)

On-site

USD 140,000 - 180,000

Full time

25 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical coverage
Dental & vision
401(k) plan with match
Flexible time off
Free snacks

Job summary

Realtor.com is seeking a Senior Site Reliability Engineer to join the newly formed Operations Excellence organization. You will contribute to the reliability and observability of platform infrastructure serving millions of users, including EKS, Skyway, Frontdoor, and Apollo GraphQL integrations.

You will implement best practices, drive chaos engineering, and collaborate with cross‑functional teams while supporting cost optimization and security initiatives.

Qualifications

  • 8+ years in Site Reliability Engineering, DevOps, or Infra Engineering
  • Bachelor’s degree or equivalent experience
  • 3+ years hands‑on AWS (EKS, EC2, RDS, S3, CloudWatch, IAM) and Kubernetes
  • Proficient in Python/Go/Java with IaC experience (Terraform, CloudFormation)
  • Production experience with observability tools (NewRelic, Datadog, Prometheus) and distributed systems
  • Experience with CI/CD platforms and GitOps workflows (CircleCI, Argo CD, Jenkins)

Responsibilities

  • Implement and maintain highly available AWS infrastructure (EKS, Fargate, multi-region)
  • Ensure reliability of critical services (Skyway CI/CD, Frontdoor, GraphQL)
  • Monitor SLOs/SLIs and participate in architectural reviews for reliability and cost efficiency
  • Build and maintain observability dashboards and error budgets
  • Participate in on‑call rotations and incident reviews
  • Contribute to cost optimization and FinOps initiatives

Skills

SRE expertise
AWS
Kubernetes
Python/Go/Java
Observability tools
CI/CD & GitOps

Education

Bachelor's degree

Tools

AWS EKS
Terraform
CloudFormation
Grafana/Prometheus
New Relic

Job description

Description

Recognized as the No. 1 site trusted by real estate professionals, Realtor.com® has been at the forefront of online real estate for over 25 years, connecting buyers, sellers, and renters with trusted insights and expert guidance to find their perfect home. Through its robust suite of tools, Realtor.com® not only makes a significant impact on the real estate industry at large, but for consumers, navigating the biggest purchase they will make in their life, by providing a user experience that is easy to use, easy to understand, and most of all, easy to make decisions.

Description

Recognized as the No. 1 site trusted by real estate professionals, Realtor.com® has been at the forefront of online real estate for over 25 years, connecting buyers, sellers, and renters with trusted insights and expert guidance to find their perfect home. Through its robust suite of tools, Realtor.com® not only makes a significant impact on the real estate industry at large, but for consumers, navigating the biggest purchase they will make in their life, by providing a user experience that is easy to use, easy to understand, and most of all, easy to make decisions. Join us on our mission to empower more people to find their way home by breaking barriers to entry, making the right connections, and building confidence through expert guidance. We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization, reporting to the Director, Operations Excellence. This role will contribute to the reliability, observability, and operational excellence of our platform infrastructure serving millions of users. As a Senior SRE, you will be a strong technical contributor who implements best practices, solves complex problems, and enables our 600+ engineers to deliver exceptional customer experiences. You will work on critical platform systems including EKS infrastructure, Skyway (CI/CD), Frontdoor (Tyk API Gateway), Pantheon (Apollo GraphQL Federation), and our observability stack, while contributing to chaos engineering practices and cost optimization initiatives with measurable ROI.

What You’ll Do
  • Implement and maintain highly available AWS infrastructure including EKS clusters, Fargate (ECS), and multi-region architectures
  • Support reliability of critical services: Skyway (CI/CD), Frontdoor (Tyk), Pantheon (Apollo GraphQL), and supporting infrastructure
  • Monitor SLIs, SLOs, and error budgets for Tier 1/2/3 systems; participate in architectural reviews for reliability and cost-efficiency
  • Implement reliability patterns including circuit breakers, graceful degradation, and automated failover
Observability & Cost Optimization
  • Implement observability solutions using NewRelic for APM, distributed tracing, metrics, and logging for rapid troubleshooting
  • Build dashboards and alerts that reduce MTTD and MTTR; contribute to observability standards across teams
  • Identify infrastructure cost optimization opportunities and implement FinOps practices including rightsizing and resource lifecycle management
  • Support cost-conscious architecture decisions and CI/CD spend optimization (CircleCI, Argo CD)
Chaos Engineering & Incident Response
  • Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing
  • Participate in game day exercises and disaster recovery simulations; create runbooks and automation for resilience
  • Participate in on-call rotation for critical systems; conduct post-incident reviews and implement improvements
  • Support incident response processes and contribute to System Health Scorecard
Technical Contribution
  • Contribute as a strong technical individual contributor to the Operations Excellence team
  • Collaborate with Platform Engineering, Quality Engineering, and product teams on reliability initiatives
  • Support security initiatives including AWS Secrets Manager migration and compliance requirements (SOC 2, PCI, GDPR)
  • Contribute to Developer Experience metrics and platform adoption goals
  • May provide technical guidance to junior team members
What You’ll Bring
  • 8+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering with demonstrated success improving system reliability
  • Bachelor’s degree or equivalent experience
  • 3+ years hands‑on experience with AWS (EKS, EC2, RDS, S3, CloudWatch, IAM) and Kubernetes including cluster management
  • Proficient programming skills (Python, Go, or Java) with infrastructure automation and Infrastructure as Code experience (Terraform, CloudFormation)
  • Production experience with observability tools (NewRelic, Datadog, Prometheus, Grafana, Splunk) and distributed systems
  • Experience with CI/CD platforms and GitOps workflows (CircleCI, Argo CD, Jenkins); on‑call rotation and incident response
  • Preferred: Exposure to chaos engineering tools, API Gateway technologies (Tyk/Kong), GraphQL federation (Apollo), cost optimization initiatives, FinOps principles
  • Technical Skills
  • Cloud & Infrastructure: AWS (EKS, Fargate, Lambda, VPC, Route53, CloudFront), Kubernetes, Docker, Istio Service Mesh
  • CI/CD & GitOps: Argo CD, CircleCI, Jenkins, GitHub Actions
  • Observability: NewRelic - APM, distributed tracing, metrics & logging; Splunk - logging
  • IaC & Automation: Terraform, CloudFormation, Helm, Kustomize, Python/Go/Bash
  • Platform Services: Tyk Gateway, Apollo GraphQL, AWS Secrets Manager, Vault
  • Incident Management: OpsGenie, PagerDuty, ServiceNow
  • Professional Qualities
  • Strong communication skills with ability to explain technical concepts to diverse audiences
  • Collaborative approach working across engineering, product, and business teams
  • Self‑motivated with ability to solve complex problems within established practices and policies
  • Data‑driven decision making with customer‑centric approach and empathy for developer experience
How We Work

We balance creativity and innovation on a foundation of in‑person collaboration. For most roles, our employees work four or more days in our offices, where they have the opportunity to collaborate in‑person, adding richness to our culture and knitting us closer together.

How We Reward You
  • Inclusive and Competitive medical, Rx, dental, and vision coverage
  • Family forming benefits
  • 13 Paid Holidays
  • Flexible Time Off
  • 8 hours of paid Volunteer Time off
  • Immediate eligibility into Company 401(k) plan with 3.5% company match
  • Tuition Reimbursement program for degreed and non‑degreed programs
  • 1:1 personalized Financial Planning Sessions
  • Student Debt Retirement Savings Match program
  • Free snacks and refreshments in each office location

Do the best work of your life at Realtor.com®

Here, you’ll partner with a diverse team of experts as you use leading‑edge tech to empower everyone to meet a crucial goal: finding their way home. And you’ll find your way home too. At Realtor.com®, you’ll bring your full self to work as you innovate with speed, serve our consumers, and champion your teammates. In return, we’ll provide you with a warm, welcoming, and inclusive culture; intellectual challenges; and the development opportunities you need to grow.

Diversity is important to us, therefore, Realtor.com® is an Equal Opportunity Employer regardless of age, color, national origin, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, marital status, status as a disabled veteran and/or veteran of the Vietnam Era or any other characteristic protected by federal, state or local law. In addition, Realtor.com® will provide reasonable accommodations for otherwise qualified disabled individuals.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Staff - DevOps
Sr. Staff - DevOps

News Corporation • Austin (TX)

On-site
USD 180,000 - 230,000
Staff Engineer, Cloud
Staff Engineer, Cloud

Realtor.com • Austin (TX)

On-site
USD 170,000 - 260,000
Medical, Rx, dental, vision
401(k) with company match
Tuition Reimbursement
+3
Software Eng, Sr Staff - DevOps
Software Eng, Sr Staff - DevOps

News Corporation • Austin (TX)

On-site
USD 120,000 - 150,000
Sr DevOps Engineer
Sr DevOps Engineer

News Corporation • Austin (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Medical, Rx, dental, and vision
Flexible Time Off
13 Paid Holidays
+3
Senior Software Engineer
Senior Software Engineer

News Corporation • Austin (TX)

On-site
USD 120,000 - 180,000
Medical coverage
401(k) match
Flexible time off
+4
Full Stack Software Engineer
Full Stack Software Engineer

Realtor.com • Austin (TX)

On-site
USD 120,000 - 180,000
Health coverage
Family benefits
Paid holidays
+5
Senior Manager of Microservices
Senior Manager of Microservices

Realtor.com • Austin (TX)

On-site
USD 130,000 - 160,000
Inclusive medical, Rx, dental, and vision coverage
Flexible Time Off
Tuition Reimbursement
Full Stack Software Engineer
Full Stack Software Engineer

News Corporation • Austin (TX), Northern (KY)

Hybrid
USD 110,000 - 160,000
Health benefits
401(k) with company match
Tuition reimbursement
+3
Staff Software Engineer, Backend
Staff Software Engineer, Backend

Realtor.com • Austin (TX)

Hybrid
USD 120,000 - 150,000
Inclusive and Competitive medical, Rx, dental, and vision coverage
Flexible Time Off
Tuition Reimbursement program
Software Engineer
Software Engineer

Realtor.com • Austin (TX)

On-site
USD 95,000 - 140,000
401(k) plan with company match
Tuition Reimbursement
13 Paid Holidays
+3