Senior Site Reliability Engineer

Unifonic

United States

Remote

USD 140,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Unifonic seeks a Senior Infrastructure Engineer for the Production Operations (Live) team to enhance reliability and scale across AWS, OCI, and OpenStack environments. You will own production services uptime, participate in on-call rotations, and drive improvements in MTTD/MTTR while advancing observability and automation.

You will work with Kafka, Redis, MySQL/PostgreSQL, and containerized services, mentoring juniors and shaping disaster recovery strategies.

Qualifications

  • 8+ years of hands-on production experience in SRE/DevOps or cloud engineering.

Responsibilities

  • Own the reliability, uptime, and scalability of critical production services 24/7.
  • Participate in on-call rotation, respond to incidents, troubleshoot live issues, and lead post-incident analysis.
  • Build robust runbooks, escalation paths, and improve MTTD/MTTR.
  • Ensure observability through SLO/SLI definitions, monitoring and capacity planning.
  • Automate operational tasks using Terraform, Helm, Jenkins, Tekton, or GitLab CI/CD.
  • Manage Kubernetes clusters (EKS, OKE, Rancher RKE2) and production databases.
  • Collaborate with cross-functional teams to strengthen SRE culture and security.

Skills

SRE experience
AWS
OCI
OpenStack
Kubernetes
Kafka
RabbitMQ
Redis
MySQL
PostgreSQL
Python
Go
Terraform
Helm
CI/CD
Linux
On-call

Education

Bachelor's degree in CS/Engineering
Master's degree (preferred)

Tools

Docker
Jenkins
Tekton
GitLab CI/CD
Prometheus
Grafana
CloudWatch

Job description

Proudly voted a Great Place to Work®, we are a dynamic startup in the CPaaS (Communication Platform as a Service) space that is revolutionizing the way businesses communicate. Our team is made up of 500 energetic and passionate Unifones who are dedicated to delivering the best possible experience to 5000+ customer-centric companies. We pride ourselves on our fun and collaborative work environment, where creativity and new ideas are constantly encouraged. As shareholders in the business, we’re so much more than a group of passionate communicators. We are Unifones. Join our team and be a part of something big! Meet the team! Our Engineering team is responsible for designing, developing, and maintaining the systems and technologies that drive Unifonic ’s solutions. We work closely with other departments to ensure our products and services meet the needs of our customers. If you are passionate about technology and are excited about working on cutting-edge communication and engagement solutions, we want you on our team. As a Senior Infrastructure Engineer in the Production Operations (Live) team you will be responsible for enhancing system reliability, scalability, and resilience. As part of our elite SRE team, you'll drive continuous improvement across our cloud infrastructure and ensure the consistent high performance of our distributed messaging platforms. Help us shape the future of communication by:

  • Owning the reliability, uptime, and scalability of critical production services 24/7.
  • Participating in the on-call rotation to respond to incidents, troubleshoot live production issues, and lead post-incident analysis 24/7.
  • Building robust operational playbooks, escalation paths, and improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Ensuring operational excellence by proactively detecting and addressing reliability risks through SLO monitoring, chaos testing, and capacity planning.
  • Automating operational tasks to minimize human intervention.
  • Being available at night during the usual non-working hours of the rest of the team according to the on-call schedule is a MUST.
  • Architecting, implementing, and managing infrastructure across AWS, Oracle Cloud Infrastructure (OCI), and OpenStack environments.
  • Optimizing cloud resources to balance performance, security, and cost-efficiency.
  • Managing Kubernetes clusters (EKS, OKE, Rancher RKE2), ensuring scalability, availability, and robust performance.
  • Deploying advanced containerization strategies and troubleshooting.
  • Managing and optimizing high-performance messaging and caching systems including Kafka, RabbitMQ, and Redis.
  • Ensuring efficient, reliable message and data delivery critical to Unifonic 's SMS and distributed systems.
  • Managing and optimizing production-grade MySQL and PostgreSQL databases.
  • Ensuring high availability, performance tuning, backups, and recovery processes for critical databases.
  • Leading the planning and execution of comprehensive disaster recovery strategies.
  • Developing and maintaining robust business continuity plans.
  • Implementing advanced observability solutions (Prometheus, Grafana, CloudWatch).
  • Defining, measuring, and enforcing Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in alignment with SRE best practices.
  • Proactively identifying issues, minimizing downtime, and enhancing system transparency.
  • Driving automation initiatives using Terraform, Helm, Jenkins, Tekton or GitLab CI/CD.
  • Streamlining deployment pipelines and reduce manual intervention through innovative automation.
  • Integrating security best practices into infrastructure and application layers.
  • Performing regular audits ensuring compliance and robust security posture.
  • Collaborating with cross-functional teams (engineering, product, QA) to foster SRE culture.
  • Mentoring junior engineers, enhancing team capabilities and promoting knowledge sharing.
Requirements

What you'll bring:

  • Bachelor's or master's degree in computer science, Engineering, or a related technical field.
  • 8+ years of hands‑on production experience in SRE, DevOps, or cloud engineering roles.
  • Strong expertise in AWS, OCI, OpenStack environments.
  • Deep understanding of Kubernetes ecosystems (EKS, OKE, Rancher RKE2).
  • Proven experience with Kafka, RabbitMQ, Redis, and distributed messaging and caching systems.
  • Solid experience managing MySQL and PostgreSQL in production environments.
  • Expert-level scripting and automation skills (Python, Bash, Go).
  • Advanced proficiency with Helm, Terraform, and modern CI/CD toolchains.
  • Demonstrable experience with Linux system administration and troubleshooting.
  • Being available at night during the usual non-working hours of the rest of the team according to the on‑call schedule is a MUST.
Craft and Toolkit
  • Distributed Systems & Architecture — Scalability, fault tolerance, consistency models, microservices.
  • Cloud Platforms — Hands-on with Amazon Web Services, Google Cloud Platform, or Microsoft Azure.
  • Infrastructure as Code — Terraform, AWS CloudFormation.
  • Containers & Orchestration — Docker, Kubernetes.
  • CI/CD & Automation — Jenkins, GitHub Actions, GitLab CI.
  • Observability — Prometheus, Grafana, ELK Stack.
  • Incident & Reliability — RCA, postmortems, SLIs/SLOs, MTTR reduction.
Character Traits
  • Analytical thinking and problem-solving – approach problems clearly, use data, and find solutions.
  • Ownership and accountability – take responsibility and follow projects through to completion.
  • Communication – explain ideas clearly and listen to others across teams.
  • Collaboration – work well with others and support team goals.
  • Adaptability and learning – adjust to change and keep learning new skills or tools.
  • Mentorship and knowledge sharing – help others grow and share what you know.
  • Resilience – stay calm under pressure and handle setbacks constructively.
  • Quality and attention to detail – do work carefully and strive for improvement.
  • Advocacy and innovation – encourage best practices, efficiency, and new ideas.
  • AI mindset and utilization – Leve
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

Unifonic • United States

Remote
USD 120,000 - 180,000
Competitive salary and bonus
Share scheme (we are all owners!)
30 holiday days after first year
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Senior IT Reliability & Automation Lead
Senior IT Reliability & Automation Lead

First Horizon Bank • Memphis (TN)

On-site
USD 120,000 - 180,000
Product Operations Lead
Product Operations Lead

Unifonic, Inc. • United States

On-site
USD 120,000 - 160,000
Competitive salary and bonus
Unifonic share scheme
30 holiday days after first year
+1
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

DriveWealth • Chicago (IL)

Hybrid
USD 160,000 - 230,000
Health & wellness packages
Unlimited vacation
Remote/hybrid options
+1
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • United States

On-site
USD 150,000 - 210,000