Senior Site Reliability Engineer

Omilia

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Omilia is seeking a Senior Site Reliability Engineer to own the reliability of production and pre-production environments, build observability solutions, and drive automation across cloud-based platforms.

You will monitor, respond to incidents, contribute to runbooks, collaborate with software and cloud teams, and help embed a culture of reliability, performance, and continuous improvement.

Qualifications

  • Bachelor’s degree or MS in Engineering or equivalent.
  • Experience operating container orchestration clusters (Kubernetes, Docker Swarm).
  • Production services at scale experience.
  • Experience with ELK, AWS, Grafana/Prometheus stack.
  • Strong scripting skills (Bash, Python or Go).
  • Excellent communication and teamwork.

Responsibilities

  • Ensure platform reliability and availability across production environments through monitoring and automation.
  • First incident response, contribute to problem management and root cause analysis.
  • Support development teams to embed reliability into the delivery lifecycle.
  • Develop runbooks and operational documentation; automate tasks.
  • Collaborate to design observability solutions with metrics, logs, and dashboards.
  • Participate in on-call rotations and improve alert quality and processes.
  • Champion reliability, performance, and continuous improvement across teams.

Skills

Cloud experience
Kubernetes
Observability
Automation
Incident response

Education

Bachelor's or MS in Engineering or equivalent

Tools

Prometheus
Grafana
ELK
AWS
Python

Job description

We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they will collaborate with team members to develop automation strategies, monitoring & alerting, and ensuring overall platform reliability. Your goal will be to become an integral part of the team, making every challenge of the platform – your own challenge, and solving them accordingly.

Responsibilities
  • Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.

  • First response for incidents, contribute to problem management and root cause analysis.

  • Supporting the development team’s effort towards reliability, creating a solid reliability culture within the development lifecycle.

  • Develop troubleshooting documentation for production support resources.

  • Collaborate with Engineering teams to develop optimised and productive runbooks, operational documentation and automation of operational tasks.

  • Collaborate with development and cloud engineering teams to embed reliability and performance into the software delivery lifecycle.

  • Design, implement, and evolve observability solutions (metrics, logs, traces, dashboards) using tools such as Prometheus, Grafana, and ELK.

  • Participate in on-call rotations and continuously improve alert quality and response processes.

  • Champion a culture of reliability, performance, and continuous improvement across teams.

  • Bachelor’s Degree or MS in Engineering or equivalent.

  • Experience in operating at least one container orchestration cluster (Kubernetes, Docker Swarm).

  • Experience developing or maintaining software for production services at scale.

  • Experience with ELK.

  • Experience with AWS.

  • Experience with Grafana/Prometheus stack.

  • Strong scripting skills (Bash, Python or Go).

  • Excellent communication skills.

  • Thinking out of the box and anticipating challenges. It is imperative we are not simply reactive; we must expect challenges and question technologies, procedures and thinking already in place. You will be expected to constantly review and challenge at all levels.

  • Versatility. We work with agile/lean methods. We’d much rather iterate and learn than assume we know all the answers.

  • Being a team player. You don’t (always) work in isolation and are excited by the thought of using your team whilst involving product, experience design, engineering, and more in the process.

Will be considered as a plus:
  • Telephony knowledge (SIP, VoIP);

  • Experience in Linux Administration (RedHat, CentOS, AL);

  • Working knowledge in Configuration Management tools (Terraform, Ansible);

  • Experience with TCP/IP and general networking concepts;

  • RDBMS knowledge (MySQL, Postgres);

  • NoSQL knowledge (Redis).

  • Fixed compensation;

  • Long-term employment with the working days vacation;

  • Development in professional growth (courses, training, etc);

  • Being part of successful cutting-edge technology products that are making a global impact in the service industry;

  • Proficient and fun-to-work-with colleagues;

  • Apple gear.

Omilia is proud to be an equal opportunity employer and is dedicated to fostering a diverse and inclusive workplace. We believe that embracing diversity in all its forms enriches our workplace and drives our collective success. We are committed to creating an environment where everyone feels welcomed, valued, and empowered to contribute their unique perspectives without regard to factors such as race, color, religion, gender, gender identity or expression, sexual orientation, national origin, heredity, disability, age, or veteran status, all eligible candidates will be given consideration for employment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Solution Consultant, Iberia
Solution Consultant, Iberia

Omilia • United States

Remote
USD 120,000 - 190,000
Apple gear
Staff DevOps Engineer
Staff DevOps Engineer

Omnissa • California (MO)

Hybrid
USD 206,000 - 343,000
Health insurance
401k with matching
Paid time off
+2
Conversational AI Application Developer (English and French Speaker)
Conversational AI Application Developer (English and French Speaker)

United States Digital Space LLC • United States

Remote
USD 110,000 - 150,000
Apple gear
Professional development
Generous vacation
Enterprise Account Executive
Enterprise Account Executive

Omilia • Northern (KY)

On-site
USD 120,000 - 180,000
Fixed compensation
Vacation days
Professional growth opportunities
+2
Enterprise Account Executive
Enterprise Account Executive

Omilia • United States

On-site
USD 140,000 - 190,000
Apple gear
Career development opportunities
Solution Consultant Director, EMEA
Solution Consultant Director, EMEA

Omilia • United States

Remote
USD 140,000 - 190,000
Apple gear
DevOps Lead, Cloud Automation & Reliability
DevOps Lead, Cloud Automation & Reliability

Omnissa • Atlanta (GA)

On-site
USD 134,000 - 280,000
Employee ownership
Health insurance
401k with matching
+3
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
DevOps Lead, Cloud Automation & Reliability
DevOps Lead, Cloud Automation & Reliability

Omnissa, LLC in • Atlanta (GA)

On-site
USD 134,000 - 280,000
Employee ownership
Health insurance
401k with matching
+4
Principal Reliability Engineer
Principal Reliability Engineer

GoToFoods • Augusta (GA)

Hybrid
USD 150,000 - 200,000