Senior Site Reliability Engineer

N-iX

United States

Hybrid

USD 140,000 - 170,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Remote or hybrid options
Competitive salary
Education reimbursement

Job summary

N-iX is seeking a Senior Site Reliability Engineer to join our international team, delivering reliable, scalable systems for an innovative hospitality company. You’ll work across Azure and traditional data centers, driving automation, monitoring, and resilience while collaborating with a cross-functional group of engineers.

You will apply strong Kubernetes and cloud experience, CI/CD discipline, and software development skills to ensure uptime, performance, and rapid improvement across

Qualifications

  • Proven experience in SRE with Azure and multi-cloud is preferred.
  • Strong Kubernetes and cloud-native architecture background.
  • Experience with CI/CD and release orchestration.
  • Unix/Linux administration knowledge is a plus.
  • Familiar with monitoring using Azure Monitor and Application Insights.
  • DevOps mindset with automation and reliability focus.

Responsibilities

  • Develop and improve the whole lifecycle of services.
  • Establish and improve monitoring to reduce outages.
  • Automate systems to improve reliability and velocity.
  • Lead designs of major components to enhance availability and latency.
  • Perform post-incident analysis with a focus on continuous improvement.

Skills

Kubernetes
Azure Cloud
CI/CD
Unix/Linux
Distributed systems
Programming: Python/Go/Java
Service Mesh

Education

Bachelor's in CS

Tools

Terraform
Chef
Azure Monitor
Application Insights

Job description

We are looking for a Senior Site Reliability Engineer who is interested in an opportunity to work for an innovative hospitality company with cutting edge technologies, with new development activities and challenges ahead. Our international team members share a common desire to develop brilliant products on reliable and resilient systems, along with their own skills. We run our services in Azure and traditional data centers. Take a chance to make a valuable contribution and enhance your professional skills. About the job: Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault tolerant systems. SRE ensures that cloud services - both our internally critical and our externally visible systems - have reliability, uptime appropriate to customer needs, and a fast rate of improvement. Additionally, SREs will keep an ever watchful eye on our systems' capacity and performance. On the SRE team, you'll have the opportunity to manage the complex challenges of scale that are unique to the project while using your expertise in coding, algorithms, complexity analysis, and large-scale system design. You will provide scalable, reliable, durable, and secure services using a customer-first approach while innovating technically. You will understand our customer needs and how we can meet them.

Responsibilities
  • Develop and improve the whole lifecycle of services
  • Establish and improve monitoring capabilities to reduce outage frequency and duration
  • Create sustainable systems through automation and uplifts
  • Develop and scale systems sustainably through mechanisms such as automation, and evolve systems by pushing for changes that improve reliability and velocity
  • Lead designs of major software components, systems, and features to improve the availability, scalability, latency, and efficiency of our services
  • Analyze and support services before they go live via system design consulting, developing software platforms and frameworks, capacity planning
  • Conduct post-incident analysis and reviews with an attitude of continuous improvement
Requirements
  • Ideally, strong experience in Azure Services and capabilities, but other cloud services (AWS, Google Cloud Platform etc.) will be considered
  • Confidence and strong experience with Kubernetes
  • Recent and fluent Terraform and (Chef platform experience nice to have)
  • Extensive expertise in software development/testing, development operations, and site reliability engineering
  • Experience of Unix/Linux administration - an appreciation of systems internals (e.g., filesystems, system calls) is a bonus
  • Experience with Continuous Integration and Deployment (CI/CD) and release orchestration and Configuration Management of VMs
  • Cloud-agnostic approach, with flexibility to work across various cloud platforms
  • Experience programming in one or more of the following languages: C#,, C++, Java, Python, JavaScript, Go, Perl, or Ruby
  • Nice to have: Bachelor's degree in Computer Science, similar technical field of study, or equivalent practical experience
  • Experience in distributed systems, storage systems, or databases
  • Experience designing, analyzing, and troubleshooting large-scale distributed systems
  • Systematic problem-solving approach, combined with excellent communication skills and a sense of ownership and drive
  • Experience in configuring application monitoring with Azure Monitor and Application Insight
  • Experience with Service Mesh
  • Previous experience as a DevOps engineer is preferred
We offer
  • Flexible working format - remote, office-based or flexible
  • A competitive salary and good compensation package
  • Personalized career growth
  • Professional development tools (mentorship program, tech talks and trainings, centers of excellence, and more)
  • Active tech communities with regular knowledge sharing
  • Education reimbursement
  • Memorable anniversary presents
  • Corporate events and team buildings
  • Other location-specific benefits *not applicable for freelancers Originally posted on Himalayas
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE (Site Realiability Engineer)
SRE (Site Realiability Engineer)

STRATIS Cloud Tech Solutions INC • Arkansas

On-site
USD 110,000 - 150,000
Competitive salary
Growth and learning opportunities
Friendly, collaborative team
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

On-site
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000