SRE/DevOps

NDEAVOUR CONSULTING

United States

Hybrid

USD 120,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote Office
Parking Space
Fun Office Space
Health Insurance
Holidays
Personal Development
Employee Referral Programme
Social Events
Family Insurance
Multisport Card

Job summary

Mobile Wave Solutions is seeking a seasoned Site Reliability Engineer to own the health, performance, and delivery of the client infrastructure. You will operate at the intersection of engineering, operations, and quality, measuring system behavior, automating toil, and enabling fast, reliable production paths.

You will set concrete reliability targets, define best practices, and apply AI/AIOps to monitor, alert, and respond intelligently while leading incident postmortems and driving cross‑team

Qualifications

  • 5+ years of experience in an SRE, DevOps, or Platform Engineering role, with a track record of owning systems end to end.
  • Strong grasp of reliability engineering fundamentals: SLIs/SLOs, error budgets, and reducing toil.
  • Hands‑on experience designing modernising and operating CI/CD pipelines.

Responsibilities

  • Own the infrastructure. Build, maintain, and scale the systems our product runs on, with reliability and cost‑efficiency as first‑class concerns
  • Measure the non‑functional. Define and track SLIs, SLOs, and error budgets for availability, latency, throughput, and scalability. Make system behavior visible and quantifiable to the whole team
  • Automate relentlessly. Identify and eliminate toil. Replace manual operational work with code, infrastructure‑as‑code, and self‑healing systems
  • Build a seamless delivery process. Design, maintain, and improve CI/CD pipelines so engineers can ship safely and frequently with fast feedback
  • Collaborate across functions. Partner closely with software engineers and QA to embed reliability and quality early – through testing strategy, deployment practices, and shared ownership of production
  • Apply AI to operations. Use AI/AIOps to automate remediation, surface anomalies, reduce alert noise, and improve the signal quality of monitoring, reporting, and on‑call
  • Lead incident response. Drive blameless postmortems and turn incidents into systemic improvements
  • Set the standard. Define reliability and operational best practices, and mentor engineers across the team to raise the bar

Skills

SRE ownership
SLIs/SLOs
CI/CD pipelines
IaC (Terraform)
Cloud platforms
Kubernetes
Python/Go
Observability (NewRelic/Prometheus)

Tools

Pulumi
CloudFormation
OpenTelemetry
NewRelic
Datadog

Job description

Mobile Wave Solutions is a professional services company specializing in software development as a service. We are committed to delivering scalable, high-quality software solutions that meet our clients’ evolving needs. With a growing team of over 120 engineers and a mission to empower businesses globally, we provide expert teams to deliver robust solutions and drive innovation.

Role Overview

We’re looking for a Site Reliability Engineer to own the health, performance, and delivery of the infrastructure of one of our client.

You’ll sit at the intersection of engineering, operations, and quality measuring how their systems actually behave, automating away manual work, and making their path from commit to production fast and dependable.

This is a hands‑on role for a seasoned engineer who treats operations as an engineering problem and sets the bar for others. You’ll define what “reliable” means for the systems in concrete, measurable terms, establish the practices the team operates by, and use modern tooling, including AI/AIOps, to monitor, alert, and respond intelligently rather than reactively.

Key Responsibilities
  • _ Own the infrastructure _. Build, maintain, and scale the systems our product runs on, with reliability and cost‑efficiency as first‑class concerns

  • _ Measure the non‑functional. _ Define and track SLIs, SLOs, and error budgets for availability, latency, throughput, and scalability. Make system behavior visible and quantifiable to the whole team

  • _ Automate relentlessly _. Identify and eliminate toil. Replace manual operational work with code, infrastructure‑as‑code, and self‑healing systems

  • _ Build a seamless delivery process. _ Design, maintain, and improve CI/CD pipelines so engineers can ship safely and frequently with fast feedback

  • _ Collaborate across functions _. Partner closely with software engineers and QA to embed reliability and quality early – through testing strategy, deployment practices, and shared ownership of production

  • _ Apply AI to operations _. Use AI/AIOps to automate remediation, surface anomalies, reduce alert noise, and improve the signal quality of monitoring, reporting, and on‑call

  • Lead incident response. Drive blameless postmortems and turn incidents into systemic improvements

  • _ Set the standard _. Define reliability and operational best practices, and mentor engineers across the team to raise the bar

Qualifications
  • 5+ years of experience in an SRE, DevOps, or Platform Engineering role, with a track record of owning systems end to end

  • Strong grasp of reliability engineering fundamentals: SLIs/SLOs, error budgets, and reducing toil

  • Hands‑on experience designing modernising and operating CI/CD pipelines

  • Solid infrastructure‑as‑code skills (Terraform / Pulumi / CloudFormation)

  • Experience with cloud platforms (GCP / AWS / Azure) and container orchestration (Kubernetes / Docker)

  • Proficiency in at least one programming/scripting language for automation (Python / Go / Bash)

  • Strong observability experience: metrics, logging, tracing, and alerting (NewRelic, Prometheus / Grafana / Datadog / OpenTelemetry)

  • Applied experience using AI/AIOps to automate, measure, report, and alert – anomaly detection, intelligent alerting, noise reduction, or automated remediation

  • A collaborative mindset and comfort working alongside engineers and QA toward shared reliability goals

  • Demonstrated technical leadership—mentoring engineers, driving cross‑team initiatives, and influencing engineering practices

You would impress us if you have
  • Experience introducing AIOps tooling into an existing observability stack

  • Background in performance and load testing.

  • Familiarity with security and compliance practices (SOC 2 / ISO 27001 / GDPR)

Our Benefits
  • Remote Office – Option to work remotely or hybrid

  • Parking Space – Free parking available

  • Fun Office Space – Game zone and relaxation area

  • Health Insurance – Private health insurance, including dental care

  • Holidays – 5 extra days after your 1st and 5th year with us

  • Personal Development – Company‑sponsored training and development

  • Employee Referral Programme – Competitive bonus for successful referrals

  • Social Events – Celebrating success together

  • Family Insurance – Add insurance coverage for a family member

  • Multisport Card – Fully covered sports pass

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Weekend Site Reliability Engineer
Weekend Site Reliability Engineer

Sporty Group • United States

Remote
USD 120,000 - 180,000
Remote first
Bonuses (quarterly)
28 days leave
+4
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Site Reliability Engineer
Site Reliability Engineer

Seek Now • Atlanta (GA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Health, dental, and vision coverage
401(k) with company match
+1
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2