Senior Site Reliability Engineer

ICIMS

Mexico

On-site

PHP 9,130,000 - 11,564,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

iCIMS is a leading enterprise hiring platform trusted by thousands of organizations in over 200 countries. We are seeking a senior SRE/Platform Reliability leader to drive observability standards, incident management, and reliability across enterprise services.

You will mentor teams, define architectures, and partner with Engineering, Cloud, and Product to improve customer outcomes. You will own the technical roadmap for telemetry, SLIs/SLOs, and cost governance while guiding cross-functional

Qualifications

  • Bachelor’s degree or equivalent in a technical field.
  • Equivalent combination of education and experience will be considered.
  • Cloud certifications (AWS, Azure, or Google Cloud) are preferred.

Responsibilities

  • Provide strategic technical direction for an SRE team across regions.
  • Own the technical strategy and roadmap for observability and reliability capabilities.
  • Define reference architectures, patterns, and standards for production.
  • Lead incident management and post-incident reviews with durable actions.
  • Mentor engineers and raise observability practices across teams.
  • Oversee enterprise-wide incident response and on-call procedures.
  • Develop runbooks and emergency procedures for critical services.
  • Work with Product/Engineering to improve customer impact through telemetry.

Skills

SRE leadership
Incident management
Mentoring
Cross-team collaboration

Education

Bachelor’s degree in Computer Science, Engineering, IS or related
Equivalent combination of education and experience
Cloud certifications (AWS, Azure or Google Cloud)

Tools

Grafana
Prometheus
Sumo Logic
New Relic
Terraform

Job description

Success Metrics
  • Customer Impact: Reduced MTTD/MTTR and improved customer experience through faster detection, diagnosis, and recovery
  • Observability Adoption: Measurable adoption of common logging, metrics, tracing, alerting, dashboarding, and service ownership standards across critical services
  • Reliability Engineering: Expanded use of meaningful SLIs, SLOs, and error budgets to drive service health and engineering priorities
  • Signal Quality: Reduction in noisy, duplicate, and non-actionable alerts while improving coverage of critical customer journeys and dependencies
  • Tooling & Cost Efficiency: Improved observability platform efficiency through governance, consolidation, telemetry optimization, and reduced tool sprawl
  • Cross-functional Adoption: Strong partnership with Product and Engineering teams that translates standards into measurable production adoption
About Us

ICIMS is a leading enterprise hiring platform that combines the scale and reliability of enterprise software with the transformative power of AI. Thousands of organizations across more than 200 countries and territories trust ICIMS to find and hire the people who shape their future and drive their business forward. Powered by insights from billions of hiring interactions, continuous AI innovation, and a highly extensible platform, ICIMS helps organizations turn talent acquisition into a competitive advantage. For more than 25 years, ICIMS has delivered end-to-end hiring solutions that improve recruiting efficiency, reduce costs and create exceptional candidate experiences.

ICIMS helps solve one of the biggest challenges businesses face today: building a workforce that can adapt, scale, and perform in an increasingly competitive and unpredictable talent market. We uniquely do that by combining enterprise-grade hiring technology, AI-powered insights and automation, and connected talent experiences to help organizations improve hiring outcomes while driving measurable impact.

Responsibilities
Technical Leadership
  • Provide strategic technical direction for a team of 5+ SRE engineers across one or more geographic regions (US, Ireland, or India)
  • Own the technical strategy and roadmap for enterprise observability and reliability capabilities in partnership with SRE, Engineering, Cloud, and Product teams
  • Define reference architectures, engineering patterns, and standards that teams can consistently apply in production
  • Drive architecture reviews and technical decision-making for complex observability, reliability, scalability, and performance challenges
  • Provide hands‑on technical mentorship and guidance, raising observability and SRE engineering practices across teams
  • Participate in enterprise-wide incident management, ensuring rapid detection, response, restoration, and prevention of recurring issues
  • Improve incident detection and triage through actionable telemetry, service health views, dependency context, and well‑designed alerting
  • Develop and maintain runbooks, emergency response procedures, and operational readiness practices for critical services
  • Lead root cause and post‑incident reviews, ensuring clear documentation and implementation of durable corrective actions
  • Participate in 24/7 on‑call and escalation procedures and serve as a senior technical leader with Engineering and Incident Management during critical incidents
Observability Strategy & Standards
  • Establish and evolve enterprise standards for logs, metrics, traces, alerting, dashboards, instrumentation, and service ownership
  • Champion OpenTelemetry‑first, vendor‑neutral instrumentation patterns with consistent context, correlation, naming, and metadata across services
  • Implement meaningful SLIs, SLOs, error budgets, and service health views that connect technical signals to customer impact
  • Drive practical adoption and governance of observability standards, measuring coverage, signal quality, and operational effectiveness across teams
Platform Reliability, Automation & Tooling
  • Design and operate scalable observability platforms and telemetry pipelines using technologies such as Grafana, Prometheus, Sumo Logic, New Relic, and cloud‑native services
  • Lead observability platform modernization, migration, and consolidation while maintaining coverage and controlling ingestion, retention, cardinality, and overall tooling cost
  • Use infrastructure‑as‑code, automation, self‑service patterns, and automated remediation to make reliability practices repeatable and reduce operational overhead
  • Monitor and optimize multi‑cloud infrastructure and core services across AWS, Azure, and GCP for reliability, performance, capacity, and operational efficiency
Qualifications
  • Bachelor’s degree in computer science, Engineering, Information Systems, or related technical field
  • Equivalent combination of education and experience will be considered
  • Cloud certifications (AWS, Azure, or Google Cloud)
Technical Experience
  • 8+ years in SRE, DevOps, Infrastructure Engineering, or Observability Engineering roles with 4+ years in senior technical positions
  • Proven hands‑on experience designing, implementing, and operating observability capabilities at scale in large enterprise SaaS or cloud production environments
  • Deep experience across logging, metrics, distributed tracing, alerting, dashboards, and OpenTelemetry, with platforms such as Grafana, Prometheus, Sumo Logic, and New Relic
  • Strong multi‑cloud and cloud‑native experience, including AWS, containers, Kubernetes/ECS, Linux, and distributed application architectures
  • Experience designing scalable telemetry pipelines and managing sampling, retention, cardinality, data quality, and cost tradeoffs in high‑volume environments
Leadership & Communication
  • Proven track record creating technical standards and successfully driving them from architecture into consistent production adoption across engineering teams
  • Experience serving as a senior technical leader during critical incidents and complex cross‑team reliability initiatives
  • Strong communication and influencing skills with engineers, architects, product leaders, and senior stakeholders
  • Demonstrated ability to mentor technical teams, build alignment across organizational boundaries, and lead through influence
SRE & Operations
  • Demonstrated success implementing SRE principles in large‑scale production environments, including practical use of SLIs, SLOs, and error budgets
  • Strong background in incident management, root cause analysis, operational readiness, automation, and continuous reliability improvement
  • Experience with ITIL frameworks and tools and with establishing service‑level expectations for enterprise SaaS products
Preferred
  • Experience leading large‑scale observability platform migrations, consolidation initiatives, and telemetry cost/governance programs
  • Infrastructure‑as‑code expertise with Terraform or CloudFormation; authentication and identity management systems knowledge is a plus
EEO Statement

iCIMS is a place where everyone belongs. We celebrate diversity and are committed to creating an inclusive environment for all employees. Our approach helps us to build a winning team that represents a variety of backgrounds, perspectives, and abilities. So, regardless of how your diversity expresses itself, you can find a home here at iCIMS. We prohibit discrimination and harassment of any kind based on race, color, religion, national origin, sex (including pregnancy), sexual orientation, gender identity, gender expression, age, veteran status, genetic information, disability, or other applicable legally protected characteristics. If you’d like to request an accommodation due to a disability, please contact us at careers@icims.com.

Competitive health and wellness benefits include medical insurance (employee and dependent family members), personal accident and group term life insurance, bonding and parental leave, lifestyle spending account reimbursements, wellness services offerings, sick and casual/emergency days, paid holidays, tuition reimbursement, retirals (PF - employer contribution) and gratuity. Benefits and eligibility may vary by location, role, and tenure. Learn more here: https://careers.icims.com/benefits

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Cloud Site Reliability Engineer, CX
Cloud Site Reliability Engineer, CX

Nice • Manila

On-site
PHP 669,600 - 892,800
Senior Site Reliability Engineer (AWS)
Senior Site Reliability Engineer (AWS)

broadridge • Philippines

On-site
PHP 1,000,000 - 2,400,000
Hybrid work model
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8, Inc. • Manila

On-site
PHP 900,000 - 1,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia • Philippines

On-site
PHP 1,000,000 - 1,800,000
Fixed compensation
Vacation leaves
Professional development opportunities
+3
Senior Site Reliability Engineer AWS
Senior Site Reliability Engineer AWS

Broadridge • Manila

Hybrid
PHP 1,200,000 - 2,400,000
Senior Site Reliability Engineer (AWS)
Senior Site Reliability Engineer (AWS)

Broadridge • Metro Manila

On-site
PHP 1,200,000 - 2,400,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AIPI Acquire Intelligence Philippines Inc. • Taguig

On-site
PHP 1,000,000 - 1,500,000