3510- Site Reliability Engineer II

Innovaccer

Dallas (TX)

On-site

USD 110,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous Paid Time Off: 22 days per year plus company holidays
Best-in-Class Parental Leave
Comprehensive insurance coverage

Job summary

A leading healthcare technology firm is seeking a Site Reliability Engineer-II to enhance its modern cloud infrastructure. Responsibilities include managing deployment processes, ensuring service availability, and leading operational reviews. The ideal candidate will possess strong cloud provider expertise, Kubernetes experience, and excellent problem-solving skills. This role offers unique challenges in a fast-paced environment in Dallas, Texas, with competitive benefits.

Qualifications

  • 4-7 years in production engineering, site reliability, or related roles.
  • Strong hands-on experience with at least one cloud provider.
  • Excellent judgment and analytical thinking.

Responsibilities

  • Take ownership of SRE pillars like Deployment and Scalability.
  • Lead production rollouts using CI/CD pipelines.
  • Manage autoscaling and perform triage in production.

Skills

Cloud provider expertise (AWS, Azure, GCP)
Kubernetes
Linux
Python scripting
Observability tools (ElasticSearch, Prometheus, etc.)
CI/CD pipelines knowledge (Jenkins, ArgoCD)
Analytical thinking
Problem-solving skills

Tools

ElasticSearch
Prometheus
Kafka
Postgres
Snowflake

Job description

Overview

We at Innovaccer are looking for a Site Reliability Engineer-II to build secured modern healthcare cloud infrastructure and a massive data stack and aim to write everything as code.

Responsibilities
  • Take ownership of SRE pillars: Deployment, Reliability, Scalability, Service Availability (SLA/SLO/SLI), Performance, and Cost
  • Lead production rollouts of new releases and emergency patches using CI/CD pipelines while continuously improving deployment processes
  • Establish robust production promotion and change management processes with quality gates across Dev/QA teams
  • Roll out a complete observability stack across systems to proactively detect and resolve outages or degradations
  • Analyze production system metrics, optimize system utilization, and drive cost efficiency
  • Manage autoscaling of the platform during peak usage scenarios
  • Perform triage and RCA by leveraging observability toolchains across the platform architecture
  • Reduce escalations to higher-level teams through proactive reliability improvements
  • Participate in the 24x7 OnCall Production Support team
  • Lead monthly operational reviews with executives covering KPIs such as uptime, RCA, CAP (Corrective Action Plan), PAP (Preventive Action Plan), and security/audit reports
  • Operate and manage production and staging cloud platforms, ensuring uptime and SLA adherence
  • Collaborate with Dev, QA, DevOps, and Customer Success teams to drive RCA and product improvements
  • Implement security guidelines (e.g., DDoS protection, vulnerability management, patch management, security agents)
  • Manage least-privilege RBAC for production services and toolchains
  • Build and execute Disaster Recovery plans and actively participate in Incident Response
  • Work with a cool head under pressure and avoid shortcuts during production issues
  • Collaborate effectively across teams with excellent verbal and written communication skills
  • Build strong relationships and drive results without direct reporting lines
  • Take ownership, be highly organized, self-motivated, and accountable for high-quality delivery
Qualifications
  • 4-7 years in production engineering, site reliability, or related roles
  • Solid hands-on experience with at least one cloud provider (AWS, Azure, GCP) with automation focus (certifications preferred)
  • Strong expertise in Kubernetes and Linux
  • Proficiency in scripting/programming (Python required)
  • Observability is very critical for the scale of our systems and ability to find insights/behavior, detect problem/failures. Looking for leads to drive this charter spanning across logs, metrics, mesh, tracing etc.
  • Knowledge of CI/CD pipelines and toolchains (Jenkins, ArgoCD, GitOps)
  • Familiarity with persistence stores (Postgres, MongoDB), data warehousing (Snowflake, Databricks), and messaging (Kafka)
  • Exposure to monitoring/observability tools such as ElasticSearch, Prometheus, Jaeger, NewRelic, etc
  • Proven experience in production reliability, scalability, and performance systems
  • Experience in 24x7 production environments with process focus
  • Familiarity with ticketing and incident management systems
  • Security-first mindset with knowledge of vulnerability management and compliance
  • Advantageous: hands-on experience with Kafka, Postgres, and Snowflake
  • Excellent judgment, analytical thinking, and problem-solving skills
  • Ability to quickly identify and drive optimal solutions within constraints
  • Lead least privilege based RBAC for various production services and tool chains
  • Able to perform with cool head under pressure situations without taking any shortcuts
  • Collaboration with solid verbal and oral communication skills are very critical to this role. Strong cross-functional collaboration skills, relationship building skills, and ability to achieve results without direct reporting relationships
  • Ability to quickly identify and drive to the optimal solution when presented with a series of constraints
  • Excellent judgment, analytical thinking, and problem-solving skills
  • Self-motivated individual that possesses excellent time management and organizational skills
  • Strong sense of personal responsibility and accountability for delivering high quality work.
Benefits
  • Generous Paid Time Off: 22 days per year plus company holidays
  • Best-in-Class Parental Leave
  • Recognition & Rewards: monetary incentives and company-wide recognition
  • Comprehensive Insurance Coverage: medical, dental, vision; 100% company-paid disability and basic life insurance

Innovaccer Inc. is an equal opportunity employer. We celebrate diversity and are committed to fostering an inclusive workplace where all employees feel valued and empowered regardless of protected characteristics. Innovaccer Inc. participates in the E-Verify program to confirm employment eligibility of all newly hired employees based out of the U.S. and employed by Innovaccer Inc.

Disclaimer: Innovaccer does not charge fees or require payment from individuals or agencies for securing employment with us. We do not guarantee job spots or engage in any financial transactions related to employment. If you encounter any posts or requests asking for payment or personal information, report them to our HR department at px@innovaccer.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

3823-Forward Deployed Engineer
3823-Forward Deployed Engineer

Innovaccer Analytics • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Generous Paid Time Off
Best-in-Class Parental Leave
Recognition & Rewards
+1
Director Product-AI Platform
Director Product-AI Platform

Innovaccer • San Francisco (CA)

On-site
USD 190,000 - 260,000
Generous Paid Time Off
Best-in-Class Parental Leave
Recognition & Rewards
+1
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Socket.dev • California (MO)

On-site
USD 150,000 - 230,000
Senior Software Engineer - Streaming (Kafka)
Senior Software Engineer - Streaming (Kafka)

Innovaccer Inc. • Northern (KY)

Hybrid
USD 140,000 - 180,000
Generous PTO
Parental Leave
Rewards program
+1
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Innovaccer • San Francisco (CA)

On-site
USD 150,000 - 230,000
Generous PTO benefits
Parental Leave
Rewards & Recognition
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Limelight Health • Austin (TX)

On-site
USD 152,000 - 195,000
Competitive salary
Stock options
Health benefits
+2
Principal Artificial Intelligence Engineer
Principal Artificial Intelligence Engineer

Innovaccer • San Francisco (CA)

On-site
USD 210,000 - 300,000
PTO 20 days/year
Parental leave
Rewards & Recognition
+1
Principal Software Engineer - Data Platform (Iceberg/Trino)
Principal Software Engineer - Data Platform (Iceberg/Trino)

Innovaccer Inc. • Northern (KY)

Hybrid
USD 180,000 - 260,000
Generous paid time off
Parental leave
Recognition & rewards
+1
Senior DevOps Engineer
Senior DevOps Engineer

Transformcap • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Flexible hours
+1
4195- Software Development Engineer-III Backend (Gravity)
4195- Software Development Engineer-III Backend (Gravity)

Innovaccer Analytics • United States

On-site
USD 140,000 - 190,000
Generous leave
Parental leave
Sabbatical
+3