Application SRE (DevOps)

ELLKAY

Elmwood Park (NJ)

Hybrid

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Remote work options
Hybrid work environment
Medical, Dental, and Vision benefits
401k w/ matching
Paid time off
Career growth opportunities
Flexible working hours
Hybrid/remote options for certainRoles

Job summary

ELLKAY is seeking an Application Site Reliability Engineer (SRE) with strong DevOps experience to improve the reliability, scalability, and performance of our applications. The role drives reliability standards, observability, and automation across teams to enable faster, safer releases.

The SRE will lead complex incident responses, partner with development teams, and implement best practices to ensure resilient software delivery in a hybrid work environment.

Qualifications

  • Minimum 7 years of SRE experience with strong DevOps background.
  • Solid understanding of Windows and Linux systems and networking.
  • Hands-on cloud experience (AWS, Azure, GCP) and container orchestration.
  • Proficiency with CI/CD tools and Infrastructure as Code.
  • Strong scripting and observability skills; incident management experience.

Responsibilities

  • Own application reliability, availability, performance, and scalability in production and non-production environments.
  • Design, build, and maintain CI/CD pipelines for deployments.
  • Automate infrastructure provisioning with Infrastructure as Code.
  • Monitor health with metrics, logs, and traces; define SLIs, SLOs, and error budgets.
  • Lead incident response and RCA, ensuring corrective actions are completed.
  • Improve resilience via capacity planning, tuning, and fault tolerance.
  • Partner with development teams to meet reliability and scalability objectives.
  • Reduce manual toil through automation and self-healing solutions.
  • Act as incident commander for Sev1/Sev2 situations when needed.

Skills

SRE/DevOps experience
Windows & Linux
Cloud platforms (AWS, Azure, GCP)
Docker & Kubernetes
CI/CD (Jenkins, GitHub Actions)
Infrastructure as Code (Terraform, in/
Scripting (Python, Bash)
Observability (Prometheus, Grafana)
Incident management & RCA

Tools

Docker
Kubernetes
Jenkins
GitHub Actions
Terraform
CloudFormation
ARM
Prometheus
Grafana
ELK
Datadog

Job description

ELLKAY started out providing connectivity solutions to laboratories and within a few years, grew to also provide data management solutions to ambulatory organizations. ELLKAY is now a trusted data management partner in five healthcare segments. ELLKAY’s solutions continue to serve laboratories and ambulatory practices and have expanded to empower hospitals and health systems, healthcare IT vendors, ambulatory practices, health plans, and other healthcare organizations with cutting-edge technologies and solutions that drive their growth and interoperability strategies.

Today, ELLKAY remains true to our core values, building strong partner relationships and offering unparalleled service and support while providing innovative, scalable solutions to the challenges our customers face in today’s data-rich world.

ELLKAY's experience, customer-focused approach, and reputation for innovation, speed, and accuracy differentiate ELLKAY as a premier partner for your interoperability needs and data management strategy.

Job Description

We are looking for an Application Site Reliability Engineer (SRE) with strong DevOps experience to improve the reliability, scalability, and performance of our applications.

The Application Site Reliability Engineer will serve as a technical contact responsible for driving the reliability, performance, and operational maturity of our application ecosystem. This role works across multiple teams to support scalable systems, establish reliability standards, improve observability, and implement automation that reduces operational effort. The SRE will lead complex incident responses, work with engineering teams in best practices, and influence architectural decisions to ensure resilient, high-quality software delivery.

You will help define reliability standards, reduce operational toil, and ensure smooth production operations while enabling faster and safer releases.

Essential Duties & Responsibilities
  • Own application reliability, availability, performance, and scalability in production and non-production environments
  • Design, build, and maintain CI/CD pipelines for application deployments
  • Automate infrastructure provisioning and configuration using Infrastructure as Code
  • Monitor application health using metrics, logs, and traces; define SLIs, SLOs, and error budgets
  • Lead incident response, root-cause analysis (RCA), ensuring corrective and preventive actions are completed and communicated.
  • Improve system resilience through capacity planning, system tuning, and fault tolerance
  • Partner with development teams to ensure services meet reliability, performance, and scalability objectives.
  • Reduce manual operational effort through automation and self-healing solutions
  • Serve as a point of contact for critical Sev1/Sev2 incidents, leading incident command when required.
Qualifications
  • Strong experience as an SRE, DevOps Engineer, or Production Support Engineer
  • Solid understanding of Windows, Linux/Unix systems and networking fundamentals
  • 7 years of experience as an SRE
  • Hands-on experience with cloud platforms such as AWS, Azure, or GCP
  • Experience with containerization and orchestration tools like Docker and Kubernetes
  • Proficiency in CI/CD tools such as Jenkins, GitHub Actions, , or similar
  • Experience with Infrastructure as Code tools like Terraform, CloudFormation, or ARM
  • Strong scripting skills in Python, Bash, or similar languages
  • Experience with monitoring and observability tools (Prometheus, Grafana, ELK, Datadog, etc.)
  • Understanding of reliability concepts such as SLAs, SLOs, and incident management
Preferred Qualifications
  • Experience supporting microservices-based architectures
  • Knowledge of security best practices in cloud and DevOps environments
  • Experience with configuration management tools (Ansible, Chef, or Puppet)
  • Exposure to chaos engineering or resilience testing practices
Soft Skills
  • Strong problem-solving and troubleshooting skills
  • Ability to work calmly during incidents and high-pressure situations
  • Clear communication and collaboration with cross-functional teams
  • Ownership mindset with a focus on continuous improvement
What We Offer
  • Opportunity to work on highly available, business-critical applications
  • Collaborative engineering culture with strong DevOps and SRE practices
  • Competitive compensation and benefits
  • Learning and growth opportunities in cloud, automation, and reliability engineering
Benefits

ELLKAY offers a comprehensive and competitive benefit package that starts day one!

  • Medical, Dental, and Vision benefits
  • Employer-paid Life and LTD
  • 401k w/ matching – once eligibility is met
  • Work/life balance
  • Paid Volunteer Program
  • Flexible working hours
  • Generous FTO
  • Remote work options
  • Employee Discounts
  • Parental Leave
  • Gym membership / Exercise class stipends
Our Awesome Culture Includes
  • Working with talented, collaborative, and friendly people who love what they do
  • Professional growth within
  • Innovation environment
  • On site in HQ Free daily lunches
Awards
  • 2025 Top Workplaces Employee Appreciation
  • 2025 Top Workplaces Employee Well-Being
  • 2025 Top Workplaces Professional Development
  • 2025 Top Workplaces Leadership
  • 2026 USA Today Top Workplaces
  • 2026 NJ Top Workplaces
Additional Information

At ELLKAY, we are committed to operating as a hybrid workplace, offering employees flexibility in how they structure their time between in-office and remote work. We recognize the significance of fostering connections, collaboration, and creativity within our office culture and its positive impact on our business. Our philosophy of operating as a hybrid workplace underscores our dedication to enabling employees to tailor work-life balance to their individual preferences. For those who do not live within 40 miles of one of our offices, we are open to considering remote work for candidates whose skills and experience strongly align with the role. While we prioritize a hybrid work environment for most roles, we understand the importance of flexibility and are open to remote work for specific positions and specialized skill sets.

For more information on our company, visit www.ELLKAY.com.

ELLKAY LLC is a Smoke‑Free Workplace.

ELLKAY, LLC provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Application SRE
Application SRE

Ellkay, Llc • Elmwood Park (NJ)

Hybrid
USD 90,000 - 110,000
Medical, Dental, and Vision benefits
401k with matching
Generous paid time off (FTO)
+2
Principal Platform Engineer
Principal Platform Engineer

ELLKAY • United States

Hybrid
USD 170,000 - 230,000
Medical benefits
401k with matching
Remote work options
+2
Principal Platform Engineer
Principal Platform Engineer

Ellkay,-LLC • United States

Hybrid
USD 160,000 - 180,000
Remote work options
401k with matching
Hybrid work model
+1
Sr. Staff Platform Engineer
Sr. Staff Platform Engineer

Ellkay, Llc • United States

On-site
USD 140,000 - 190,000
Senior Software Engineer
Senior Software Engineer

ELLKAY • United States

Hybrid
USD 140,000 - 210,000
Medical benefits
Dental benefits
Vision benefits
+5
Staff Engineer - Data Platform & APIs
Staff Engineer - Data Platform & APIs

ELLKAY • United States

Hybrid
USD 135,000 - 180,000
Medical, Dental, and Vision benefits
Employer-paid Life and LTD
401k w/ matching
+7
Sr. Director, Delivery Governance & Operations
Sr. Director, Delivery Governance & Operations

ELLKAY • United States

Hybrid
USD 180,000 - 240,000
Remote work options
Hybrid work model
401k with company match
+2
Staff Engineer - Data Platform & APIs
Staff Engineer - Data Platform & APIs

Socket.dev • United States

Hybrid
USD 150,000 - 210,000
Medical, Dental, and Vision benefits
Remote work options
401k w/ matching
+2
Staff Engineer - Data Platform & APIs
Staff Engineer - Data Platform & APIs

Ellkay, Llc • Northern (KY)

Hybrid
USD 180,000 - 200,000
Medical, Dental, Vision benefits
Employer-paid Life and LTD
401k with matching
+2
Staff CI/CD Engineer
Staff CI/CD Engineer

Ellkay, Llc • United States

On-site
USD 130,000 - 190,000