Site Reliability Engineer

OneStream Software LLC

Northern (KY)

Hybrid

USD 114,000 - 148,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

OneStream Software LLC is seeking a Site Reliability Engineer to ensure platform reliability, performance, and service availability. Responsibilities include implementing observability solutions and automating infrastructure deployments, alongside a small team to maintain reliable systems.

The ideal candidate will have extensive experience with cloud infrastructure, including Azure services and container orchestration. A passion for technology and ability to mentor others are crucial for this role. The position requires collaboration with product and engineering teams, contributing to a high-availability service environment.

Qualifications

  • Proven work experience as a Site Reliability Engineer or similar role.
  • 6+ years of cloud infrastructure and software development experience.
  • 2+ years Azure Kubernetes Services experience with hands-on skills.

Responsibilities

  • Implement application/infrastructure observability solutions.
  • Participate in regular On-Call rotations and post-mortem reports.
  • Proactively partner with Product and Engineering teams.

Skills

Cloud infrastructure experience
Automation scripting (PowerShell, Bash, Python)
Container orchestration (Kubernetes, AKS)
Infrastructure-as-Code (Terraform, CloudFormation)
APM and observability tools

Education

BS/BA in computer science or related field

Tools

Dynatrace
Azure DevOps
Git

Job description

Gross Annual Base Salary: USD 114,000 - 148,000

Additional variable compensation and benefits may apply. Total compensation is based on experience, skills, and location using objective, job-related criteria.

Summary

As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant, and highly available. If you enjoy staying at the forefront of technology and automating infrastructure deployments, then this is the job for you. This vital role within Cloud Services requires knowledge and experience designing, implementing, and monitoring scalable and secure cloud services. The employee is expected to work well in a small team and is willing to share responsibilities with other team members as needed. You will interact with internal staff, managers, and customers to implement and maintain operations. A passion for technology and learning, and the ability to grow others are vital for success in this role.

Primary Duties and Responsibilities
  • Implement application/infrastructure observability solutions to ensure desired application availability, reliability, and performance.
  • Participate in regular On-Call rotations and share details related to incidents and their resolution through post-mortem reports and regular review meetings.
  • Proactively partner with Product and Engineering teams to identify, develop, deploy, and maintain reliable systems and services.
  • Influence and create new designs, architectures, standards, and methods for large-scale systems.
  • Sustain a high level of reliability for key services and automated systems.
  • Automate processes to improve reliability, performance, and availability.
  • Update technical documentation, workflows, and knowledge base articles.
  • Provide feedback in pull requests and peer coding reviews.
  • Implement codified automated solutions that build integrations between Dynatrace, Azure DevOps and Jira.
  • Solid knowledge in focused areas of OneStream Software.
  • Ability to mentor others in several technical areas.
  • Understanding practical use of SOC/FedRAMP controls to assist Compliance and Security teams.
Required Education and Experience
  • BS/BA in computer science, engineering, or technology-related field (or equivalent work experience).
  • Proven work experience as a Site Reliability Engineer or in a similar role.
  • 6+ years of cloud infrastructure and software development experience.
  • 2+ years hands on experience of Azure Kubernetes Services (AKS) with container-based deployment skills or other platforms such as OpenShift, GKS, EKS.
  • Advanced understanding of APM and observability tools such as Dynatrace, AppInsights, DataDog, Log Analytics, New Relic, Prometheus and Grafana.
  • Advanced understanding of Infrastructure-as-Code (IaC) concepts and tooling (Terraform, CloudFormation templates, Bicep or ARM templates) on Microsoft Azure, Amazon Web Services (AWS), or Google Cloud Platform (GCP).
  • Deep knowledge of Configuration Management/Orchestration utilities such as Ansible, PowerShell DSC, Chef, and Puppet.
  • Advanced understanding of cloud concepts including elasticity, security, and identity management.
  • Well versed familiarity with Agile Development methodologies utilizing Jira or Azure DevOps Boards.
  • 6+ years of hands‑on experience with the following technologies, tools, and concepts:
    • Automating processes using PowerShell, Bash, CLI, REST APIs, Python, ARM Templates or other scripting languages.
    • Comfortable leveraging source control tools such as Git, Azure DevOps, or GitHub.
    • Knowledge of container orchestration platforms such as Kubernetes, OpenShift, AKS, GKS or helm.
    • Microsoft Azure, Amazon Web Services (AWS) or Google Cloud (GCP).
Preferred Education and Experience
  • Experience working for a cloud service provider (CSP), managed service provider (MSP), or SaaS provider.
  • 6+ years of relevant Azure experience deploying and managing leveraging Infrastructure-as-Code (IAC) concepts.
  • Experience with Microsoft and .NET (.NET, C#, SQL).
  • Experience writing efficient and reliable code in a development environment.
  • Debian, Ubuntu, Alpine or other distributions of the Linux operating systems.
  • Deep knowledge and understanding of containerized applications, with special attention to reliability and monitoring of those containerized applications.
Knowledge, Skills, and Abilities
  • Deal well with ambiguous/undefined problems.
  • Ability to self-motivate and work independently.
  • Strong organizational and prioritization skills.
  • Ability to find and apply effective solutions to emerging problems and challenges.
  • Strong attention to detail.
  • Comfortable communicating with all levels of management and engineering.
  • Ability to get up to speed quickly with modern technologies and services.
  • Ability to multitask on a variety of projects.
Travel
  • Travel Requirement: Travel is not expected to exceed 5%.

All candidates must be legally authorized to work for any company in the country where this position is located without sponsorship.

OneStream is an Equal Opportunity Employer.

Equal Opportunity Employer

This employer is required to notify all applicants of their rights pursuant to federal employment laws. For further information, please review the Know Your Rights notice from the Department of Labor.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Site Reliability Engineer - Cloud & Observability
Remote Site Reliability Engineer - Cloud & Observability

OneStream Software • Birmingham (MI)

On-site
USD 114,000 - 148,000
Vision insurance
Medical insurance
Life insurance
+2
Senior Site Reliability Engineer - Cloud & Observability
Senior Site Reliability Engineer - Cloud & Observability

OneStream Software LLC • Northern (KY)

Hybrid
USD 114,000 - 148,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Plano (TX)

On-site
USD 152,000 - 192,000
Discretionary incentive eligible
Benefits eligible
Senior SysOps Engineer
Senior SysOps Engineer

OneStream Software • Rochester (MI)

On-site
USD 120,000 - 149,000
Medical Insurance
Dental Insurance
Vision Insurance
+3
Site Reliability Engineer
Site Reliability Engineer

Axle • Frederick (MD)

On-site
USD 140,000 - 155,000
Paid Time Off
401K match
Educational Benefits
+5
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Plano (TX)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Paid time off
Access to resources and support
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Charlotte (NC)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Paid time off
Discretionary incentive eligibility