Site Reliability Engineer

OnBoard Group

Northern (KY)

Hybrid

USD 120,000 - 150,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Fully remote work with equipment
Competitive benefits: medical, dental,
401K with company match
Paid PTO & holidays
Diversity and inclusion programs

Job summary

OnBoard Group seeks a Cloud Operations Engineer III to lead observability and automation for a multi-product SaaS platform. You will own the Datadog observability practice, define SLIs, and build dashboards and alerts while guiding modernization from legacy stacks.

You will also advance security controls, runbooks, and incident response, collaborating across teams. The role emphasizes automation, IaC, and cross-functional leadership to improve reliability and reduces toil in a fast-paced, remote

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent practical experience.
  • 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps for production SaaS systems.
  • Hands-on with a modern observability platform – Datadog preferred – including APM, logs, dashboards, monitors, and SLOs.
  • Strong scripting in PowerShell; ability to write safe tooling for others.
  • Knowledge of containers/Kubernetes; autoscaling, cluster upgrades, pod diagnostics.
  • Production experience with Azure services (AKS, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, Entra ID) or similar cloud.
  • Experience with infrastructure-as-code and CI/CD (Bicep/Terraform; Helm/Kustomize; Azure DevOps).
  • Incident response experience in customer-facing environments; on-call and post-incident reviews.

Responsibilities

  • Own the Datadog platform across all products and environments, including instrumentation and tagging.
  • Instrument services for APM, tracing, logs, and synthetic monitoring; close gaps with engineering teams.
  • Maintain dashboard/SLO catalog and define SLIs for critical journeys; drive prioritization with Eng/Prod.
  • Design high-signal alerting with owners and runbooks; reduce noise and duplication.
  • Develop automation in PowerShell, Python, Bash for provisioning, diagnostics, remediation, and reporting.
  • Extend IaC estate (Bicep, Terraform, Helm, Kustomize) for reproducible environments across regions.
  • Convert runbooks to automated workflows: upgrades, rotations, provisioning, data purges, access.

Skills

Datadog
Observability
PowerShell
Python
Bash
Kubernetes
Azure
CI/CD
Incident response
Communication
Teamwork

Education

Bachelor's degree in CS/IT or related field

Tools

Bicep
Terraform
Helm
Kustomize
Azure DevOps

Job description

Reports to: Manager,Cloud Operations

Location: Remote - United States

Position Summary

The Cloud Operations Engineer III is a senior member of thecloud operationsteam, responsible for the reliability, observability, performance, and operational security of our multi-product SaaS platform. This role owns our Datadog observability practice — instrumentation standards, dashboards, SLOs, monitors, and alert routing — and leads the migration off our legacy monitoring stack. The idealcandidateis aproactiveproblem-solverwhothrives indynamic,evolvingenvironmentsand workseffectivelyacrossdepartmentsto address complexchallenges.They haveexperiencepartneringwith cross-functionalteams tounderstandanddocumentrequirements,thentranslatingthose needsintomeaningfuldashboardsthatimproveservicevisibility(Observability) and support informeddecision-making.They are passionate about automation,process improvement,andeliminatingunnecessarymanualeffort.Theyconfidentlyproposebetterapproacheswhenopportunitiesfor improvementarise.

Key Responsibilities

Observability and Datadog Ownership

  • Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration.
  • Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner withengineeringteams to close instrumentation gaps in both legacy and modern codebases.
  • Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product.
  • Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook.

Automation and Toil Elimination

  • Develop andmaintainautomation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting.
  • Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free.
  • Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning.

Security, Documentation, and Mentorship

  • Implement andmaintainplatform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning.
  • Author andmaintainrunbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response.

Skills and Experience Needed

  • Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
  • 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering forproductionSaaS systems.
  • Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs.
  • Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safelyoperate.
  • Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures.
  • Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud.
  • Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm orKustomize; Azure DevOps preferred).
  • Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews.
  • Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation.
  • Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals.
  • Strongcommunication,and teamwork skills, including the ability to work effectively with legacy systems and their constraints.
  • Nice tohave:experience migrating from a legacy monitoring stack to aconsolidatedobservability platform; relevant Azure, Kubernetes, or Datadog certifications.

Accountability

AI Curiosity & Innovation

Business Acumen

Customer Focus

Dealing with Ambiguity

Decision Making

Driving for Results

Initiating Action

Technical/Professional Knowledge

About the Company:

Boards set the standard for what organizations can achieve. At OnBoard, our board management software helps boards function at a higher level so every organization can make a bigger difference in the world.

Launched in 2011, today, OnBoard serves as the board intelligence platform for more than 5,000 organizations and their 12,000 boards and committees in 60 countries worldwide. With customers in higher education, nonprofit, healthcare systems, government, and enterprise business, OnBoard is the leading board management provider.

OnBoard has grown from a class project at Purdue University in West Lafayette, Indiana in 2003 into the world’s leading board management software platform today. Backed by JMI Equity and the acquisitions of eScribe and Govenda, OnBoard is positioned to become the industry leader in Board Management and Meeting Solutions for private and public sector entities.

Benefits and Perks:

  • Fully remote work with company provided equipment (laptop, software, etc.)
  • Employment with a growing, casual, fun, philanthropic minded company
  • US Based Employees
    • Comprehensive, high-quality medical/prescription drug plan options, as well as dental and vision plan offerings.
    • An employer contribution to your Health Savings Account (HSA) if you participate in a High Deductible Healthcare Plan.
    • Medical Flexible Spending Accounts available.
    • Dependent Care Flexible Spending Accounts available.
    • Basic life insurance in the amount of $50,000 or 1 X's your salary (whichever is higher) .
    • Short and long-term disability and Accidental Death and Dismemberment benefits at no cost to you.
    • 401K Retirement Savings Plan with automatic enrollment at the first of the month following 60 days of employment at 5% to help you secure your financial freedom. We offer a generous company match that starts on the first of the month following 60 days of employment. The company match is dollar for dollar on the first 3% of your pay that you contribute and $0.50 on the dollar on the next 2%, for a total match of 4%.
    • Paid Time Off (PTO)/Holiday
  • CAN Based Employees
    • Employer paid Life and Accidental Death Insurance
    • Contribution to Health Care Spending Account
    • Dependent Life Insurance
    • Optional Life Insurance
    • LTD Insurance
    • Drug and Paramedical Coverage
    • Dental Insurance
    • Vision Insurance
    • EAP
  • AUS Based employees
    • Superannuation rate of 12%
    • Monthly stipend of $400 AUD to purchase private medical insurance
  • UK Based Employees (via EPG)
    • Pension - Aegon
      • Passageways/OnBoard contributes 8% of the employee's basic salary
      • Employees can contribute up to 100% of salary subject to max limits
      • Enrolled from Day 1 of employment
    • Private Medical Insurance
    • Life Assurance
    • Income Protection
    • Critical Illness
    • Employee Assistance Programme
    • Serious Illness Benefit
    • Help@Hand
    • Cashplan

Diversity Statement - Culture of Togetherness:

AtOnBoard, our mission is to encourage and celebrate a culture of togetherness. We acknowledge that uniqueness is powerful, and we welcome, foster, and appreciate all. Diversity, Equity, and Inclusiveness fuelthe Pathfinder atmosphere and all our efforts. Our power is in our people and we Pledge 1% to give back to our communities and across the globe.

OnBoardis an equal opportunity employer and committed to a diverse and inclusive working environment.We do not discriminate based on race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status.

Interview Transparency & Technology DisclosureWe use video/audio recordings and artificial intelligence (AI) tools during our interview process to transcribe responses, evaluate skills, and streamline evaluations. Your data is processed securely and handled in line with our Privacy Policy and local data protection laws

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer - AI
Principal Software Engineer - AI

Socket.dev • United States

Remote
USD 160,000 - 210,000
Fully remote work with equipment
Company equipment provided
Competitive medical/dental/vision
+2
Principal Software Engineer - AI
Principal Software Engineer - AI

OnBoard Group • Northern (KY)

Hybrid
USD 140,000 - 200,000
Senior Manager, Marketing Operations
Senior Manager, Marketing Operations

OnBoard • United States

On-site
USD 120,000 - 180,000
Fully remote work with equipment
Health, dental, vision plans
401(k) with company match
+3
Senior Manager, Marketing Operations
Senior Manager, Marketing Operations

OnBoard Group • Northern (KY)

Hybrid
USD 120,000 - 180,000
Fully remote work with equipment
Health, dental, vision benefits
401K with company match
+1
Sales Development Representative
Sales Development Representative

OnBoard Group • Indianapolis (IN), Northern (KY)

Hybrid
USD 50,000 - 85,000
Fully remote work with equipment
Medical, dental & vision coverage
401K with company match
+1
Sales Development Representative
Sales Development Representative

OnBoard • Indianapolis (IN)

On-site
USD 60,000 - 90,000
Fully remote work with equipment
Health benefits
401(k) plan
+1
Director, Brand
Director, Brand

OnBoard Group • Northern (KY)

Hybrid
USD 140,000 - 200,000
Fully remote
Health plan
HSA available
+3
Marketing Operations Specialist New
Marketing Operations Specialist New

OnBoard Group • Northern (KY)

Hybrid
USD 65,000 - 90,000
Fully remote work with equipment
Health insurance options
401K retirement plan with companyMatch
+1
Manager I, Engineering - Cloud FinOps
Manager I, Engineering - Cloud FinOps

United States Digital Space LLC • New York (NY)

Hybrid
USD 200,000 - 250,000
Stock equity
ESPP
Professional development
+5
Senior Enablement Program Manager, TS
Senior Enablement Program Manager, TS

Datadog • New York (NY)

Hybrid
USD 116,000 - 155,000
Generous benefits
Stock options
Career development
+2