Site Reliability Engineer II

Kalmbach Feeds

Bellaire Gardens (OH)

On-site

USD 120,000 - 170,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Kalmbach Feeds is building a small SRE team to shape the next generation of our hybrid platform, spanning company data centers and Microsoft Azure. You’ll own architecture decisions and drive standards as we deploy production RKE2 clusters with Rancher, AKS, and disaster‑recovery plans.

You’ll work hands‑on from rack to cloud, debugging outages, improving observability, and partnering with infrastructure teams. Bring practical curiosity and a bias toward automation.

Qualifications

  • 3+ years in SRE, platform, systems engineering, or DevOps with hands-on Kubernetes production experience.
  • Hands-on data-center, servers, virtualization, storage, and networking experience.
  • Strong Linux fundamentals; TCP/IP, DNS, TLS are essential.
  • Experience with monitoring, logging, alerting, and incident response; calm during outages.
  • Scripting in Bash, Python, or Go; Git; familiarity with CI/CD, GitOps, or IaC.

Responsibilities

  • Operate production and non-production Kubernetes clusters, including upgrades, node lifecycle, networking, ingress, DNS, certificates, and capacity.
  • Troubleshoot control‑plane, scheduling, CoreDNS, and memory issues; improve observability and runbooks.
  • Support data-center hardware, virtualization, storage, and networking; collaborate with infra teams on networking and storage.
  • Assist with disaster recovery, backups, and connectivity between data centers and cloud; document DR plans.
  • Join on-call rotation and reduce toil through automation and IaC.

Skills

Kubernetes ops
SRE / platform eng
Linux fundamentals
Networking DNS TLS
Automation / IaC
Scripting Bash Python Go
Git & CI/CD

Tools

RKE2/Rancher
AKS
Terraform
Ansible

Job description

Join a small SRE team at the start of an exciting build: we are shaping the next generation of our hybrid platform across company data centers and Microsoft Azure—not simply maintaining someone else’s established setup. You’ll have direct ownership and a real voice in the architecture, standards, and technologies we put into production. The platform will include production RKE2 clusters managed with Rancher, Azure Kubernetes Service (AKS), hybrid workloads, and disaster recovery, while storage, networking, and data-platform designs are still open to influence. This hands‑on role spans racks to cloud, and what you help design will become what the company runs. You need not know every tool on day one, but should learn quickly and bring a practical, curious approach.

What You’ll Do
  • Operate production and non-production RKE2/Rancher Kubernetes clusters, including upgrades, node lifecycle, networking, ingress, DNS, certificates, and capacity; troubleshoot control‑plane, scheduling, CoreDNS, and memory (OOM) issues, and improve observability, alerting, and runbooks.
  • Run Kubernetes on physical HPE servers and virtual machines in company data centers, including hardware, firmware, RAID, and out‑of‑band management; partner with infrastructure teams on Cisco networking, SAN/NVMe/object storage, failure domains, and capacity planning.
  • Support AKS, Azure Container Registry (ACR), and connectivity between data centers and Azure. Implement, test, and document disaster‑recovery plans, and verify that backups can be restored.
  • Join the on‑call rotation; investigate and respond to incidents, escalating to the Staff SRE when appropriate. Reduce recurring toil through automation, infrastructure as code (IaC), safer CI/CD, SLIs/SLOs, and attention to single points of failure.
What We’re Looking For
  • 3+ years in SRE, platform, systems engineering, or DevOps, with hands‑on experience operating and troubleshooting Kubernetes in production.
  • Hands‑on data‑center, colocation, or equivalent experience with servers, virtualization, storage, and networking.
  • Strong Linux fundamentals and working knowledge of TCP/IP, DNS, and TLS.
  • Experience with monitoring, logging, alerting, and incident response; clear communication and a calm, methodical approach during outages.
  • Scripting experience in Bash, Python, or Go; Git experience; and familiarity with CI/CD, GitOps, or IaC practices using any toolset.

Helpful but not required:RKE2/Rancher; AKS in a hybrid environment; Kyverno/OPA; 25/100GbE or Cisco Nexus; SAN/NVMe/object storage; stateful data platforms on Kubernetes; GPU/AI workloads; GitOps pipelines; Terraform or Ansible.

Why Join Us

Take meaningful ownership of a platform being built for its next chapter. You’ll help shape it from the ground up, influence foundational decisions, and see your work become the systems the company relies on—from physical servers to cloud. If you want to build, improve, and own real infrastructure rather than inherit a ticket queue, this is your opportunity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II - Build a Hybrid Kubernetes Platform
Site Reliability Engineer II - Build a Hybrid Kubernetes Platform

Kalmbach Feeds • Bellaire Gardens (OH)

On-site
USD 120,000 - 170,000
Infrastructure/Cloud DevOps - SRE
Infrastructure/Cloud DevOps - SRE

Bayside Solutions • Cupertino (CA)

On-site
USD 150,000 - 230,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer
Site Reliability Engineer

Veritas Search Group • Tustin (CA)

On-site
USD 140,000 - 190,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000