Get more replies from employers
Send a job-specific resume in minutes.
Highbrow LLC is looking for an experienced SRE DevOps Engineer based in Overland Park, KS. The ideal candidate should have 4–9 years of experience in SRE or DevOps, with strong expertise in Kubernetes, incident troubleshooting, and automation.
The role involves resolving technical incidents, improving tooling, and mentoring junior engineers. Candidates with additional skills in cloud platforms and security will be preferred.
Location: Overland Park, KS / Atlanta, GA / Frisco, TX (Onsite)
Advanced Incident Troubleshooting & Resolution
Expectation: Diagnose and resolve escalated incidents that L1 cannot handle, often across multiple layers (infrastructure, application, network).
Example: For an API outage, identify if the root cause is in Kubernetes pod networking, APIgateway misconfig, or backend DB latency — and apply fixes.
Kubernetes & Container Orchestration Expertise
Expectation: Comfortable with deployments, scaling, networking, and debugging cluster‑level issues.
Example: Troubleshoot why pods are pending by checking node capacity, taints/tolerations, and cluster autoscaler logs.
Automation & Scripting (Python, Go, Bash, Ansible, Terraform)
Expectation: Write scripts and automation to reduce manual toil, enhance monitoring, and improve incident resolution speed.
Example: Develop a Python script to automatically collect pod and system logs when a service crashes.
Observability & Monitoring Tooling
Expectation: Deep understanding of monitoring, alerting, tracing, and logging systems.
Example: Build Prometheus alert rules to detect DB query spikes; configure Grafana dashboards for API latency.
CI/CD & Infrastructure as Code (IaC)
Expectation: Familiarity with GitOps workflows, CI/CD pipelines, and infrastructure provisioning.
Example: Enhance Jenkins pipeline to add automated smoke tests before promoting Kubernetes deployments.
Database Troubleshooting (SQL & NoSQL)
Expectation: Identify performance bottlenecks, connection issues, and basic tuning opportunities.
Example: Run queries to detect slow‑running SQL statements causing latency in an application.
Incident Management & RCA
Expectation: Act as incident commander for escalated issues, lead bridge calls, and produce Root Cause Analyses.
Example: After a WAF misconfiguration causes downtime, lead the investigation, document the timeline, and propose preventive actions.
Mentorship & Runbook Improvement
Expectation: Coach L1 engineers, refine runbooks, and introduce new automated workflows.
Example: Update a runbook to add automated Kubernetes log collection instead of manual steps.
Cloud Platform Engineering (AWS, Azure, GCP)
Expectation: Hands‑on skills in provisioning, scaling, and securing cloud workloads.
Example: Diagnose why an AWS ALB is misrouting traffic after a deployment.
Security & WAF Management
Expectation: Understand WAF rules, common attacks (SQL injection, XSS), and how to apply fixes.
Example: Investigate false positives in WAF logs and adjust rule sets with security teams.
Capacity & Performance Engineering
Expectation: Anticipate scaling needs, tune resource utilization, and propose optimizations.
Example: Identify that a Kubernetes deployment is CPU‑throttled and adjust HPA (Horizontal Pod Autoscaler) configs.
Automation Platform Integration (AIOps, ChatOps)
Expectation: Integrate AI/ML‑powered tools for anomaly detection and auto‑remediation.
Example: Implement a ChatOps bot that runs predefined Kubernetes troubleshooting commands in Slack.
Cross‑Platform Expertise (Hybrid Infra)
Expectation: Experience supporting both on‑prem and cloud environments seamlessly.
Example: Compare latency patterns between on‑prem DBs and cloud‑hosted APIs to identify bottlenecks.