Production Operations Engineer

DryvIQ

United States

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A rapidly growing software company in the United States is seeking a Production Operations / SRE professional who will ensure reliability and security across hybrid environments. The role demands at least 5 years of experience in DevOps or SRE, with proficiency in Kubernetes, Helm, and Docker. Responsibilities include deployment automation and monitoring. Competitive pay range is $120,000 - $160,000 annually, and the position supports a hybrid work model.

Qualifications

  • 5+ years in SRE, DevOps, or Production Ops roles supporting hybrid or on-prem software delivery.
  • Minimum 3 years working with Fortune 500 companies implementing or maintaining enterprise software.
  • Expertise with Kubernetes, Helm, and Docker in mixed cloud environments (Azure AKS, AWS EKS, on-prem K3s).
  • Solid understanding of network security and Linux administration.
  • Strong scripting and automation skills, especially related to K8s.
  • Familiarity with CI/CD pipelines.

Responsibilities

  • Understand, deploy and maintain Helm charts and CI/CD workflows.
  • Monitor uptime, capacity, and performance across distributed clusters.
  • Participate in release readiness and hardening cycles.
  • Integrate static/dynamic security scanning and image-signing pipelines.
  • Extend monitoring to include customer-facing telemetry.

Skills

Kubernetes
Helm
Docker
Bash
Python
PowerShell
YAML / Terraform
Network security
Linux administration

Tools

CI/CD pipelines (GitHub Actions, TeamCity, Argo CD, Flux)
Prometheus
Grafana
Azure Monitor
CloudWatch

Job description

Production Operations / SRE

Ensures reliability, security, and consistency across DryvIQ’s hybrid environments—spanning Azure, AWS, and on-prem customer installations (note: not a traditional multi‑tenant SaaS environment). This role bridges engineering and operations, owning deployment automation, monitoring, and incident response for mission‑critical data‑management workloads.

Base pay range

$120,000.00/yr - $160,000.00/yr

About the Company

DryvIQ is a rapidly growing, venture‑backed software company headquartered in the Ann Arbor tech cluster with a 90% remote workforce across all U.S. time zones. We help enterprises safeguard their most sensitive documents and content through intelligent, data‑driven visibility and synchronization.

We value curiosity, technical excellence, and collaboration. Our culture is strictly merit-based — We believe in recognizing and rewarding contributions based on impact and outcomes. Our culture values collaboration, initiative, and growth over politics or tenure, creating an environment where everyone has a fair chance to succeed.

We also embrace pragmatic AI adoption: we use AI tools (including AI‑assisted code generation) to speed up development and improve quality, but we are not zealots about it. We have no intention of replacing humans and believe the best solutions come from human creativity, experience, and judgment.

Key Responsibilities
  • Deployment & Automation
    • Understand, deploy and maintain Helm charts, and CI/CD workflows for AKS, EKS, and on‑prem Kubernetes (K3s or RKE2) in customer environments.
    • Standardize customer deployments (private cloud / air‑gapped) using reproducible manifests and configuration validation tooling.
    • Maintain our single‑node and multi‑node install processes; improve installer packaging.
  • Environment Reliability
    • Monitor uptime, capacity, and performance across distributed clusters (migration, scan, OLAP DB node groups).
    • Implement proactive alerting (Prometheus, Grafana, Azure Monitor, CloudWatch) and ensure runbooks exist for all major services.
    • Coordinate with customer IT/security teams to handle firewall, proxy, and credential configurations safely and consistently.
  • Release & Incident Management
    • Participate in release‑readiness and hardening cycles; validate new images and helm charts before customer rollout.
    • Lead incident response for production issues—triage, communicate status, and drive post‑incident reviews and root‑cause documentation.
    • Track reliability metrics (MTTR, deployment success rate, change‑failure rate) and feed insights back into engineering planning.
  • Security & Compliance
    • Integrate static/dynamic security scanning (GitHub Advanced Security / CodeQL / Dependabot) and image‑signing pipelines.
    • Ensure secrets, credentials, and certificates are rotated and stored per corporate security standards.
    • Support ISO / SOC2 audit evidence collection (CCR change control, deployment logs, access reviews).
  • Tooling & Observability
    • Extend monitoring to include customer‑facing telemetry where allowed; maintain log shipping and retention policies.
    • Contribute to internal dashboards showing environment health, install duration, and customer success metrics.
  • Work closely with Dev/QA/Support to reproduce issues in controlled environments and publish fixes or workarounds.
  • Provide training and documentation for Services and Support engineers deploying or maintaining on‑prem instances.
  • Champion “build‑to‑run” culture—drive automation, resiliency testing, and feedback loops between engineering and field ops.
Required Experience
  • 5+ years in SRE, DevOps, or Production Ops roles supporting hybrid or on‑prem software delivery.
  • Minimum 3 years working with Fortune 500 companies implementing or maintaining enterprise software.
  • Expertise with Kubernetes, Helm, and Docker in mixed cloud environments (Azure AKS, AWS EKS, on‑prem K3s).
  • Solid understanding of network security (proxies, TLS, VPN, firewalls) and Linux administration.
  • Strong scripting and automation skills (Bash, Python, PowerShell, YAML / Terraform) especially as it relates to K8s.
  • Familiarity with CI/CD pipelines (GitHub Actions, TeamCity, Argo CD or Flux).
  • Comfort working directly with enterprise customer admins and security teams.
Success Indicators (How you will be measured)
  • 95 %+ installation success on first attempt across customer environments
  • Measurable reduction in install/upgrade time and CRI (Customer Raised Issues) related to configuration or infrastructure
  • Clear, actionable runbooks for all critical services
  • Improved observability and automation coverage across all deployment models
Seniority level

Mid‑Senior level

Employment type

Full‑time

Job function

Management and Manufacturing

Industries

Software Development

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

HTEC Group • United States

Hybrid
USD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Director of Software Engineering (Infrastructure)
Director of Software Engineering (Infrastructure)

ServiceTitan • United States

On-site
USD 180,000 - 240,000
Flexible time off
Fully paid medical, dental, and vision
HSA, FSA, dependent care FSA
+6
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Platform Engineer
Platform Engineer

Relativity Space • Long Beach (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision coverage
401(k)
Generous parental leave
+1
Manager, DevOps Engineering
Manager, DevOps Engineering

Plexusworldwidellc • Scottsdale (AZ)

On-site
USD 120,000 - 150,000
Competitive medical plans
401(k) with company match
Personalized health coaching
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
DevOps Engineer
DevOps Engineer

Union Technologies • Dallas (TX)

On-site
USD 110,000 - 140,000