Senior Platform Operations Engineer, Infrastucture

Viome

Bellevue (WA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Viome is seeking a Senior Platform Operations Engineer, Infrastructure to own our Azure-based service platform and simplify it through robust Unix-based systems engineering.

The role consolidates services onto Linux VMs with systemd supervision, Apache/HAProxy/Nginx routing, and scripted deployment with rollback. Strong emphasis on immutable artifacts and clean retirement of stale resources.

Qualifications

  • 7+ years operating production Unix/Linux systems with systemd and Apache/HAProxy.
  • Strong shell plus one scripting language with deploy/rollback tooling experience.
  • Ability to recover undocumented services and map real dependencies.
  • Experience with message queues: Azure Service Bus, SQS, RabbitMQ.
  • PostgreSQL/MySQL ops, replicas, access control, diagnostics.
  • Kubernetes and cloud networking to operate AKS clusters and Azure firewall/DNS.
  • Migration experience with zero-downtime cutover and rollback tests.
  • Track record of reducing surface area by consolidating services.

Responsibilities

  • Plan and execute zero-downtime migrations from AKS/containers to VM-based hosting.
  • Build deployment scripts for the target platform (bash-based) from build to rollback.
  • Maintain host hygiene: patching, log rotation, and centralized logging.
  • Manage external integrations, webhook delivery, and idempotency during migrations.
  • Operate networking, load balancing, DNS, and security during the transition.
  • Keep observability stack healthy (OpenTelemetry, ELK, Grafana) for the new platform.
  • Coordinate cross-cloud dependencies with AWS for DNS/edge routing and egress allowlists.

Skills

Unix/Linux operations
Systemd
Shell scripting
Reverse engineering
Kubernetes
Cloud networking
DNS and routing
Observability tooling
Network security basics

Tools

Apache
HAProxy
Nginx
systemd
Nagios
OpenTelemetry
ELK
Grafana
Azure Service Bus
SQS
RabbitMQ
Terraform
Kubernetes (AKS)
Istio
Kyverno

Job description

At Viome, we are driven by a singular mission: to help people live a healthy, disease-free life. This mission guides our actions, fuels our passion, and shapes the impact we aim to have in the world. Our core values - Be Bold, Be Collaborative, Be Frugal, and Grow Continuously - underpin our approach to achieving this goal. If you are motivated by the idea of working in an environment that prioritizes bold innovation, teamwork, efficient resource use, and continuous learning, all towards promoting health and preventing disease, join us in our journey to transform lives and create a healthier future for all.

We are looking for a Senior Platform Operations Engineer, Infrastructure to take ownership of Viome’s Azure-based service platform with a clear mandate: make it dramatically simpler through robust Unix-based systems engineering.

This is a consolidation role for a seasoned systems administrator. The work involves migrating a suite of abstracted services onto a deliberately stable target platform: Linux VMs, systemd-supervised services, Apache, HAProxy, and Nginx routing. You must possess strong Unix knowledge to operate and eventually transform the current stack safely. We are looking for a traditional operations background where success is measured by the removal of unnecessary layers and a genuine bias toward simplicity and host-level stability.

Simplification & Migration (core mandate)
  • Plan and execute zero-downtime migrations of services from AKS/containers to VM-based hosting: systemd unit authoring and service supervision (restart policy, resource limits, sandboxing), Apache, HAProxy and/or nginx as reverse proxy and TLS terminator, certificate automation.
  • Build and own the deployment scripts for the target platform: build → test → archive → ship → symlink-flip → health-check → rollback, scripted in bash.
  • Preserve two non-negotiable invariants while simplifying everything else: immutable, commit-traceable artifacts and scripted rollback.
  • Inventory and retire stale infrastructure: dormant deployments, unused DNS records, orphaned firewall rules, unpinned image tags.
  • Maintain host hygiene for consolidated services: OS patching discipline, runtime vendoring, log rotation, centralized log aggregation.
  • Familiarity with UptimeKuma, Nagios or similar.
  • Operate and migrate data stores: PostgreSQL and/or MySQL.
Network and Infrastructure Operations
  • Operate and troubleshoot workloads during the transition, focusing on the networking layer, load balancing, and core platform services.
  • Administer the hub network: Azure Firewall rules, VPN gateways, VNet peering, public and private DNS zones and reason about a packet’s full path from public IP to service.
  • Support the existing release process and network and infrastructure operations until each service is migrated to the new VM-based standard.
  • Keep the observability stack healthy (OpenTelemetry, ELK, Grafana, uptime and cost monitoring) and carry its essentials forward to the simplified platform.
  • Coordinate cross-cloud dependencies with AWS: DNS/edge routing, queue consumers, and egress IP allowlists.
  • L2+ operations support; manage runbooks for external L0, L1 support.
External Integrations & Security
  • Own the external integrations most at risk during migration: e-commerce and subscription platforms, messaging/notification providers, and clinical/health-data partners — webhook delivery, signature verification, idempotency, and retry semantics.
  • Raise the security baseline as you consolidate: secrets management, webhook authentication, least-privilege network access, and data-retention hygiene. Findings from an internal review are ready for you to remediate; the instinct to spot and close this class of issue — and not create more — is part of the job.
Required
  • 7+ years operating production Unix/Linux systems, with deep knowledge of systemd, process supervision; Apache, HAProxy
  • Strong shell plus one scripting language (bash/Python/PHP) with a track record of building deploy and rollback tooling, not just using it.
  • Demonstrated reverse-engineering ability: taking ownership of an undocumented production service and recovering its real dependencies and failure modes.
  • Message-queue literacy: Azure Service Bus, SQS, RabbitMQ, or equivalent — ordering, lock/ack semantics, dead-letter handling.
  • PostgreSQL operations: replicas/clones, connection proxying, network-restricted access, production diagnostics.
  • Working proficiency with Kubernetes and cloud networking — enough to operate private AKS clusters, an Istio-style ingress layer, and Azure firewall/DNS during the transition. We will onboard you on our specifics; deep specialization is not required.
  • Migration experience: consolidating or re-platforming production services with zero-downtime cutover and tested rollback.
  • A demonstrable record of reducing system surface area — services consolidated, infrastructure retired.
  • Clear written communication for runbooks, migration plans, and cross-functional coordination.
Strongly Preferred
  • Azure networking at landing-zone depth: hub-spoke VNets, Azure Firewall, Private Link/Private DNS, VPN gateways.
  • Administration of external-dns, cert-manager, and policy engines such as Kyverno.
  • ELK and OpenTelemetry pipeline operations.
  • Webhook-heavy integration experience (e-commerce platforms such as Shopify, subscription billing, marketing/notification platforms, or healthcare data exchanges).
  • Experience in a regulated or health-data environment: PHI handling, audit trails, least-privilege network design.
  • Terraform or equivalent IaC for cloud network and compute resources.
  • Enough AWS to manage cross-cloud seams (Route53, SQS, egress allowlists).
  • AI minded; proficient with Claude or similar.
  • First 30 days: Trace and fix a production issue end-to-end (DNS → firewall → ingress → service → database) with guidance; produce a true-dependency inventory for one candidate service.
  • First 90 days: First service migrated off AKS to the VM/systemd platform with a passing rollback drill; firewall rules, DNS zones, and cluster add-ons documented and reproducible; stale-resource inventory complete and retirement underway.
  • First 6 months: Migration cadence established with multiple services consolidated; release and rollback drills routine on both platforms; measurable reduction in infrastructure footprint and spend; security-baseline remediations closed.

We are an equal opportunity employer and do not discriminate based on race, color, religion, sex, national origin, age, disability, or any other legally protected status. We value diversity, equity, and inclusion, and are committed to creating a workplace that reflects these values.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Operations Engineer, Infrastucture
Senior Platform Operations Engineer, Infrastucture

Viome Life Sciences • Bellevue (WA)

On-site
USD 130,000 - 150,000
Senior Platform Engineer - Zero-Downtime Infra Migrations
Senior Platform Engineer - Zero-Downtime Infra Migrations

Viome Life Sciences • Bellevue (WA)

On-site
USD 130,000 - 150,000
Senior DevOps Engineer (Azure)
Senior DevOps Engineer (Azure)

Motion Recruitment Partners LLC • Chelmsford (MA)

Hybrid
USD 120,000 - 180,000
Hybrid work schedule
Professional development opportunities
401(k) match
VP, Product Platform & Engineering
VP, Product Platform & Engineering

Verramobility • Mesa (AZ)

On-site
USD 150,000 - 210,000
Senior Engineer, Platform Infrastructure
Senior Engineer, Platform Infrastructure

Vts • New York (NY)

On-site
USD 160,000 - 200,000
Competitive salary
Comprehensive health benefits
401(k) plan
+2
Senior Platform Engineer (Cloud Workloads)
Senior Platform Engineer (Cloud Workloads)

Veeam Software • San Jose (CA)

On-site
USD 158,000 - 295,000
Unlimited paid time off
Paid parental leave
Medical, dental, and vision coverage
+2
Infrastructure and Platform Team Lead
Infrastructure and Platform Team Lead

VectorVMS • Raleigh (NC)

Hybrid
USD 140,000 - 190,000
Hybrid work model
Leadership scope in Infra/Platform
Senior Linux Systems Engineer
Senior Linux Systems Engineer

Veilant • Tysons (VA)

On-site
USD 100,000 - 130,000
Flexible PTO
401k match
Medical, dental, and vision insurance
+2
Director of Platform Engineering & Operations, Hands-on, Onsite in Charlotte, NC
Director of Platform Engineering & Operations, Hands-on, Onsite in Charlotte, NC

Gina’s Tech Jobs - IT Recruiting Agency • Charlotte (NC)

On-site
USD 130,000 - 160,000
Medical insurance
Dental
Vision
+2
Sr. Platform Infrastructure Engineer
Sr. Platform Infrastructure Engineer

Vantor Inc. • Westminster (CO)

On-site
USD 128,000 - 187,000
401(k) with company match
Mental health resources
Student loan repayment assistance
+2