Senior DevOps Engineer, Reliability & Platform Operations

MedTrainer

United States

Remoto

MXN 781.000 - 1.004.000

Tempo pieno

26 ore fa
Candidati tra i primi
Generatore di candidature

Trasforma questa posizione in un colloquio — un curriculum e una lettera di presentazione creati in base a ciò questo datore di lavoro sta cercando.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

100% remote work
Major Medical Insurance
Home office stipend
English classes
Gym discounts
Savings plan
Paid time off

Descrizione del lavoro

MedTrainer, a leading healthcare compliance platform, seeks a Senior DevOps Engineer, Reliability & Platform Operations, to improve cloud infrastructure, AKS, CI/CD, databases, and middleware. You will work with DevOps, software delivery, security, and product teams to reduce production risk and raise observability.

You will provide design guidance, automation, and reliability standards while teams maintain application delivery readiness and implementation.

Competenze

  • 8+ years of experience in DevOps, SRE, platform engineering, or related roles.
  • Strong English communication skills.
  • Hands-on experience with Azure and production Kubernetes/AKS environments.
  • Experience designing or improving CI/CD, IaC, configuration management, observability, and deployment automation at scale.
  • Strong production troubleshooting across cloud infrastructure, Linux, containers, databases, and middleware.

Mansioni

  • Improve reliability engineering practices across production systems with SLIs/SLOs and postmortems.
  • Lead reliability improvements across cloud platforms, AKS, CI/CD pipelines, and databases.
  • Operate and improve AKS clusters, including upgrades and autoscaling.
  • Provide guidance on Kubernetes deployment patterns and best practices.
  • Enhance delivery reliability with GitHub Actions workflows and automated patterns.
  • Build automation and self-service capabilities to reduce toil and drift.
  • Maintain Infrastructure as Code and configuration management practices.
  • Own Ansible-based configuration management for OS and services.
  • Support Azure as primary platform; AWS/GCP experience preferred.
  • Improve observability across metrics, logs, and dashboards.
  • Support incident response with triage and postmortems.
  • Mentor engineers on reliability best practices and standards.

Conoscenze

DevOps
Site Reliability
Kubernetes AKS
GitHub Actions
Azure Cloud
Python
Linux
Automation
Security Best Practices

Formazione

Bachelor's degree in Computer Science or equivalent

Strumenti

Terraform
Pulumi (Python)
Ansible
New Relic
ProxySQL
RabbitMQ
MySQL

Descrizione del lavoro

Company Description

MedTrainer is the only all-in-one compliance platform purpose-built for healthcare organizations, by healthcare professionals. For over 13 years, we have helped healthcare teams remain audit-ready and confident in their regulatory compliance through continuous innovation and deep industry expertise.

MedTrainer is the only all-in-one compliance platform purpose-built for healthcare organizations, by healthcare professionals. For over 13 years, we have helped healthcare teams remain audit-ready and confident in their regulatory compliance through continuous innovation and deep industry expertise. Our cloud-based platform unifies learning management, credentialing, and compliance into a single intelligent system designed to simplify complex healthcare operations. Powered by automation and AI-driven workflows, MedTrainer enables organizations to onboard faster, streamline processes, and scale efficiently while maintaining the highest standards of compliance and workforce readiness.

Job Description

We are looking for a Senior DevOps Engineer, Reliability & Platform Operations to improve the reliability, operational stability, risk posture, and delivery effectiveness of our cloud infrastructure, Kubernetes platforms, CI/CD workflows, databases, middleware, and production systems.

In this role, you will work closely with DevOps Engineers / Software Delivery Engineers, Software Engineering, Security, Database providers, Support, Product, and leadership teams to reduce production risk, improve observability, strengthen operational standards, automate repetitive work, and enable safer paths to production.

You will provide design guidance, advanced technical support, automation, best practices, and reliability standards to DevOps Engineers / Software Delivery Engineers, while those teams retain ownership of application delivery readiness and implementation.

This is a senior technical individual contributor role for someone who has strong production ownership, deep infrastructure and platform experience, solid SRE practices, and the ability to proactively identify recurring problems and drive permanent improvements.

Responsibilities
  • Improve reliability engineering practices across production systems, including SLIs, SLOs, error budgets, postmortems, runbooks, service health indicators, and corrective-action tracking.
  • Lead reliability improvements across cloud platforms, AKS, CI/CD pipelines, application environments, databases, and supporting middleware.
  • Operate and improve AKS clusters, including upgrades, autoscaling, node pools, networking, storage, identity, security, reliability, and operational standards.
  • Provide guidance, best practices, and standards for Kubernetes application deployment patterns.
  • Improve delivery reliability through advanced GitHub Actions workflows, reusable automation, secure deployment patterns, rollback support, and pipeline optimization.
  • Build automation, self-service capabilities, reusable infrastructure patterns, and operational procedures to reduce toil, ticket-driven work, manual handoffs, and infrastructure drift.
  • Define, implement, and maintain Infrastructure as Code and configuration management practices using approved tools.
  • Own and maintain Ansible-based configuration management for OS, middleware, application-supporting services, and operational automation.
  • Support and improve Azure cloud environments as the primary platform, with AWS and GCP experience preferred.
  • Implement and improve observability across metrics, logs, traces, dashboards, alerting, APM, capacity planning, performance tuning, and incident dashboards.
  • Support incident response for production, deployment, infrastructure, database, middleware, and application reliability issues, including severity recommendation, triage coordination, stakeholder communication, mitigation support, postmortems, and corrective actions.
  • Operate with production awareness and ownership, supporting reliability improvements and production-risk reduction.
  • Recommend blocking, delaying, or rolling back releases when reliability, security, operational, or business-continuity risks are identified.
  • Understand PHP Symfony monolith behavior sufficiently to support production reliability and collaborate effectively with software engineering teams.
  • Operate, tune, monitor, back up, troubleshoot, and support MySQL, ProxySQL, and RabbitMQ directly, while coordinating with the database provider where required.
  • Strengthen security and compliance practices across infrastructure, CI/CD, and runtime environments, including secrets management, access control, dependency and image scanning, signing/provenance, CI/CD security controls, cloud security configuration, and audit support.
  • Improve backup, restore, disaster recovery, cloud cost visibility, cost control, and operational resilience across critical systems.
  • Mentor DevOps Engineers / Software Delivery Engineers through technical guidance, design review, standards definition, operational knowledge sharing, and reliability best practices.
  • Produce and maintain runbooks, troubleshooting guides, deployment procedures, reliability standards, architecture notes, and knowledge-base articles.
  • Propose technical initiatives related to reliability, platform operations, observability, automation, incident reduction, cloud maturity, and operational risk reduction.
Qualifications
  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or equivalent professional experience.
  • 8+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, cloud operations, infrastructure automation, production operations, or related roles
  • Strong written and verbal English communication skills.
  • Strong experience operating production systems where reliability, observability, incident response, automation, and operational risk reduction are core responsibilities.
  • Strong hands-on experience with Azure and production Kubernetes / AKS environments.
  • Experience designing or improving CI/CD, Infrastructure as Code, configuration management, observability, and deployment automation at scale.
  • Strong production troubleshooting skills across cloud infrastructure, Linux, containers, Kubernetes, databases, middleware, CI/CD pipelines, and application-supporting services.
  • Ability to understand application behavior and support production reliability of PHP Symfony applications.
  • Experience operating or supporting MySQL, ProxySQL, and RabbitMQ in production environments.
  • Strong understanding of secure delivery and infrastructure security practices, including secrets management, least privilege, access reviews, CI/CD security, and audit support.
  • Experience with backup, restore, disaster recovery, business continuity, capacity planning, performance tuning, and cost optimization.
  • Ability to mentor engineers, define technical standards, challenge weak practices constructively, and lead technical initiatives without formal people-management authority.
  • Fully remote work capability, including disciplined written communication, self-management, asynchronous collaboration, and reliable participation in distributed-team workflows.
Essential Technologies And/or Skills
  • Cloud platforms: Azure required; AWS and GCP preferred.
  • Kubernetes: AKS, cluster operations, troubleshooting, upgrades, autoscaling, networking, storage, identity, security, and platform standards.
  • CI/CD: GitHub Actions, reusable workflows, composite actions, environments, approvals, OIDC, self-hosted runners, rollback workflows, artifacts, caching, and pipeline security controls.
  • Infrastructure as Code: Terraform or Pulumi; Pulumi with Python preferred.
  • Configuration management: Ansible.
  • Scripting and automation: Python required; Bash preferred.
  • Containers: Docker / OCI and standalone Docker hosts.
  • Operating systems: Linux administration and troubleshooting.
  • Databases and middleware: MySQL, ProxySQL, RabbitMQ.
  • Observability: New Relic, Logz.io, Azure Metrics, ClickStack preferred, metrics, logs, traces, dashboards, APM, synthetic checks, and SLO-based alerting.
  • Reliability engineering: SLIs, SLOs, error budgets, postmortems, runbooks, toil reduction, incident response, and corrective actions.
  • Security: HashiCorp Vault, 1Password, secrets management, least privilege, access reviews, dependency scanning, container image scanning, signing/provenance, and CI/CD security controls.
  • Application reliability: PHP Symfony applications production-support context.
  • Operational resilience: backup, restore, disaster recovery, business continuity, capacity planning, and performance tuning.
  • Documentation: runbooks, troubleshooting guides, operational procedures, standards, and knowledge-base articles.
  • Working style: structured problem-solving, strong ownership, production awareness, mentoring, automation mindset, and proactive risk reduction.
Additional Information
  • 100% remote work from anywhere in Mexico.
  • Competitive monthly salary after taxes.
  • Major Medical Insurance and healthcare coverage.
  • Home office and ergonomics support, including internet and electricity.
  • Professional development opportunities, including English classes.
  • Wellness benefits such as TotalPass gym discounts.
  • Savings plan.
  • Paid time off, including personal days.
  • Collaborative, international, and growth-oriented environment.
  • Opportunity to work with modern cloud, automation, CI/CD, Kubernetes, reliability, observability, and platform technologies.
  • Collaborative engineering culture focused on automation, reliability, operational excellence, and continuous improvement.

All your information will be kept confidential according to EEO guidelines.

Salary range: $70,000 MXN - $90,000 MXN

At MedTrainer, every role contributes to building technology that supports safer, more efficient healthcare organizations. We value collaboration, ownership, and continuous improvement across all teams.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

DevOps/Cloud Engineer
DevOps/Cloud Engineer

AgileEngine • Stati Uniti

Remoto
MXN 900.000 - 1.300.000
100% remote work
Annual learning budget
Well-being programs
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Medallia • McLean (VA)

In loco
USD 129.000 - 190.000
Health benefits
401(k) matching
Paid parental leave
+1
Senior Manager, DevOps
Senior Manager, DevOps

Clinician Nexus • Stati Uniti

In loco
USD 133.000 - 222.000
Medical and dental coverage
401(k) and profit-sharing
Flexible spending accounts
+5
3510- Site Reliability Engineer II
3510- Site Reliability Engineer II

Innovaccer • Dallas (TX)

In loco
USD 110.000 - 140.000
Generous Paid Time Off: 22 days per year plus company holidays
Best-in-Class Parental Leave
Comprehensive insurance coverage
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

In loco
USD 165.000 - 215.000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Staff DevOps Engineer
Staff DevOps Engineer

Omnissa • California (MO)

Ibrido
USD 206.000 - 343.000
Health insurance
401k with matching
Paid time off
+2
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • Stati Uniti

Remoto
USD 165.000 - 215.000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+6
Senior Platform Security Engineer (100% Remote within Spain)
Senior Platform Security Engineer (100% Remote within Spain)

Docplanner • Stati Uniti

Ibrido
USD 104.000 - 150.000
Private healthcare
Gym access
English & Spanish classes
+6
Senior DevOps & Platform Reliability Engineer - Remote MX
Senior DevOps & Platform Reliability Engineer - Remote MX

MedTrainer • Stati Uniti

Remoto
MXN 781.000 - 1.004.000
100% remote work
Major Medical Insurance
Home office stipend
+4
Site Reliability Engineer
Site Reliability Engineer

Enzo Health • Lehi (UT)

In loco
USD 120.000 - 180.000
Competitive salary
Equity
401k & Insurance
+2