A healthcare technology company in Miami is seeking a Senior Software Engineer who will focus on operating Kubernetes and AWS services, designing observability systems, and improving service reliability. The role involves leading incident response and mentoring engineers in best practices. The position offers a robust benefits package, including health care, retirement plans, and paid time off.
Qualifications
Strong experience operating Kubernetes and cloud-native infrastructure in production environments.
Proficiency in AWS services, including networking and logging/monitoring tools.
Skilled in Terraform and Infrastructure as Code practices.
Deep understanding of observability tooling and incident management workflows.
Responsibilities
Design and implement robust monitoring, alerting, and observability systems.
Lead reliability reviews and incident response.
Improve service scalability and performance through architectural input.
Build and maintain automation for infrastructure management.
Skills
Kubernetes
AWS services
Terraform
Observability tooling
Coding skills
Job description
Overview
Job Title: Senior Software Engineer
Department: Engineering
Location: Miami, FL
Reports to:
Responsibilities
Design and implement robust monitoring, alerting, and observability systems across all services and infrastructure
Lead reliability reviews, incident response, and post-incident analysis—focusing on prevention, learning, and long-term improvements
Improve service scalability, fault tolerance, and performance through architectural input and systems optimisation
Build and maintain automation for infrastructure management using Terraform, and delivery pipelines using GitHub Actions
Partner with software engineers to improve the operational readiness and resilience of services, including capacity planning and runbooks
Lead initiatives to reduce operational toil through tooling, automation, and process improvement
Manage and optimise production Kubernetes and AWS environments with a focus on reliability, security, and cost-effectiveness
Contribute to security hardening efforts, including network controls, secrets management, and compliance readiness
Participate in and lead in-person stand-ups, incident reviews, and cross-team planning sessions
Share knowledge and mentor engineers on best practices in observability, incident response, and operational engineering
Qualifications
Technical Skills (Essential)
Strong experience operating Kubernetes and cloud-native infrastructure (preferably EKS on AWS) in production environments
Proficiency in AWS services, including networking, compute, IAM, and logging/monitoring tools (e.g. CloudWatch, ELB, VPC)
Skilled in Terraform and Infrastructure as Code practices
Deep understanding of observability tooling (metrics, logs, tracing) and incident management workflows
Strong coding skills for building tools, scripts, and automation
Ability to troubleshoot complex infrastructure issues and lead delivery of reliable cloud solutions
Preferred
Experience implementing SLAs, SLOs, and error budgets to guide operational priorities
Background in healthcare or other regulated industries with security and compliance requirements