Senior DevOps Engineer, Infrastructure & Reliability

Worth AI

Miami (FL)

Hybrid

USD 135,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health Plan (Medical, Dental & Vision)
401k Retirement Plan
Life Insurance
Flexible Paid Time Off
9 paid Holidays
Family Leave
Remote work option
Hybrid work (Orlando associates)
Free Food & Snacks
Wellness Resources

Job summary

Worth AI is seeking a Senior DevOps Engineer to join the Infrastructure team. You will write Terraform, tune Kubernetes workloads, and automate manual tasks, shipping infrastructure changes to production. You’ll join a small platform team with a defined roadmap and a strong voice in how work gets built.

You will own the Kubernetes platform, optimize CI/CD, and design secure networking, IAM, and backups, enabling faster, more reliable software delivery.

Qualifications

  • 8+ years in DevOps, SRE, or infrastructure engineering.
  • Proven experience operating production Kubernetes environments at scale.
  • Deep hands-on expertise with AWS infrastructure and cloud networking.
  • Strong experience building Terraform modules across large cloud environments.
  • Demonstrated ownership of CI/CD systems and improvement of DORA metrics.
  • Experience leading incident response and postmortems.
  • Understanding of distributed systems, event-driven architectures (Kafka), and PostgreSQL performance.
  • Ability to modernize legacy infra and reduce toil.
  • Track record taking scoped infra projects to production with minimal direction.
  • Proven ability to earn trust across teams and raise reliability standards.

Responsibilities

  • Write Terraform to standardize cloud provisioning and reduce drift.
  • Own and evolve Kubernetes platform (EKS or self-managed) for reliability.
  • Improve CI/CD pipelines to increase release velocity and confidence.
  • Design secure networking, IAM, and secrets management across environments.
  • Improve observability with metrics, logs, and tracing via DataDog.
  • Optimize cloud costs through rightsizing and autoscaling strategies.
  • Implement disaster recovery and multi-region resilience plans.
  • Refactor brittle infrastructure into automated, testable systems.
  • Introduce new infra tooling and drive adoption via docs and workshops.
  • Collaborate with engineering to remove CI/CD friction and improve deployments.

Skills

DevOps
Kubernetes
AWS
Terraform
CI/CD
DataDog
PostgreSQL
Kafka
Redis
Automation

Tools

Terraform
Kubernetes
GitHub Actions
DataDog
PostgreSQL
Kafka
Redis
Python
TypeScript

Job description

Worth AI, a leader in the computer software industry, is looking for a Senior DevOps Engineer to join our Infrastructure team with a singular mission: to make our systems faster, more reliable, and more resilient while making life dramatically easier for engineers shipping software.

This is a hands-on build role. You will spend most of your time writing Terraform, tuning Kubernetes workloads, automating things that are currently manual, and shipping infrastructure changes to production. You'll join a small platform team with an established roadmap and existing patterns, and a strong voice in how the work gets built.

  • Implement scalable Infrastructure-as-Code patterns using tools like Terraform to standardize cloud provisioning and reduce configuration drift.
  • Own and evolve our Kubernetes platform (EKS or self-managed), ensuring workloads are secure, scalable, and resilient by default.
  • Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase confidence in releases.
  • Design and enforce secure networking, IAM, and secrets management strategies across environments.
  • Improve observability by refining metrics, logs, and tracing using tools like DataDog, ensuring actionable insight into system health.
  • Optimize cloud cost efficiency through rightsizing, autoscaling strategies, and architectural improvements.
  • Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives.
  • Refactor brittle or manually managed infrastructure into automated, testable, and reproducible systems.
  • Introduce new infrastructure tooling or architectural shifts and drive adoption through documentation, workshops, and hands-on support.
  • Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments.
  • Communicate technical trade-offs clearly across engineering and product stakeholders, balancing speed with safety.
Technology Stack
  • Cloud & Infrastructure: AWS (EKS, RDS, MSK, S3, Lambda, IAM, VPC)
    Containerization & Orchestration: Kubernetes, ArgoCD
    Infrastructure-as-Code: Terraform
    CI/CD: GitHub Actions
    Monitoring & Observability: DataDog
    Data & Messaging: PostgreSQL, Kafka, Redis
    Languages (as needed): Bash, Python, TypeScript, JavaScript
  • 8+ years in DevOps, SRE, or infrastructure engineering.
  • Proven experience designing and operating production Kubernetes environments at scale.
  • Deep hands-on expertise with AWS infrastructure and cloud networking.
  • Strong experience building and maintaining Terraform modules across large cloud environments.
  • Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics.
  • Experience leading incident response processes and driving meaningful postmortem outcomes.
  • Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).
  • Proven ability to modernize legacy infrastructure and eliminate manual operational toil.
  • Track record of taking a scoped infrastructure project from an ambiguous starting point to production without needing daily direction.
  • Demonstrated ability to build trust across teams while raising the reliability bar.
Success Metrics
  • System Reliability: Maintain or exceed defined SLO/SLA targets with reduced incident frequency and duration.
  • Infrastructure Stability: Reduce production incidents caused by misconfiguration, manual processes, or infrastructure drift.
  • Operational Efficiency: Increase the percentage of infrastructure managed through code and automation.
  • Cost Optimization: Improve cloud cost efficiency without sacrificing reliability or performance.
Bonus Points (Nice to Have)
  • Experience coding applications
  • Experience operating high-throughput Kafka clusters (MSK or self-managed).
  • Strong background in database performance tuning (PostgreSQL, Redis).
  • Experience implementing autoscaling strategies for high-traffic systems.
  • Familiarity with service mesh technologies.
  • Experience building internal developer platforms (IDP).
  • Background in security best practices (zero-trust networking, policy-as-code).
  • Experience with multi-region or globally distributed systems.
  • Experience introducing platform-wide reliability frameworks (SLOs, error budgets, chaos testing).

All Remote Hires will be required to travel to Orlando, Florida at least twice per year for Town Halls and team collaboration, in addition to orientation in Orlando.

  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k)
  • Life Insurance
  • Flexible Paid Time Off
  • 9 paid Holidays
  • Family Leave
  • Remote
  • Hybrid work (for Orlando Associates)
  • Free Food & Snacks (Orlando)
  • Wellness Resources
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer, Infrastructure & Reliability
Senior DevOps Engineer, Infrastructure & Reliability

Worth AI • Atlanta (GA)

Hybrid
USD 140,000 - 190,000
Health insurance
401k
Life Insurance
+7
Senior DevOps Engineer, Infrastructure & Reliability
Senior DevOps Engineer, Infrastructure & Reliability

Doist • Orlando (FL)

On-site
USD 150,000 - 210,000
Health Care Plan
Retirement Plan (401k)
Life Insurance
+7
Senior DevOps Engineer, Infrastructure & Reliability
Senior DevOps Engineer, Infrastructure & Reliability

Worth AI, Inc. • Orlando (FL), Tampa (FL), Miami (FL), Atlanta (GA)

Hybrid
USD 140,000 - 170,000
Health insurance
401k
Life Insurance
+7
Senior Software Engineer, Platform & Developer Experience
Senior Software Engineer, Platform & Developer Experience

Worth AI • Jacksonville (FL)

Hybrid
USD 140,000 - 180,000
Health Care Plan
Retirement Plan
Life Insurance
+3
Senior Software Engineer, Platform & Developer Experience
Senior Software Engineer, Platform & Developer Experience

Worth AI, Inc. • Orlando (FL)

Hybrid
USD 130,000 - 170,000
Health Care Plan (Medical, Dental &amp
Retirement Plan (401k, IRA)
Life Insurance
+5
Engineering Manager
Engineering Manager

Worth AI, Inc. • Orlando (FL)

Hybrid
USD 120,000 - 150,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Flexible Paid Time Off
+1
Senior DevOps Engineer: Remote Infra & Reliability
Senior DevOps Engineer: Remote Infra & Reliability

Worth AI, Inc. • Orlando (FL), Tampa (FL), Miami (FL), Atlanta (GA)

Hybrid
USD 140,000 - 170,000
Health insurance
401k
Life Insurance
+7
Engineering Manager
Engineering Manager

Worth AI • Miami (FL)

On-site
USD 150,000 - 210,000
Health Care Plan
401k / Retirement
Life Insurance
+7
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

N-iX • Town of Poland (NY)

Hybrid
USD 140,000 - 190,000
Flexible work format
Competitive salary
Career growth
+3
Senior Integrations Engineer - Remote API & Data Pipelines
Senior Integrations Engineer - Remote API & Data Pipelines

Worth Ai • Orlando (FL)

On-site