Domain
About Fold Health
Fold Health is building an AI-poweredhealthcare technology platform designed to support Value-Based Care (VBC)programs. By combining advanced data integration, analytics, and intelligentautomation, we help providers, payers, and care teams deliver better outcomesand improve patient experiences. Our mission is to simplify healthcaretechnology, ensure interoperability, and enable innovation at scale. Join usand be part of shaping the future of AI-driven healthcare.
Role Overview
We are looking for a DevOps Engineer who willbe responsible for building, maintaining, and scaling the cloud infrastructureand delivery pipelines that power Fold Health's healthcare platform. The idealcandidate is hands-on, reliability-focused, and comfortable working across AWS,GCP, Terraform, and Prometheus in a fast-paced healthcare environment.
Responsibilities
- Design,build, and maintain highly available, secure, and scalable cloudinfrastructure across AWS and GCP.
- Manageproduction environments including Amazon ECS Fargate, RDS PostgreSQL,ElastiCache, ALB, SQS/SNS, Lambda, Step Functions, Cognito, Route 53, andsupporting GCP services.
- Optimizecloud resources for performance, reliability, availability, and costefficiency.
- Design andmaintain networking, security groups, load balancing, and disasterrecovery capabilities.
2. Infrastructure as Code & Automation
- Developand maintain reusable Infrastructure as Code using Terraform followingindustry best practices.
- Build,enhance, and maintain CI/CD pipelines to enable reliable, secure, andautomated application deployments.
- Automateinfrastructure provisioning, operational tasks, and cloud governancethrough scripting and tooling.
- Standardizeinfrastructure components, deployment patterns, and operational workflowsacross environments.
3. Observability, Monitoring & Incident Management
- Build andmaintain comprehensive monitoring, logging, and alerting solutions usingCloudWatch, Prometheus, AlertManager, Grafana, and xMatters.
- Developactionable alerting strategies to reduce alert fatigue while ensuringrapid detection of production issues.
- Createdashboards, metrics, and operational insights to improve platform healthand service reliability.
- Participatein production incident response, root cause analysis, and post-incidentreviews, driving preventive improvements.
4. Reliability, Performance & Security
- Ensureplatform reliability, scalability, and performance through proactivecapacity planning, tuning, and optimization.
- Troubleshootcomplex production issues across infrastructure, networking, databases,and distributed applications.
- Implementsecurity best practices including IAM, secrets management, encryption,vulnerability remediation, and least-privilege access.
- Supportcompliance initiatives such as HIPAA, SOC 2, and HITRUST by implementingrequired technical controls and maintaining audit readiness.
- Partnerwith software engineering teams to improve application deployment,reliability, and operational excellence.
- AdvocateDevOps best practices, automation, and Infrastructure as Code throughoutthe engineering organization.
- Continuouslyevaluate and adopt new cloud technologies, tools, and processes to improveplatform efficiency and developer experience.
- Create andmaintain technical documentation, runbooks, and operational procedures tosupport knowledge sharing and onboarding.
Requirements
Requirements
- Bachelor'sdegree in Computer Science, Engineering, Information Technology, or a relatedfield.
- Minimum 5 years ofhands-on experience in DevOps, Platform Engineering, or Site ReliabilityEngineering.
- Strongexperience with AWS services such as ECS, RDS, IAM, CloudWatch, ALB, Lambda,and networking.
- Hands-onexperience with Infrastructure as Code using Terraform and CI/CD pipelinesusing GitLab CI/CD or similar tools.
- Strongexperience with Docker and containerized application deployments.
- Workingknowledge of Kubernetes, including deployments, services, configuration,scaling, and troubleshooting.
- Familiaritywith PostgreSQL, Redis or ElastiCache, and microservices architectures.
- Proficiencyin Linux administration, Python or Bash scripting, and infrastructureautomation.
- Experiencewith monitoring and observability tools such as Prometheus, Grafana,CloudWatch, and AlertManager.
- Goodunderstanding of cloud security, IAM, networking, secrets management, andproduction troubleshooting.
- Strongproblem-solving, communication, and collaboration skills.
Good to Have
- Experiencewith Google Cloud Platform services.
- Experiencemanaging Kubernetes workloads in production using EKS, GKE, or similarplatforms.
- Experiencewith on-call management tools such as xMatters or PagerDuty.
- Knowledge ofhealthcare compliance standards such as HIPAA, HITRUST, or SOC 2.
Why Join Us?
- Opportunity to build and scalean AI-powered healthcare platform impacting millions of lives
- Work with modern cloudtechnologies in a mission-driven environment
- Collaborative and innovativeculture with strong growth opportunities