Senior Site Reliability Engineer

Carta Healthcare

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Carta Healthcare seeks a Senior Site Reliability Engineer to build and scale internal platform services, ensuring reliability and performance for applications. You will design monitoring and incident response, collaborate with software engineers to scale systems, and drive gradual improvements as the company grows globally.

The role emphasizes Python development, cloud exposure across AWS/GCP/Azure, and modern IaC and CI/CD practices.

Qualifications

  • Cloud platforms such as AWS, GCP or Azure with services like EC2, S3, RDS, Lambda; Kubernetes or similar is preferred.
  • IaC with Terraform, Ansible or CloudFormation for provisioning cloud infra.
  • Networking concepts incl. CNI, network policies; proxies and service mesh are a plus.
  • Monitoring/observability using Prometheus, Grafana, ELK or Datadog; set up and maintain monitoring.
  • Software development in Python with writing clean, scalable code.
  • API services design and maintenance; REST and/or GraphQL principles.
  • AI fluency; use of AI tools to reduce toil and build agents.
  • CI/CD experience is beneficial.

Responsibilities

  • Build and scale internal platform offerings (compute, storage, networking) for reliability and performance.
  • Design and implement monitoring, alerting, and incident response systems.
  • Collaborate with application engineers to ensure scalable designs.
  • Act as agent of change to incrementally improve systems as Carta expands globally.

Skills

Cloud platforms
IaC
Networking
Monitoring/Observability
Python
API services
AI fluency
CI/CD

Tools

Terraform
Ansible
CloudFormation
Kubernetes
Prometheus
Grafana
ELK Stack
Datadog
Python

Job description

The Problems You’ll Solve

At Carta, our employees set out on a mission to unlock the power of equity ownership for more people in more places. We believe that the problems we solve today unlock the opportunities of tomorrow. As a Senior Site Reliability Engineer, you’ll work to:

  • Build and scale our internal platform offerings (compute, storage and networking services) to ensure the reliability, and performance of our applications.
  • Design and implement monitoring, alerting, and incident response systems.
  • Collaborate with application software engineers (as needed) to guide their design and ensure it scales for what Carta needs in the long run.
  • Act as an agent of change and push boundaries to incrementally improve our systems as we expand globally.
The Team You’ll Work With

You’ll be joining the Infrastructure Engineering team at Carta. The Infrastructure Engineering team is responsible for providing secure, reliable, scalable and performant infrastructure to Carta’s customers and developers.

We are Software and Infrastructure Engineers who specialize in cloud computing, networking, systems design and architecture, storage, real time data telemetry, associated automation, tooling and processes. We possess a breadth and depth of knowledge about Carta’s infrastructure and industry wide best practices, that translates into leverage for Carta’s business.

About You

You are excited by the idea of developing scalable, reliable and efficient infrastructure that powers the entire company. We’re looking for strong communicators who enjoy collaborating to solve complex problems. Familiarity with infrastructure best practices on performance, reliability and security and their associated tools is appreciated.

Our stack is Python, Java, Terraform, gRPC, Docker, Kubernetes, Postgres, running on AWS. Come join us!

Qualifications
  • Cloud Platforms: Extensive experience with cloud services such as AWS, Google Cloud Platform, or Azure, including services like EC2, S3, RDS, and Lambda. Experience with Kubernetes or other container orchestration is preferred!
  • Infrastructure as Code (IaC): Proficient in using tools such as Terraform, Ansible, or CloudFormation for managing and provisioning cloud infrastructure.
  • Networking: Experience with networking concepts and tools, including Container Network Interface (CNI), Network policy implementations. Experience with proxies and service mesh is a big plus.
  • Monitoring and Observability: Strong knowledge of monitoring tools and practices, such as Prometheus, Grafana, ELK Stack, or Datadog, and the ability to set up and maintain comprehensive monitoring solutions.
  • Software Development: Proficiency in Python, with the ability to write efficient, maintainable, and scalable code.
  • API Services: Experience in designing, deploying, and maintaining API services, with a strong understanding of RESTful and/or GraphQL API design principles.
  • AI Fluency: You use AI tools in your own day-to-day work in addition to enabling others. You're comfortable building agents to reduce toil and expect this to be a normal part of how you operate.
  • Experience operating CI/CD and its associated best practices is also appreciated though not essential.
Disclosures

We are an equal opportunity employer and are committed to providing a positive interview experience for every candidate. If accommodations due to a disability or medical condition are needed, please connect with the talent partner via email.

Carta uses E-Verify in the United States for employment authorization. See the E-Verify and Department of Justice websites for more details.

For information on our data privacy policies, see Privacy, CA Candidate Privacy, and Brazil Transparency Report.

Please note that all official communications from us will come from an @carta.com or @carta-external.com domain. Report any contact from unapproved domains to security@carta.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer London, England, United Kingdom
Senior Site Reliability Engineer London, England, United Kingdom

Carta, Inc. • Greater London

On-site
GBP 95,000 - 135,000
Senior SRE — Global Platform & Observability
Senior SRE — Global Platform & Observability

Carta, Inc. • Greater London

On-site
GBP 95,000 - 135,000
Senior Cloud Reliability & Platform Engineer
Senior Cloud Reliability & Platform Engineer

Carta Healthcare • Greater London

On-site
GBP 90,000 - 130,000
Senior Software Engineer II London, London, United Kingdom
Senior Software Engineer II London, London, United Kingdom

Carta, Inc. • Greater London

Hybrid
GBP 90,000 - 130,000
Customer Success Manager
Customer Success Manager

Carta • Greater London

On-site
GBP 65,000 - 90,000
Paralegal, Compliance
Paralegal, Compliance

Carta Healthcare • Greater London

On-site
GBP 28,000 - 40,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Customer Success Manager
Customer Success Manager

Carta Healthcare • Greater London

Hybrid
GBP 75,000 - 115,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Symphony • Belfast City District

On-site
GBP 60,000 - 70,000
Regional specific competitive benefits
Build your own Benefits (BYOB) perk
Local events, team building, and devop
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Boston Consulting Group (BCG) • Greater London

On-site
GBP 80,000 - 120,000