Senior Site Reliability Engineer - Infra Ops

Circle

San Francisco (CA)

On-site

USD 123,000 - 205,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Circle is seeking a Senior Site Reliability Engineer to design, build, and operate the scalable platform infrastructure behind critical digital assets, AI, and app workloads. You will write and maintain services, automate repeatable workflows, and advance Kubernetes-based platforms across hybrid and public-cloud environments.

You will collaborate with platform, product, and application teams to translate workload requirements into resilient designs, taking ownership of production outcomes and

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems.
  • Deep, hands-on Kubernetes expertise: designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale.
  • Strong Terraform experience, including authoring reusable modules, managing state and environments, and delivering infrastructure changes through reviewable, automated workflows.

Responsibilities

  • Design, build, and operate Kubernetes platforms that provide secure, highly available, and scalable foundations for production services across hybrid and public-cloud environments.
  • Build infrastructure as code with Terraform, creating reusable modules, safe delivery workflows, and well-governed infrastructure changes.
  • Develop backend services, internal tools, and operational automation in Go, Python, or JavaScript/TypeScript to eliminate manual work and improve the developer experience.
  • Partner with engineering and product teams to understand workload requirements and design pragmatic solutions for reliability, performance, capacity, security, and cost.
  • Improve the production lifecycle through reliable CI/CD, deployment automation, progressive delivery, and clear operational ownership.
  • Define and evolve observability practices across metrics, logs, traces, alerting, and dashboards so teams can detect issues early and troubleshoot effectively.
  • Own production reliability by participating in on-call, incident response, root-cause analysis, and durable corrective actions.
  • Establish and maintain reliability targets through SLIs, SLOs, error budgets, capacity planning, disaster-recovery testing, and resilience improvements.
  • Embed security and compliance into platform operations, partnering with Security to protect infrastructure, workloads, and data while meeting regulatory requirements.
  • Apply AI-assisted and data-driven operational techniques to improve signal detection and automation opportunities.
  • Raise the bar through code reviews, documentation, knowledge sharing, and mentorship.
  • Mentor and support team growth, fostering collaboration and scalability.

Skills

Ownership
Strong communication
Problem solving

Tools

Kubernetes
Terraform
Go
Python
JavaScript/TypeScript
CI/CD
GitOps
Cloud platforms
Observability

Job description

Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open, global economy through digital assets, payment applications, and programmable blockchain infrastructure. Circle’s platform includes the world’s largest regulated stablecoin network anchored by USDC, Circle Payments Network for global money movement, and Arc, an enterprise-grade blockchain designed to become the Economic OS for the internet. Enterprises, financial institutions, and developers use Circle to power trusted, internet-scale financial innovation. Learn more at circle.com.

What You’ll Be Part Of

Circle is committed to visibility and stability in everything we do. As we grow as an organization, we're expanding into some of the world's strongest jurisdictions. Speed and efficiency are motivators for our success and our employees live by our company values: High Integrity, Future Forward, Multistakeholder, Mindful, and Driven by Excellence. We have built a flexible work environment where new ideas are encouraged and everyone is a stakeholder.

What You’ll Be Responsible For

As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical digital-assets, AI, and application workloads. You will bring an engineering mindset to production operations: writing and maintaining services and automation, developing reliable Kubernetes platforms, and using Terraform to make infrastructure repeatable, auditable, and easy to evolve.

You’ll work closely with platform, product, and application engineering teams to translate workload requirements into resilient technical designs across hybrid and public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed-systems problems, taking ownership of production outcomes, and raising the reliability, performance, security, and cost-effectiveness of the systems our customers depend on.

What You'll Work On
  • Design, build, and operate Kubernetes platforms that provide secure, highly available, and scalable foundations for critical production services across hybrid and public-cloud environments.
  • Build infrastructure as code with Terraform, creating reusable modules, safe delivery workflows, and well-governed infrastructure changes.
  • Develop backend services, internal tools, and operational automation in Go, Python, or JavaScript/TypeScript to eliminate manual work and improve the developer experience.
  • Partner with engineering and product teams to understand workload requirements and design pragmatic solutions for reliability, performance, capacity, security, and cost.
  • Improve the production lifecycle through reliable CI/CD, deployment automation, progressive delivery, and clear operational ownership.
  • Define and evolve observability practices across metrics, logs, traces, alerting, and dashboards so teams can detect issues early and troubleshoot effectively.
  • Own production reliability by participating in on-call, leading incident response, performing root-cause analysis, and driving blameless postmortems and durable corrective actions.
  • Establish and maintain reliability targets through meaningful SLIs, SLOs, error budgets, capacity planning, disaster-recovery testing, and resilience improvements.
  • Embed security and compliance into platform operations, partnering with Security to protect infrastructure, workloads, and data while meeting applicable regulatory requirements.
  • Apply AI-assisted and data-driven operational techniques to improve signal detection, reduce alert noise, accelerate root-cause analysis, and surface opportunities for automation.
  • Raise the bar for the team through thoughtful code reviews, documentation, knowledge sharing, and mentorship.
  • Mentor and support team growth, fostering collaboration and scalability.
What You’ll Bring To Circle (not All Required)
  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems.
  • Deep, hands-on Kubernetes expertise: designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale.
  • Strong Terraform experience, including authoring reusable modules, managing state and environments, and delivering infrastructure changes through reviewable, automated workflows.
  • Production software-development experience in Go, Python, or JavaScript/TypeScript, with the ability to build maintainable backend services, tooling, and automation—not only scripts.
  • Demonstrated success improving the reliability, performance, scalability, or cost efficiency of distributed systems in production.
  • Experience with cloud infrastructure and core networking concepts, including IAM, DNS, load balancing, routing, service networking, and secure connectivity.
  • Strong observability and troubleshooting skills using metrics, logs, traces, alerting, and incident data to diagnose complex systems.
  • Experience defining and operating against SLIs, SLOs, error budgets, incident-management processes, postmortems, and disaster-recovery practices.
  • Familiarity with CI/CD, GitOps or deployment automation, and safe rollout strategies such as canary or blue-green deployments.
  • A security-minded approach to infrastructure and a track record of partnering effectively with Security and engineering teams in regulated or high‑availability environments.
  • Clear written and verbal communication, strong ownership, and the judgment to balance speed, risk, and operational excellence.
  • Experience applying AI-assisted tooling to engineering or operations workflows is a plus.

Circle is on a mission to create an inclusive financial future, with transparency at our core. We consider a wide variety of elements when crafting our compensation ranges and total compensation packages.

Starting pay is determined by various factors, including but not limited to: relevant experience, skill set, qualifications, and other business and organizational needs. Please note that compensation ranges may differ for candidates in other locations.

Base Pay Range: $152,500 - $205,000

We are an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status, or any other protected status required by the laws in the locations where we hire. Additionally, Circle participates in the E-Verify Program in certain locations, as required by law.

Should you require accommodations or assistance in our interview process because of a disability, please reach out to accommodations@circle.com for support. We respect your privacy and will connect with you separately from our interview process to accommodate your needs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - CLO
Senior Site Reliability Engineer - CLO

Circle • San Francisco (CA)

On-site
USD 153,000 - 205,000
Senior Site Reliability Engineer - CLO
Senior Site Reliability Engineer - CLO

Omaze • San Francisco (CA)

On-site
USD 153,000 - 205,000
Senior/Staff Site Reliability Engineer
Senior/Staff Site Reliability Engineer

Circle • Portland (OR)

On-site
USD 152,000 - 205,000
Senior Site Reliability Engineer - CLO
Senior Site Reliability Engineer - CLO

circle • United States

On-site
USD 153,000 - 205,000
Senior Staff Data Engineer
Senior Staff Data Engineer

Circle • San Francisco (CA)

On-site
USD 225,000 - 290,000
Manager, Software Engineering
Manager, Software Engineering

Circle • Boise (ID)

On-site
USD 195,000 - 258,000
Manager, Software Engineering
Manager, Software Engineering

Circle • Illinois

On-site
USD 195,000 - 258,000
Manager, Software Engineering
Manager, Software Engineering

Circle • San Diego (CA)

On-site
USD 195,000 - 258,000
Manager, Software Engineering
Manager, Software Engineering

Circle • Charlotte (NC)

On-site
USD 195,000 - 258,000
Manager, Software Engineering
Manager, Software Engineering

Circle • Washington

On-site
USD 195,000 - 258,000