Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)

CrowdStrike Holdings, Inc.

New York (NY)

Hybrid

USD 140,000 - 215,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Market leader compensation
Comprehensive wellness programs
Paid parental leaves
Professional development opportunities
Diversity and inclusion programs

Job summary

CrowdStrike is seeking a Principal SRE to operate at the intersection of our Core Platform and Embedded Reliability, building foundational libraries, services, and tooling that power our AI-native Falcon platform.

You will partner with product engineering teams to drive reliability outcomes at scale, shaping architecture and influencing decisions across multiple squads while advancing observability and automation.

Qualifications

  • 10+ years building and operating distributed systems
  • 5+ years developing microservices for a SaaS product
  • Expert-level proficiency in at least one language (Go preferred)
  • Deep understanding of distributed systems, concurrency, and scaling
  • Proven experience driving reliability and architectural decisions

Responsibilities

  • Define and drive multi-year reliability roadmaps across product groups
  • Design architectural improvements to services, libraries, and platforms
  • Develop and maintain services meeting reliability and scalability demands
  • Lead initiatives around performance, cost optimization, and resilience in large-scale systems
  • Mentor engineers and establish architectural standards across the org
  • Contribute to observability, automation, and infrastructure-as-code efforts

Skills

Distributed systems
Go
Backend development
SRE
Microservices
Cloud platforms

Education

Degree in Computer Science

Tools

Kubernetes
AWS
Kafka
Elasticsearch/OpenSearch

Job description

As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed - we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.

About the Role

CrowdStrike Falcon is the industry standard in cloud‑native cybersecurity and threat hunting, processing trillions of events per day. As a Principal SRE, you will operate at the intersection of our Core Platform and Embedded Reliability charters: building the foundational libraries, services, and tooling that every product group depends on, while embedding directly with product engineering teams and their leadership to drive reliability outcomes at scale. While we embrace the SRE moniker, at CrowdStrike it means something far more service‑oriented and engineering‑heavy than traditional operations. This is hands‑on systems engineering - writing production code, re‑architecting critical systems, and eliminating entire classes of failure - not ticket management. It is far and away our most self‑driven and autonomous backend engineering role, with the freedom to move up, down, and laterally across the stack as needed. You'll join a group with bottom‑up visibility and ownership of a fast‑expanding codebase (including AI‑based Detection & Response) and a far‑reaching mandate for service resiliency. Several key feature additions and critical re‑architecture initiatives are underway: deploying core services to additional cloud providers, modularizing into reusable components, improving core libraries and frameworks, maturing observability tooling (tracing, profiling, alerting, SLOs), and automating away manual toil across all of the above. Recent examples of the team's work include introducing adaptive concurrency into our core Kafka library to implement flow control that protects database performance, resolving critical issues in leader election libraries, and building infrastructure‑as‑code tooling that eliminated manual deployment processes. At the Senior Engineer level, your influence is organizational. Product engineers and engineering leaders will come to you for guidance on architectural decisions because you've earned credibility through hands‑on work and delivered results. You will shape architectural choices that affect every feature development team and provide shared architectural components leveraged throughout the Falcon Platform — including Unified Search, Protobuf libraries, and other shared‑tier infrastructure.

Why This Role Matters

Our customers depend on us to protect their businesses from sophisticated threats, and reliability isn't optional - it’s fundamental to our mission. Your work directly impacts whether organizations around the world can defend themselves against cyberattacks. You'll work on problems that matter, at a scale few companies can match, with the autonomy to make real architectural decisions.

Location

This position requires candidates to be based in Midtown Manhattan, NY; Redmond, WA; Sunnyvale, CA; or Austin, TX. It's a hybrid role, with employees typically working 2 to 3 days per week in the office.

What You'll Do
  • Partner with engineering leadership across multiple product groups to define and drive multi‑year reliability roadmaps.
  • Design and implement architectural improvements to services, libraries, and platforms that impact teams across all of CrowdStrike.
  • Develop and maintain services that meet aggressive reliability and scalability demands.
  • Extend and build new libraries for cross‑cutting concerns spanning CrowdStrike's cloud platform, which comprises hundreds of libraries and services.
  • Lead initiatives around reliability, scalability, performance, and cost efficiency in large‑scale distributed systems.
  • Establish foundational observability practices: ensure teams instrument services properly, react to signals effectively, and leverage observability to drive automation such as continuous delivery.
  • Define and implement service‑level objectives and error budgets that drive real decision‑making and prioritization.
  • Lead performance and cost optimization efforts: profiling, bottleneck analysis, capacity planning, and cloud efficiency improvements.
  • Conduct resilience engineering: chaos experiments, failure injection, failure modeling, and designing for graceful degradation.
  • Design and implement automation and infrastructure‑as‑code to improve infrastructure reliability and eliminate manual toil.
  • Provide technical leadership during complex incidents and ensure follow‑through on retrospectives with concrete improvements that eliminate entire classes of failures.
  • Identify opportunities to extract common patterns into shared libraries and tools, or partner with platform teams on improvements benefiting multiple product groups.
  • Continuously re‑evaluate our products to improve architecture, knowledge models, developer and user experience, performance, and stability.
  • Drive strategic technical decisions and influence infrastructure and operational improvements across the organization.
  • Mentor and coach engineers, raising the technical IQ of the team and driving architectural standards across the org.
  • Use and give back to the open source community; evangelize software engineering best practices, especially as they pertain to Go.
  • Brainstorm, define, and build collaboratively with members across multiple teams as an energetic self‑starter who owns and is accountable for deliverables.
What You'll Need
  • 10+ years of experience building and operating distributed systems and service‑oriented backends at scale.
  • 5+ years developing microservices for a SaaS product in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js).
  • Expert‑level proficiency in at least one programming language, with expert‑level Go or demonstrated ability and willingness to reach expert level in Go.
  • Deep understanding of distributed systems: consensus algorithms, replication, consistency models, failure modes, and scalability patterns.
  • Proven experience scaling backend systems — sharding, partitioning, horizontal scaling, capacity planning, and performance optimization are second nature.
  • Deep understanding of multi‑threading, concurrency, and parallel processing.
  • Track record of making impactful architectural decisions at organizational scope and seeing them through to production.
  • Strong systems thinking and the ability to influence without direct authority across organizational boundaries.
  • Thorough command of engineering best practices: appropriate testing paradigms, effective peer code review, and resilient architecture.
  • Ability to thrive in a fast‑paced, test‑driven, collaborative, and iterative environment; strong team‑player orientation.
  • A desire to ship code and a love of seeing your bits run in production.
  • Degree in Computer Science, or commensurate experience in data structures, algorithms, and distributed systems.
  • Proven experience utilizing AI technologies to enhance decision‑making, streamline workflows and processes, improve efficiency and drive business outcomes.
Bonus Points
  • Experience driving reliability improvements in organizations with hundreds or thousands of microservices.
  • Deep knowledge of Kubernetes or other large‑scale orchestration systems.
  • Hands‑on experience with AWS, Cassandra, Kafka, Elasticsearch/OpenSearch, or similar large‑scale distributed technologies.
  • Experience with Google Cloud Platform (GCP).
  • Experience with Oracle Cloud Infrastructure (OCI).
  • Experience delivering or operating services across multiple cloud providers, including multi‑cloud abstraction layers, portability, and cloud‑agnostic tooling.
  • Track record of building internal platforms, developer platforms, or tools that other engineers depend on.
  • Experience with infrastructure cost optimization at scale.
  • Background in performance engineering: profiling, optimization, and identifying system bottlenecks.
  • Experience with chaos engineering or resilience testing practices.
  • History of establishing SLO/SLI frameworks and error budgets in production environments.
  • Contributions to the open source community (GitHub, Stack Overflow, technical blogging).
  • Prior experience in the cybersecurity or intelligence fields.
Benefits of Working at CrowdStrike
  • Market leader in compensation and equity awards
  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees regardless of level or role
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
  • Vibrant office culture with world class amenities
  • Great Place to Work Certified across the globe

CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program. CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions—including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay‑offs, return from lay‑off, terminations and social/recreational programs—on valid job requirements.

If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at recruiting@crowdstrike.com for further assistance.

Find out more about your rights as an applicant.

CrowdStrike participates in the E-Verify program.

Right to Work CrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location. The base salary range for this position for all U.S. candidates is $140,000 - $215,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.

CrowdStrike was founded in 2011 to fix a fundamental problem: The sophisticated attacks that were forcing the world’s leading businesses into the headlines could not be solved with existing malware-based defenses. Founder GeorgeKurtz realized that a brand new approach was needed — one that combines the most advanced endpoint protection with expert intelligence to pinpoint the adversaries perpetrating the attacks, not just the malware. There’s much more to the story of how Falcon has redefined endpoint protection but there’s only one thing to remember about CrowdStrike: We stop breaches.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Engineer - Cloud Backend (Hybrid) - Sunnyvale
Sr. Engineer - Cloud Backend (Hybrid) - Sunnyvale

CrowdStrike Holdings, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 215,000
Market leading compensation
Equity awards
Health insurance
+1
Engineering Manager, Cloud Backend - AI Detection and Response (AIDR) (Hybrid)
Engineering Manager, Cloud Backend - AI Detection and Response (AIDR) (Hybrid)

CrowdStrike Holdings, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 215,000
Equity awards
Wellness programs
Vacation & holidays
+3
Sr. Engineer, Cloud Native - AI Detection and Response (AIDR) (Hybrid, Sunnyvale)
Sr. Engineer, Cloud Native - AI Detection and Response (AIDR) (Hybrid, Sunnyvale)

CrowdStrike Holdings, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 215,000
Market compensation
Wellness programs
Paid time off
+3
Software Engineer, Product Security - Security Automation (Remote)
Software Engineer, Product Security - Security Automation (Remote)

CrowdStrike Holdings, Inc. • United States

Remote
USD 120,000 - 180,000
Equity awards
Wellness programs
Vacation & holidays
+5
Software Engineer, Product Security - Security Automation (Remote)
Software Engineer, Product Security - Security Automation (Remote)

CrowdStrike Holdings, Inc. • Northern (KY)

Remote
USD 120,000 - 180,000
Health insurance
401k
Paid time off
+4
Sr. Full Stack Engineer, Cloud Native - AI Detection and Response (AIDR) (Hybrid, Sunnyvale)
Sr. Full Stack Engineer, Cloud Native - AI Detection and Response (AIDR) (Hybrid, Sunnyvale)

CrowdStrike Holdings, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 215,000
Compensation & equity
Wellness programs
Vacation & holidays
+5
Sr. Program Manager, Engineering (Hybrid)
Sr. Program Manager, Engineering (Hybrid)

CrowdStrike Holdings, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 215,000
Equity awards
Wellness programs
Generous vacation
+3
Sr Engineer, SRE TechOps CICD (Remote)
Sr Engineer, SRE TechOps CICD (Remote)

CrowdStrike Holdings, Inc. • Northern (KY)

Remote
USD 140,000 - 215,000
Competitive compensation
Wellness programs
Generous vacation and holidays
+1
Sr Engineer, SRE TechOps CICD (Remote)
Sr Engineer, SRE TechOps CICD (Remote)

CrowdStrike Holdings, Inc. • Virginia (MN)

Hybrid
USD 140,000 - 215,000
Equity awards
Wellness programs
Vacation and holidays
+2
Engineering Manager - Core Infrastructure (Remote)
Engineering Manager - Core Infrastructure (Remote)

CrowdStrike Holdings, Inc. • Northern (KY)

Remote
USD 140,000 - 215,000
Market-leading compensation
Wellness programs
Vacation and holidays
+5