Sr. Site Reliability Engineer

Blackpoint Cyber

United States

Remote

USD 140,000 - 200,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401k plan
Discretionary Time Off

Job summary

Blackpoint Cyber is seeking a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on‑premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance.

You will work across cloud platform administration, container orchestration, data streaming, observability, and incident response—partnering with engineering teams to keep our systems reliable, secure, and efficient, and helping foster continuous improvement.

Qualifications

  • 5+ years in a Senior Site Reliability Engineer role or equivalent, focused on cloud infrastructure and automation.
  • Proficiency with Infrastructure as Code: Terraform, Terragrunt.
  • Extensive AWS knowledge for secure, scalable cloud architectures.
  • Experience with distributed data streaming: Confluent Cloud, Apache Kafka.
  • Redis for caching and Amazon RDS for relational DB management.
  • Monitoring/alerting with Prometheus, Grafana, Alert Manager, OpsGenie/PagerDuty.
  • Kubernetes administration (Helm, ArgoCD, Istio); familiarity with Kustomize.
  • Experience with feature flag systems (LaunchDarkly/PostHog) for controlled releases.

Responsibilities

  • Design, develop, and maintain scalable infrastructure using IaC (Terraform/Terragrunt).
  • Own and optimize AWS cloud environment for cost efficiency, security, and high availability.
  • Manage and optimize Kubernetes clusters (Helm, ArgoCD, Istio; Kustomize knowledge).
  • Administer and scale data streaming (Confluent Cloud, Apache Kafka).
  • Deploy, configure, and maintain Redis and relational databases (RDS).
  • Implement monitoring/alerting frameworks (Prometheus, Grafana, Alert Manager, OpsGenie/PagerDuty).
  • Enable controlled feature deployments with LaunchDarkly and PostHog.
  • Collaborate with software teams to integrate services into existing infra.
  • Diagnose complex production issues and drive reliability improvements.
  • Advance automation tooling and engineering practices for scalability.

Skills

Cloud infrastructure management
Automation
Problem-solving
Communication
Collaboration
Agile

Tools

Terraform
Terragrunt
AWS
Kubernetes
Helm
ArgoCD
Istio
Kustomize
Confluent Cloud
Apache Kafka
Redis
OpenSearch/Elasticsearch
Prometheus
Grafana
Alert Manager
PagerDuty/OpsGenie
LaunchDarkly
PostHog

Job description

Blackpoint Cyber is the leading provider of world-class cybersecurity threat hunting, detection and remediation technology. Founded by former National Security Agency (NSA) cyber operations experts who applied their learningsto bring national security-grade technology solutions to commercial customers around the world, Blackpoint Cyber is in hyper-growth mode, fueled by a recent $190m series C round.

SUMMARY

We’re hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance. You’ll work across cloud platform administration, container orchestration, data streaming, observability, and incident response — partnering with engineering teams to keep our systems reliable, secure, and efficient, and helping foster a culture of continuous improvement.

RESPONSIBILITIES
  • Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.
  • Own and optimize our AWS cloud environment, ensuring cost efficiency, security best practices, and high-availability standards.
  • Manage and optimize Kubernetes cluster environments (Helm, ArgoCD, Istio, Kustomize) to support continuous delivery and infrastructure-as-code practices.
  • Administer and scale data streaming infrastructure (Confluent Cloud, Apache Kafka) to support enterprise-level data processing.
  • Deploy, configure, and maintain Redis for caching and real-time data processing.
  • Implement and maintain monitoring, alerting, and incident response frameworks (Prometheus, Grafana, Alert Manager, OpsGenie/PagerDuty) to ensure system reliability and performance.
  • Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.
  • Partner with software development teams to ensure seamless integration of new services, applications, and features into existing infrastructure.
  • Diagnose and resolve complex system-level issues, implementing solutions that maintain high performance and maximize uptime.
  • Drive continuous improvement of automation tooling, operational processes, and engineering methodologies to enhance scalability, reliability, and maintainability.
  • Stay current on emerging SRE trends and tools, and help the team adopt relevant industry advancements and best practices.
REQUIREMENTS
  • 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial emphasis on cloud infrastructure management and automation.
  • Expertise in Infrastructure as Code (Terraform, Terragrunt) for enterprise-scale deployments.
  • Comprehensive knowledge of AWS, including designing, implementing, and maintaining secure, scalable, resilient cloud architectures.
  • Extensive hands-on experience with distributed data streaming (Confluent Cloud, Apache Kafka).
  • Proven experience with Redis for caching and Amazon RDS for relational database management.
  • Experience with enterprise search and analytics platforms (OpenSearch, Elasticsearch, ChaosSearch).
  • Proficiency designing and implementing monitoring/alerting infrastructure (Prometheus, Grafana, Alert Manager, OpsGenie/PagerDuty).
  • Practical experience with feature flag systems (LaunchDarkly/PostHog) for controlled release management.
  • Extensive experience administering production-grade Kubernetes (Helm, ArgoCD, Istio); working knowledge of Kustomize.
  • Strong problem-solving skills, with the ability to troubleshoot complex systems in production.
  • Strong communication and collaboration skills, with experience working in Agile environments.
NICE TO HAVE
  • Multi-cloud experience (Google Cloud Platform, Microsoft Azure).
  • Understanding of security frameworks and compliance standards for cloud-native/containerized environments.
  • Serverless computing and CI/CD pipeline experience (Jenkins, GitHub Actions).
  • Software development proficiency in Node.js, Python, and/or Go.

Blackpoint Cyber welcomes and encourages applications from qualified individuals of all races, colors, religions, sex, sexual orientation, gender identity or expression, national origin, age, marital status, or any other legally protected status. We are committed to equality of opportunity in all aspects of employment.

For eligible employees in the US, Blackpoint offers competitive Health, Vision, Dental, and Life Insurance plans, a robust 401k plan, Discretionary Time Off, and other minor perks. International employees receive competitive benefits in accordance with local market standards and applicable country requirements.

Blackpoint believes all employees should share in the company’s success – equity participation is available to employees globally, with program details varying by location and employment structure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

CyberPoint International • Maryland

Hybrid
USD 120,000 - 170,000
Senior MDR Analyst
Senior MDR Analyst

Blackpoint Cyber • United States

Remote
USD 110,000 - 160,000
Health insurance
Vision insurance
Dental insurance
+4
Support Engineer III
Support Engineer III

Blackpoint Cyber • United States

Hybrid
USD 90,000 - 130,000
Health insurance
Vision insurance
Dental insurance
+4
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

On-site
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
AI/ML Software Engineer
AI/ML Software Engineer

Blackpoint Cyber • United States

Remote
USD 110,000 - 180,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer
Site Reliability Engineer

Fortress Information Security, LLC • Patuxent Highland (MD)

On-site
USD 160,000 - 180,000
Medical, dental, and vision plans
401(k) match
Flexible Paid Time Off
+1
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

United States Digital Space LLC • United States

On-site
USD 145,000 - 200,000
Remote-first