Stand out for this role — generate a tailored resume and cover letter in about a minute.
Blackpoint Cyber seeks a Senior Site Reliability Engineer to design, implement, and maintain cloud and on‑prem infrastructure, CI/CD pipelines, and automation. You will work across AWS, Kubernetes, data streaming, and observability to keep systems reliable, secure, and scalable, partnering with engineering teams to drive reliability and performance.
The role emphasizes infrastructure as code, cost awareness, incident response, and continuous improvement, with opportunities to influence
Blackpoint Cyber is the leading provider of world-class cybersecurity threat hunting, detection and remediation technology. Founded by former National Security Agency (NSA) cyber operations experts who applied their learningsto bring national security-grade technology solutions to commercial customers around the world, Blackpoint Cyber is in hyper-growth mode, fueled by a recent $190m series C round.
We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance. You'll work across cloud platform administration, container orchestration, data streaming, observability, and incident response — partnering with engineering teams to keep our systems reliable, secure, and efficient, and helping foster a culture of continuous improvement.
Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.
Own and optimize our AWS cloud environment, ensuring cost efficiency, security best practices, and high-availability standards.
Manage and optimize Kubernetes cluster environments (Helm, ArgoCD, Istio, Kustomize) to support continuous delivery and infrastructure-as-code practices.
Administer and scale data streaming infrastructure (Confluent Cloud, Apache Kafka) to support enterprise-level data processing.
Deploy, configure, and maintain Redis for caching and real-time data processing.
Implement and maintain monitoring, alerting, and incident response frameworks (Prometheus, Grafana, Alert Manager, Grafana Cloud OpsGenie/PagerDuty) to ensure system reliability and performance.
Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.
Partner with software development teams to ensure seamless integration of new services, applications, and features into existing infrastructure.
Diagnose and resolve complex system-level issues, implementing solutions that maintain high performance and maximize uptime.
Drive continuous improvement of automation tooling, operational processes, and engineering methodologies to enhance scalability, reliability, and maintainability.
Stay current on emerging SRE trends and tools, andtools and help the team adopt relevant industry advancements and best practices.
Blackpoint Cyber welcomes and encourages applications from qualified individuals of all races, colors, religions, sex, sexual orientation, gender identity or expression, national origin, age, marital status, or any other legally protected status. We are committed to equality of opportunity in all aspects of employment.
For eligible employees in the US, Blackpoint offers competitive Health, Vision, Dental, and Life Insurance plans, a robust 401k plan, Discretionary Time Off, and other minor perks. International employees receive competitive benefits in accordance with local market standards and applicable country requirements.
Blackpoint believes all employees should share in the company’s success – equity participation is available to employees globally, with program details varying by location and employment structure.