Staff Site Reliability Engineer

TENEX.AI

Sarasota (FL)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TENEX.AI is seeking a Staff Site Reliability Engineer to scale and secure our AI‑driven cybersecurity platform. You will drive resilient infrastructure, automate operations, and partner with cross‑functional teams to embed reliability from concept to production.

The role emphasizes on‑premise readiness with Monday–Thursday onsite work and Friday WFH, offering a competitive salary and benefits in a fast‑growing startup environment.

Qualifications

  • 10+ years in SRE, DevOps, or Software/Systems Engineering.
  • Public cloud expertise (AWS, GCP, or Azure) and Kubernetes (EKS/GKE).
  • Infrastructure as Code with Terraform, Pulumi, or similar tools.
  • Observability with monitoring, logging, tracing stacks (Prometheus, Grafana, ELK, Datadog).
  • Distributed systems with microservices and event-driven architectures.

Responsibilities

  • Design, build, and maintain scalable, secure infrastructure for the AI-driven platform.
  • Develop internal tooling and automation for deployment, incident response, and capacity planning.
  • Monitor performance, identify bottlenecks, optimize for low-latency AI workloads.
  • Lead incident response, post-mortems, and implement long-term reliability solutions.
  • Manage IaC across cloud environments (AWS, GCP).
  • Collaborate with engineering, product, and security to bake reliability into the lifecycle.

Skills

SRE
DevOps
Cloud infrastructure
Observability
Distributed systems

Education

Bachelor’s or Master’s in CS/Engineering

Tools

Terraform
Kubernetes
Prometheus/Grafana
Datadog/ELK

Job description

Company Overview

TENEX is an AI-native, automation-first, built-for-scale Managed Detection and Response (MDR) provider. We are a force multiplier for defenders, helping organizations enhance their cybersecurity posture through advanced threat detection, rapid response, and continuous protection. Our team is composed of industry experts with deep experience in cybersecurity, automation, and AI-driven solutions. Backed by leading investors, we are rapidly growing and seeking top talent to join our mission of revolutionizing the AI‑Native MDR landscape.

We’re a fast‑growing startup backed by industry experts and top‑tier investors led by Crosspoint Capital Partners and also backed by Shield Capital, DTCP (formerly Deutsche Telekom Capital Partners), Deepwork Capital, and the Florida Opportunity Fund. Seed round led by Andreessen Horowitz (a16z). As an early employee, you’ll play a meaningful role in defining and building our culture. Get in on the ground floor. We’re a small but well‑funded team that just raised a substantial round – joining now comes with limited risk and unlimited upside.

As a Staff Site Reliability Engineer at TENEX, you will be a key technical driver responsible for ensuring the scalability, reliability, and performance of our AI‑driven cybersecurity platform. You will play a crucial role in designing resilient infrastructure, automating operational workflows, and shaping the future of our production environments while collaborating across engineering teams to drive technical excellence.

Culture is one of the most important things at TENEX.AI—explore our culture deck at culture.tenex.ai to witness how we embody it, prioritizing the irreplaceable collaboration and community of in‑person work.

Location: This role will require Monday - Thursday onsite in any of our locations. WFH Friday.

Job Responsibilities

  • System Resilience: Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI‑native cybersecurity platform.
  • Automation & Tooling: Develop internal tooling and automation to streamline deployment processes, incident response, and capacity planning.
  • Performance Engineering: Monitor system performance and proactively identify bottlenecks, optimizing infrastructure for low‑latency, high‑throughput AI workloads.
  • Incident Management: Lead incident response efforts, conduct post‑mortems, and implement long‑term solutions to prevent recurring reliability issues.
  • Infrastructure as Code (IaC): Manage infrastructure via code, driving consistency, auditability, and scalability across our cloud environments (e.g., AWS, GCP).
  • Cross‑Functional Collaboration: Partner with sibling Engineering teams, Product, and Security teams to ensure reliability is baked into our development lifecycle from concept to production.
Required Skills & Qualifications
SRE & Infrastructure Expertise
  • Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at scale.
  • Cloud Infrastructure: Deep expertise in public cloud environments (AWS, GCP, or Azure) and managing services such as Kubernetes (EKS/GKE), networking, and storage.
  • Infrastructure as Code: Extensive experience with tools like Terraform, Pulumi, or similar technologies to manage complex infrastructure deployments.
  • Observability: Hands‑on experience with monitoring, logging, and tracing stacks (e.g., Prometheus, Grafana, ELK, Datadog) to drive data‑informed reliability decisions.
  • Distributed Systems: Solid understanding of microservices architecture, distributed databases, and event‑driven systems.
Soft Skills
  • Communication: Clear, concise communication skills and a bias for collaborative problem‑solving.
  • Leadership Alignment: Proven track record of guiding multi‑stakeholder initiatives and influencing engineering practices across teams.
  • Analytical Rigor: Strong problem‑solving, debugging, and analytical skills, especially in high‑pressure environments.
Nice‑to‑have
  • Domain Background: Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure.
  • AI/ML Infrastructure: Experience supporting infrastructure for large‑scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization).
  • Startup Mentality: Background driving high‑impact engineering initiatives in high‑growth startups or enterprise SaaS.
  • Strong familiarity with Agentic Workflows such as Agno, Temporal, etc..
Education & Certifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • Relevant certifications (CKA/CKAD, AWS/GCP Professional Cloud Architect, etc.) are a plus.
Why Join Us?
  • Opportunity to work with cutting‑edge AI‑driven cybersecurity technologies and Google SecOps solutions.
  • Collaborate with a talented and innovative team focused on continuously improving security operations and system reliability.
  • Competitive salary and benefits package.
  • A culture of growth and development, with opportunities to expand your knowledge in AI, cybersecurity, and emerging technologies.

If you're passionate about building resilient infrastructure, scaling AI systems, and working at the intersection of reliability and security, we encourage you to apply!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

TENEX.AI • United States

Hybrid
USD 140,000 - 200,000
AI/ML Engineer
AI/ML Engineer

TENEX.AI • Sarasota (FL)

Hybrid
USD 90,000 - 120,000
Cutting-edge AI technologies
Collaborative team atmosphere
Competitive salary and benefits
+1
Senior AI/ML Engineer
Senior AI/ML Engineer

TENEX.AI • Sarasota (FL)

Hybrid
USD 120,000 - 150,000
Competitive salary and benefits package
Opportunities for growth and development
Cutting-edge AI technologies
Principal AI Engineer
Principal AI Engineer

Tenex • Sarasota (FL), Kansas City (MO), San Jose (CA)

Hybrid
USD 180,000 - 240,000
Senior AI/ML Engineer
Senior AI/ML Engineer

TENEX.AI • United States

Hybrid
USD 120,000 - 150,000
Competitive salary and benefits package
Opportunities for growth and development
Cutting-edge AI-driven technologies
Software Engineer II
Software Engineer II

TENEX.AI • San Jose (CA)

Hybrid
USD 100,000 - 130,000
Competitive salary and benefits package
Opportunity for growth in AI and cybersecurity
AI/ML Engineer II
AI/ML Engineer II

Tenex • Sarasota (FL), Kansas City (MO), San Jose (CA)

On-site
USD 120,000 - 150,000
AI/ML Engineer
AI/ML Engineer

TENEX.AI • United States

Hybrid
USD 100,000 - 150,000
Competitive salary
Benefits package
Growth and development opportunities
AI/ML Engineer II
AI/ML Engineer II

TENEX.AI • San Jose (CA)

Hybrid
USD 120,000 - 150,000
Competitive salary and benefits
Culture of growth and development
Opportunity to work with cutting-edge technologies
Staff Software Engineer
Staff Software Engineer

TENEX.AI • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Competitive salary and benefits package
Opportunity for growth and learning
Collaborative culture focused on development