- As a Staff Software Engineer on the Agent Platform team, you will help secure tens of millions of devices across Windows, Linux, and macOS while processing billions of security events every day as part of SentinelOne’s Endpoint Protection product line
- You’ll help drive customer-critical incident response across the Agent Platform, rapidly build context across multiple platform services, and partner with engineering teams to continuously improve the platform’s reliability, operability, and scalability as a whole
- You will build and evolve the high-throughput, highly available services responsible for policy, configuration, and command distribution to millions of agents worldwide, while contributing to the agent platform protocols that enable other SentinelOne teams to deliver new security capabilities safely and at scale
- This role offers broad technical ownership, the opportunity to work across service boundaries, and the ability to influence the design and reliability of critical platform capabilities used throughout the Agent Platform at SentinelOne. You’ll solve some of the company’s most complex production challenges while helping shape the future of our platform
- Drive rapid response to customer-critical incidents by diagnosing, triaging, and resolving complex production issues that span multiple services across the Agent Platform
- Lead systematic root-cause analysis and translate incident learnings into durable reliability, scalability, observability, and operability improvements
- Quickly build a systems-level understanding of unfamiliar services and codebases, collaborating closely with engineering teams to resolve complex cross-service issues
- Design, develop, test, document, deploy, and operate large-scale, high-volume, low-latency distributed systems processing millions of events per second
- Understand, maintain, and continuously improve existing services through refactoring, feature development, and architectural enhancements
- Maintain application stability and data integrity by monitoring key metrics, improving operational visibility, and continuously strengthening the codebase
- Translate business and functional requirements into robust, scalable, and operable technical solutions
- Partner closely with engineering teams across SentinelOne to solve cross-functional problems, influence technical direction, and deliver solutions that scale
- Continuously evaluate and adopt technologies that improve the platform’s scalability, reliability, and operational excellence
- Your Toolkit:
- Our backend services are primarily developed in with Go and Python also playing important roles
- We use gRPC, REST, GraphQL, and Kafka for service communication
- Our data layer includes Redis, PostgreSQL, Cassandra, ClickHouse, and our own columnar time-series database
- Our services run across multiple AWS and GCP regions on Kubernetes, using Docker, GitHub, and ArgoCD
- We provide engineers with modern AI-powered tools to improve both engineering productivity and software quality
Benefits
- Medical, dental, and vision coverage
- Employee assistance program
- Gym reimbursement
- Incentive-based challenges
- Mental health and mindfulness
- Unlimited Time Off
- Grandparent Leave
- Volunteer Time Off
- Paid Sick Time
- Paid Holidays
- 16 weeks Gender-Neutral Parental Leave
- Restricted Stock Unit Program
- Flexible Spending Accounts
- Life Insurance
- Short and Long Term Disability Insurance
- 401K
- Team building activities
- Celebrations and social gatherings
- Community volunteering events
- Global all hands and local town hall events
We are looking for an excellent software engineer with a reliability-first mindset who thrives on solving complex distributed systems challenges in production and turning those learnings into durable platform improvements and long-term engineering ownership Experience with messaging systems and data platforms such as Kafka, PostgreSQL, Redis, Cassandra, ClickHouse, or similar technologies Strong hands-on experience designing, building, and operating large-scale distributed systems, with a deep understanding of failure modes, performance trade-offs, resilience patterns, and operational excellence A high degree of ownership, autonomy, curiosity, and the ability to drive ambiguous technical problems to successful outcomes Experience in an enterprise SaaS or cybersecurity software company is highly desirable Excellent communication and collaboration skills, with the ability to work effectively across engineering teams, Product, Technical Account Managers, and other stakeholders, influence technical direction, and mentor fellow engineers A reliability-first mindset with a proven ability to solve complex production incidents, perform systematic root-cause analysis, and implement durable engineering solutions that prevent recurrence Experience with AWS, GCP, or similar cloud platforms, as well as Docker, Helm, and Kubernetes Demonstrated ability to quickly understand unfamiliar distributed systems, navigate large codebases, and troubleshoot complex failures across service boundaries 8+ years of professional backend software engineering experience with deep expertise in at least one of Go or Python