An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Onebrief is hiring a Site Reliability Engineering Manager to lead the SRE team within Infrastructure & Security. You will ensure our deployments are reliable, secure, and well supported across on-prem DoD and AWS environments.
You will guide incident response, own priorities, and coordinate with platform engineering, application engineering, security, and customer success to deliver operational excellence. Expect a mix of hands-on leadership and strategic planning.
Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.
Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.
Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.
We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.
Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.
This role requires regularly working on-site at customer locations.
If you are not currently within commuting distance, you must be willing to relocate. Onebrief provides relocation assistance.
We're hiring a Site Reliability Engineering Manager to lead our SRE team within Infrastructure & Security. You'll work closely with platform engineering, application engineering, security, and customer success to ensure Onebrief's mission-critical deployments are reliable, secure, and well supported across on-prem DoD and AWS environments.
You'll lead a team whose work spans customer-facing operations, infrastructure, observability, automation, and application reliability. Most of the team focuses on deploying and operating Onebrief in demanding customer environments.
You'll own the team's priorities, planning, execution, and development. A significant part of this role is coordinating work: understanding demand, balancing capacity, sequencing tasks, managing dependencies, and keeping commitments realistic as customer needs change. You'll help the team deliver immediate operational support while making steady progress on improvements that reduce future support demands.
You'll bring the technical grounding to evaluate risks, ask useful questions, and guide decisions. Your engineers will own technical implementation and lead incident response. You'll provide direction, remove blockers, and create the conditions for them to succeed.
You care deeply about reliability and understand the challenges of operating software in environments where connectivity, access, and deployment options can be constrained. You treat infrastructure and operability as products that deserve clear ownership, thoughtful design, and continuous improvement.
You're an effective people manager who sets clear expectations, gives useful feedback, and helps engineers grow. You build accountability through clear priorities and meaningful ownership, and you recognize when your team needs direction, support, or room to solve a problem.
You're comfortable managing a changing workload. You can turn competing requests into an actionable plan, account for operational interruptions, and explain what the team can commit to with its available capacity. You surface tradeoffs early and work with stakeholders to make deliberate decisions about scope and timing.
You bring calm and structure when priorities shift or incidents occur. You support engineers leading the response, help resolve escalations, and coordinate with customer-facing partners. You build a culture where engineers can surface risks early and examine failures honestly.
You have the technical judgment to help the team determine whether a recurring problem needs an infrastructure change, better automation, an application fix, or a clearer process. You bring the right people together to address it and ensure they have time to follow through.
You should have practical experience operating production systems and enough breadth to guide engineers working across these areas:
We value depth in relevant areas and the ability to guide specialists across the rest.