Senior Technical Duty Officer, Cloud Ops

Linuxconfig

Redwood City, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Box is seeking a Senior Technical Duty Officer to lead live-site incidents and drive rapid resolution for critical cloud services in Redwood City, CA. You will own incident bridges, design automated tooling, and partner with SRE teams to reduce MTTR while improving observability and resiliency across Box's cloud stack.

You will guide change review processes, mentor colleagues, and shape the evolution of incident and problem management in a fast-paced, AI-powered environment.

Qualifications

  • 5+ years in SRE, production operations, reliability engineering, or equivalent high-scale internet/SaaS operations.
  • Experience leading or co-leading major production incidents.

Responsibilities

  • Own and direct live-site Critical and Blocker incidents from identification to recovery.
  • Lead incident bridges, coordinate resources, and drive swift mitigation.
  • Improve incident tooling and automation to reduce overhead and time to mitigate.
  • Collaborate with SRE and engineering to deepen knowledge of dependencies and failure modes.
  • Lead daily change reviews in Jira to evaluate and minimize risk.
  • Provide day-to-day technical expertise for 24x7 environments and policy decisions.
  • Lead projects to improve tools and processes for resiliency and observability.
  • Represent GTOC/NOC in problem management and readiness reviews.

Skills

Incident Command
SRE Leadership
Python automation
Observability
SLIs/SLOs
Communication
Cloud: GCP/AWS
Kubernetes

Tools

Prometheus
Grafana
Chronosphere
SignalFx
Catchpoint
PagerDuty
Jira
Terraform

Job description

Senior Technical Duty Officer, Cloud Ops

Redwood City, CA, United States

WHAT IS BOX?

Box (NYSE:BOX) is the leader in Intelligent Content Management. Our platform enables organizations to fuel collaboration, manage the entire content lifecycle, secure critical content, and transform business workflows with enterprise AI. We help companies thrive in the new AI-first era of business. Founded in 2005, Box simplifies work for leading global organizations, including JLL, Morgan Stanley, and Nationwide. Box is headquartered in Redwood City, CA, with offices across the United States, Europe, and Asia.

By joining Box, you will have the unique opportunity to continue driving our platform forward. Content powers how we work. It’s the billions of files and information flowing across teams, departments, and key business processes every single day: contracts, invoices, employee records, financials, product specs, marketing assets, and more. Our mission is to bring intelligence to the world of content management and empower our customers to completely transform workflows across their organizations. With the combination of AI and enterprise content, the opportunity has never been greater to transform how the world works together and at Box you will be on the front lines of this massive shift.

WHY BOX NEEDS YO

Box is looking for a Senior Technical Duty Officer - a Senior Incident Commander with a strong Site Reliability Engineering (SRE) skillset and proficient Python for tooling and automation. The GTOC (Global Technical Operations Center) team protects the customer experience, and responsible for detecting customer-impacting failures early, drive rapid mitigation when they occur, coordinate production change risk, and continuously improve how Box runs incidents.

In this role, you will be leading critical and blocker incidents and driving them to swift resolution, while leveraging your software engineering expertise to design and build the tools, automation, and processes that define the next generation of cloud operations. You will shape how we improve manageability, observability, and resiliency of our most critical services, and drive down MTTx so that when incidents happen, we respond faster and smarter than ever before.

WHAT YOU'LL DO

  • Own and direct live-site Critical and Blocker (and other high severity) incidents from identification and escalation through mitigation and recovery.
  • Triage, refine, and verify the problem and customer-impact statements. Organize the incident bridge, establish clear swim lanes, coordinate SME’s resources, and lead cross-functional Incident Bridges to mitigate issues and restore service quickly.
  • Improve Incident Platform tooling - Optimize templates, automate repetitive response steps, and develop tooling to reduce operational overhead and time to mitigate.
  • Partner with SRE and engineering teams to deepen shared knowledge of Box's dependencies, Tier 1 journeys, failure modes. Work together to implement secure, automated workflows that close operability gaps.
  • Lead daily change reviews of planned changes in Jira, partnering collaboratively with engineering teams to evaluate and minimize change risk.
  • Provide day to day technical expertise and experience to the organization to address issues in globally diverse, 24x7 environments - influencing from policy and procedural decisions to key architectural and tooling insights to improve Box's Incident, Change, and Problem Management engineering capabilities.
  • Lead projects to improve tools and processes related that enhance overall site resiliency, service manageability and observability.
  • Represent GTOC/NOC in problem management, readiness reviews, and reliability in cross functional forums.
  • Turn incident learnings into actionable engineering asks (such as SLOs, signal quality, rollback readiness, dependency docs) to reduce customer impact.
  • Leverage observability APIs, PagerDuty/Jira, and internal services to make response workflows measurable and repeatable.
  • Continuously mentor and uplift the team by running tabletop exercises and improving runbooks.

WHO YOU ARE

We are an AI-first company. This means you approach your work with a growth mindset and find ways to leverage AI to help make faster, smarter decisions that will 10X your impact at Box.

  • 5+ years in SRE, production operations, reliability engineering, or equivalent high-scale internet/SaaS operations with repeated experience leading or co-leading major production incidents.
  • Demonstrated Incident Commander / Technical Duty Officer (or equivalent) skill: calm under uncertainty, clear communication, ability to delegate, and judgment on when to escape vs when to dig deeper.
  • Strong SRE fundamentals: SLIs/SLOs and error budgets (practical use), observability (metrics, logs, traces), golden signals, blameless postmortems, toil reduction, and “you build it, you run it” partnership with product teams.
  • Proficient in Python for automation and tooling (not just one-off scripts): readable code, APIs, packaging or service-style tools others can run, and comfort reviewing others’ automation.
  • Networking literacy sufficient for real incidents (DNS, TLS, load balancing, HTTP, basic routing/firewall concepts).
  • Experience with cloud environments (GCP preferred; AWS/Azure also valued) and container/orchestration concepts (Kubernetes or equivalent).
  • Excellent written and verbal communication: exec-ready impact statements, precise Slack/bridge facilitation, and documentation that others can trust under stress.
  • Proven ability to coach and uplift others — training, mentoring, runbook quality, drills.

Preferred skills

  • Experience in a 24x7 NOC / GTOC / follow-the-sun operations center, or partnering closely with one.
  • Hands-on with Prometheus-compatible observability (e.g. Chronosphere, Grafana), distributed tracing(e.g. SignalFx), synthetic monitoring (e.g. Catchpoint), and PagerDuty (or similar).
  • Experience improving incident tooling (ChatOps, workflow automation, status/comms templates, drill frameworks).
  • Familiarity with change management under pressure (emergency changes, break-glass, CAB/async approval patterns).
  • Additional languages or tooling useful for ops automation (Go, shell, Terraform, CI/CD).
  • Prior work mapping dependency graphs, service catalogs, or Tier 1 journey documentation.

Box lives its values, with community and in-person collaboration being a core part of our culture. Boxers are expected to work from their assigned office a minimum of 3 days per week.Your Recruiter will share more about how we work and company culture during the hiring process.

At Box, we believe unique and diverse experiences benefit our culture, our products, our customers, our company, and our world. We aim to recruit a passionate, high-performing workforce that reflects the world we live in.

EQUAL OPPORTUNITY

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, disability, and any other protected ground of discrimination under applicable human rights legislation. Box strives to respect the dignity and independence of people with disabilities and is committed to giving them the same opportunity to succeed as all other employees. Inclusiveness is core to our culture at Box, and we strive to ensure you get the most from your interview experience.

Box makes reasonable accommodations for applicants with disabilities. If a reasonable accommodation is needed to participate in the job application or interview process, please complete this form . Reasonable accommodations may include scheduling adjustments, document dictation and beyond.

Notice to applicants in Los Angeles: Box, Inc and its related branches will consider for employment, qualified applicants with criminal histories in a manner consistent with the Los Angeles Fair Chair Ordinance. The Fair Chance Ordinance is provided here .

Notice to applicants in San Francisco: Box, Inc and its related branches will consider for employment, qualified applicants with criminal histories in a manner consistent with the San Francisco Fair Chair Ordinance. The Fair Chance Ordinance is provided here .

For details on how we protect your information when you apply, please see our Personnel Privacy Notice. If you are a California-resident, please read our California Applicant & Candidate Privacy Notice here .

Box is committed to fair and equitable compensation practices. Actual base salary (or OTE if commissionable role) is dependent upon factors such as: knowledge, skill level, experience, and work location. This role is also eligible for equity and benefits. For more information, check out our benefits and perks .

Here at Box our goal is to fully leverage and engage the unique talents/capabilities of diverse teams who feel like they belong. We provide a safe space to contribute your unique skills, ideas, and insights. A welcoming space is provided to learn, try new things, grow your career, test your thinking, and evolve. Our goal is to build diverse teams with broad representation that matches our customers and the world at large. We also want to make sure that you feel welcomed and supported, connected, included, and are treated fairly and with respect. We also ask that you do the same for others.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Technical Duty Officer, Cloud Ops
Senior Technical Duty Officer, Cloud Ops

Box • Redwood City (CA)

On-site
USD 187,000 - 234,000
Senior Software Engineer, Microservice Foundation
Senior Software Engineer, Microservice Foundation

Box • Town of Texas (WI)

Hybrid
USD 199,000 - 248,000
Senior Software Engineer, Microservice Foundation
Senior Software Engineer, Microservice Foundation

Box • Redwood City (CA)

On-site
USD 180,000 - 260,000
Equity
Benefits
Product Support Specialist
Product Support Specialist

Box • Chicago (IL)

Hybrid
USD 55,000 - 75,000
Associate Technical Consultant
Associate Technical Consultant

Box • Town of Texas (WI)

Hybrid
USD 98,000 - 112,000
Software Engineer III, Network Platform
Software Engineer III, Network Platform

Box • Redwood City (CA)

On-site
USD 165,000 - 207,000
Senior Engineering Manager, Cloud Ops
Senior Engineering Manager, Cloud Ops

Box • Town of Texas (WI)

Hybrid
USD 253,000 - 316,000
Equity
Benefits
Hybrid work model
Senior Organizational Development Manager
Senior Organizational Development Manager

Box • Chicago (IL)

On-site
USD 120,000 - 180,000
Enterprise Sales Engineer
Enterprise Sales Engineer

Box • Redwood City (CA)

Hybrid
USD 140,000 - 210,000
Senior Organizational Effectiveness Manager
Senior Organizational Effectiveness Manager

Box • Redwood City (CA)

On-site
USD 147,000 - 171,000