About the Role
Lead the Platform Engineering organization at OCC, overseeing Platform Engineering Governance & Compliance, Site Reliability Engineering, Cloud Strategy & Architecture, and Metrics & Reporting. The role requires scaling a mature SRE practice, driving cloud architectural standards, ensuring platform health visibility, and managing FinOps and SecOps domains as Product Manager.
Responsibilities
- Scale and mature the SRE practice: establish error budgets, SLOs, SLAs, incident response frameworks, and reliability standards across all platform services.
- Define and enforce on‑call models, blameless post‑mortems, and corrective action tracking to drive continuous improvement.
- Partner with Platform Foundation teams (Kubernetes, Kafka, FinOps/Security) to embed reliability principles and reduce toil through automation.
- Serve as Product Manager for FinOps and SecOps domains, owning product vision, prioritization, and stakeholder alignment for governance tooling.
- Maintain a governance framework that ensures adherence to incident, problem, risk, and audit standards.
- Own end‑to‑end PE compliance processes, ensuring timely resolution of incidents, problems, risks, and audit observations.
- Collaborate with Risk, Compliance, and Security functions to identify gaps, drive remediation, and maintain compliance posture.
- Define and execute multi‑year cloud architecture strategy aligned to growth, scalability, regulatory compliance, and cost optimization.
- Establish cloud architectural standards, reference architectures, and governance frameworks for landing zones, identity, and network patterns.
- Guide cloud‑native architecture decisions, including containerization, IaaS/PaaS adoption, disaster recovery, and multi‑region patterns.
- Own metrics and reporting: build dashboards, define KPIs, and enable data‑driven prioritization across engineering and executive audiences.
- Coordinate with Platform Engineering Program Management to scope, sequence, and resource platform initiatives.
- Build and retain a high‑performing team of engineering managers and staff, fostering a culture of automation, full‑stack ownership, and transparency.
- Oversee audit findings remediations, risk mitigation, and budget adherence for assigned areas.
- Manage work/life balance and overall team delivery quality.
Qualifications
- Proven executive‑level leadership of SRE, cloud engineering, or platform reliability in a regulated environment.
- Demonstrated ability to build and scale SRE practices, including SLO/SLA frameworks, on‑call models, error budgets, and incident response programs.
- Deep expertise in cloud architecture strategy and governance, with enterprise‑wide standards experience.
- Strong cross‑functional partnership with Program/Product Management; ability to translate platform capabilities into roadmaps.
- Experience serving in a Product Manager role for technical domains such as FinOps, SecOps, or platform tooling.
- Experience establishing and managing governance and compliance frameworks, overseeing incidents, problem management, risk items, and audit obligations.
- Ability to design and maintain metrics and reporting frameworks that provide visibility into platform health, performance, and compliance.
- Exceptional written and verbal communication skills for executive‑level insight translation.
- Proven leadership of high‑performing, highly technical teams through accountability, coaching, and clear ownership.
- Agile/Scrum experience with strong prioritization and deadline management.
- Preferred: production change control process experience; audit and compliance work in regulated industries; CIS, NIST exposure.
- Preferred: financial services or similarly regulated industries experience.
- Required Technical Skills: SRE tooling (Prometheus, Grafana, PagerDuty, Datadog, etc.); cloud platforms (AWS, Azure, or GCP) and multi‑cloud exposure; container orchestration (Kubernetes, Kafka); CI/CD tooling (GitHub Actions, Jenkins, etc.); metrics/reporting platform design; FinOps principles and cloud cost governance; SecOps tooling and security governance practices.
- Preferred Technical Skills: GRC tooling (ServiceNow, Archer, etc.).
- Required Education: Bachelor’s degree in a technical discipline or equivalent experience.
- Required Experience: 15+ years in cloud engineering, platform reliability, or infrastructure roles, with 5+ years in senior engineering leadership.
- Certificates: AWS Solutions Architect Associate or higher strongly desired; other cloud or SRE certifications preferred.
Benefits
- Hybrid work environment with up to 2 days per week remote.
- Tuition reimbursement.
- Student loan repayment assistance.
- Technology stipend for remote work.
- Generous PTO, parental leave, and 401(k) with employer match.
- Competitive health benefits (medical, dental, vision).
All employees may be eligible for a discretionary bonus. Compensation will be determined by skills, experience, and education, and will be aligned with OCC’s pay equity policies.
OCC is an Equal Opportunity Employer. We are committed to diversity, equity, and inclusion. OCC provides equal employment opportunities to all employees and applicants for employment without regard to race, color, national origin, citizenship status, sex, sexual orientation, gender identity or expression, disability, age, marital status, religion, veteran status, or any other characteristics protected by applicable federal, state, or local laws.