AIOPs Observability/SRE Lead

Altera Corporation

San Jose, Northern (CA, KY)

Hybrid

USD 187,000 - 271,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Altera Corporation in San Jose, CA is seeking an Onsite AIOps Observability/SRE Lead to build and guide a reliability engineering team across critical IT and engineering systems. You will shape scalable monitoring, incident response, and automation strategies to improve platform stability and performance.

You will partner with IT, security, and cloud teams to modernize networks, enable hybrid cloud, and ensure secure, reliable connectivity.

Qualifications

  • Bachelor’s degree in CS/IT/Engineering with 10+ years in enterprise networking.
  • 10+ years in site reliability engineering or DevOps focused on platform reliability.
  • Proven leadership of engineering teams in a senior/manager capacity.
  • Deep understanding of SRE principles: SLIs, SLOs, error budgets, reliability frameworks.
  • Strong observability, incident management, blameless post-mortem culture.
  • Experience driving automation and toil-reduction across infra and platforms.
  • Knowledge of Azure/AWS and containerized environments.
  • Excellent stakeholder communication for cross-functional collaboration.

Responsibilities

  • Lead and grow the SRE team incl. hiring, mentoring, and developing reliability capabilities.
  • Define SRE practice including SLIs, SLOs, error budgets, and reliability targets across platforms.
  • Drive automation initiatives to eliminate toil and improve reliability.
  • Oversee incident management, blameless post-mortems, and improvement programs.
  • Collaborate with cloud, infra, security, and external service providers on reliability.
  • Establish observability standards: monitoring, alerting, logging, tracing.
  • Manage on-call processes, escalation, and 24x7 operations.
  • Report on platform reliability, SLO compliance, and maturity to IT leadership.
  • Partner with teams to design reliable network connectivity solutions.
  • Document AIOps architectures, configurations, standards, and procedures.
  • Provide technical leadership to Observability and operations teams.
  • Identify opportunities for automation and modernization of ops.

Skills

SRE leadership
Cloud platforms (Azure, AWS)
Observability
Incident management
Automation

Education

Bachelor’s degree in Computer Science, IT, Engineering

Job description

## AIOPs Observability/SRE LeadApply: San Jose, California, United States: Full time: Posted Yesterday: R03209# **Job Details:**### ## **Job Description:****Onsite Requirement:** This position requires regular in-office work and is an onsite role based in San Jose, CA. Candidates must be able to work onsite in San Jose, CA. **About Altera**At AlteraTM, our independence as the world’s largest pure-play FPGA solutions provider gives us the focus, speed, and agility to innovate without compromise. With more than four decades of industry-leading FPGA expertise, our singular mission is to deliver high-performance, flexible FPGA solutions that enable customers to solve their most complex computing challenges. As Altera continues to evolve and scale as an independent company, our IT organization is transforming its global infrastructure to support the needs of the business, engineering organizations, laboratories, and employees around the world. **About the Role**We are seeking a Site Reliability Engineering Manager to lead the SRE function. This role builds and manages a team of reliability engineers who drive platform stability, observability, automation, and operational excellence across the semiconductor company's critical IT and engineering systems. In this role, you will define and implement scalable, secure, and high-performance monitoring architectures that support both current and future business requirements. You will work closely with IT, Engineering, Integration teams, security teams, and external service providers to modernize network infrastructure, enable hybrid cloud connectivity, and ensure a smooth transition from existing environments to the future-state network. The ideal candidate brings deep expertise in enterprise SRE architecture and transformation, strong hands-on knowledge of routing and switching technologies, and experience integrating cloud environments such as AWS and Azure with large-scale on-premises infrastructure. **Key Responsibilities*** Lead and grow the SRE team including hiring, mentoring, and developing reliability engineering capabilities.* Define SRE practice including SLIs, SLOs, error budgets, and reliability targets across critical platforms.* Drive automation initiatives to eliminate toil and improve platform reliability and scalability.* Oversee incident management, blameless post-mortems, and systematic reliability improvement programs.* Collaborate with cloud, infrastructure, and application teams to embed reliability into platform design.* Establish observability standards including unified monitoring, alerting, logging, and tracing strategies.* Manage on-call processes, escalation procedures, and team wellbeing for 24x7 operations.* Report on platform reliability, SLO compliance, and operational maturity to IT leadership.* Collaborate with Integration teams, IT infrastructure teams, cybersecurity teams, and external service providers to design and implement reliable connectivity solutions.* Evaluate network technologies and solutions and provide technical recommendations based on business requirements, scalability, performance, security, and cost.* Ensure network transformation initiatives align with applicable security, regulatory, and compliance requirements.* Partner with cybersecurity teams to incorporate appropriate security controls, segmentation, firewall policies, VPN connectivity, and access controls into network architecture.* Develop and maintain comprehensive documentation for AIOps architectures, topologies, configurations, standards, policies, migration plans, and operational procedures.* Establish architecture standards, design principles, and best practices that promote consistency, scalability, reliability, and operational efficiency.* Provide technical leadership and guidance to Observability engineering and operations teams throughout architecture, implementation, migration, and optimization activities.* Identify opportunities for continuous improvement, automation, standardization, and modernization of network operations.* Troubleshoot and provide architectural guidance for complex Observability services* Stay current with emerging enterprise SRE, cloud networking, automation, and security technologies and assess their applicability to Altera’s environment. **Salary Range**The pay range below is for Bay Area California only. Actual salary may vary based on a number of factors including job location, job-related knowledge, skills, experiences, trainings, etc. We also offer incentive opportunities that reward employees based on individual and company performance. **$187,000 - $270,700 USD** We use artificial intelligence to screen, assess, or select applicants for the position. Applicants must be eligible for any required U.S. export authorizations. #LI-MD1### ## **Qualifications:****Minimum Qualifications*** Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field, with 10+ years of professional experience in enterprise networking.* 10+ years of experience in site reliability engineering or DevOps with a focus on platform reliability.* Proven experience leading technical engineering teams in a senior or management capacity.* Deep understanding of SRE principles including SLIs, SLOs, error budgets, and reliability frameworks.* Strong background in observability, incident management, and blameless post-mortem culture.* Experience driving automation and toil reduction programs across infrastructure and platform teams.* Knowledge of cloud platforms (Azure, AWS) and containerized environments.* Excellent communication skills for stakeholder reporting and cross-functional collaboration. **Preferred Qualifications*** Experience establishing SRE practices in semiconductor, HPC, or EDA-dependent environments.* Familiarity with AIOps platforms and AI-assisted incident management tooling.* Knowledge of chaos engineering practices and resiliency testing frameworks.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AIOPs Observability/SRE Lead
AIOPs Observability/SRE Lead

191 Altera Corporation • San Jose (CA)

On-site
USD 187,000 - 271,000
Senior Director, Silicon Hardware Engineering - Boards, Packaging, & Reliability
Senior Director, Silicon Hardware Engineering - Boards, Packaging, & Reliability

Altera Corporation • San Jose (CA)

On-site
USD 232,000 - 335,000
AIOps Observability & SRE Lead
AIOps Observability & SRE Lead

191 Altera Corporation • San Jose (CA)

On-site
USD 187,000 - 271,000
Lead AIOps Observability & SRE | Onsite in San Jose
Lead AIOps Observability & SRE | Onsite in San Jose

Altera Corporation • San Jose (CA), Northern (KY)

Hybrid
USD 187,000 - 271,000
AI Lead Architect - Silicon Design Execution
AI Lead Architect - Silicon Design Execution

Altera Corporation • San Jose (CA)

Hybrid
USD 206,000 - 303,000
Senior Director, Silicon Hardware Engineering - Boards, Packaging, & Reliability
Senior Director, Silicon Hardware Engineering - Boards, Packaging, & Reliability

191 Altera Corporation • San Jose (CA)

On-site
USD 232,000 - 335,000
Virtualization & Cloud Platform Engineer
Virtualization & Cloud Platform Engineer

Altera Corporation • San Jose (CA)

Hybrid
USD 149,000 - 213,000
Senior Director, Silicon Hardware Engineering - Boards, Packaging, & Reliability
Senior Director, Silicon Hardware Engineering - Boards, Packaging, & Reliability

altera • San Jose (CA)

On-site
USD 250,000 - 360,000
Senior Design Automation Engineer - Front End Design Verification
Senior Design Automation Engineer - Front End Design Verification

Altera Corporation • San Jose (CA)

On-site
USD 188,000 - 271,000
Senior Network Architect
Senior Network Architect

191 Altera Corporation • San Jose (CA)

On-site
USD 187,000 - 271,000