Staff Site Reliability Engineer

WEX Inc.

United States

On-site

USD 121,000 - 151,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Retirement savings plan
Paid time off
Tuition reimbursement

Job summary

WEX Inc. is seeking a Staff Site Reliability Engineer (SRE) to lead reliability strategy across a diverse platform.

You will architect systems for availability, scale, and observability while mentoring engineers and guiding AI-enabled automation initiatives. You will drive incident management, capacity planning, and cost optimization, collaborating with cross-functional teams to reduce toil and improve performance.

Qualifications

  • 8+ years of experience in large-scale system reliability.
  • Proven expertise in system architecture, cloud platforms, and automation frameworks.
  • Deep knowledge of Kubernetes, service meshes, and distributed tracing.
  • Experience with monitoring and logging platforms (Grafana, ELK, Splunk).

Responsibilities

  • Architect and oversee mission-critical systems with focus on availability and scalability.
  • Define and enforce SRE practices and operational standards.
  • Lead cross-functional initiatives to improve reliability, performance and efficiency at scale.
  • Mentor engineers on production-grade AI-enabled reliability solutions.

Skills

Large-scale reliability
Kubernetes
Service meshes
Distributed tracing
Incident management
AI/agent tooling
Cloud platforms
Automation frameworks
Observability

Tools

Grafana
ELK stack
Splunk
OpenTelemetry
Jaeger
Prometheus

Job description

About the Role

We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative impact across WEX’s platform reliability and operational excellence. This is a particularly exciting time to be part of the SRE function at WEX. Our diverse product ecosystem supports a wide array of customer businesses and generates rich, complex telemetry across applications, infrastructure, and platforms. Ensuring these systems are scalable, observable, and resilient is critical to unlocking business value and customer success. As a Staff SRE, you will play a pivotal role in shaping the reliability engineering strategy at WEX. You’ll architect and lead efforts that improve availability, performance, and efficiency at scale, driving initiatives across observability, automation, incident management, problem management, capacity planning, and performance optimization. You’ll be hands‑on in building foundational tooling and frameworks while also acting as a multiplier, mentoring engineers, aligning cross‑functional teams, and influencing platform decisions with a strong reliability lens. You’ll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high‑TOIL work, integrating safely into our AI ecosystem, and establishing security and operational guardrails so intelligent automation is trustworthy, measurable, and scalable. Our team embraces agile development, a strong product mindset, and modern engineering practices, including AI‑assisted operations and intelligent automation. You’ll take on some of the most complex, high‑impact challenges at WEX, supported by a team of highly skilled engineers and technical leaders invested in your success and growth. If you’re a senior technical leader passionate about building reliable systems, leading through influence, and making a meaningful impact with AI‑enabled operations, this is a fantastic opportunity for you.

What You’ll Do

Architect and oversee the implementation of mission‑critical systems with a focus on availability, scalability, and operational excellence. Define and enforce SRE best practices and operational standards across engineering and platform teams. Lead cross‑functional initiatives to enhance system reliability, performance, and efficiency at scale. Serve as a technical advisor for engineering leadership on reliability, architecture, and operational risk. Develop capacity planning and load testing strategies that proactively identify and mitigate scalability risks. Design self‑healing and auto‑recovery mechanisms that reduce manual intervention during failures. Drive cloud cost optimization and budgeting initiatives without compromising reliability. Design, build, and govern AI agents and reusable skills that automate operational workflows and reduce TOIL. Evaluate and integrate AI ecosystems, including models, agent frameworks, orchestration, tooling interfaces, and evaluation practices, into SRE and platform workflows. Apply AI security and governance controls, including least‑privilege tool access, secure data and prompt handling, auditability, and safe automation boundaries. Lead AI‑enabled initiatives for incident response, runbook automation, anomaly detection, and capacity/performance insights, with clear measurement of TOIL reduction and reliability outcomes. Mentor engineers on production‑grade agentic solutions and help embed AI into day‑to‑day reliability practices.

What You’ll Bring

8+ years of experience with a focus on large‑scale system reliability. Expertise in system architecture, cloud platforms, and automation frameworks. Deep knowledge of Kubernetes, service meshes, and distributed tracing. Experience with monitoring and logging platforms (Grafana, ELK stack, Splunk, etc.). Knowledge of containerization and orchestration (Docker, Kubernetes). Experience designing high‑availability, fault‑tolerant architectures. Strong understanding of database reliability engineering (MySQL, PostgreSQL, NoSQL), plus networking, databases, and storage architectures. Excellent incident command and crisis management skills. Hands‑on experience building AI agents and skills/tools that integrate with operational systems (APIs, observability, ticketing, CI/CD). Working knowledge of AI ecosystems and agent architectures, including orchestration, tool calling, context/memory, evaluation, and human‑in‑the‑loop patterns. Practical understanding of AI security and governance for production use, secure permissions, data leakage prevention, secrets handling, and guarded autonomous actions. Demonstrated ability to reduce TOIL with AI by automating repetitive operational work and delivering measurable efficiency and reliability gains.

Nice to Have

Experience with multi‑region and multi‑cloud deployments. Deep expertise in scalable microservices and event‑driven architectures. Strong experience with advanced observability tools (OpenTelemetry, Jaeger, Prometheus). Leadership in driving large‑scale SRE transformations. Experience designing and developing AI agents, skills, and copilots for SRE/platform engineering, including evaluation and safe rollout practices. Familiarity with enterprise agent platforms, skill registries, and observability for AI/agent workflows. Ability to influence engineering culture and process improvements, including adoption of AI‑assisted operations under change control, safety, and audit requirements.

The base pay range represents the anticipated low and high end of the pay range for this position. Actual pay rates will vary and will be based on various factors, such as your qualifications, skills, competencies, and proficiency for the role. Base pay is one component of WEX's total compensation package. Most sales positions are eligible for commission under the terms of an applicable plan. Non‑sales roles are typically eligible for a quarterly or annual bonus based on their role and applicable plan. WEX's comprehensive and market competitive benefits are designed to support your personal and professional well‑being. Benefits include health, dental and vision insurances, retirement savings plan, paid time off, health savings account, flexible spending accounts, life insurance, disability insurance, tuition reimbursement, and more. For more information, check out the "About Us" section. Pay Range: $120,600.00 - $150,900.00 WEX is a global commerce platform that helps businesses solve for operational complexities like employee benefits, managing and mobilizing fleets, and streamlining payments. With over 6,500 employees, we work with large and small companies in more than 200 countries and territories, and can tailor our services to meet the unique needs of their businesses. We hire people who share our passion for continuous innovation and client service that is unparalleled in the industry. Offering comprehensive and market competitive benefits, our offerings are designed to support your personal and professional well‑being. If you’re looking for a growing career - come be part of WEX today.

WEX is an equal opportunity employer committed to diversity and inclusion in the workplace. All qualified applicants will receive consideration for employment without regard to sex, race, color, age, national origin, religion, sexual orientation, gender identity, protected veteran status, disability or other protected status. WEX promotes a drug‑free workplace. Qualified individuals with a disability have the right to request a reasonable accommodation. If you require a reasonable accommodation as a result of your disability at any point in the job application process, please submit your request through our Reasonable Accommodation Request Form. This form is for accommodation requests only and cannot be used to inquire about the status of applications.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

WEX • Chicago (IL)

On-site
USD 121,000 - 151,000
Health insurance
Dental insurance
Vision insurance
+7
SRE Architect, AI-Powered Reliability
SRE Architect, AI-Powered Reliability

WEX, Inc. • Seattle (WA)

On-site
USD 200,000 - 251,000
Health, dental, vision insurance
Retirement plan
Paid time off
+1
SRE Architect, AI-Powered Reliability
SRE Architect, AI-Powered Reliability

WEX, Inc. • Dallas (TX)

On-site
USD 200,000 - 251,000
SRE Architect, AI-Powered Reliability
SRE Architect, AI-Powered Reliability

WEX, Inc. • Chicago (IL)

On-site
USD 200,000 - 251,000
Health insurance
Retirement savings plan
Paid time off
+1
Senior Software Engineer, Corporate Payments
Senior Software Engineer, Corporate Payments

WEX Inc. • Minneapolis (MN)

On-site
USD 122,000 - 146,000
SRE Architect, AI-Powered Reliability
SRE Architect, AI-Powered Reliability

WEX Inc. • Portland (ME)

On-site
USD 200,000 - 251,000
Health insurance
Dental and vision insurances
Retirement savings plan
+1
Sr Manager, Identity Operations & Engineering
Sr Manager, Identity Operations & Engineering

WEX Inc. • United States

On-site
USD 152,000 - 173,000
Health Insurance
Retirement Savings Plan
Paid Time Off
+6
Sr. Engineering Manager, Web Experience
Sr. Engineering Manager, Web Experience

WEX Inc. • Portland (ME)

Hybrid
USD 150,000 - 210,000
Health insurance
Dental insurance
Vision insurance
+7
Sr. Engineering Manager, Web Experience
Sr. Engineering Manager, Web Experience

WEX Inc. • United States

On-site
USD 180,000 - 240,000
Health, dental and vision insurances
Retirement savings plan
Paid time off
+5
Principal Software Engineer, Semantic Data Services
Principal Software Engineer, Semantic Data Services

WEX Inc. • United States

On-site
USD 201,000 - 250,000
Health insurance
Retirement plan
Paid time off
+2