Senior Software Engineer- Product Reliability Engineering

Visa

Austin (TX)

On-site

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
401(k)
FSA/HSA
Paid time off
Wellness program

Job summary

Visa is seeking a Software Development Engineer on the Product Reliability Engineering (PRE) team to design, build, and automate large-scale infrastructure and data platforms. You will write Python, manage AI-enabled tooling, and contribute to reliable, real-time payment systems.

This role emphasizes automation, observability, scalability, and security across cloud-native and hybrid environments, with opportunities to shape reliability practices and ship agentic AI-powered tooling.

Qualifications

  • Bachelor's degree with 2+ years of relevant professional experience, OR an advanced degree with at least 2 years of relevant experience, OR 5+ years of relevant work experience.
  • Hands-on software engineering or automation experience using Python, Java, Go, JavaScript/TypeScript, or a comparable language.
  • Experience supporting Linux-based systems and troubleshooting distributed applications or infrastructure.
  • Hands-on experience with one or more logging or search platforms such as Splunk, ClickHouse, OpenSearch, or Elasticsearch.
  • Hands-on experience with metrics and visualization technologies such as Prometheus, Thanos, Grafana, or Bosun.
  • Experience developing backend services, APIs, command-line tools, integrations, or operational automation.
  • Experience with cloud platforms, preferably AWS or GCP, and with cloud-native architecture and services.
  • Experience with containers and orchestration technologies such as Docker, Kubernetes, or equivalent enterprise container platforms.
  • Experience with infrastructure as code, configuration management, CI/CD pipelines, Git, automated testing, and deployment tooling.
  • Understanding of telemetry pipelines, log collection and parsing, metrics collection, alerting, dashboards, data retention, access controls, and platform integrations.
  • Understanding of distributed systems, scalability, high availability, disaster recovery, performance tuning, capacity planning, and reliability engineering.
  • Experience in production incident troubleshooting, root cause analysis, problem management, and implementing preventive remediation.
  • Experience with vulnerability remediation, secure configuration, certificate management, patching, upgrades, and software lifecycle management.
  • Ability to translate user and platform requirements into maintainable engineering solutions and clear technical documentation.
  • Strong problem-solving, communication, and collaboration skills, with the ability to work effectively across globally distributed teams.

Responsibilities

  • Design, develop, and deploy end-to-end automation for deployment pipelines, infrastructure provisioning, platform operations, and release orchestration across complex production environments.
  • Write clean, production-grade Python (and Go or Bash where it counts) to eliminate toil, reduce operational risk, and improve reliability and scalability of critical engineering workflows.
  • Design and implement reusable frameworks for release scheduling, validation, rollback, reporting, and configuration management that support the software delivery lifecycle.
  • Drive automation initiatives that improve engineering efficiency, standardization, and operational excellence across teams.
  • Design, build, operate, and continuously improve relational database platforms supporting critical payment systems and high-volume transaction processing.
  • Contribute to architecture decisions, platform enhancements, and engineering solutions that improve scalability, resiliency, and performance.
  • Lead database health and lifecycle operations including upgrades, patching, backup and recovery strategies, and platform modernization efforts.
  • Analyze and optimize database performance through index tuning, execution plan analysis, replication monitoring, and capacity management.
  • Develop automation for database operations, configuration management, and schema deployments using tools such as Ansible, Liquibase, and CI/CD pipelines.
  • Build proactive monitoring, observability, and reporting solutions that identify reliability risks before they impact production services.
  • Design and build GenAI-powered engineering solutions that automate deployment orchestration, operational workflows, release governance, and platform management.
  • Integrate LLM-driven capabilities into observability, incident response, troubleshooting, and developer productivity workflows to improve operational effectiveness.
  • Evaluate and implement emerging AI, automation, and machine learning technologies that improve reliability, efficiency, and engineering velocity.
  • Contribute to agentic automation strategies that help evolve PRE into an increasingly intelligent and autonomous engineering organization.
  • Design and build dashboards, alerts, telemetry pipelines, and health indicators using Prometheus, Grafana, Splunk, or ELK to provide visibility across globally distributed systems.
  • Analyze platform performance, reliability, utilization, and availability data to identify trends and implement long-term improvements.
  • Lead troubleshooting efforts across infrastructure, applications, databases, and platform services, performing root cause analysis and driving durable corrective actions.
  • Design and implement self-healing, automated remediation, and auto-scaling capabilities that improve system resilience and reduce operational overhead.
  • Design and implement highly available, scalable infrastructure solutions that support business-critical payment systems operating at global scale.
  • Ensure platforms and services meet security, compliance, governance, and resiliency requirements across cloud-native and hybrid environments.
  • Drive vulnerability remediation, configuration hardening, patch management, and security automation efforts to improve platform security posture.
  • Partner with engineering teams to build reliability and security practices directly into the software development lifecycle.
  • Partner with software engineers, product managers, platform teams, and global PRE peers to design, deliver, and operate reliable engineering solutions.
  • Participate in architecture reviews, design discussions, code reviews, and technical planning activities, contributing engineering expertise and best practices.
  • Create and maintain technical documentation, runbooks, operational procedures, and engineering standards that improve team effectiveness and knowledge sharing.
  • Participate in on-call rotations and incident response activities, driving operational improvements and helping teams learn from production events.
  • Take ownership of assigned initiatives from design through implementation, deployment, and operational support while continuously seeking opportunities to improve systems and processes.

Skills

Python
Go
JavaScript/TypeScript
Linux
Splunk/OpenSearch/Elasticsearch
Prometheus/Grafana
AWS/GCP
Docker/Kubernetes
CI/CD
IaC/Config Mgmt

Education

Bachelor's degree

Tools

Ansible
Liquibase
CI/CD pipelines

Job description

About Us

Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.

At Visa, you'll have the opportunity to create impact at scale - tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.

Join Visa and do work that matters - to you, to your community, and to the world. Progress starts with you.

Job Description

Every time someone taps, swipes, or clicks to pay- Visa infrastructure makes it happen in milliseconds, across 200+ countries. As a Software Development Engineer on the Product Reliability Engineering (PRE) team, you won't just watch those systems run- you'll be one of the engineers building, automating, and evolving them.

PRE is not a traditional ops team. We are a software engineering organization that treats infrastructure as code, reliability as a product, and automation as a strategic advantage. You'll write Python, build agentic AI tools, manage data platforms, and contribute to the distributed systems that process billions of real-time transactions. From day one, you are an engineer- and from day one, your work matters.

If you are endlessly curious about how large-scale systems stay resilient, obsess over elegant automation, and want to launch your career at the intersection of AI, infrastructure, and global financial technology - this role was built for you.

Build Automation That Scales
  • Design, develop, and deploy end-to-end automation for deployment pipelines, infrastructure provisioning, platform operations, and release orchestration across complex production environments.
  • Write clean, production-grade Python (and Go or Bash where it counts) to eliminate toil, reduce operational risk, and improve the reliability and scalability of critical engineering workflows.
  • Design and implement reusable frameworks for release scheduling, validation, rollback, reporting, and configuration management that support the software delivery lifecycle.
  • Drive automation initiatives that improve engineering efficiency, standardization, and operational excellence across teams.
Manage & Evolve Data Platforms
  • Design, build, operate, and continuously improve relational database platforms supporting critical payment systems and high-volume transaction processing.
  • Contribute to architecture decisions, platform enhancements, and engineering solutions that improve scalability, resiliency, and performance.
  • Lead database health and lifecycle operations including upgrades, patching, backup and recovery strategies, and platform modernization efforts.
  • Analyze and optimize database performance through index tuning, execution plan analysis, replication monitoring, and capacity management.
  • Develop automation for database operations, configuration management, and schema deployments using tools such as Ansible, Liquibase, and CI/CD pipelines.
  • Build proactive monitoring, observability, and reporting solutions that identify reliability risks before they impact production services.
Ship Agentic AI & ML-Powered Tools
  • Design and build GenAI-powered engineering solutions that automate deployment orchestration, operational workflows, release governance, and platform management.
  • Integrate LLM-driven capabilities into observability, incident response, troubleshooting, and developer productivity workflows to improve operational effectiveness.
  • Evaluate and implement emerging AI, automation, and machine learning technologies that improve reliability, efficiency, and engineering velocity.
  • Contribute to agentic automation strategies that help evolve PRE into an increasingly intelligent and autonomous engineering organization.
Own Observability & Platform Health
  • Design and build dashboards, alerts, telemetry pipelines, and health indicators using tools such as Prometheus, Grafana, Splunk, or ELK to provide visibility across globally distributed systems.
  • Analyze platform performance, reliability, utilization, and availability data to identify trends and implement long-term improvements.
  • Lead troubleshooting efforts across infrastructure, applications, databases, and platform services, performing root cause analysis and driving durable corrective actions.
  • Design and implement self-healing, automated remediation, and auto-scaling capabilities that improve system resilience and reduce operational overhead.
Engineer for Reliability & Security
  • Design and implement highly available, scalable infrastructure solutions that support business-critical payment systems operating at global scale.
  • Ensure platforms and services meet security, compliance, governance, and resiliency requirements across cloud-native and hybrid environments.
  • Drive vulnerability remediation, configuration hardening, patch management, and security automation efforts to improve platform security posture.
  • Partner with engineering teams to build reliability and security practices directly into the software development lifecycle.
Collaborate, Learn & Grow Fast
  • Partner with software engineers, product managers, platform teams, and global PRE peers to design, deliver, and operate reliable engineering solutions.
  • Participate in architecture reviews, design discussions, code reviews, and technical planning activities, contributing engineering expertise and best practices.
  • Create and maintain technical documentation, runbooks, operational procedures, and engineering standards that improve team effectiveness and knowledge sharing.
  • Participate in on-call rotations and incident response activities, driving operational improvements and helping teams learn from production events.
  • Take ownership of assigned initiatives from design through implementation, deployment, and operational support while continuously seeking opportunities to improve systems and processes.

Visa requires at least 3 days in office, expectations of these days will be confirmed by your Hiring Manager.

Qualifications
Basic Qualifications
  • Bachelor's degree with 2+ years of relevant professional experience, OR an advanced degree with at least 2 years of relevant experience, OR 5+ years of relevant work experience.
Preferred Qualifications
  • Hands-on software engineering or automation experience using Python, Java, Go, JavaScript/TypeScript, or a comparable language.
  • Experience supporting Linux-based systems and troubleshooting distributed applications or infrastructure.
  • Hands-on experience with one or more logging or search platforms such as Splunk, ClickHouse, OpenSearch, or Elasticsearch.
  • Hands-on experience with metrics and visualization technologies such as Prometheus, Thanos, Grafana, or Bosun.
  • Experience developing backend services, APIs, command-line tools, integrations, or operational automation.
  • Experience with cloud platforms, preferably AWS or GCP, and with cloud-native architecture and services.
  • Experience with containers and orchestration technologies such as Docker, Kubernetes, or equivalent enterprise container platforms.
  • Experience with infrastructure as code, configuration management, CI/CD pipelines, Git, automated testing, and deployment tooling.
  • Understanding of telemetry pipelines, log collection and parsing, metrics collection, alerting, dashboards, data retention, access controls, and platform integrations.
  • Understanding of distributed systems, scalability, high availability, disaster recovery, performance tuning, capacity planning, and reliability engineering.
  • Experience in production incident troubleshooting, root cause analysis, problem management, and implementing preventive remediation.
  • Experience with vulnerability remediation, secure configuration, certificate management, patching, upgrades, and software lifecycle management.
  • Ability to translate user and platform requirements into maintainable engineering solutions and clear technical documentation.
  • Strong problem-solving, communication, and collaboration skills, with the ability to work effectively across globally distributed teams.
Information for US Applicants

Visa has a comprehensive benefits package for which this position may be eligible that includes Medical, Dental, Vision, 401(k), FSA/HSA, Life Insurance, Paid Time Off, and Wellness Program.

Work Hours

Varies upon the needs of the department.

Travel Requirements

This position requires travel 5-10% of the time.

Mental/Physical Requirements

This position will be performed in an office setting. The position will require the incumbent to sit and stand at a desk, communicate in person and by telephone, frequently operate standard office equipment, such as telephones and computers.

Visa is an EEO Employer

Qualified applicants will receive consideration for employment without regard to race, color religion, sex, national origin, sexual orientation, gender identity, disability or protect veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with the EEOC guidelines and applicable local law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Development Engineer- Product Reliability Engineering
Software Development Engineer- Product Reliability Engineering

Visa • Austin (TX)

On-site
USD 140,000 - 170,000
Medical Insurance
401(k) Plan
Paid Time Off
Senior Software Engineer
Senior Software Engineer

VISA • Bellevue (WA)

Hybrid
USD 111,000 - 172,000
Medical
Dental
Vision
+5
Software Engineer- Operations and Infrastructure
Software Engineer- Operations and Infrastructure

Visa • Austin (TX)

On-site
USD 100,000 - 150,000
Benefits package
Staff SW Engineer
Staff SW Engineer

Visa • Foster City (CA)

On-site
USD 146,000 - 234,000
Sr. SW Engineer
Sr. SW Engineer

VISA • Foster City (CA)

On-site
USD 123,000 - 191,000
Medical
Dental
Vision
+5
Sr. Software Engineer
Sr. Software Engineer

Visa • Austin (TX)

Hybrid
USD 112,000 - 204,000
Medical benefits
401(k)
Wellness Program
Senior Software Engineer- PaaS Engineering
Senior Software Engineer- PaaS Engineering

Visa • Austin (TX)

On-site
USD 111,000 - 172,000
Medical, Dental, Vision
401(k)
Paid Time Off
+1
Software Engineer, Sr. Consultant Level
Software Engineer, Sr. Consultant Level

Visa • Bellevue (WA)

Hybrid
USD 163,000 - 260,000
Lead SW Engineer
Lead SW Engineer

Visa • Foster City (CA)

On-site
USD 192,300 - 307,600
Director, Software Engineering
Director, Software Engineering

Visa • Austin (TX)

On-site
USD 187,000 - 299,000