Staff Software Engineer, Platform Infrastructure

WebHosting

Ireland

Remote

EUR 110,000 - 150,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Pantheon is seeking a Staff Software Engineer focused on Platform Infrastructure from Ireland, with a strong SRE orientation. You will standardize reliability practices across 15+ engineering teams and help evolve the internal Service Foundation platform with Go, GCP, Terraform, and Grafana-centric observability.

You will influence across PIE and the Internal Platform Group, driving security, compliance, and scalable patterns while mentoring engineers and shaping the reliability culture of

Qualifications

  • Proven experience delivering large-scale platform reliability improvements.
  • Hands-on experience with GCP managed services.
  • Strong DevOps and incident response background.
  • Ability to mentor teams and raise the technical bar.

Responsibilities

  • Establish SRE as a discipline with clear goals and reliable practices.
  • Set observability standards using metrics, traces, and logs.
  • Design and maintain service templates and infrastructure patterns.
  • Default to GCP primitives and managed services to reduce complexity.
  • Drive security, compliance, and IAM/network controls.
  • Mentor engineers and promote a culture of reliability.
  • Operate in a full DevOps lifecycle including on-call rotation.
  • Lead with clear communication to technical and non-technical stakeholders.

Skills

SRE principles
Observability focus
Technical leadership
Communication

Tools

Terraform
Prometheus
OpenTelemetry
Grafana
Cloud Run
GKE
Cloud SQL
IAP
IAM

Job description

Staff Software Engineer, Platform Infrastructure
Job Description

About Pantheon

Pantheon brings together exceptional builders who take ownership, drive meaningful impact, and shape what’s next for the web. We power more than 300,000 websites globally for organizations including Google, Princeton, Salesloft, Clorox, and the United Nations. Every day, thousands of developers and marketers use our WebOps platform to build, iterate, and scale WordPress, Drupal, and Next.js sites that reach billions of people worldwide. As an employer, we operate with the same philosophy that drives our product: foundation and freedom. Pantheon is a vibrant, remote-forward team of experts who care deeply about their craft and results. Here, you take ownership of work that matters, contribute alongside exceptional people, and see the impact you create.

The Role

The Platform Infrastructure Engineering (PIE) team is the foundation that lets every engineering team at Pantheon move fast, safely. We build the patterns, templates, and tools that power a self-service developer experience – and we’re responsible for the reliability and observability standards Pantheon’s entire internal platform depends on.

We’re looking for a Staff Software Engineer with a strong SRE orientation to take our platform reliability practice to the next level. Today, reliability is distributed and inconsistent across teams. This role exists to change that – establishing Site Reliability Engineering as a first-class discipline within our group and setting the patterns that 15+ engineering teams will use to ship quickly and safely for years to come.

In the short term, the focus will be on building and finalizing our new internal Service Foundation platform – meeting strict compliance requirements while making it easier, more reliable, and quicker to ship. Long term, as that foundation matures, the focus shifts toward SRE best practices: formalizing SLO/SLI frameworks, incident response culture, and reliability engineering across the organization. You’ll be a technical voice not just within PIE, but across the broader Internal Platform Group – working with adjacent teams to drive alignment and raise the bar.

The new platform runs on Go, GCP primitives (Cloud Run primarily, with GKE being strategically used), Terraform, and a Grafana-centered observability stack.

Pantheon’s core values are Trust, Teamwork, Passion, and Customers First. Within engineering, we value collaboration, character, autonomy, and a no-blame culture.

What You Will Do

  • Establish SRE as a discipline.Define and drive adoption of SLO/SLI frameworks, reliability standards, and incident response practices across PIE and the broader Internal Platform Group. Build a quieter, more reliable internal platform.
  • Set the observability standard.Own and evolve our observability patterns – Prometheus metrics, OpenTelemetry tracing, and structured logging – so every team has clear, actionable signals. Observability is a first class citizen here.
  • Build and maintain the paved path.Design and deliver the service templates, Terraform modules, and infrastructure patterns that make success easy: Go APIs, CLIs, Cloud Run services, and GKE workloads that come with reliability and observability built in from day one.
  • Default to GCP primitives.Contribute to key strategic initiatives – migrating to GCP Secret Manager, standardizing on GitHub Actions and Cloud Build, moving to Cloud SQL – by establishing GCP-native managed services as the default for scaling. Help the organization move away from self-hosted complexity as the default, and toward primitives that reduce operational burden.
  • Drive security and compliance.Apply security-first thinking to networking and identity – IAP, IAM, VPC design – and support compliance posture across SOC 2, ISO 27001, PCI DSS, and other frameworks.
  • Mentor the team.Guide engineers in designing and implementing high-impact reliability and infrastructure work, raising the technical bar across PIE and adjacent teams.
  • Own the full lifecycle.Operate in a full DevOps model – development, testing, operations, and support for the systems you build.
  • Participate in on-call.After an initial ramp period (typically 3-6 months), join on-call rotation and actively contribute to reducing its burden through better automation, runbooks, and reliability work.

What You Need to Succeed

  • Site Reliability Engineering principles:deep understanding of SLOs, SLIs, error budgets, toil reduction, and how to operationalize these practices in a team that hasn’t had them before.
  • Cloud infrastructure:hands-on expertise with GCP services – Cloud Run, GKE, Cloud SQL, GCP Secret Manager, IAP/IAM, networking – with a strong bias toward managed and serverless over self-hosted.
  • Observability:practical experience designing and implementing all three observability pillars (metrics, traces, logs) using Prometheus, OpenTelemetry, and structured logging. Grafana is our first-class observability platform – practical experience in turning signals into actionable reliability improvements, not just pretty graphs.
  • Infrastructure as code:strong Terraform skills, including module design for reusability and adoption across teams.
  • Security mindset:security is foundational, not a checkbox – IAM least-privilege, IAP, network controls, secrets management, and compliance requirements inform how you build.
  • Technical leadership:experience setting technical direction for a platform or team, translating ambiguous reliability goals into concrete architecture, and influencing engineers who don’t report to you.
  • Platform engineering philosophy:you think in patterns and self-service – you’d rather give teams a great pattern or template than solve their specific problem for them.
  • Communication:clear and direct – you can articulate reliability risk, architectural decisions, and incident postmortems to both engineers and non-technical stakeholders.

Working At Pantheon From Ireland

This role is based in Ireland and can be performed remotely within the country. Pantheon has a distributed engineering culture – you’ll collaborate primarily with teams in North America and Europe, which means some scheduling flexibility is expected for cross-timezone standups and incident response. Pantheon complies with all applicable Irish employment law including statutory leave entitlements, and compensation is benchmarked to the Irish market.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Platform Infrastructure at Pantheon
Staff Software Engineer, Platform Infrastructure at Pantheon

Sharkey • Ireland

Remote
EUR 120,000 - 150,000
Equity plan
28 days vacation
Private medical and dental
+5
Staff Security Engineer – Security Operations
Staff Security Engineer – Security Operations

WebHosting • Dublin

Hybrid
EUR 120,000 - 180,000
Staff Platform Reliability Engineer
Staff Platform Reliability Engineer

Sharkey • Ireland

Remote
EUR 120,000 - 150,000
Equity plan
28 days vacation
Private medical and dental
+5
Staff Platform Reliability Engineer (SRE/Observability)
Staff Platform Reliability Engineer (SRE/Observability)

WebHosting • Ireland

Remote
EUR 110,000 - 150,000
Staff Security Engineer – Security Operations – Ireland
Staff Security Engineer – Security Operations – Ireland

Webhosting • Ireland

On-site
EUR 150,000 - 190,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Okta • Ireland

Hybrid
EUR 140,000 - 190,000
Work from home opportunities
Health + Wellness
Financial Benefits
+4
Staff Security Engineer - Security Operations
Staff Security Engineer - Security Operations

Pantheon • Ireland

On-site
GBP 120,000 - 180,000
Industry competitive compensation and
Equity plan
28 days of holiday
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Okta • Ireland

Hybrid
EUR 110,000 - 150,000
Work from home opportunities
Health + Wellness
Financial Benefits
+4
Site Reliability Engineer
Site Reliability Engineer

Fulcrum Digital Inc • Dublin

On-site
EUR 110,000 - 140,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Dublin

On-site
EUR 92,000 - 127,000