Senior Site Reliability Engineer

Tracksuit Limited

Auckland

On-site

NZD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive compensation
ESOP
Wellness benefits
Parental leave
L&D budget
Flexible working
Office in Auckland

Job summary

Tracksuit is seeking a Senior Site Reliability Engineer to shape reliability, observability and security across our Auckland-based platform. You will guide cloud infrastructure, guardrails for agents, and golden paths to empower product teams.

Bring deep experience with IaC, AWS, Datadog, and incident management, plus strong communication to balance reliability with feature delivery. This is an office-based role with a collaborative, fast-paced startup culture.

Qualifications

  • Extensive hands-on IaC experience with Terraform/Terragrunt or CDK.
  • Production familiarity with AWS (or GCP/Azure) and ECS.
  • Strong observability skills using Datadog, tracing and structured logging.
  • Security-minded with least privilege, credential handling and incident readiness.
  • Experience working with agents and tooling (Claude Code or equivalent).
  • Ability to communicate health and risks to technical and non-technical stakeholders.
  • Collaborative, product-minded and customer-focused.

Responsibilities

  • Build and maintain cloud infrastructure with resilient defaults.
  • Provide self-serve observability: monitoring, alerting, logging, tracing and automation.
  • Make incidents easier to handle with runbooks and post-incident reviews.
  • Define identity/least-privilege model for non-human actors (agents, CI, MCP).
  • Document infrastructure designs and runbooks for both humans and bots.
  • Set safety/guardrails for shipping with minimal gatekeeping.
  • Oversee cloud spend visibility and cost-controls.
  • Coach engineers in reliability practices and align with product goals.
  • Lead incident command on real production incidents and scale the platform.

Skills

Infrastructure as Code
Cloud platforms
Observability
Security best practices
Agentic systems
Collaboration
Product-minded communication

Tools

Terraform
Terragrunt
CDK
ECS
AWS
GCP/Azure
CI/CD pipelines
Python
TypeScript
Bash
Datadog
Claude Code

Job description

Tracksuit exists to help marketers prove their brand building is working. We give teams the data they need to make smarter decisions, convince stakeholders, defend budgets, and track their progress. The brand tracking industry is dominated by 100-page reports, static data and big price-tags. We're doing things differently by being built for the modern marketer: always-on, accessible, and approachable. We're now tracking more than 1,000 brands across 25 countries globally. With offices in Auckland, Sydney, London and New York City, we're scaling fast with a brilliant team of collaborative and ambitious humans. Our culture is defined by "high care, high performance". We strive to be the best and look after each other while we do it.

Are you our next Senior Site Reliability Engineer?

We’re on the lookout for a Senior Site Reliability Engineer to join Tracksuit, based in our Auckland office.

You’ll set how reliability, observability and security work at Tracksuit, from the AWS infrastructure underneath the platform through to the agents, MCP servers and model-backed features running on top. Teams build and run their own services, and the SRE team makes it straightforward for them to do that well, through the golden paths, guardrails, tooling and defaults that make the reliable way the easy way, plus the support to lean on when something does break.

  • Platform-wide impact: You set how reliability, observability and operational readiness work across the whole platform, and every team shipping on it feels the difference.
  • Agentic systems in production: You’ll set the guardrails, identity and blast-radius controls for agents operating against real systems and data, and work out what could go wrong before it does.
  • Golden paths: You’ll make the platform legible to agents as well as people, through golden paths, MCP servers, skills and documentation that both can actually use.
  • Startup pace with scaleup reach: Five years in, we’re working with over 1,000 brands across AU, NZ, USA and the UK, with plenty of scaling still ahead.
  • Build and maintain the cloud infrastructure and paved paths teams ship on, so resilience, security and scalability come as defaults rather than as decisions each team makes on its own.
  • Give teams observability they can self-serve: monitoring, alerting, logging and tracing, plus automation for provisioning, deployments and the operational work nobody should be doing by hand.
  • Make incidents easier to handle: the tooling, runbooks and practice that let whoever is closest to the problem debug it quickly, and post-incident reviews that turn into actual changes.
  • Set the identity and least-privilege model for non-human actors, including agents, CI and MCP servers, so teams can give agents real access without real risk.
  • Document infrastructure designs and operational procedures in a form both people and agents can act on, including machine-readable runbooks.
  • Set the standard for what is safe to ship, and make it easy to meet through guardrails and checks in the pipeline rather than through gatekeeping.
  • Give teams visibility and controls over cloud and inference spend, so the cost of what they run is something they can see and act on.
  • Coach engineers in reliability practices, and work with Engineering and Product to balance feature delivery against platform needs.
  • You’ve run incident command on real production incidents, and you’ve owned a platform through a meaningful scaling step.
  • Deep infrastructure skills: Infrastructure as Code (Terraform, Terragrunt, CDK), containers (ECS), cloud platforms (AWS preferred, or GCP/Azure), CI/CD, and scripting in Python, TypeScript or Bash.
  • Strong on observability: Datadog or similar, distributed tracing and structured logging, and a solid grasp of incident management methodology.
  • Security-minded: Networking and cloud architecture fundamentals, plus an understanding of the security model for agentic systems, including credential handling, least privilege and prompt injection.
  • Works with agents: You use coding agents in your own work (Claude Code or equivalent) and have a point of view on where they help and where they don’t.
  • Product-oriented: You navigate technical complexity with business outcomes in mind, and can clearly communicate system health, risks and trade-offs to technical and non-technical people.
  • Kind: This is a collaborative role, so we ultimately want a great, supportive team player.
  • Bonus points if you’re passionate about marketing, design, and building exceptional user experiences.

Some of the tools we use: AWS, ECS, Terraform and Terragrunt, GitHub Actions, Datadog, Claude Code, Linear, Notion, Postgres, DynamoDB and Snowflake.

We’re a tight-knit, supportive, and ambitious team, driven to empower companies to use brand to drive success. Our culture thrives on complete transparency, trust, learning, and constant development and improvement.

Underpinning the experience are our great benefits, including:
  • Compensation: Competitive market rate remuneration, which is reviewed twice annually. Our radically transparent compensation policy ensures that salaries are fair across the entire team.
  • Employee Share Option Program (ESOP): So that everyone on the team has a share in Tracksuit’s success.
  • Progressive health and wellness benefits: Including an annual wellness bonus, access to a premium EAP platform, and 6 weeks of paid annual leave.
  • Generous parental benefits: 12 weeks’ paid parental leave for either caregiver, additional sick leave for IVF, gradual return to work.
  • A $1000 personal L&D budget for each Trackstar, plus additional growth opportunities including mentorships, speaking engagements, and travel.
  • Flexible working: We have beautiful offices in Auckland, Sydney, London, and New York. We are an office first work environment, however we understand that life can get messy. We provide flexibility in our approach to working options and will work with you to find the right balance and approach to WFH/in-office work.
  • Most importantly, when you join, you’ll receive an epic Tracksuit which reflects our vibe. We are built for speed and comfort, we're fun and informal, and we're practical and ready for anything.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tracksuit • Auckland

On-site
NZD 150,000 - 190,000
ESOP
Health & wellness
Parental leave
+3
Engineering Manager
Engineering Manager

Monograph • Auckland

Hybrid
NZD 160,000 - 190,000
Competitive remuneration
Employee Share Option Program
Annual wellness bonus
+3
Senior Full Stack Engineer
Senior Full Stack Engineer

Tracksuit • Auckland

On-site
NZD 100,000 - 130,000
Competitive market rate remuneration
Employee Share Option Program (ESOP)
Annual wellness bonus
+2
Head of Product
Head of Product

Tracksuit • Auckland

Hybrid
NZD 100,000 - 130,000
ESOP for everyone
Transparent salaries
L&D budget
+3
Senior Data Scientist
Senior Data Scientist

Tracksuit • Auckland

On-site
NZD 90,000 - 120,000
Competitive salary review twice annually
Employee Share Option Program
Progressive health and wellness benefits
+2
Head of Product
Head of Product

black.ai • Auckland

Hybrid
NZD 205,000 - 270,000
Employee Stock Ownership Plan (ESOP)
Transparent salaries
Learning & Development budget
+4
Senior Analytics Engineer
Senior Analytics Engineer

black.ai • Auckland

On-site
NZD 135,000 - 208,000
Competitive salary
ESOP program
Wellbeing benefits
+3
Senior Analytics Engineer
Senior Analytics Engineer

Tracksuit Limited • Auckland

On-site
NZD 140,000 - 190,000
6 weeks paid leave
Annual wellness bonus
Employee stock options (ESOP)
+1
Senior Analytics Engineer
Senior Analytics Engineer

Tracksuit • Auckland

On-site
NZD 140,000 - 180,000
Equity
6 weeks leave
Wellbeing bonus
+3
Head of Finance
Head of Finance

Tracksuit • Auckland

Hybrid
NZD 180,000 - 240,000
ESOP
Competitive salary
Wellness benefits
+4