DevOps and Site Reliability Engineer

BigGeo Inc.

Calgary

Hybrid

CAD 90,000 - 130,000

Full time

19 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

BigGeo Inc. is seeking a DevOps and Site Reliability Engineer to design, automate, secure, and operate the infrastructure powering the Spatial Cloud. You will own reliability, deployment systems, observability, and incident response while collaborating across software, platform, data, and product teams.

You will build infrastructure as a product, automate everything, and enable engineers to move faster without sacrificing reliability. Calgary-based role with on-site work.

Qualifications

  • 4+ years of experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering.
  • Strong experience operating cloud infrastructure in production.
  • Experience building Infrastructure-as-Code solutions.
  • Experience managing Kubernetes and containerized workloads.
  • Knowledge of CI/CD pipelines and deployment automation.
  • Strong Linux systems administration skills.
  • Experience with monitoring, logging, and observability platforms.
  • Understanding of networking, security, and distributed systems concepts.
  • Experience supporting production incident response and troubleshooting.
  • Experience participating in an on-call rotation for production systems.
  • Strong communication and cross-functional collaboration skills.

Responsibilities

  • Design, deploy, and maintain scalable cloud infrastructure.
  • Build and manage Infrastructure-as-Code solutions.
  • Develop and optimize CI/CD pipelines for engineering teams.
  • Operate Kubernetes-based production environments.
  • Improve system reliability, performance, and fault tolerance.
  • Implement monitoring, observability, and alerting strategies.
  • Manage cloud networking, security, and access controls.
  • Automate operational processes and infrastructure workflows.
  • Lead incident response, root-cause analysis, and blameless postmortems.
  • Establish reliability standards, SLOs, and operational metrics.
  • Perform capacity planning for stateful and resource-intensive workloads.
  • Reduce operational toil through automation and elimination of recurring issues.
  • Collaborate with engineering teams to improve deployment velocity.
  • Evaluate and integrate AI-powered operational and automation tools.
  • Contribute to platform architecture decisions and infrastructure strategy.

Skills

Cloud infrastructure
Kubernetes
CI/CD
IaC
Observability
Security
On-call
Linux
Networking
GitOps

Tools

Terraform
Terragrunt
Pulumi
GitHub Actions
GitLab CI
ArgoCD
Flux
Jenkins
Prometheus
Grafana
Datadog
OpenTelemetry
ELK/OpenSearch
IAM
Secrets management
Vulnerability management
Cloud security tooling

Job description

DevOps and Site Reliability Engineer
About BigGeo

BigGeo is the Spatial Cloud. We help companies manage and access the world’s spatial data. Any size, any slice, any insight. Delivered in seconds.

We’re building something that hasn’t existed before: a new layer of the internet where the “where” and “when” behind every decision is instantly clear, programmable, and actionable. Our platform removes the complexity that has kept spatial data locked in silos for decades, and replaces it with speed, precision, and control.

We’re a Calgary-based company, early and moving fast, with real customers, real infrastructure, and a clear point of view on where the world is going.

Why BigGeo Exits and Why People Build Here

Most companies are spatially blind. They know what their data says, but not where or when things actually happen. That gap costs real money, creates real risk, and limits what AI can actually do in the physical world. BigGeo exists to close that gap. We’re not building another tool. We’re building the rails that connect the planet’s moving data to the systems that run the world. That’s a big problem, and it takes people who care about doing things right, not just fast.

People build here because:

  • The problem is real and the category is open. We’re not competing for the middle of an existing market. We’re defining a new one. Your work shapes what the category becomes.
  • Your fingerprints are on the architecture. We’re at the stage where the decisions you make today become the foundation tomorrow. What you ship matters.
  • We run on clarity, not politics. We move with purpose. No bureaucratic drag, just a team that agrees on the mission and gets to work.
  • You’ll grow fast because the problems are hard. Spatial data at scale is a genuinely difficult domain. If you want to be stretched, you’ll be stretched.
  • We’re building for longevity. We’re not chasing hype cycles. We’re building infrastructure, the kind that compounds in value over time and earns the trust of the companies that depend on it.
The Role

BigGeo is seeking a DevOps & Site Reliability Engineer to design, automate, secure, and operate the infrastructure that powers the Spatial Cloud.

This role is responsible for building reliable cloud infrastructure, deployment systems, observability platforms, security controls, and operational tooling that enable engineering teams to deliver production software with confidence.

The role combines DevOps and Site Reliability Engineering responsibilities. You will build the systems that deliver software to production, and you will own the reliability of what runs there — service level objectives, observability, capacity planning, incident response, and the on-call practice that supports them. We combine these deliberately: engineers who build delivery systems make better reliability decisions when they also operate what they ship.

You will work closely with software engineers, platform engineers, data engineers, and product teams to ensure BigGeo’s systems remain scalable, resilient, secure, and highly available as the platform grows.

This role is ideal for someone who enjoys building infrastructure as a product, automating everything possible, and creating systems that allow engineering teams to move faster without sacrificing reliability.

What You Will Build and Own

  • Cloud infrastructure supporting BigGeo production environments
  • Infrastructure-as-Code frameworks and deployment pipelines
  • Kubernetes clusters and container orchestration platforms
  • CI/CD systems supporting engineering delivery workflows
  • Monitoring, logging, alerting, and observability platforms
  • Security automation and compliance controls
  • Reliability engineering practices and operational standards
  • Service level objectives and reliability measurement frameworks
  • Incident response processes and on-call practices
  • Disaster recovery and business continuity capabilities
  • Cost optimization frameworks across cloud environments
  • Internal developer platforms and operational tooling
Key Responsibilities

Core Responsibilities

  • Design, deploy, and maintain scalable cloud infrastructure
  • Build and manage Infrastructure-as-Code solutions
  • Develop and optimize CI/CD pipelines for engineering teams
  • Operate Kubernetes-based production environments
  • Improve system reliability, performance, and fault tolerance
  • Implement monitoring, observability, and alerting strategies
  • Manage cloud networking, security, and access controls
  • Automate operational processes and infrastructure workflows
  • Lead incident response, root-cause analysis, and blameless postmortems
  • Establish reliability standards, SLOs, and operational metrics
  • Perform capacity planning for stateful and resource-intensive workloads
  • Reduce operational toil through automation and elimination of recurring issues
  • Collaborate with engineering teams to improve deployment velocity
  • Evaluate and integrate AI-powered operational and automation tools
  • Contribute to platform architecture decisions and infrastructure strategy

Reliability and On-Call

Reliability is treated as a core engineering responsibility at BigGeo rather than a separate function. This role carries meaningful ownership of it.

Reliability engineering

  • Define and maintain service level objectives for critical services
  • Build observability that makes system behavior measurable and actionable
  • Perform capacity planning for both growth and failure scenarios
  • Drive continuous improvement through postmortems and reliability reviews

On-call

  • Production alerting is automated and routed by severity
  • On-call responsibility is currently shared across the engineering team and is being formalized into a structured rotation as the team grows
  • You will participate in that rotation and help define escalation paths, response expectations, and handoff practices
  • New team members shadow incidents before taking primary responsibility
  • On-call scheduling and supporting policies are being established as part of this work

Technology and Tools

The exact stack will evolve as the platform grows, but experience with many of the following technologies is expected:

Cloud Platforms

  • AWS, Azure, or Google Cloud Platform

Infrastructure and Automation

  • Terraform
  • Terragrunt
  • Pulumi
  • Infrastructure-as-Code frameworks
  • GitOps workflows

Containers and Orchestration

  • Docker
  • Kubernetes
  • Helm

CI/CD

  • GitHub Actions
  • GitLab CI
  • ArgoCD
  • Flux
  • Jenkins

Observability

  • Prometheus
  • Grafana
  • Datadog
  • OpenTelemetry
  • ELK/OpenSearch

Reliability and Incident Management

  • SLO and SLI frameworks
  • Incident management and paging tools
  • Status and escalation workflows

Security

  • IAM
  • Secrets management
  • Cloud security tooling
  • Vulnerability management

Collaboration and Operations

  • Slack
  • Google Workspace
  • Monday
  • Linear

AI Tooling

  • Modern AI assistants and development tools
  • AI-supported operational automation
  • AI-enhanced monitoring and troubleshooting workflows
What You Bring

Required Experience

  • 4+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering
  • Strong experience operating cloud infrastructure in production environments
  • Experience building Infrastructure-as-Code solutions
  • Experience managing Kubernetes and containerized workloads
  • Knowledge of CI/CD pipelines and deployment automation
  • Strong Linux systems administration skills
  • Experience with monitoring, logging, and observability platforms
  • Understanding of networking, security, and distributed systems concepts
  • Experience supporting production incident response and troubleshooting
  • Experience participating in an on-call rotation for production systems
  • Strong communication and cross-functional collaboration skills

Preferred Experience

  • Experience supporting large-scale data platforms
  • Experience with geospatial or location-based systems
  • Experience building internal developer platforms
  • Experience with multi-cloud environments
  • Familiarity with high-volume data processing systems
  • Experience implementing reliability engineering practices
  • Experience defining SLOs, SLIs, and error budgets
  • Experience establishing or improving on-call and incident response processes
  • Knowledge of database operations and performance tuning
  • Experience working within startup or high-growth technology environments
  • Experience integrating AI systems into operational workflows

How This Role Contributes to the Spatial Cloud

The Spatial Cloud depends on infrastructure that is reliable, scalable, secure, and globally available.

As a DevOps & Site Reliability Engineer, you will build and operate the foundational systems that allow BigGeo to manage and access the world’s spatial data at massive scale. Your work directly impacts platform reliability, developer productivity, customer experience, and the company’s ability to deliver spatial insights in seconds.

The systems you build become critical components of the infrastructure powering the next generation of spatial computing.

Work Environment and Collaboration

BigGeo operates as a highly collaborative, AI-native startup environment.

This is an on-site role based in Calgary, Alberta. Candidates must be located in the Calgary area or willing to relocate.

Team members are expected to take ownership, move quickly, communicate clearly, and contribute beyond traditional role boundaries when needed. You will work closely with engineering, product, data, and leadership teams while helping establish the operational foundation of a category-defining company.

Success in this role requires curiosity, initiative, systems thinking, and a desire to build infrastructure that enables others to do their best work.

You will have significant influence over how BigGeo scales its platform, operations, and engineering capabilities as we continue defining the Spatial Cloud category.

Success Metrics

Success in this role means quickly developing a deep understanding of BigGeo’s production environment, taking meaningful ownership of platform reliability, and improving the systems that allow engineering teams to deploy and operate the Spatial Cloud safely at scale.

First 30 Days — Understand the Environment and Establish Operational Context

By the end of your first 30 days, you will:

  • Develop a working understanding of BigGeo’s cloud infrastructure, production architecture, Kubernetes environments, networking, deployment systems, and Infrastructure-as-Code.
  • Understand the current CI/CD workflows and how software moves from development through deployment into production.
  • Review existing monitoring, logging, alerting, and observability coverage for critical services.
  • Understand current production reliability expectations, incident response practices, escalation paths, and the developing on-call model.
  • Shadow production incidents and participate in troubleshooting alongside engineering team members before taking primary on-call responsibility.
  • Identify initial reliability, security, infrastructure, deployment, or operational risks and document recommended priorities.
  • Establish working relationships with software, platform, data, and product teams and understand the infrastructure requirements of their workloads.
  • Begin using AI-assisted tools for infrastructure analysis, troubleshooting, automation, documentation, and operational workflows.

Success at 30 days: You can explain how BigGeo’s production infrastructure operates, how software reaches production, how system health is measured, where the most important operational risks exist, and how you will begin improving reliability and developer operations.

First 60 Days — Take Ownership and Improve Reliability

By the end of your first 60 days, you will:

  • Independently own meaningful components of BigGeo’s cloud infrastructure, Kubernetes environments, deployment systems, or observability stack.
  • Deliver at least one material improvement to Infrastructure-as-Code, CI/CD, deployment automation, observability, or operational tooling.
  • Establish or materially improve SLOs and SLIs for critical production services, creating clearer visibility into reliability and service health.
  • Improve monitoring and alerting so critical production conditions are actionable while unnecessary or low-value alerts are reduced.
  • Participate actively in the on-call rotation with appropriate escalation support and demonstrate effective production troubleshooting and incident response.
  • Lead or materially contribute to root-cause analysis and blameless postmortems, ensuring recurring issues result in clear corrective actions.
  • Identify recurring operational toil and automate at least one high-value manual infrastructure or operational workflow.
  • Assess capacity requirements for critical stateful or resource-intensive workloads and identify scaling or failure risks.
  • Strengthen infrastructure security, access controls, secrets management, or vulnerability management where gaps are identified.
  • Use AI-assisted operational workflows to accelerate troubleshooting, infrastructure development, analysis, or automation without compromising reliability or security.

Success at 60 days: You are independently operating important parts of BigGeo’s infrastructure, contributing effectively to production response, and delivering measurable improvements to reliability, automation, observability, and engineering delivery.

First 90 Days — Strengthen the Operational Foundation of the Spatial Cloud

By the end of your first 90 days, you will:

  • Own the reliability and operational health of defined production infrastructure and services, with clear SLOs, monitoring, alerting, and response expectations.
  • Demonstrate measurable improvement in at least one key operational area such as deployment reliability, deployment velocity, incident frequency, recovery time, alert quality, infrastructure automation, or operational toil.
  • Help establish a structured on-call practice with clear escalation paths, response expectations, handoff practices, and supporting documentation.
  • Establish repeatable incident response and postmortem practices that convert production failures into system and process improvements.
  • Improve infrastructure resilience through capacity planning, fault-tolerance improvements, recovery planning, or elimination of identified single points of failure.
  • Advance disaster recovery and business continuity capabilities for critical infrastructure and workloads.
  • Deliver reusable infrastructure automation or internal developer tooling that allows engineering teams to deploy and operate services with less manual intervention.
  • Identify cloud cost or resource-efficiency opportunities and implement practical improvements without compromising reliability or performance.
  • Establish clear operational documentation and runbooks for critical infrastructure, deployment, incident, and recovery workflows.
  • Contribute informed recommendations to platform architecture and infrastructure strategy based on direct production experience.
  • Demonstrate that infrastructure improvements are creating leverage across the engineering organization by making production systems easier to deploy, observe, troubleshoot, recover, and scale.

Success at 90 days: You are operating as a trusted owner of BigGeo’s production reliability. Engineering teams can ship with greater confidence, critical systems are more observable and resilient, operational practices are becoming repeatable, and infrastructure is better prepared to support the continued scale of the Spatial Cloud.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps and Site Reliability Engineer
DevOps and Site Reliability Engineer

BigGeo • Calgary

On-site
CAD 110,000 - 160,000
Lead Spatial Cloud Solutions
Lead Spatial Cloud Solutions

BigGeo Inc. • Calgary

Hybrid
CAD 150,000 - 190,000
Data Partner Manager
Data Partner Manager

BigGeo Inc. • Calgary

Hybrid
CAD 90,000 - 140,000
Client Success Manager
Client Success Manager

BigGeo • Calgary

On-site
CAD 70,000 - 110,000
Head of Engineering
Head of Engineering

BigGeo • Calgary

On-site
CAD 120,000 - 160,000
Director, Global DevOps, Site Reliability, & Infrastructure (Vancouver, B.C. or Austin, TX)
Director, Global DevOps, Site Reliability, & Infrastructure (Vancouver, B.C. or Austin, TX)

Bitkernel Technology Inc • Vancouver

On-site
CAD 160,000 - 240,000
Flexible work options
Office and Operations Coordinator
Office and Operations Coordinator

RGIT Australia • Calgary

On-site
CAD 45,000 - 75,000
Technical Product Owner
Technical Product Owner

BigGeo Inc. • Calgary

Hybrid
CAD 110,000 - 150,000
Senior DevOps
Senior DevOps

Quartermaster Inc. • Toronto

Hybrid
CAD 160,000 - 215,000
30 days PTO annually
Health, dental and wellness benefits
Tech allowance
+1
Technical Product Owner
Technical Product Owner

Kibbi Technologies Inc. • Calgary

On-site
CAD 110,000 - 150,000