Senior DevOps Engineer, Agentic Automation & Live Games

Socket.dev

Toronto

Hybrid

CAD 140,000 - 180,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Big Viking Games is seeking a Senior DevOps Engineer for Agentic Automation & Live Games in Toronto. This hybrid role requires in-office presence three days per week and focuses on modernizing live production infrastructure, building agentic automation, and improving reliability and security across AWS, serverless components, and CI/CD pipelines.

The ideal candidate has extensive experience with cloud platforms, IaC, container orchestration, and GitOps, with a passion for automation that

Qualifications

  • 7+ years of experience in DevOps, infrastructure engineering, or similar roles.
  • Experience delivering live production systems with uptime and reliability requirements.
  • 1+ year building with agentic coding tools, tool-calling systems, or automation that takes real action.

Responsibilities

  • Design, build, and operate agentic tooling that performs real DevOps work.
  • Maintain integration layer for safe tooling actions (APIs, MCP-style servers, webhooks).
  • Automate recurring maintenance tasks and SOPs into unattended workflows.
  • Establish guardrails: least-privilege, dry-run, approvals, logging, audit trails, rollback plans.
  • Use agentic tools to accelerate IaC authoring, migrations, incident analysis, and documentation.
  • Improve how the engineering team uses automation and AI-assisted tooling.

Skills

AWS
Docker
Kubernetes
Terraform
CI/CD
GitOps
Automation tooling
Security best practices
Observability

Tools

ArgoCD
Terraform

Job description

Senior DevOps Engineer, Agentic Automation & Live Games

Toronto, ON

Hybrid, 3 days per week in office

Full-time

Department: Engineering

Reports to: Engineering Leadership

Compensation range: CAD $140,000 to $180,000

About Big Viking Games

Big Viking Games is a Canadian gaming company focused on building, operating, and growing long-standing online game communities. Our games have entertained players for years, supported by loyal audiences, live operations, evolving content systems, product innovation, and deep player-driven economies.

Our flagship titles, YoWorld and FishWorld, have served millions of players over their lifetime. These are enduring live-service virtual worlds with rich in-game economies, virtual goods, social interaction, and long-term player engagement at their core.

We are entering a new phase of modernization and growth, with a focus on stronger infrastructure, better automation, practical AI adoption, improved reliability, stronger security practices, and scalable systems that help our games and teams perform at a higher level.

About the Role

Big Viking Games is hiring a Senior DevOps Engineer, Agentic Automation & Live Games to modernize the infrastructure behind our live-service games and help reinvent how those systems are operated.

This is a hands‑on senior infrastructure role for someone fluent in modern cloud operations: AWS, containers, Infrastructure as Code, CI/CD, observability, production reliability, incident response, security, and automation.

But this is not a traditional DevOps role.

We are looking for someone who can build automation that does real work, not just scripts that run or dashboards that summarize. The right person has started using agentic coding tools, tool-calling systems, API integrations, workflow automation, or MCP‑style tooling to safely diagnose, provision, remediate, monitor, or maintain infrastructure.

Our games run on mature systems with real players, real revenue, real constraints, and real consequences. There is meaningful room to automate how they are operated, but the work must be done with discipline. Uptime, data integrity, least‑privilege access, rollback paths, auditability, and production safety matter.

The defining trait for this role is self-direction. Given a backlog, you improve how the work gets done. Left to your own judgment, you find repetitive operational work nobody has flagged, decide what is worth automating, and build it safely.

This is a hybrid role based in Toronto, with an expectation of working in office three days per week. Live‑service games require operational awareness outside regular business hours, including periodic on‑call and incident response availability.

What You’ll Do

Build Agentic Automation for Infrastructure Work

  • Design, build, and operate agentic tooling that performs real DevOps work, including diagnosis, remediation, provisioning, routine maintenance, and operational follow‑up.
  • Build and maintain the integration layer that lets tooling act safely on our systems, including API integrations, MCP‑style servers, webhook‑driven orchestration, serverless functions, and permission‑controlled automation.
  • Convert manual runbooks, SOPs, recurring maintenance tasks, and repetitive operational chores into automation that can run unattended where appropriate.
  • Establish guardrails that make automated action against production defensible, including least‑privilege scopes, dry‑run modes, approval paths, logging, audit trails, rollback plans, and clear escalation rules.
  • Use agentic coding tools to accelerate infrastructure work, including IaC authoring, migration scripts, incident analysis, log analysis, documentation, and operational troubleshooting.
  • Help improve how the wider engineering team uses automation and AI‑assisted tooling safely and effectively.
Own and Modernize Live Production Infrastructure
  • Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms supporting our live games, data systems, internal tools, and operational workflows.
  • Drive infrastructure modernization while maintaining uptime for live games with active player communities.
  • Implement and maintain Infrastructure as Code using Terraform, CloudFormation, CDK, Pulumi, or similar tools.
  • Improve CI/CD pipelines, release workflows, deployment reliability, and environment management so teams can ship safely and frequently.
  • Operate containerized workloads and GitOps‑based deployment patterns.
  • Improve local development, build, test, deploy, and production support workflows.
  • Help manage cloud usage, resource tagging, environment efficiency, and infrastructure cost discipline.
Reliability, Observability, and Security
  • Rebuild and improve observability across the stack, including logging, metrics, alerting, dashboards, and operational visibility.
  • Pay particular attention to early detection of silent failures, pipeline failures, data freshness issues, and degraded production behavior.
  • Maintain and monitor data pipelines between game source databases, Snowflake, and downstream analytics and reporting systems.
  • Own secrets and credential lifecycle management across platforms, including API key rotation, access controls, environment variable governance, and least‑privilege practices.
  • Support incident response, root cause analysis, remediation planning, and post‑incident improvements.
  • Automate the parts of incident response and remediation that repeat.
  • Create documentation, runbooks, and SOPs that are executable wherever possible, not just descriptive.
  • Improve backup, restore, disaster recovery, access review, and production‑readiness practices.
What You Bring
  • 7+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, platform engineering, or a similar role.
  • Strong hands‑on experience designing, maintaining, and improving production infrastructure.
  • Experience supporting live production systems where uptime, reliability, data integrity, and performance matter.
  • Experience working with mature or legacy systems that predate modern cloud‑native patterns.
  • 1+ year building with agentic coding tools, tool‑calling systems, workflow automation, or infrastructure automation that takes real action.
  • Experience with systems such as agents wired into pipelines, MCP‑style integrations, automated diagnosis, automated remediation, provisioning workflows, or production‑safe infrastructure tooling.
  • Experience identifying repetitive operational work and turning it into reliable automation.
Core Technical Skills
  • Strong hands‑on experience with AWS or similar cloud platforms.
  • Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar, including shared state and team‑based workflows.
  • Experience with containerized applications, especially Docker.
  • Experience with container orchestration such as Kubernetes, ECS, EKS, or similar, including debugging real production issues.
  • Experience with GitOps and declarative deployment tools such as ArgoCD, Flux, or equivalent.
  • Experience with CI/CD tooling, version control, deployment automation, and modern release workflows.
  • Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
  • Experience establishing observability, including choosing tooling, defining alerts, tuning alert noise, and creating useful dashboards.
  • Experience with relational databases such as MariaDB, MySQL, or Postgres, including replication, backup, verified restore, and schema changes against systems that stay online.
  • Security‑aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least‑privilege practices.
How You Work
  • Self‑starting. You identify the work rather than waiting for it to be assigned, and you can explain why one problem matters more than another.
  • Proactive about toil. You notice repetitive work and treat it as a defect to be engineered away, not a cost of doing business.
  • Inventive but pragmatic. You reach for novel approaches where they genuinely help and recognize when a simple script is the better answer.
  • Strong problem‑solving skills and the ability to investigate complex infrastructure or production issues, including silent failures and data pipeline outages.
  • Comfortable working directly with engineers to improve build, deploy, and operational workflows.
  • Practical ownership mindset with the ability to prioritize, execute, and close loops.
  • Strong communication with technical and non‑technical stakeholders.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer: Agentic Automation for Live Games
Senior DevOps Engineer: Agentic Automation for Live Games

Socket.dev • Toronto

Hybrid
CAD 140,000 - 180,000
Technical Director, Live Games
Technical Director, Live Games

Socket.dev • Toronto

Hybrid
CAD 180,000 - 220,000
Retirement plan
Health, dental, vision coverage
Wellness spending account
+4
Principal Engineer, Agentic Products & Workflows
Principal Engineer, Agentic Products & Workflows

Worky • Toronto

On-site
CAD 175,000 - 200,000
Group Retirement Savings Plan matching
Comprehensive benefits package
Health and Wellness spending account
+3
Principal Engineer, Agentic Products & Workflows
Principal Engineer, Agentic Products & Workflows

Big Viking Games Inc. • Toronto

On-site
CAD 175,000 - 200,000
Employee Stock Option Plan
Health and dental benefits
Vacation and wellness days
+3
Principal Engineer, Agentic Products & Workflows
Principal Engineer, Agentic Products & Workflows

Big Viking Games • Toronto

On-site
CAD 175,000 - 200,000
Stock options (Employee Stock Option)
Health, dental, and vision coverage
Wellness spending account
+3
Principal Game Engineer
Principal Game Engineer

WorkinGames • Toronto

On-site
CAD 175,000 - 200,000
Group Retirement Savings Plan
Health, dental, and vision coverage
Vacation days
DevOps Engineer
DevOps Engineer

Code Wizards • Vancouver

On-site
USD 56,000 - 86,000
Senior DevOps Engineer
Senior DevOps Engineer

Quest Global • Vancouver

On-site
CAD 100,000 - 120,000
401(k) matching
Health insurance
Dental insurance
+5
Principal Engineer, Agentic Products & Workflows
Principal Engineer, Agentic Products & Workflows

United States Digital Space LLC • Toronto

On-site
CAD 175,000 - 200,000
Group retirement savings plan matching
Health, dental, vision benefits
Wellness spending account
+3
Senior DevOps
Senior DevOps

Quartermaster inc. • Toronto

Hybrid
CAD 160,000 - 215,000
30 days of PTO annually
Health, dental, and wellness benefits
Tech allowance benefit
+1