Lead Application Reliability Engineer

Innodata

Toronto

On-site

CAD 150,000 - 180,000

Full time

24 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Innodata is seeking a hands-on Lead Application Support Engineer to sustain and improve enterprise applications running on Google App Engine and microservices. The role focuses on availability, incident response, troubleshooting, and incremental feature enhancements rather than building new apps from scratch.

The ideal candidate will rapidly understand a complex microservices landscape, restore service during incidents, and deliver safe, scalable updates across test, pre-production, and

Qualifications

  • 10+ years in Application Support, Engineering, Cloud or related roles.
  • Strong hands-on experience with GCP and Google App Engine.
  • Experience with microservices, REST/gRPC, and distributed tracing.
  • Proficient in CI/CD tooling and deployments across environments.

Responsibilities

  • Provide production support and maintenance for enterprise apps on GAE.
  • Lead triage, diagnosis, and restoration of user-impacting incidents.
  • Troubleshoot latency, errors, and inter-service failures with root-cause analysis.
  • Design and deliver feature enhancements based on user/business needs.
  • Manage deployments, canary releases, and rollback plans across environments.
  • Collaborate with cross-functional teams (Engineering, QA, Data, Infra).

Skills

GCP experience
Google App Engine
Microservices
CI/CD pipelines
SQL/cloud databases
Incident management
Communication

Tools

Cloud Build
Jenkins
GitHub Actions
Terraform
Cloud Run/Kubernetes
Apigee/API Gateway

Job description

Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

We are looking for a hands-on Lead Application Support Engineer to support, maintain, and enhance business-critical enterprise applications built and running on Google App Engine (GAE) and microservices.

The role is focused on application availability, production support, incident response, troubleshooting, and continuous feature enhancement rather than building a new application from the ground up. The ideal candidate can quickly understand an existing microservices-based application landscape, restore service when users are impacted, and deliver incremental enhancements safely across test, pre-production, and production environments.

Experience supporting large-scale, business-critical applications in complex enterprise technology environments is preferred.

What You’ll Own:
  • Provide production support and maintenance for enterprise applications hosted on Google App Engine, including standard and flexible environments.
  • Act as a first point of contact for user-impacting incidents: triage, diagnose, restore service, and drive issues to closure within agreed SLAs/SLOs.
  • Troubleshoot application errors, failed requests, latency and performance degradation, service-to-service failures, configuration issues, quota/scaling limits, and dependency or integration failures.
  • Design, develop, and deliver feature enhancements and functional improvements to existing applications based on user and business needs.
  • Support applications built on microservices architecture, including service boundaries, APIs/contracts, inter-service communication, authentication, and failure/retry behavior.
  • Own build, release, and deployment activities across development, test, pre-production, and production environments with appropriate validation, approvals, and rollback plans.
  • Manage App Engine deployments including versions, traffic splitting/migration, canary and staged rollouts, rollbacks, service configuration, and scaling settings.
  • Perform root-cause analysis for recurring production issues and implement sustainable fixes.
  • Build and maintain monitoring, alerting, logging, dashboards, and error reporting using Cloud Monitoring, Cloud Logging, Error Reporting, and Cloud Trace.
  • Support platform, framework, library, dependency, and runtime upgrades while maintaining stability, supportability, and compliance.
  • Support IAM, service accounts, access controls, secrets management, and operational governance.
  • Participate in change management, release-readiness reviews, and on-call/rotational support as required.
  • Collaborate with client and cross-functional Application Engineering, Product, QA, Data Engineering, Infrastructure, and Platform teams.
  • Adapt to established client-specific engineering, security, review, change-management, and operational processes.
  • Create and maintain technical documentation, operational runbooks, troubleshooting guides, deployment procedures, and support playbooks.
  • Service decomposition and ownership: understand service boundaries, upstream/downstream dependencies, and ownership of behavior or data.
  • APIs and contracts: REST and gRPC/protobuf interfaces, versioning, backward compatibility, and contract-change impacts.
  • Inter-service communication: synchronous calls, asynchronous/event-driven messaging (Pub/Sub, Cloud Tasks), idempotency, retries, timeouts, backoff, and circuit breaking.
  • Service-to-service authentication and authorization using service accounts, identity/tokens, and least-privilege access.
  • Distributed troubleshooting using logs, distributed tracing, and correlation IDs to isolate failures across services.
  • Failure modes at scale, including cascading failures, partial outages, hot spots, quota exhaustion, and graceful degradation.
  • Independent deployability and coordination of multi-service releases when required.
  • Service-level observability, including dashboards, SLIs/SLOs, alerting, and error budgets.
You’ll Thrive in This Role If You Have:
  • 10+ years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related role.
  • Strong hands-on experience supporting, maintaining, and enhancing production applications on Google Cloud Platform (GCP).
  • Hands-on experience with Google App Engine, including deployment, configuration, scaling, versioning, and troubleshooting.
  • Solid understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging.
  • Practical experience with deployments and promotions across multiple environments, including release validation, rollback, and change control.
  • Working knowledge of SQL and application data stores such as Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery.
  • Good understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations.
  • Experience with CI/CD pipelines and automated build/deployment tooling such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI.
  • Experience troubleshooting complex production environments and performing root-cause analysis under time pressure.
  • Ability to quickly understand existing systems, codebases, services, configurations, and client-specific tools and workflows.
  • Strong communication and collaboration skills, including communication during user-impacting incidents.
Preferred Qualifications:
  • Experience supporting large-scale internal or enterprise applications with demanding availability, reliability, and performance requirements.
  • Ability to quickly learn and operate within enterprise-specific application frameworks, deployment tooling, and support processes.
  • Cloud Run, GKE, Cloud Functions, or other GCP application services.
  • Apigee / API Gateway, load balancing, and API management.
  • Pub/Sub, Cloud Tasks, Cloud Scheduler, and asynchronous/event-driven patterns.
  • Infrastructure as Code, particularly Terraform.
  • SRE practices including SLIs/SLOs, error budgets, incident management, and blameless postmortems.
  • Containerization with Docker and Kubernetes fundamentals.
  • Frontend or full-stack experience supporting user-facing web applications.
  • Looker / Tableau / BI platforms or reporting integrations.
  • Supporting AI/ML, GenAI, or LLM-based applications running on GCP, including Vertex AI.

The expected salary range for this position is $150,000 - $180,000 CAD per year, based on experience, skills, and qualifications.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Application Reliability Engineer
Application Reliability Engineer

Innodata Inc. • Canada

Hybrid
CAD 80,000 - 150,000
Application Reliability Engineer
Application Reliability Engineer

Innodata • Toronto

On-site
CAD 80,000 - 150,000
Data Platform Engineer (GCP)
Data Platform Engineer (GCP)

Innodata Inc. • Canada

On-site
CAD 80,000 - 150,000
Senior GCP App Engine Reliability Engineer
Senior GCP App Engine Reliability Engineer

Innodata Inc. • Canada

On-site
CAD 150,000 - 180,000
Practice Customer Developer, Data Management, Google Cloud
Practice Customer Developer, Data Management, Google Cloud

Google Inc. • Toronto

On-site
CAD 194,000 - 198,000
Top Customer Solutions Developer, High Touch Support, Compute, Google Cloud
Top Customer Solutions Developer, High Touch Support, Compute, Google Cloud

Google Inc. • Southwestern Ontario

On-site
CAD 138,000 - 141,000
Equity
Benefits
Outcome Customer Developer, Google Cloud Platform, Google Cloud
Outcome Customer Developer, Google Cloud Platform, Google Cloud

Google • Toronto, Montreal (administrative region)

On-site
CAD 194,000 - 198,000
Top Customer Solutions Developer II, Networking, Google Cloud (Mandarin)
Top Customer Solutions Developer II, Networking, Google Cloud (Mandarin)

Google • Southwestern Ontario

Hybrid
CAD 170,000 - 174,000
Manager, Technical Solutions, Cloud Applied AI
Manager, Technical Solutions, Cloud Applied AI

Google • Southwestern Ontario

On-site
CAD 248,000 - 255,000
Health insurance
Dental insurance
Vision insurance
+8
Outcome Customer Developer, Google Cloud Platform, Google Cloud
Outcome Customer Developer, Google Cloud Platform, Google Cloud

United States Digital Space LLC • Toronto

On-site
CAD 194,000 - 198,000
Equity
Bonus target
Benefits