Site Reliability Engineer (Google Cloud Platform)

Grid Dynamics

Wrocław

On-site

PLN 180,000 - 280,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Flexible schedule
Sports
Well-equipped office
Professional development

Job summary

Grid Dynamics is seeking a Site Reliability Engineer to lead a cloud platform migration with emphasis on observability and FinOps for a large financial transformation.

You will own the cloud platform layer end‑to‑end, delivering visibility into performance and infrastructure spend across pods and services. Collaborating with a nearshore team of senior engineers, you will instrument distributed systems using OpenTelemetry, GKE, Cloud Run, Terraform, and multi‑language code in Kotlin/Java.

Qualifications

  • 4–6 years of software/DevOps experience.
  • 3+ years of hands-on cloud infrastructure management in production.
  • Deep expertise in GCP with GKE and Cloud Run.
  • Proficient in Python for tooling and automation.
  • Terraform expertise across multi-environment setups and modules.
  • Read/modify Kotlin and Java backend codebases.
  • Fluent written English for async collaboration and reviews.
  • Strong autonomy and ability to navigate large codebases in enterprise settings.

Responsibilities

  • Own cloud platform layer end-to-end, with visibility into performance and spend.
  • Architect and implement a cloud usage and cost attribution dashboard.
  • Instrument distributed tracing and OpenTelemetry across Kubernetes environments.
  • Develop tooling for database performance monitoring and load testing.
  • Build scalable monitoring, alerting, and SLO tracking for services.
  • Manage infrastructure changes using Terraform and enterprise standards.
  • Read and extend platform code across Kotlin/Java without impacting business logic.

Skills

OpenTelemetry
Python
Terraform
GCP
Kubernetes
Java
Kotlin
English
Observability
Autonomy

Tools

Cloud Pub/Sub
Cloud Run

Job description

We are looking for a Site Reliability Engineer to join a high-stakes global tech ecosystem and drive the delivery of a critical enterprise platform migration to the cloud. Your core mission will be to architect, build, and productionalize the observability and cost intelligence (FinOps) layer for a massive, multi-year financial platform transformation. You will take end-to-end ownership of the cloud platform layer, giving internal stakeholders full visibility into platform behavior, performance, and infrastructure spend. Working alongside a nearshore team of senior engineers, you will solve highly complex architectural challenges in a production-grade, distributed system.

Responsibilities
  • End-to-End Infrastructure & FinOps Ownership: Architect and implement a cloud usage and cost attribution dashboard, providing detailed per-pod and per-service cost breakdown using cloud billing APIs and internal FinOps hubs.
  • Advanced Observability & Tracing: Instrument end-to-end distributed tracing using OpenTelemetry, configuring collectors within Kubernetes environments and exporting traces to cloud monitoring systems utilizing RED metrics.
  • Performance Engineering & Stress Testing: Write custom tooling from scratch to deliver database performance monitoring, load testing, and trend analysis for critical underlying storage layers.
  • Monitoring & Alerting Automation: Build and deploy scalable production monitoring, custom alerting policies, and SLO tracking for containerized and serverless services.
  • Infrastructure as Code: Independently manage, write, and apply infrastructure modifications using Terraform, working within established enterprise repository standards, modules, and environment state management.
  • Cross-Language Codebase Extension: Read, debug, and extend existing platform code across a diverse stack including Kotlin, Java, and Python to seamlessly integrate technical metrics without disrupting business logic.
  • Quality & Release Assurance: Implement rigorous unit testing with high code coverage for all newly developed monitoring tools to comply with strict enterprise quality gates and sign-offs.
Requirements
  • Experience: 4 to 6 years of professional software or DevOps engineering experience, with at least 2 to 3 years of hands-on cloud infrastructure management in production.
  • Advanced Cloud Infrastructure: Deep operational proficiency with Google Cloud Platform (GCP), specifically with managing and configuring workload-level alerting on Google Kubernetes Engine (GKE) and Cloud Run.
  • Observability & OpenTelemetry: Proven track record of building observability solutions in distributed systems, using OpenTelemetry (both auto and manual instrumentation) alongside distributed tracing and profiling tools.
  • Strong Automation Scripting: Intermediate-to-advanced fluency in Python for writing custom test tooling, metrics integration scripts, and backend automation from scratch.
  • Solid Infrastructure as Code: Strong proficiency in Terraform, including experience with multi-environment setups, workspaces, and corporate module standards.
  • Polyglot & JVM Familiarity: Practical ability to read, understand, and modify existing backend codebases written in Kotlin and Java.
  • Crucial Non-Technical Skills: Extreme technical autonomy to resolve blockers independently, rapid onboarding skills into large unfamiliar codebases, and fluent written English for async alignment and pull requests.
  • Process Alignment: Ability to thrive in a highly regulated enterprise environment with strict peer reviews, robust documentation requirements, and formal deployment procedures.
Nice to have
  • Domain Knowledge: Previous experience working within financial services, fintech, investment banking, or other highly regulated industries.
  • Enterprise Streaming Tools: Working knowledge of cloud messaging systems (such as Cloud Pub/Sub) utilized for inter-service communication.
  • Advanced Storage Engines: Familiarity with high-throughput distributed database architectures, such as Google Cloud Bigtable.
  • Systems Languages Awareness: Ability to read or debug foundational code written in low-level systems languages like Rust or C++ during multi-stack production deployments.
We offer
  • Opportunity to work on bleeding-edge projects
  • Work with a highly motivated and dedicated team
  • Competitive salary
  • Flexible schedule
  • Benefits package - medical insurance, sports
  • Corporate social events
  • Professional development opportunities
  • Well-equipped office
About Us

Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation. A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI, supported by profound expertise and ongoing investment in data, analytics, cloud & DevOps, application modernization and customer experience. Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (Google Cloud Platform)
Site Reliability Engineer (Google Cloud Platform)

Grid Dynamics • Warszawa

On-site
PLN 180,000 - 240,000
Medical insurance
Sports
Professional development
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Grid Dynamics • Wrocław

On-site
Competitive salary
Flexible schedule
Medical insurance
+4
Senior DevOps Engineer
Senior DevOps Engineer

Grid Dynamics • Wrocław

On-site
PLN 200,000 - 280,000
Competitive salary
Flexible schedule
Medical insurance
+4
Senior DevOps Engineer
Senior DevOps Engineer

Grid Dynamics • Województwo pomorskie

On-site
PLN 180,000 - 280,000
Medical insurance
Sports benefits
Flexible schedule
+2
Senior DevOps Engineer
Senior DevOps Engineer

Grid Dynamics • Warszawa

On-site
PLN 180,000 - 300,000
Competitive salary
Flexible schedule
Medical insurance
+2
Senior DevOps Engineer
Senior DevOps Engineer

Grid Dynamics • Kraków

On-site
PLN 180,000 - 260,000
Flexible schedule
Competitive salary
Medical insurance
+4
Senior ML engineer
Senior ML engineer

Grid Dynamics • Województwo pomorskie

On-site
PLN 180,000 - 280,000
Flexible schedule
Competitive salary
Medical insurance
+4
Senior ML engineer
Senior ML engineer

Grid Dynamics • Kraków

On-site
PLN 180,000 - 240,000
Senior ML engineer
Senior ML engineer

Grid Dynamics • Wrocław

Hybrid
PLN 300,000 - 420,000
Flexible schedule
Medical insurance
Professional development
+2
Senior ML engineer
Senior ML engineer

Grid Dynamics • Warszawa

On-site
PLN 200,000 - 320,000
Medical insurance
Sports benefits
Flexible schedule
+1