Site Reliability Engineer with GCP

MSRcosmos LLC

Dallas (TX)

On-site

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

MSRcosmos LLC is seeking a Site Reliability Engineer with strong GCP and Kubernetes expertise to ensure reliability, performance, and scalability of both on-premise and cloud-based systems. The role emphasizes reducing Google Cloud costs while maintaining uptime and efficiency.

You will design and implement cloud infrastructure, automate tasks with Terraform and Ansible, monitor services with Prometheus/Grafana, and collaborate with development and operations teams to plan capacity and respond

Qualifications

  • Experience with Google Cloud Platform (GCP) and Kubernetes.
  • Experience with automation and configuration management tools.
  • Proficiency in monitoring and incident response.
  • Ability to collaborate across teams and document procedures.

Responsibilities

  • Ensure reliability and uptime of critical services.
  • Design, implement, and manage cloud infrastructure on GCP.
  • Develop automation scripts to improve efficiency.
  • Monitor and respond to incidents to minimize downtime.
  • Collaborate with development and operations teams to improve reliability and performance.
  • Perform capacity planning and performance tuning to handle growth.
  • Create and maintain documentation for configurations, processes, and procedures.

Skills

GCP fundamentals
System reliability
Monitoring concepts
Cost optimization

Tools

Kubernetes
Terraform
Ansible
Puppet
Jenkins
GitLab CI
Networking basics

Job description

Job Description
  • We are looking for a talented Site Reliability Engineer (SRE) with a strong background in Google Cloud Platform (GCP) and kubernetes. The ideal candidate will be responsible for ensuring the reliability, performance, and scalability of our on-premise and cloud-based systems along with focus on reducing costs for Google Cloud.
  • System Reliability: Ensure the reliability and uptime of critical services and infrastructure.
  • Google Cloud Expertise: Design, implement, and manage cloud infrastructure using Google Cloud services.
  • Automation: Develop and maintain automation scripts and tools to improve system efficiency and reduce manual intervention.
  • Monitoring and Incident Response: Implement monitoring solutions and respond to incidents to minimize downtime and ensure quick recovery.
  • Collaboration: Work closely with development and operations teams to improve system reliability and performance.
  • Capacity Planning: Conduct capacity planning and performance tuning to ensure systems can handle future growth.
  • Documentation: Create and maintain comprehensive documentation for system configurations, processes, and procedures.
Skills
  • Experience with database technologies (SQL & no-SQL - AlloyDB (PostgreSQL), DataBricks, Firestore, BigQuery, etc.).
  • Familiarity with Google BI and AI/ML tools (Looker, BigQuery ML, Vertex AI, etc.).
  • Experience with automation tools (Terraform, Ansible, Puppet).
  • Familiarity with CI/CD pipelines and tools (Azure pipelines Jenkins, GitLab CI, etc.).
  • Knowledge of networking concepts and protocols. (Service mesh experience a plus).
  • Experience with monitoring tools (Prometheus, Grafana, etc.).
Preferred Certifications
  • Google Cloud Professional DevOps Engineer
  • Google Cloud Professional Cloud Architect
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • New Jersey

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

JPS Tech Solutions • San Jose (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer - GCP & Automation Focus
Site Reliability Engineer - GCP & Automation Focus

Insight Global • United States

On-site
USD 100,000 - 125,000
Site Reliability Engineer - GCP
Site Reliability Engineer - GCP

TechDigital Group • San Jose (CA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

NexTier Completion Solutions Inc. • Houston (TX)

On-site
USD 110,000 - 150,000
GCP SRE | Kubernetes & Cloud Reliability Engineer
GCP SRE | Kubernetes & Cloud Reliability Engineer

MSRcosmos LLC • Dallas (TX)

On-site
USD 120,000 - 180,000
GCP & Kubernetes SRE: Reliable, Scalable Cloud Ops
GCP & Kubernetes SRE: Reliable, Scalable Cloud Ops

Compunnel, Inc. • New Jersey

On-site
USD 120,000 - 150,000
Site Rel Eng III, GCP
Site Rel Eng III, GCP

Optimum • Plano (TX)

On-site
USD 140,000 - 180,000
GCP Site Reliability Engineer
GCP Site Reliability Engineer

TechDigital Group • Frisco (TX)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer – Google Distributed Cloud Edge (Edge SRE)
Senior Site Reliability Engineer – Google Distributed Cloud Edge (Edge SRE)

CoSourcing Partners Inc. • Chicago (IL)

Hybrid
USD 150,000 - 190,000