Site Reliability Engineer (Only W2)

MSRcosmos LLC

Mahwah (NJ)

On-site

USD 120,000 - 150,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

MSRcosmos LLC in Mahwah, NJ invites an experienced Site Reliability Engineer to strengthen the reliability and scalability of our on-premise and cloud-based systems. You will focus on Google Cloud Platform, Kubernetes, automation, cost optimization, and incident response to keep services performant and resilient.

You will collaborate with development and operations teams, implement monitoring with Prometheus/Grafana, design cloud infrastructure with GCP services, and contribute to capacity

Qualifications

  • Strong background in Google Cloud Platform (GCP) and Kubernetes.
  • Experience designing reliable, scalable systems and incident response.
  • Familiarity with cost optimization in cloud environments.

Responsibilities

  • Ensure reliability and uptime of critical services and infrastructure.
  • Design, implement, and manage cloud infrastructure using Google Cloud services.
  • Develop automation scripts and tools to improve system efficiency.
  • Implement monitoring solutions and respond to incidents to minimize downtime.

Skills

GCP
Kubernetes
Automation
Monitoring
CI/CD

Tools

Terraform
Ansible
Puppet
Azure Pipelines
Jenkins
GitLab CI

Job description

Position: Site Reliability Engineer with GCP (Only W2)

Job Description:

We are looking for a talented Site Reliability Engineer (SRE) with a strong background in Google Cloud Platform (GCP) and kubernetes. The ideal candidate will be responsible for ensuring the reliability, performance, and scalability of our on-premise and cloud-based systems along with focus on reducing costs for Google Cloud.

System Reliability: Ensure the reliability and uptime of critical services and infrastructure.

Google Cloud Expertise: Design, implement, and manage cloud infrastructure using Google Cloud services.

Automation: Develop and maintain automation scripts and tools to improve system efficiency and reduce manual intervention.

Monitoring and Incident Response: Implement monitoring solutions and respond to incidents to minimize downtime and ensure quick recovery.

Collaboration: Work closely with development and operations teams to improve system reliability and performance.

Capacity Planning: Conduct capacity planning and performance tuning to ensure systems can handle future growth.

Documentation: Create and maintain comprehensive documentation for system configurations, processes, and procedures.

Skills:

Experience with database technologies (SQL & no-SQL - AlloyDB (PostgreSQL), DataBricks, Firestore, BigQuery, etc).

Familiarity with Google BI and AI/ML tools (Looker, BigQuery ML, Vertex AI, etc).

Experience with automation tools (Terraform, Ansible, Puppet).

Familiarity with CI/CD pipelines and tools (Azure pipelines Jenkins, GitLab CI, etc).

Knowledge of networking concepts and protocols. (Service mesh experience a plus).

Experience with monitoring tools (Prometheus, Grafana, etc).

Preferred Certifications:

Google Cloud Professional DevOps Engineer

Google Cloud Professional Cloud Architect

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer with GCP
Site Reliability Engineer with GCP

MSRcosmos LLC • Dallas (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • New Jersey

On-site
USD 120,000 - 150,000
Senior SRE – GCP
Senior SRE – GCP

DMS Vision Inc • Dallas (TX)

On-site
USD 120,000 - 190,000
Senior Site Reliability Engineer – Google Distributed Cloud Edge (Edge SRE)
Senior Site Reliability Engineer – Google Distributed Cloud Edge (Edge SRE)

CoSourcing Partners - Enterprise-AI and IT Services Company • Chicago (IL)

Hybrid
USD 120,000 - 150,000
Senior Site Reliability Engineer – Google Distributed Cloud Edge (Edge SRE)
Senior Site Reliability Engineer – Google Distributed Cloud Edge (Edge SRE)

CoSourcing Partners Inc. • Chicago (IL)

Hybrid
USD 150,000 - 190,000
GCP SRE | Kubernetes & Cloud Reliability Engineer
GCP SRE | Kubernetes & Cloud Reliability Engineer

MSRcosmos LLC • Dallas (TX)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

NexTier Completion Solutions Inc. • Houston (TX)

On-site
USD 110,000 - 150,000
Staff SRE
Staff SRE

Selby Jennings • Chicago (IL)

Hybrid
USD 140,000 - 180,000
Site Reliability Engineer: Build Reliable, Scalable Systems
Site Reliability Engineer: Build Reliable, Scalable Systems

Google • Sunnyvale (CA)

On-site
USD 256,000 - 300,000
Site Rel Eng III, GCP
Site Rel Eng III, GCP

Optimum • Plano (TX)

On-site
USD 140,000 - 180,000