A complete application in a minute — tailored resume and cover letter, ready to send.
MSRcosmos LLC in Mahwah, NJ invites an experienced Site Reliability Engineer to strengthen the reliability and scalability of our on-premise and cloud-based systems. You will focus on Google Cloud Platform, Kubernetes, automation, cost optimization, and incident response to keep services performant and resilient.
You will collaborate with development and operations teams, implement monitoring with Prometheus/Grafana, design cloud infrastructure with GCP services, and contribute to capacity
Position: Site Reliability Engineer with GCP (Only W2)
Job Description:
We are looking for a talented Site Reliability Engineer (SRE) with a strong background in Google Cloud Platform (GCP) and kubernetes. The ideal candidate will be responsible for ensuring the reliability, performance, and scalability of our on-premise and cloud-based systems along with focus on reducing costs for Google Cloud.
System Reliability: Ensure the reliability and uptime of critical services and infrastructure.
Google Cloud Expertise: Design, implement, and manage cloud infrastructure using Google Cloud services.
Automation: Develop and maintain automation scripts and tools to improve system efficiency and reduce manual intervention.
Monitoring and Incident Response: Implement monitoring solutions and respond to incidents to minimize downtime and ensure quick recovery.
Collaboration: Work closely with development and operations teams to improve system reliability and performance.
Capacity Planning: Conduct capacity planning and performance tuning to ensure systems can handle future growth.
Documentation: Create and maintain comprehensive documentation for system configurations, processes, and procedures.
Skills:
Experience with database technologies (SQL & no-SQL - AlloyDB (PostgreSQL), DataBricks, Firestore, BigQuery, etc).
Familiarity with Google BI and AI/ML tools (Looker, BigQuery ML, Vertex AI, etc).
Experience with automation tools (Terraform, Ansible, Puppet).
Familiarity with CI/CD pipelines and tools (Azure pipelines Jenkins, GitLab CI, etc).
Knowledge of networking concepts and protocols. (Service mesh experience a plus).
Experience with monitoring tools (Prometheus, Grafana, etc).
Preferred Certifications:
Google Cloud Professional DevOps Engineer
Google Cloud Professional Cloud Architect