Site Reliability Engineer

Macromill

Bengaluru

On-site

INR 1,500,000 - 2,300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Macromill is looking for a Site Reliability Engineer to own the end-to-end platform and cloud infrastructure. You will design, build, and operate scalable systems, automate CI/CD pipelines, and implement IaC with Terraform. The role emphasizes reliability, security, and developer productivity across teams.

You will collaborate with engineering on platform usability, monitor systems, and drive SRE best practices, including SLIs/SLOs, error budgets, and incident response.

Qualifications

  • 4+ years of experience building and operating production systems at scale.
  • Experience with cloud platforms (GCP, AWS, or Azure).
  • Proficiency in Infrastructure as Code (Terraform or similar).
  • Experience with containers and orchestration (Docker, Kubernetes) in production.
  • Strong understanding of Linux systems, networking, and troubleshooting.
  • Experience with CI/CD pipelines and automation.
  • Experience with monitoring and observability tools (metrics, logs, tracing).
  • Understanding of SLIs, SLOs, and service reliability practices.
  • Experience in incident handling, debugging, and root cause analysis.
  • Programming experience in Go or a similar language, plus scripting (bash).
  • Security best practices in cloud environments.

Responsibilities

  • Own and manage end-to-end platform, including infrastructure, networking, CI/CD, observability, and operations.
  • Design, build, and operate scalable, secure systems on cloud platforms.
  • Develop and maintain CI/CD pipelines and automation for faster deployments.
  • Implement and manage IaC using Terraform.
  • Improve reliability, availability, and performance through monitoring and optimization.
  • Build and enhance observability to reduce incidents and improve debugging.
  • Lead incident management, troubleshooting, and root cause analysis.
  • Automate infrastructure and operational processes to reduce toil.
  • Support service onboarding and migration to modern platform architecture.
  • Ensure security best practices across infrastructure and applications.
  • Collaborate with developers to improve developer experience and platform usability.

Skills

Cloud platforms
CI/CD automation
Observability & monitoring
Incident management
Infrastructure as Code
Linux & networking
Distributed systems
Go programming
Security best practices
DevOps mindset
Automation & toil reduction
Platform engineering
Mentoring/leadership
Open-source contributions
Japanese stakeholder collaboration

Tools

Terraform
Docker
Kubernetes

Job description

As a Site Reliability Engineer, you will own and manage the end-to-end platform and infrastructure at Macromill. This includes cloud infrastructure, networking, CI/CD, observability, and platform tooling. You will work closely with engineering teams to build a reliable, scalable, secure, and cost-efficient platform while improving developer productivity and system performance.

Responsibilities:
  • Own and manage the end-to-end platform, including infrastructure, networking, CI/CD, observability, and operations.
  • Design, build, and operate scalable, reliable, and secure systems on cloud platforms.
  • Develop and maintain CI/CD pipelines and automation to enable faster and safer deployments.
  • Implement and manage Infrastructure as Code (IaC) using tools like Terraform.
  • Improve system reliability, availability, and performance through monitoring, alerting, and optimization.
  • Build and enhance observability (logging, metrics, tracing) to reduce incidents and improve debugging.
  • Lead incident management, troubleshooting, and root cause analysis for production systems.
  • Automate infrastructure and operational processes to reduce manual work and toil.
  • Support service onboarding and migration to modern platform architecture.
  • Ensure security best practices across infrastructure and applications.
  • Collaborate with developers to improve developer experience and platform usability.
  • Drive adoption of SRE, DevOps, and engineering best practices across teams.
  • Define and maintain SLIs, SLOs, and SLAs to ensure service reliability.
Requirements:
  • 4+ years of experience building and operating production systems at scale.
  • Strong hands‑on experience with cloud platforms (GCP, AWS, or Azure).
  • Experience owning and managing end-to-end infrastructure and platform systems.
  • Proficiency in Infrastructure as Code (Terraform or similar).
  • Experience with containers and orchestration (Docker, Kubernetes) in production.
  • Solid understanding of distributed systems and microservices architecture.
  • Strong knowledge of Linux systems, networking (TCP/IP, DNS), and troubleshooting.
  • Experience with CI/CD pipelines and automation.
  • Experience with monitoring and observability tools (metrics, logs, tracing).
  • Understanding of SLIs, SLOs, and service reliability practices.
  • Experience in incident handling, debugging, and root cause analysis.
  • Ability to design systems and write technical design documents.
  • Programming experience in Go or a similar language, along with scripting (e. g., bash).
  • Basic understanding of security best practices in cloud environments.
  • Strong ownership mindset and ability to work across teams.
  • Experience with large‑scale or high‑traffic systems.
  • Hands‑on experience with SRE practices such as SLO‑driven operations and error budgets.
  • Understanding of platform engineering and building internal developer platforms.
  • Experience with cost optimization and performance tuning at scale.
  • Experience with agile development practices
  • Contributions to open‑source projects or active participation in technical communities.
  • Experience in mentoring engineers or leading technical initiatives.
  • Experience in applying AI in the SRE field.
  • Proven experience working with Japanese stakeholders and products.
  • Strong ownership of systems with a focus on reliability and performance.
  • Proactive in identifying problems and driving end-to-end improvements.
  • Thinks in terms of automation, scalability, and reducing toil.
  • Comfortable working in ambiguous environments.
  • Communicates well and collaborates across teams.
  • Continuously learns and improves engineering practice.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Sr. Site Reliability Engineer I
Sr. Site Reliability Engineer I

MetLife • Hyderabad

Hybrid
INR 1,500,000 - 2,300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Technologies Pvt. Ltd. • Pune District

On-site
INR 900,000 - 1,400,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior DevOps Engineer
Senior DevOps Engineer

Rwindia • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Lead
Senior Site Reliability Lead

Generac • Pune District

On-site
INR 3,000,000 - 6,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Five9 • Bengaluru

On-site
INR 4,000,000 - 7,000,000