At Bajaj Broking, high availability isn't just an internal KPI- it's the backbone of every trade, order execution, and portfolio update across millions of retail investors. During high-volatility market opens, our systems process massive spikes in concurrent transactions. Zero downtime is our baseline.
We are looking for a Lead Site Reliability Engineer (SRE) with 8-10 years of engineering experience to own the resilience, scalability, and performance of our core trading platforms. You will bridge the gap between software engineering and systems operations—building self-healing infrastructure, automating away operational toil, and maintaining ultra-low latency infrastructure.
What You'll Do
- Architect for Scale: Design, deploy, and manage self-healing, high-concurrency cloud infrastructure and multi-tenant Kubernetes clusters.
- Eliminate Toil: Drive an "Automation-First" culture using Terraform, Ansible, and Python/Go to automate infrastructure provisioning, auto-scaling, and failovers.
- Master Observability: Build full-stack observability pipelines (Prometheus, Grafana, Datadog, ELK) to capture high-cardinality metrics, tracing, and log aggregation before issues hit production.
- Define Reliability Standards: Establish and enforce SLIs, SLOs, and Error Budgets across microservices teams to strike the right balance between rapid deployment and platform stability.
- Incident Commander & RCA: Lead high-severity incident responses, conduct blameless Root Cause Analyses (RCAs), and build long-term engineering fixes to eliminate recurring failure modes.
- CI/CD Optimization: Optimize continuous deployment pipelines (GitHub Actions, GitLab CI, ArgoCD) for zero-downtime, continuous release cycles.
- Disaster Recovery (DR) & Chaos Engineering: Conduct chaos testing and design multi-region disaster recovery protocols to guarantee business continuity under any scenario.
What We're Looking For
- 8-10 years of hands-on experience in SRE, Platform Engineering, or Cloud Infrastructure handling mission-critical, large-scale systems.
- Containerization & Orchestration: Deep expertise with Kubernetes (EKS/GKE/Self-hosted) and Docker in production.
- Infrastructure as Code (IaC): Advanced hands-on mastery of Terraform and modular automation tools.
- Cloud Mastery: Strong expertise across AWS, Azure, or GCP core services, networking (VPCs, BGP, DNS, Load Balancers), and IAM security models.
- Coding & Scripting: Strong programming ability in Python, Go, or Bash to build custom operators, internal tools, and automation scripts.
- Observability Expert: Practical experience setting up distributed tracing, APM, and automated alerting frameworks.
Nice to Have
- Fintech Experience: Prior exposure to high-frequency trading platforms, broking systems, payment gateways, or banking backends.
- Service Mesh: Exposure to Istio or Linkerd.
- Certifications: CKA/CKAD, AWS Solutions Architect Professional, or GCP Cloud Engineer.
Equal Opportunity Employer
At Bajaj Broking, we build teams solely on technical merit, problem-solving ability, and shared commitment to technical excellence.