About LotusFlare
LotusFlare is a provider of cloud-native SaaS products based in the heart of Silicon Valley. Founded by the team that helped Facebook reach over one billion users, LotusFlare was founded to make affordable mobile communications available to everyone on Earth.
Today, LotusFlare focuses on designing, building, and continuously evolving a digital commerce and monetization platform that delivers valuable outcomes for enterprises. Our platform, Digital Network Operator® (DNO™) Cloud, is licensed to telecommunications service providers and supports millions of customers globally.
LotusFlare also powers Nomad, the leading eSIM travel product providing data connectivity in over 190 countries.
The Role: Engineering at the Intersection of Systems and Software
Site Reliability Engineering (SRE) at LotusFlare combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. We are "our engineers' engineers"—we love taking things apart to rebuild them better and keep our networks running to ensure users have the fastest experience possible.
We are looking for a proficient Software Engineer to lead and redefine our site reliability, observability, and monitoring strategy for a flagship product suite that supports millions of users and global enterprise partners.
Key Responsibilities
- Lead a team of Software/Systems Engineers on mission‑critical projects and be directly responsible for system uptime.
- Lead by example: mentor the team, establish technical standards, and build a culture of intellectual curiosity and blamelessness.
- Manage on‑call rotations and establish sustainable incident response patterns for a global service.
- Manage the complete life‑cycle of services—from inception and design through deployment, operation, and refinement.
- Conduct system design consultations and perform launch reviews to ensure reliability is "baked in" from day one.
- Design, write, and deliver software to improve the availability, scalability, and efficiency of LotusFlare’s services.
- Optimize existing systems and eliminate "toil" through automation, creating automated telemetry systems that proactively identify regressions.
- Lead distributed tracing initiatives by correlating frontend exceptions (Sentry) with backend logs and traces (VictoriaLogs/OpenSearch).
- Deep‑dive into application code (TypeScript/Lua/Scala) to fix bottlenecks at the source rather than masking them with infrastructure.
- Define and refine SLIs/SLOs that truly reflect user experience and revenue impact, balancing aggressive feature delivery with long‑term system stability.
Job Requirements
- Strong background in Computer Science fundamentals, algorithms, complexity analysis, and large‑scale system design.
- Experience managing and scaling high‑volume systems (OpenSearch, VictoriaLogs) with a focus on indexing strategies and high‑volume log ingestion.
- Proficiency in TypeScript, Scala, and Lua; comfortable reading and debugging service codebases to trace request flows end‑to‑end.
- Hands‑on expertise with Grafana and Sentry (including performance profiling).
- Experience operating in environments where downtime has significant financial consequences.
- Desire to mentor, lead projects, and eventually transition into people management.
- Passion for problem‑solving in an open, collaborative, and risk‑tolerant environment.
Benefits
- Competitive salary package.
- Paid lunch (in the office).
- Comprehensive medical insurance.
- Yearly performance bonus.
- Training, workshops, and a defined path toward Engineering Management.
Recruitment Process
- HR Interview: culture fit and background alignment.
- Technical Rounds (4–5 levels): deep dives into coding/algorithms, system design, observability patterns, and a leadership strategy session.
Website
www.lotusflare.com