Own architecture for high-scale, high-availability products - Cruise, domains, and business email - with agent-first contracts and fail-closed trade-offs.
Work mode Remote (India)
Experience 10+ years
Compensation ₹45-65 LPA
Employment Full-time
About the role
You will own how BytesBrains’ products fit together under real load: Cruise (our LLM API gateway), domains, and business email - plus the surfaces humans and AI agents both call. You have designed and operated high-scale, high-availability systems before, and you turn product intent into durable boundaries, contracts, and failure modes - not slide decks. The work matters because we treat agents as first-class customers: every design must be machine-usable, cost-aware, and fail closed when the supplier or the budget says no.
What you will do
- Design end-to-end architectures for multi-product SaaS that must stay available under load - tenancy, trust boundaries, and how callables, gateways, and UIs share one contract
- Set the bar for high availability and operability: failure domains, retries, degradations, observability, and runbooks that a human or agent can act on
- Write and review architecture decisions (APIs, auth, isolation, secrets, billing/ledger) so trade-offs and refusals stay explicit
- Shape OpenAI-compatible and reseller-facing APIs so agents get structured errors they can branch on, never prose-only failures
- Guide integrations with external suppliers (LLM providers, OpenSRS, email hosts) using environment-scoped credentials, rate limits, and fail-closed switches
- Review critical designs and PRs for security, cost, scale, and “absence means unknown” - not silent defaults
- Raise the quality of ADRs, diagrams, and handoffs so the next person can ship without tribal knowledge
What you will need
- 10+ years building production software, including ownership of architecture on high-scale, high-availability products or services
- Proven system design under load: tenancy, authn/authz, idempotency, async workflows, capacity, and clear API contracts (OpenAPI or equivalent)
- Hands-on experience keeping production systems available - SLOs, incident response, graceful degradation, and supplier or dependency failure
- Strong judgement on cloud backends (e.g. Firebase/GCP or similar): rules, secrets isolation, and what must never sit on a customer-readable document
- Ability to write precise technical prose and lead design reviews without hiding risk behind jargon
- Comfort collaborating in a small remote team that ships TypeScript/Node services and modern web front ends
Nice to have
- LLM gateways, agent tooling, or OpenAI-compatible proxy experience (routing, spend controls, multi-provider fallback)
- Registrar, hosted-email, or other reseller API integrations at production volume
- Next.js (or similar) + Cloud Functions (or equivalent serverless) in production
- Public technical writing, ADRs, or open-source architecture work
How we hire
- A 10-minute introductory call.
- A practical task: you commit work to a public GitHub repository and share the pull request.
- A 25-minute final conversation.
What we offer
- Ownership of real product architecture on systems that must stay up - not advisory theatre
- Direct impact on Cruise and the platforms agents can buy from and call
- Remote-first India work with a small, high-context team
- Compensation in the published ₹45-65 LPA band, negotiated in the final round within that commitment