Staff Software Engineer — Platform & Distributed Systems The Problem

AiFA Labs

Hyderabad

On-site

INR 3,000,000 - 6,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AiFA Labs in Hyderabad is seeking a Staff Software Engineer to transform an AI platform by instrumenting a production system to capture signals, stream data to a learning layer, and enable continual learning without downtime.

You will own event-driven design across 50+ API endpoints, instrument frontend (React) and backend, and drive observability with traces, metrics, and structured logs. This role requires leadership and deep experience in distributed systems.

Qualifications

  • 8+ years software engineering with staff-level depth
  • Distributed systems experience with CAP and eventual consistency
  • Event-driven architecture experience with Kafka, RabbitMQ, or similar at scale
  • Full-stack capability: backend Python and frontend Next.js/React
  • Production transformation experience: refactor, migrate, instrument systems while serving traffic

Responsibilities

  • Design and implement event emission across 50+ API endpoints
  • Define schemas capturing full context and why events happen
  • Build reliable event pipelines with exactly-once semantics where it matters
  • Handle backpressure, failures, and replay in streaming pipelines
  • Instrument React frontend to capture user behavior and signals
  • Ensure end-to-end event flow: UI → API → Event Bus → Storage
  • Apply strangler pattern for incremental migration and feature flags
  • Define patterns and review designs for event-driven consistency
  • Document decisions and trade-offs for future maintainers

Skills

Python
Next.js
React
Kafka
MongoDB
PostgreSQL
Redis
Docker
Kubernetes
Azure
Event-driven
Distributed systems

Tools

OpenTelemetry

Job description

Staff Software Engineer – Platform & Distributed Systems

Location: Hyderabad

Experience: 7+ Years

Type: Full-time

Notice Period: Immediate to 15 Days

Level: Staff Engineer (L6/ Consultant/ Architect Equivalent)

We have a working AI platform. Now we need to make it learn. Every AI output, every user interaction, every accept/reject/modify decision contains signal. Right now, that signal is lost. Your job: instrument a production system to capture everything, stream it to a learning layer, and close the feedback loop — without breaking what’s already working for customers.

This is surgery on a moving train. You’ll touch frontend, backend, event pipelines, and data flows. You’ll design schemas that capture context without bloat. You’ll refactor services to emit events at scale. You’ll do it incrementally, behind feature flags, with zero downtime.

Not a rewrite. A transformation for Scale, Efficiency, Reliability, & continual learning.

What You’ll Own
Event-Driven Architecture
  • Design and implement event emission across 50+ API endpoints
  • Define schemas that capture full context (not just what happened, but why)
  • Build reliable event pipelines (Kafka) with exactly-once semantics where it matters
  • Handle backpressure, failures, and replay
Full-Stack Instrumentation
  • Instrument React frontend to capture user behavior (actions, timing, implicit signals)
  • Build low-friction feedback components that users actually use
  • Ensure end-to-end event flow: UI -> API -> Event Bus -> Storage
System Transformation
  • Apply strangler pattern to extract services without disruption
  • Implement feature flags for incremental rollouts
  • Design for observability from day one (traces, metrics, structured logs)
  • Migrate historical data to new event schemas
Technical Leadership
  • Define patterns that the rest of engineering will follow
  • Review designs and code for event-driven consistency
  • Document decisions and trade-offs for future maintainers
Why This Role
  • Transform, don’t build from scratch — Harder than greenfield. You’ll make a production system smarter while it’s serving customers.
  • Full ownership — You’ll make architectural decisions, not just implement tickets.
  • AI platform scale — Your instrumentation directly impacts how the platform learns. Bad events = bad AI. Good events = compounding intelligence.
  • Staff-level scope — Cross-team influence, technical direction, patterns that scale.
  • Modern stack, real problems — Kafka, Kubernetes, event sourcing, distributed systems. Not legacy maintenance — system evolution.
You Are
Required
  • 8+ years software engineering—Staff-level depth. You’ve made architectural mistakes and learned from them.
  • Distributed systems experience — You understand CAP theorem trade-offs, eventual consistency, and when strong consistency actually matters.
  • Event-driven architecture — You’ve built or significantly contributed to event-driven systems. Kafka, RabbitMQ, or similar at scale.
  • Full-stack capability — Fluent in backend (Python) AND frontend (Next.js/React). Can move between layers without context-switching pain.
  • Production transformation experience — You’ve refactored, migrated, or instrumented systems while they were serving real traffic. Strangler pattern, feature flags, incremental rollouts.
Strong Plus
  • Observability implementation (OpenTelemetry, Datadog, Prometheus)
  • AI/ML platform experience (LLM applications, inference pipelines)
  • Event sourcing or CQRS patterns
  • High-scale analytics instrumentation (Segment, Amplitude, custom pipelines)
  • Microservices decomposition from monoliths
Tech Stack (Important)
  • Python
  • MongoDB
  • React
  • Kafka
  • PostgreSQL
  • Redis
  • Docker
  • Kubernetes
  • Azure
  • Next.js
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer II - Platform & Integrations (Backend) Engineering - Development Hyderabad
Senior Software Engineer II - Platform & Integrations (Backend) Engineering - Development Hyderabad

Seismic • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Staff AI Engineer
Staff AI Engineer

Pattern • Pune District

Hybrid
INR 2,000,000 - 3,000,000
Staff Software Engineer - Python, React & Cloud
Staff Software Engineer - Python, React & Cloud

Blue Yonder • Hyderabad

Hybrid
INR 1,500,000 - 2,500,000
Lead / Staff Engineer–AI Platform | Python | React | AWS
Lead / Staff Engineer–AI Platform | Python | React | AWS

ULTISOURCE • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Staff Engineer - Backend Platform
Staff Engineer - Backend Platform

Prismforce • Pune District

On-site
INR 2,000,000 - 3,000,000
Full Stack Lead Engineer
Full Stack Lead Engineer

YottaFlex AI Technologies Inc • Hyderabad

On-site
INR 3,500,000 - 5,500,000
Competitive Hyderabad market rates
Lead high-impact client engagements
Structured L&D budget
+2
Staff Software Engineer - AI Platform
Staff Software Engineer - AI Platform

Addepar • Pune District

On-site
INR 1,800,000 - 2,500,000
Senior Software Engineer - Backend Focused (Full Stack)
Senior Software Engineer - Backend Focused (Full Stack)

First Object • Dadri, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Staff Engineer- Full Stack
Staff Engineer- Full Stack

Michael Page • Bengaluru Urban

On-site
Competitive salary package
Attractive holiday leave policies
Collaborative work culture
+1
Staff Software Engineer - Data Platform
Staff Software Engineer - Data Platform

United States Digital Space LLC • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000