Now Hiring**Location:**Indonesia**Job Type:**Full Time## Job DescriptionWe are seeking an autonomous Senior Backend Engineer to build and scale our real-time AI avatar platform.You will own systems that transform user voice input into GPU-rendered, lip-synced avatar video streamed in real time, covering lipsync inference, conversational orchestration, voice-cloning/TTS, and GCP deployment.| What You'll Do> Real-Time Streaming: Develop and scale WebSocket/Protobuf services for low-latency audio and lip-synced video streaming.> AI Orchestration: Build and enhance FastAPI/asyncio services coordinating LLM, TTS/voice-cloning, and lipsync services.> GPU Inference: Manage PyTorch/CUDA inference pipelines, including GPU allocation, warm pools, batching, and multi-user concurrency.> CPU Avatar Rendering: Develop a CPU-only avatar rendering path to support lower-cost, scalable sessions where full GPU lipsync is not required.> API & Integration: Design and maintain REST, WebSocket, and Protocol Buffer APIs and service contracts.> Cloud & Reliability: Deploy and operate Docker-based services on GCP Cloud Run/GCE and GPU cloud providers, focusing on health checks, scaling, cold starts, reliability, and cost.> System Design: Collaborate with frontend engineers on API contracts and end-to-end data flow.> Session & Data: Work with Redis for session/state management and GCS/Vertex AI for storage and RAG-based retrieval.| What You'll Bring> 5+ years software engineering experience in backend or distributed systems.> Strong expertise in Python 3.10+, asyncio, type hints, and FastAPI or equivalent async frameworks.> Hands-on experience with WebSockets, streaming protocols, concurrency, and low-latency systems.> Strong knowledge of REST APIs and Protocol Buffers/gRPC-style contracts.> Experience with microservices, inter-service communication, and session-state management such as Redis.> Hands-on experience with GCP, including Cloud Run, GCE, Cloud Storage and Artifact Registry.> Proficiency in Docker, Linux administration, deployment and troubleshooting.> Understanding of GPU inference, including PyTorch/CUDA, GPU memory, batching, and concurrency.| Nice to Have> ML inference serving experience with PyTorch, ONNX Runtime, model warm-loading and batching.> LLM/GenAI API experience, including Vertex AI, Gemini and RAG.> Knowledge of audio/video streaming, encoding, buffering and backpressure.> Experience with GPU cloud providers such as RunPod/Vast.ai and cost/latency optimisation.> Firebase or equivalent authentication experience.> CI/CD pipeline experience.> Work on production-scale AI systems for real-world retail applications.> Take end-to-end ownership in a highly autonomous environment.> Join a fast-moving team that values innovation, technical excellence and clean code.