Get more replies from employers
Send a job-specific resume in minutes.
SYNC is seeking an MLOps + DevOps Engineer to build and operate an AI-native platform at the intersection of ML systems, backend infrastructure, and distributed workflows. You will ensure models, agents, APIs, and end-to-end workflows run reliably, scale predictably, and stay observable in production.
You will design cloud architectures, manage ML platforms on AWS, and implement CI/CD pipelines beyond basic GitHub Actions, while ensuring security and governance across services.
We are building an AI-native platform where agentic systems, backend services, and real workflows operate together. This role sits at the intersection of ML Systems, Backend infrastructure and Distributed system operations. You will be responsible for ensuring that models, agents, APIs and workflows run reliably, scale predictably and remain observable end to end. This is not a traditional ML/DevOps role, but it is about operating intelligent systems in production.
1. Design and implement cloud architectures supporting AI/ML workloads and production-grade systems
2. Build and manage ML platforms using AWS services including EC2, ECS/EKS, Lambda, S3, RDS, VPC, and IAM
3. Leverage AWS Bedrock, SageMaker, or similar managed AI services for model training and deployment
4. Use Infrastructure-as-Code tools such as Terraform, CloudFormation, or CDK to automate cloud provisioning
5. Work with vector databases (Milvus, Pinecone, Weaviate) and graph databases (Neo4j) to support retrieval-based and knowledge-driven AI solutions
6.1.Deploy and manage Small and Medium Language Models (SLMs)
6.2.Manage external LLM integrations
6.3. Build pipeline for model versioning, evaluation and fine-tuning
6.4. Support RAG systems, embeddings and Vector database infra
7. Agentic System Runtime
7.1.Enable execution of multi-agent workflows
7.3.Ensure consistency, fault-tolerance and latency control
8. Design and operate event-driven backend infrastructure – Apache Kafka (or equivalent)
9. Handle async workflows, retries, ordering, idempotency and enable reliable communication between backend and AI
10.Own Kubernetes cluster design, scaling strategies and workload isolation