We are seeking a highly skilled Azure Cosmos DB Database Administrator (DBA) to manage, optimize, and support enterprise-scale distributed databases. This role will be responsible for ensuring high availability, performance, security, and cost efficiency of Cosmos DB deployments across mission-critical applications.
Key Responsibilities
1. Database Administration & Operations
- Manage and support Azure Cosmos DB environments
- Monitor health, performance, and availability of Cosmos DB accounts, containers, and workloads
- Perform capacity planning, throughput (RU/s) management, and partition key optimization
- Ensure SLA adherence for availability, latency, and throughput
2. Performance & Cost Optimization
- Analyze query performance and optimize indexing policies, partition strategies, and data models
- Right-size RU/s provisioning and implement autoscale strategies to reduce costs
- Identify and remediate hot partitions and inefficient query patterns
- Implement caching strategies and optimize SDK usage where applicable
3. Reliability, Backup & Disaster Recovery
- Configure and manage multi-region replication and failover policies
- Validate backup and restore processes (continuous backup or periodic modes)
- Define and maintain DR runbooks aligned to RTO/RPO requirements
- Participate in DR tests and resilience validation exercises
4. Security & Compliance
- Implement and enforce RBAC, managed identities, and network security controls (Private Endpoints, VNet integration)
- Ensure encryption (at-rest and in-transit) and compliance with enterprise security standards
- Support audit and compliance requirements (SOX, SOC, etc.)
5. Automation & DevOps Integration
- Automate provisioning and management using Infrastructure-as-Code (Ansible)
- Integrate Cosmos DB operations into CI/CD pipelines
- Develop scripts/runbooks for monitoring, alerting, and remediation
6. Monitoring & Observability
- Leverage Azure Monitor, Log Analytics, and Application Insights for telemetry
- Build dashboards and alerts for proactive issue detection
- Partner with Observability platforms (e.g., Dynatrace) for full-stack visibility
- Work closely with application teams on access patterns, and performance tuning
- Provide L3/L4 support for production incidents and root cause analysis (RCA)
- Contribute to architecture reviews and cloud migration initiatives
Required Skills & Experience
Core Database Expertise
- Strong hands-on experience with Azure Cosmos DB
- Deep understanding of:
- Indexing and query optimization
- RU/s consumption and scaling models
- Experience with NoSQL data modeling and distributed systems
- Proficiency in:
- Deep knowledge of Microsoft Azure
- Experience with:
- High availability and failover strategies
- Networking (VNet, Private Endpoints, NSGs)
Nice to Have / Secondary Skills
Other NoSQL / Distributed Databases
- Hands-on experience with:
- MongoDB (native)
- DynamoDB or equivalent distributed stores
- Understanding of:
- CAP theorem, eventual consistency, and distributed design trade-offs