About The Team
TDO need to own any issues related to site availability and drive through the resolution and have the accountability to make necessary changes to fix issues and bring the sites backup online and make sure customer experience is seamless, execute recovery levers during outages, configuration mishaps, and DR situations. Leveraging technical experience to keep critical systems running through any event.
The Technical Duty Officer (TDO) is responsible for the availability and performance of our global sites. The TDO will take command and control of major incidents focusing on restoration by identifying and coordinating with appropriate resources through all the phases of triage, restoration and validation. Technically you will understand the full end‑to‑end stack and use this knowledge to detect and lead a team through incident response. Excellent judgement is crucial as you will provide final approval on site changes and hold critical switches for functionality of the site.
You will ensure all documentation surrounding the major incidents are accurate and communication with the leadership team is clear and complete. Your ability to continuously challenge yourself and develop a strong network with peers and stakeholders cross‑functionally will see you exceed in this role. Our goal is to protect the customer experience and deliver outstanding levels of availability.
What You’ll Do
- Architecture Acumen: Requires knowledge of architectural principles, systems and environment behavior, architectural styles, patterns and plans, architectural standards, non‑functional system performance parameters, technology strategy.
- Defect Management and Troubleshooting: Knowledge of defect life‑cycle process, defect tracking tools and methodologies, defect reporting, regression testing, root cause analysis and corrective action. Track registered issues for the product/solution and prioritize them for resolution.
- DevOps Orientation: Knowledge of different operating systems, software maintenance tools and techniques, application monitoring tools and techniques, debugging tools, mock screen, pseudocodes, reverse engineering, traceability matrix, system performance, security, integration, data migration and accessibility, design methodologies.
- Agentic AI Framework: Managing human‑in‑the‑loop systems where AI agents plan and execute tasks while ensuring alignment with business goals, security, and safety.
- Prompt Engineering and Evaluation: Proficiency in designing prompts that enable agents to act reliably and setting up evaluation frameworks for monitoring performance.
- Requirement and Scoping Analysis: Knowledge of traceability matrix, risk analysis methodologies, cost analysis, business objectives, classification of requirements, user stories.
- Strong incident management skills with relevant experience in an enterprise organization.
- Methodical and systematic problem‑solving approach combined with a solid awareness of ownership, initiative and drive.
- Experience investigating, analysing and troubleshooting large‑scale enterprise systems.
- Understanding of Unix/Linux systems from kernel to shell and beyond, including system libraries, file systems, and client‑server protocols.
- Experience working with and developing enterprise monitoring/tooling solutions like Grafana, Prometheus, Kibana, Splunk, Graphite, Dynatrace, Catchpoint.
- Working knowledge of one or more cloud technologies such as Azure, GCP and OpenStack.
- Excellent verbal and written communication skills.
- Excellent judgement in decision making.
- Strong focus on collecting and inferring metrics.
What You’ll Bring
- 10‑14 years in an infrastructure, systems, engineering or development environment delivering operational excellence to highly complex distributed systems.
- Bachelor’s Degree in Computer Science or a related field, or relevant work experience of 10+ years.
- Experience and exposure working in a 24/7 operations support environment.
- Working and technical expertise in Kubernetes and microservice architectures.
- Experience administering Unix/Linux in a production environment.
- Ability to supervise the Site Reliability Operations team, mentor and provide guidance.
- Working knowledge of BASH, Python, AI or other scripting languages.
- Utilize AI‑powered monitoring and anomaly detection tools to predict potential failures and resource bottlenecks before they impact users.
- Ensure the reliability, performance, and scalability of infrastructure specifically designed for AI/ML workloads.
- Networking knowledge and understanding of network concepts such as different protocols (TCP/IP, UDP, ICMP, etc.), MAC addresses, IP packets, DNS, OSI layers, and load balancing.
Benefits
Beyond a great compensation package, you can receive incentive awards for your performance. Other perks include a host of best‑in‑class benefits such as maternity and parental leave, PTO, health benefits, and more.
Equal Opportunity Employer
Walmart, Inc., is an Equal Opportunities Employer – By Choice. We believe we are best equipped to help our associates, customers and the communities we serve live better when we really know them. That means understanding, respecting and valuing unique styles, experiences, identities, ideas and opinions – while being inclusive of all people.
Minimum Qualifications
- Option 1: Bachelor’s degree in computer science, computer engineering, computer information systems, software engineering, or related area and 4 years’ experience in software engineering or related area.
- Option 2: 6 years’ experience in software engineering or related area.
Preferred Qualifications
- Master’s degree in Computer Science, Computer Engineering, Computer Information Systems, Software Engineering, or related area and 2 years’ experience in software engineering or related area.