(IND) STAFF, SOFTWARE ENGINEER

Walmart Global Tech India

Chennai District

On-site

INR 2,500,000 - 3,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Incentive awards
Health benefits
PTO

Job summary

Walmart Global Tech India is seeking a Technical Duty Officer responsible for the availability and performance of global sites. The role requires 10-14 years in infrastructure or engineering environments. You will lead incident response, manage a team, and ensure system reliability. Preferred qualifications include a Bachelor's in Computer Science and experience in a 24/7 operations support environment. The position offers a broad range of benefits, including health benefits and PTO.

Qualifications

  • 10-14 years in infrastructure or engineering environments.
  • Experience in a 24/7 operations support environment.
  • Exposure to large-scale enterprise systems.

Responsibilities

  • Manage availability and performance of global sites.
  • Lead incidents focusing on restoration and documentation.
  • Oversee Site Reliability Operations team.

Skills

Incident Management
Kubernetes
DevOps
Problem-solving
Unix/Linux Systems
Cloud Technologies
Scripting Languages
Networking Concepts
Monitoring Tools
AI Systems

Education

Bachelor’s Degree in Computer Science or related field
Master’s degree in Computer Science or related field

Job description

About The Team

TDO need to own any issues related to site availability and drive through the resolution and have the accountability to make necessary changes to fix issues and bring the sites backup online and make sure customer experience is seamless, execute recovery levers during outages, configuration mishaps, and DR situations. Leveraging technical experience to keep critical systems running through any event.

The Technical Duty Officer (TDO) is responsible for the availability and performance of our global sites. The TDO will take command and control of major incidents focusing on restoration by identifying and coordinating with appropriate resources through all the phases of triage, restoration and validation. Technically you will understand the full end‑to‑end stack and use this knowledge to detect and lead a team through incident response. Excellent judgement is crucial as you will provide final approval on site changes and hold critical switches for functionality of the site.

You will ensure all documentation surrounding the major incidents are accurate and communication with the leadership team is clear and complete. Your ability to continuously challenge yourself and develop a strong network with peers and stakeholders cross‑functionally will see you exceed in this role. Our goal is to protect the customer experience and deliver outstanding levels of availability.

What You’ll Do
  • Architecture Acumen: Requires knowledge of architectural principles, systems and environment behavior, architectural styles, patterns and plans, architectural standards, non‑functional system performance parameters, technology strategy.
  • Defect Management and Troubleshooting: Knowledge of defect life‑cycle process, defect tracking tools and methodologies, defect reporting, regression testing, root cause analysis and corrective action. Track registered issues for the product/solution and prioritize them for resolution.
  • DevOps Orientation: Knowledge of different operating systems, software maintenance tools and techniques, application monitoring tools and techniques, debugging tools, mock screen, pseudocodes, reverse engineering, traceability matrix, system performance, security, integration, data migration and accessibility, design methodologies.
  • Agentic AI Framework: Managing human‑in‑the‑loop systems where AI agents plan and execute tasks while ensuring alignment with business goals, security, and safety.
  • Prompt Engineering and Evaluation: Proficiency in designing prompts that enable agents to act reliably and setting up evaluation frameworks for monitoring performance.
  • Requirement and Scoping Analysis: Knowledge of traceability matrix, risk analysis methodologies, cost analysis, business objectives, classification of requirements, user stories.
  • Strong incident management skills with relevant experience in an enterprise organization.
  • Methodical and systematic problem‑solving approach combined with a solid awareness of ownership, initiative and drive.
  • Experience investigating, analysing and troubleshooting large‑scale enterprise systems.
  • Understanding of Unix/Linux systems from kernel to shell and beyond, including system libraries, file systems, and client‑server protocols.
  • Experience working with and developing enterprise monitoring/tooling solutions like Grafana, Prometheus, Kibana, Splunk, Graphite, Dynatrace, Catchpoint.
  • Working knowledge of one or more cloud technologies such as Azure, GCP and OpenStack.
  • Excellent verbal and written communication skills.
  • Excellent judgement in decision making.
  • Strong focus on collecting and inferring metrics.
What You’ll Bring
  • 10‑14 years in an infrastructure, systems, engineering or development environment delivering operational excellence to highly complex distributed systems.
  • Bachelor’s Degree in Computer Science or a related field, or relevant work experience of 10+ years.
  • Experience and exposure working in a 24/7 operations support environment.
  • Working and technical expertise in Kubernetes and microservice architectures.
  • Experience administering Unix/Linux in a production environment.
  • Ability to supervise the Site Reliability Operations team, mentor and provide guidance.
  • Working knowledge of BASH, Python, AI or other scripting languages.
  • Utilize AI‑powered monitoring and anomaly detection tools to predict potential failures and resource bottlenecks before they impact users.
  • Ensure the reliability, performance, and scalability of infrastructure specifically designed for AI/ML workloads.
  • Networking knowledge and understanding of network concepts such as different protocols (TCP/IP, UDP, ICMP, etc.), MAC addresses, IP packets, DNS, OSI layers, and load balancing.
Benefits

Beyond a great compensation package, you can receive incentive awards for your performance. Other perks include a host of best‑in‑class benefits such as maternity and parental leave, PTO, health benefits, and more.

Equal Opportunity Employer

Walmart, Inc., is an Equal Opportunities Employer – By Choice. We believe we are best equipped to help our associates, customers and the communities we serve live better when we really know them. That means understanding, respecting and valuing unique styles, experiences, identities, ideas and opinions – while being inclusive of all people.

Minimum Qualifications
  • Option 1: Bachelor’s degree in computer science, computer engineering, computer information systems, software engineering, or related area and 4 years’ experience in software engineering or related area.
  • Option 2: 6 years’ experience in software engineering or related area.
Preferred Qualifications
  • Master’s degree in Computer Science, Computer Engineering, Computer Information Systems, Software Engineering, or related area and 2 years’ experience in software engineering or related area.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

(IND) SENIOR, SOFTWARE ENGINEER
(IND) SENIOR, SOFTWARE ENGINEER

Walmart Global Tech India • Chennai District

On-site
INR 80,000 - 110,000
Incentive awards
Health benefits
Maternity and parental leave
(IND) SENIOR, SOFTWARE ENGINEER
(IND) SENIOR, SOFTWARE ENGINEER

Walmart Global Tech India • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Incentive awards for performance
Maternity and parental leave
Health benefits
+1
Senior Technology Operations
Senior Technology Operations

Walmart Global Tech India • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Incentive awards
Maternity and parental leave
Health benefits
+1
STAFF, SOFTWARE ENGINEER
STAFF, SOFTWARE ENGINEER

Walmart Global Tech India • Chennai District

On-site
INR 2,500,000 - 4,000,000
Incentive awards
Health benefits
Maternity and parental leave
+1
STAFF, DATA ENGINEER
STAFF, DATA ENGINEER

Walmart Global Tech India • Chennai District

On-site
INR 1,500,000 - 2,000,000
Incentive awards
Maternity and parental leave
Health benefits
+1
STAFF, SOFTWARE ENGINEER (Machine Learning)
STAFF, SOFTWARE ENGINEER (Machine Learning)

Walmart • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Maternity/Parental leave
Health benefits
Paid time off
(IND) GROUP DIRECTOR, SOFTWARE ENGINEERING
(IND) GROUP DIRECTOR, SOFTWARE ENGINEERING

Walmart Global Tech India • Bengaluru

On-site
INR 3,000,000 - 5,000,000
Incentive awards
Health benefits
PTO
+1
PRINCIPAL, SOFTWARE ENGINEER
PRINCIPAL, SOFTWARE ENGINEER

Walmart • Chennai District

On-site
INR 2,000,000 - 3,000,000
Incentive awards
Maternity and parental leave
Health benefits
(IND) SENIOR, SOFTWARE ENGINEER
(IND) SENIOR, SOFTWARE ENGINEER

Walmart Global Tech India • Bengaluru

On-site
INR 1,500,000 - 2,200,000
Incentive awards
Health benefits
Parental leave
SENIOR, SOFTWARE ENGINEER
SENIOR, SOFTWARE ENGINEER

Walmart Global Tech India • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Incentive awards
Maternity and parental leave
Health benefits
+1