Senior Technical Lead - DevOps

TechDigital Group

Bellevue (WA)

On-site

USD 80,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

An established industry player is seeking a skilled support engineer to enhance system stability and performance. In this dynamic role, you will provide critical consulting services, troubleshoot large-scale distributed systems, and support applications running on Kubernetes and cloud platforms. Your expertise will be vital in leading root cause analysis sessions and ensuring high availability and disaster recovery readiness. Join a collaborative team that values innovation and problem-solving, where your contributions will have a significant impact on operational excellence and system reliability.

Qualifications

  • Expertise in troubleshooting large-scale distributed systems.
  • Strong experience with Kubernetes and cloud technologies.

Responsibilities

  • Provide 24X7 on-call support and troubleshoot customer issues.
  • Lead root cause analysis sessions and ensure documentation is updated.

Skills

Analyzing large-scale distributed systems
Kubernetes
Gloo
AWS
Apigee API Gateway
REST, SOAP, and GraphQL API support
Git
GitLab
Docker
Unix/Linux operating systems
Scripting knowledge
Networking, routing, and TLS/SSL

Tools

Git
GitLab
Docker
Postman
Splunk
App Dynamics
Imperva WAF
CI/CD tools

Job description

- Provide consulting services for improved system stability, availability, performance and reliability.

- Assist in determining the impact of operational issues and provide input into their resolution via data extraction and quantification.

- Work through day-to-day support issues, ensure effective and timely resolution of issues in production environment, troubleshoot customer impacting issues.

- Support multiple applications, specifically running Kubernetes, Gloo, AWS, Apigee, PCF, GCP/Java based systems in an enterprise environment.

- Support Gloo running on Kubernetes, Apigee opdk and saas, Grafana, Prometheus, Cassandra, Postgres, Spring Boot or Java based applications running on Kubernetes, PCF, and Java application servers.

- Apply GitOps principles to manage infrastructure and application configurations.

- Apply monitoring and create complex alerts and dashboards for production systems.

- Provide capacity analysis and tuning analysis for Apigee and Java applications hosted on LINUX and container platform.

- Available to provide 24X7 on-call support on a rotating basis with other team members.

- Lead efforts in troubleshooting, recovery, and root cause investigation.

- Perform analysis of user requirements and problems to automate or improve systems and review system capabilities, workflow, and scheduling limitations.

- Able to follow and develop detailed work plans, schedules, project estimates, resource plans, and status reports.

- Facilitate HA (High Availability) / DR (Disaster Recovery) exercises to ensure that the team is fully prepared for any event.

- Lead root cause analysis sessions to understand what causes issues in Production and come up with RCA Report along with solutions that will prevent them from happening in the future.

- Ensure documentation is created and remains updated for any related work.

- Strong understanding of UNIX operating systems and any scripting language.

- Forecast and plan for a rapidly growing environment.

- Evaluate new software product and service solutions.

Skill Requirements:
  1. Expertise in analyzing and troubleshooting large-scale distributed systems.
  2. Strong experience with Kubernetes – Container Orchestration Tool, Gloo, AWS, Apigee API Gateway.
  3. Experience with REST, SOAP, and GraphQL API support.
  4. Experience with tools like: Git, Gitlab, Docker, Postman, Splunk, App Dynamics, Imperva WAF and CI/CD tools.
  5. Good experience in GitOps process, performance measurement tuning, capacity planning and management, contingency, and disaster recovery.
  6. Good understanding and strong experience with Unix/Linux operating systems.
  7. Ability to debug, optimize code, and automate routine tasks.
  8. Systematic problem-solving approach coupled with effective communication skills.
  9. Strong scripting knowledge and experience.
  10. Good understanding of networking, routing, and TLS/SSL.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, DevOps
Manager, DevOps

1 O.C. Tanner Company • Salt Lake City (UT)

On-site
USD 120,000 - 150,000
Lead Full Stack Developer
Lead Full Stack Developer

Compunnel, Inc. • Atlanta (GA)

On-site
USD 100,000 - 130,000
Java DevOps Engineer
Java DevOps Engineer

NLB Services • Buffalo (NY)

On-site
USD 80,000 - 100,000
Senior DevOps Engineer
Senior DevOps Engineer

FutureRecruit.net • City of White Plains (NY)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Dallas (TX)

On-site
USD 90,000 - 120,000
Senior DevOps Engineer (Kubernetes, Docker, Jenkins)
Senior DevOps Engineer (Kubernetes, Docker, Jenkins)

CatchProbe Intelligence Technologies • San Francisco (CA)

Remote
USD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

Relanto • Fremont (CA)

On-site
USD 120,000 - 150,000
Senior Engineer, Systems Engineering
Senior Engineer, Systems Engineering

ICE • Jacksonville (FL)

On-site
USD 110,000 - 150,000
Senior Application Development Advisor
Senior Application Development Advisor

Jobtailor • Colorado

Hybrid
USD 130,000 - 170,000
Staff Software Engineer - Cloud Platform
Staff Software Engineer - Cloud Platform

UKG • Lowell (MA)

On-site
USD 130,000 - 180,000