Sr. Site Reliability Engineer

Optomi

Fort Worth, Arlington (TX, TX)

Hybrid

USD 89,544 - 103,320

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Vision insurance

Job summary

A leading tech solutions firm is seeking a Senior Site Reliability Engineer for a 6-month contract based in Fort Worth, Texas. This role involves enhancing infrastructure, improving system performance, and optimizing cost-effectiveness in a hybrid work setting. The ideal candidate has over 5 years of experience with Docker and Kubernetes, along with strong analytical skills. Competitive pay range of $65-$75/hr offered.

Qualifications

  • 5+ years of experience as a Site Reliability Engineer or DevOps Engineer.
  • Proven track record of delivering highly available systems.
  • Deep expertise in Docker, Kubernetes, and AWS.

Responsibilities

  • Identify performance improvements in system availability and scalability.
  • Collaborate with software engineering and DevOps teams.
  • Monitor system performance using Datadog.

Skills

Docker
Kubernetes
AWS
Observability tools
Analytical skills

Tools

Datadog
Terraform
Ansible
Argo CD

Job description

Get AI-powered advice on this job and more exclusive features.

This range is provided by Optomi. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.

Base pay range

$65.00/hr - $75.00/hr

Direct message the job poster from Optomi

Senior Cloud and Infrastructure Recruiter at Optomi

Sr. Site Reliability Engineer

*6-month contract with extensions*

Hybrid: 4x a week onsite in Plano, TX

Optomi, in partnership with our premier client in the manufacturing industry, is seeking a Senior Site Reliability Engineer to join a dynamic and fast-paced engineering team. The ideal candidate will have extensive experience managing large-scale, microservice-based systems, ensuring high availability, and implementing best practices in reliability engineering. This role involves close collaboration with development and operations teams to enhance infrastructure, improve system performance, and optimize cost-effectiveness.

Responsibilities:

  • Proactively identify performance improvements in responsiveness, availability, and scalability.
  • Establish and drive adoption of best practices in observability, monitoring, and incident response.
  • Lead incident response efforts and perform post-mortem analyses to prevent recurrence.
  • Collaborate with Software Engineering and DevOps teams to design, implement, and maintain scalable and reliable systems using Kubernetes, Docker, and Istio.
  • Monitor system performance and troubleshoot issues using tools like Datadog.
  • Implement and tune Horizontal Pod Autoscalers (HPAs) for efficient resource utilization.
  • Develop and maintain automation for deployment, monitoring, and incident response.
  • Support GitOps-based deployment strategies using Argo CD.
  • Implement deployment strategies such as A/B testing, canary deployments, and traffic mirroring.
  • Mentor junior engineers and foster team knowledge sharing.
  • Coordinate with global SRE teams to support effective on-call rotations.
  • Utilize Helm for application deployment and configuration management.
  • Implement and manage cloud infrastructure, particularly AWS services including Load Balancers and routing for high-traffic systems.
  • Participate in on-call rotations and provide production system support.

Required Qualifications:

  • 5+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Software Engineer in production environments.
  • Proven track record of delivering highly available and scalable systems.
  • Strong analytical, troubleshooting, and decision-making skills.
  • Deep expertise in Docker, Kubernetes, and Istio.
  • Advanced knowledge of AWS infrastructure and services.
  • Experience with Argo CD and GitOps deployment practices.
  • Proficiency in observability tools such as Datadog, ELK, Grafana, AppDynamics, or Prometheus.
  • Experience implementing progressive delivery techniques (A/B, Canary, Blue/Green, traffic mirroring).
  • Familiarity with Infrastructure as Code and automation tools (e.g., Terraform, Ansible).
  • Experience delegating tasks and mentoring junior team members.
  • Excellent understanding of system interdependencies and holistic architecture.
  • Strong communication and collaboration skills.
  • High level of organization, independence, and attention to detail.
  • Bonus: Proficiency in Golang or Rust.
Seniority level
  • Seniority level
    Mid-Senior level
Employment type
  • Employment type
    Full-time
Job function
  • Job function
    Information Technology
  • Industries
    Motor Vehicle Manufacturing and Software Development

Referrals increase your chances of interviewing at Optomi by 2x

Inferred from the description for this job

Medical insurance

Vision insurance

Get notified when a new job is posted.

Sign in to set job alerts for “Site Reliability Engineer” roles.
Software Developer / Junior Software Engineer

Irving, TX $107,120.00-$160,680.00 1 week ago

Backend Associate Developer/Developer, IT Applications

Dallas, TX $125,000.00-$170,000.00 6 days ago

Dallas, TX $90,000.00-$180,000.00 2 weeks ago

Plano, TX $112,000.00-$130,000.00 2 weeks ago

Plano, TX $112,000.00-$130,000.00 2 weeks ago

Dallas, TX $140,000.00-$160,000.00 1 month ago

Plano, TX $64,300.00-$98,300.00 5 days ago

Dallas, TX $70,000.00-$82,000.00 2 weeks ago

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Signature IT World Inc • Washington

On-site
USD 120,000 - 200,000
Medical insurance
Vision insurance
Child care support
+3
Site Reliability Engineer
Site Reliability Engineer

Felix Recruitment • Dallas (TX)

On-site
USD 110,000 - 175,000
Medical insurance
Vision insurance
401(k)
Site Reliability Engineer
Site Reliability Engineer

Optomi • Town of Florida (NY)

On-site
USD 145,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Interactive Resources - iR • Austin (TX)

On-site
USD 126,000 - 223,000
Medical insurance
Vision insurance
401(k)
Site Reliability Engineer
Site Reliability Engineer

Prestige Staffing • Atlanta (GA)

On-site
USD 130,000 - 150,000
Vision insurance
401(k)
Paid maternity leave
+2
*REMOTE* Cloud Platform Engineer (Azure, Kubernetes, Terraform)
*REMOTE* Cloud Platform Engineer (Azure, Kubernetes, Terraform)

Optomi • United States

Remote
USD 100,000 - 180,000
Fully Remote
Medical insurance
Vision insurance
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment • Atlanta (GA)

On-site
USD 93,000 - 158,000
Data Engineer
Data Engineer

Optomi • Dallas (TX)

On-site
USD 120,000 - 140,000
Site Reliability Engineer (10+ Years) (Need Local only)
Site Reliability Engineer (10+ Years) (Need Local only)

Ampstek • Bellevue (WA)

On-site
USD 130,000 - 150,000
Senior DevOps Engineer
Senior DevOps Engineer

Xoriant • Fort Worth (TX), Arlington (TX)

On-site
USD 107,000 - 161,000