Technical Lead Site Reliability & Performance Engineer

ECLARO

Addison (TX)

On-site

USD 140,000 - 190,000

Full time

45 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401k Retirement Savings Plan
Commuter Check Pretax Benefits
Medical, Dental & Vision Insurance

Job summary

ECLARO seeks a Technical Lead Site Reliability and Performance Engineer for its client in Addison, TX. The role provides technical leadership across reliability, performance, security, observability, and operational management of enterprise platforms and modern web applications.

The position targets a senior engineer or architect who mentors teams, shapes technical direction, and leads modernization efforts while remaining hands-on.

Qualifications

  • 8+ years supporting enterprise applications, cloud platforms, infrastructure services, or web technologies.
  • 3+ years serving as a Technical Lead, Senior Engineer, Architect, or equivalent.
  • Strong experience with AWS, Terraform, Kubernetes/EKS, DNS, CDN, AWS WAF, SSL/TLS, observability platforms, CI/CD, and SRE practices.
  • Experience troubleshooting large-scale customer-facing web applications.

Responsibilities

  • Provide technical leadership across Platform Engineering and Site Reliability Engineering functions.
  • Establish engineering standards, operational best practices, and reliability objectives.
  • Lead technical decision-making for cloud infrastructure, observability platforms, domain services, infrastructure automation, application security, and operational tooling.
  • Mentor engineers and provide technical coaching across multiple disciplines.
  • Drive technical roadmaps and continuous improvement initiatives.

Skills

AWS
Terraform
Kubernetes/EKS
DNS
CDN
SSL/TLS
WAF
Observability
SRE practices
CI/CD

Tools

Terraform
Kubernetes
AWS WAF
DNS/CDN tools

Job description

Technical Lead Site Reliability and Performance Engineer

Job Number: 26-01385

Ready to take your career to the next level with a company that values innovation, leadership, and personal growth? ECLARO is looking for a Technical Lead Site Reliability and Performance Engineer for our client in Addison, TX.

ECLARO’s client is an industry-leading organization that has built a reputation for excellence through its dedication to quality products, customer satisfaction, and employee development. Join a team where your contributions can help drive meaningful results on a global scale.

Position Overview
  • Highly experienced Technical Lead - SRE and Platform Engineering to provide technical leadership for the reliability, performance, security, observability, and operational management of enterprise platforms and modern web applications.
  • This role is ideal for a senior engineer, technical lead, or architect who enjoys solving complex technical challenges, mentoring engineers, and influencing technical direction while remaining hands-on.
  • The position offers a clear growth path into a future Technical Manager - SRE and Platform Engineering role as organizational needs and leadership responsibilities expand.
  • The Technical Lead will serve as a senior technical leader for Site Reliability Engineering (SRE), monitoring and observability, domain portfolio management, cloud platform operations, infrastructure automation, web application security, and application availability.
  • The role will work closely with engineering teams in Dallas, Europe, and Asia Pacific to establish consistent operational standards, improve platform reliability, and drive technology modernization across Client's global technology landscape.
  • Play a critical role in ensuring the reliability, performance, security, and operational health of Client's externally facing digital platforms through ownership of key platform services including observability, domain services, DNS, certificate management, CDN technologies, web application firewalls, cloud platform infrastructure, and Infrastructure-as-Code (IaC) solutions.
  • This role offers a unique opportunity to work as part of a globally distributed engineering organization supporting mission-critical platforms and digital experiences used across international business operations.
  • Collaborate closely with engineering and operational teams located in Dallas, Texas (Corporate Headquarters), Europe Region, and Asia Pacific Region.
  • Through these partnerships, you will gain exposure to diverse technologies, operating models, and global business challenges while helping drive platform reliability, observability, security, and operational excellence across multiple markets.
  • This position provides the opportunity to work with globally distributed engineering teams operating in a follow-the-sun support model, influence platform strategy across multiple regions, partner with technical leaders worldwide, and develop toward a future Technical Manager role.
  • Provides a clear progression path into a Technical Manager - SRE and Platform Engineering role for candidates demonstrating leadership, operational ownership, mentoring, infrastructure automation expertise, and strategic influence.
  • Applicants must be authorized to work in the United States on a full-time basis without the need for current or future visa sponsorship.
Responsibilities
  • Technical Leadership:
    • Provide technical leadership across Platform Engineering and Site Reliability Engineering functions.
    • Establish engineering standards, operational best practices, and reliability objectives.
    • Lead technical decision-making for cloud infrastructure, observability platforms, domain services, infrastructure automation, application security, and operational tooling.
    • Mentor engineers and provide technical coaching across multiple disciplines.
    • Drive technical roadmaps and continuous improvement initiatives.
    • Evaluate, promote, and help operationalize emerging engineering capabilities, including AI-assisted development tools, coding agents, Infrastructure-as-Code automation, and other technologies that improve engineering productivity, quality, and speed of delivery.
  • Site Reliability Engineering (SRE):
    • Lead enterprise reliability initiatives focused on availability, scalability, resiliency, performance, and operational excellence.
    • Define and drive adoption of SLOs, SLIs, Error Budgets, Incident Management, and Root Cause Analysis.
    • Drive automation initiatives that reduce operational overhead and improve service reliability.
  • Monitoring & Observability Platform Ownership:
    • Serve as the technical owner for the enterprise monitoring and observability platform.
    • Support APM, infrastructure monitoring, synthetic monitoring, Real User Monitoring (RUM), centralized logging, and distributed tracing.
    • Define dashboards, alerting standards, operational metrics, and reporting.
  • Domain Portfolio & DNS Management:
    • Lead governance of the company's global domain portfolio.
    • Manage registrations, renewals, DNS services, certificate lifecycle management, and related vendor relationships.
    • Ensure domain-related services remain secure, compliant, and highly available.
  • Cloud Platform Engineering, Automation & Edge Services:
    • Provide technical leadership for AWS infrastructure, Kubernetes/EKS, CDN, DNS, SSL/TLS, load balancing, Web Application Firewalls (WAF), and edge security services.
    • Design and support Infrastructure-as-Code solutions using Terraform.
    • Establish standards for cloud provisioning, automation, and environment consistency.
  • Infrastructure as Code & Terraform:
    • Design, maintain, and optimize Terraform modules and deployment pipelines.
    • Promote automated provisioning, version control, testing, and infrastructure governance.
    • Drive reduction of manual deployment activities through automation.
  • Web Application Security & WAF Administration:
    • Administer and optimize AWS WAF and comparable WAF technologies.
    • Manage WAF rules, rate limiting, bot protection, IP reputation controls, and application-layer threat mitigation.
    • Partner with Information Security to improve web application protection capabilities.
  • Modern Web Application Support & Troubleshooting:
    • Serve as a senior escalation point for complex production issues.
    • Lead troubleshooting across client-side and server-side technologies.
    • Diagnose issues involving browser behavior, APIs, DNS, CDN, WAF, load balancing, networking, cloud infrastructure, and application performance.
    • Drive reliability, resiliency, and end-user experience improvements.
  • Global Collaboration:
    • Work closely with engineering teams across North America, Europe, and Asia Pacific.
    • Participate in technical reviews, architecture discussions, operational planning, and knowledge sharing.
    • Help evolve a follow-the-sun operating model.
  • On-Call & Operational Support Responsibilities:
    • Participate in scheduled on-call rotations supporting critical platforms and services.
    • Provide leadership during major incidents and after-hours escalations.
    • Support maintenance, upgrades, deployments, and disaster recovery activities.
Required Qualifications
  • 8+ years supporting enterprise applications, cloud platforms, infrastructure services, or web technologies.
  • 3+ years serving as a Technical Lead, Senior Engineer, Architect, or equivalent.
  • Strong experience with AWS, Terraform, Kubernetes/EKS, DNS, CDN, AWS WAF, SSL/TLS, observability platforms, CI/CD, and SRE practices.
  • Experience troubleshooting large-scale customer-facing web applications.
Preferred Qualifications
  • Experience managing global domain portfolios.
  • Experience with enterprise observability platforms and global support models.
  • Experience developing enterprise Infrastructure-as-Code frameworks and reusable Terraform modules.
  • Experience leveraging AI-assisted development tools and coding agents such as GitHub Copilot, Microsoft Copilot, Claude Code, Cursor, Amazon Q Developer, or similar technologies to accelerate software delivery, infrastructure automation, troubleshooting, and operational efficiency.
  • AWS, Terraform, Kubernetes, SRE, networking, security, or cloud certifications.
Benefits
  • 401k Retirement Savings Plan administered by Merrill Lynch
  • Commuter Check Pretax Commuter Benefits
  • Eligibility to purchase Medical, Dental & Vision Insurance through ECLARO

Equal Opportunity Employer: ECLARO values diversity and does not discriminate based on Race, Color, Religion, Sex, Sexual Orientation, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status, in compliance with all applicable laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Platform SRE Engineer #11145
Cloud Platform SRE Engineer #11145

ECCO Select • Dallas (TX)

Hybrid
USD 110,000 - 150,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
SRE and Platform Engineer
SRE and Platform Engineer

YOH Services LLC • Addison (TX)

Hybrid
USD 77,000 - 110,000
Medical, Prescription, Dental & Vision
Health Savings Account (HSA)
Life & Disability Insurance
+2
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Charles Schwab Corporation • Southlake (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Castleton Commodities International • Stamford (CT)

On-site
USD 120,000 - 150,000
Comprehensive medical and dental benefits
Tuition assistance and reimbursement
Employee wellness programs
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Empower • United States

Hybrid
USD 114,000 - 165,300
Medical, dental, vision and life
401(k) with company match
Tuition reimbursement
+3
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Cloud Analytics Technologies, LLC • Atlanta (GA)

On-site
USD 140,000 - 210,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Oaks (PA)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare coverage
401(k) matching
Tuition reimbursement
+1
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

Verint Systems, Inc. • Frankfort (KY)

Hybrid
USD 65,000 - 90,000