Systems Engineering Associate - GovCloud [Salesforce National Security]

Salesforce

Virginia (MN)

On-site

USD 140,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Salesforce is seeking an engineering candidate to join the Site Reliability organization. Working closely with counterparts in the Infrastructure and R&D organizations, this team monitors cloud service availability and repairs issues quickly to protect customers.

You will detect and resolve incidents within minutes by supervising services, reacting to problems, and proactively addressing issues before they affect performance or availability.

Qualifications

  • Active TS/SCI with polygraph clearance required
  • Related technical degree required
  • Systems engineering experience in enterprise scale internet service engineering or support role
  • Expertise in TCP/IP related technologies
  • Expertise in CLI enterprise support of Unix variants (Linux/Solaris/BSD) with strong Linux/UNIX knowledge (RHEL/Solaris)
  • Solid understanding of monitoring security systems and administration
  • Good interpersonal skills (Written and Oral)
  • Experience in Incident Management and ITIL service operations
  • Experience in working in a 24/7 team managing large data centers
  • Be available to work shift work if required (1400 - 2200 x 4 days a week)
  • Experience provisioning, operating, and running AWS/C2S based infrastructure and systems
  • Understand and have experience with writing scripts in Python, Go, or other languages

Responsibilities

  • Keep customer-facing services available at top performance by maintaining the health of the supporting systems
  • Incident management: act in key support roles during major incidents (Sev0/Sev1) and participate in RCAs
  • Problem management: participate in RCAs and hand them off to Global Solutions
  • Ensure work aligns with internal compliance policy and directives
  • Collaborate with other staff to solve technical issues and drive improvements
  • Lead and mentor team members to stay on top of industry innovation and technology
  • Operate effectively in a fast-paced environment and balance multiple priorities
  • Automate detection and resolution of recurring production issues
  • Help create and improve processes to reduce operations and engineering toil

Skills

US Citizenship
TS/SCI with polygraph
System engineering
TCP/IP
Linux/UNIX
Incident Management
ITIL
Scripting (Python/Go)
AWS
24/7 operations
Security monitoring

Education

Related technical degree

Tools

AWS
Kubernetes
Jenkins
Spinnaker
Puppet
Chef

Job description

Job Category

Software Engineering

Job Details
About Salesforce

Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn't a buzzword - it's a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.

Ready to level-up your career at the company leading workforce transformation in the agentic era? You're in the right place! Agentforce is the future of AI, and you are the future of Salesforce.

PLEASE NOTE: Qualification for this job is contingent upon acceptable results from a background investigation as well as your obtaining and maintaining a TS/SCI with Polygraph is required for this role.

Salesforce is seeking an engineering candidate to join the Site Reliability organization. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a team of engineers monitoring cloud service availability and ready to swiftly repair any service-impacting issues. Seven days a week, 24 hours a day, the Site Reliability team keeps the Salesforce cloud and our customers protected. As a member of the Site Reliability team, you will be responsible for the primary task of detecting and resolving incidents within minutes. This objective is met by supervising the services, reacting to problems, and proactively addressing issues before they affect performance or availability.

The team contributes to the customer and Salesforce by securing data through monitoring, automation, self-healing and resiliency initiatives, destructive testing, and game day exercises. The incumbent in this role would demonstrate a solid focus on tactical operations, as well as large-scale production engineering and orchestration.

Your Impact
  • Keep the customer-facing services available at top performance by maintaining the constant health of the supporting systems.
  • Incident management - Act in key support roles during major incidents e.g. Sev0, Sev1. Also, participate in the technical review of the incident for problem management
  • Problem Management - populate and participate in RCAs and hand them off to the Global Solutions team
  • Ensuring that work carried out by the Site Reliability team is performed in such a way as to stay in sync with the company's internal compliance policy and directives
  • Passionate about solving technical issues and customer concerns with other technical staff as required.
  • Work with and lead other members of the team in staying on top of key industry innovation and technology, and assist in team development growth
  • Ability to operate in the fast paced environment and solve sophisticated issues quickly successfully balance multiple priorities
  • Work to automate detection and resolution of recurring issues in the production environment
  • Help create and improve current processes to reduce operations and engineering toil
Required Skills
  • U.S. citizen (U.S. born or naturalized) who does not hold dual citizenship. You must have an active TS/SCI with polygraph position with the U.S. federal government or other clearances as deemed appropriate for the role
  • A related technical degree required
  • Systems engineering experience in enterprise scale internet service engineering or support role
  • Expertise in TCP/IP related technologies (networking protocols, network programming, etc.)
  • Expertise in CLI enterprise support of Unix variants (Linux/Solaris/BSD) as well as strong Linux/UNIX knowledge with significant exposure to Red Hat Enterprise Linux and Solaris
  • Solid understanding of monitoring security systems and administration
  • Good interpersonal skills (Written and Oral)
  • Past experience in Incident Management and good understanding of ITIL service operations
  • Experience in working in a 24/7 team managing large data centers
  • Be available to work shift work if required (1400 - 2200 x 4 days a week).
  • Experience provisioning, operating, and running AWS/C2S based infrastructure and systems
  • Understand and have experience with writing scripts in Python, Go, or other languages
Preferred Skills
  • Prior Chef/Puppet or automated deployment experience
  • Prior Jenkins/Bamboo/Spinnaker pipeline execution experience
  • Experience in supporting and maintaining a monitoring and alert systems
  • Experience in supporting and maintaining Java applications
  • Hands on experience configuring and running AWS (Amazon Web Services), using the CLI/SDKs
  • Experience managing systems monitoring and alerts.
  • Certifications in Linux+, RedHat and AWS
  • Experience in supporting and leading Kubernetes based applications and services
  • Familiar with Agile Process and DevOps
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, GovCloud 24x7
Site Reliability Engineer, GovCloud 24x7

Engg • Denver (CO)

On-site
USD 120,000 - 180,000
Systems Engineering Associate - GovCloud [Salesforce National Security]
Systems Engineering Associate - GovCloud [Salesforce National Security]

109 Computable Insights LLC • Herndon (VA)

On-site
USD 111,000 - 122,000
Competitive salary range: $111,000 – $122,000 annually
Time-off programs
Medical, dental, vision benefits
+1
Systems Engineering Associate - GovCloud [Salesforce National Security]
Systems Engineering Associate - GovCloud [Salesforce National Security]

salesforce.com, inc. • Herndon (VA)

On-site
USD 111,000 - 122,000
Medical, Dental, Vision insurance
Mental health support
401(k)
+2
Site Reliability Engineer, GovCloud 24x7
Site Reliability Engineer, GovCloud 24x7

salesforce.com, inc. • Burlington (MA)

On-site
USD 114,000 - 125,000
Time off programs
Medical insurance
Dental insurance
+6
Computer Security Incident Report Analyst with TS/SCI Clearance [Salesforce National Security]
Computer Security Incident Report Analyst with TS/SCI Clearance [Salesforce National Security]

Salesforce • Virginia (MN)

On-site
USD 120,000 - 150,000
Site Reliability Engineer, GovCloud 24x7
Site Reliability Engineer, GovCloud 24x7

salesforce.com, inc. • McLean (VA)

On-site
USD 114,000 - 125,000
Health insurance
401(k)
Employee stock purchase program
+1
Site Reliability Engineer, GovCloud 24x7
Site Reliability Engineer, GovCloud 24x7

salesforce.com, inc. • Denver (CO)

On-site
USD 114,000 - 125,000
Medical
Dental
Vision
+3
Technical Support Engineer - Sales Cloud
Technical Support Engineer - Sales Cloud

Salesforce • United States

On-site
USD 90,000 - 130,000
Site Reliability Engineer, GovCloud 24x7
Site Reliability Engineer, GovCloud 24x7

Salesforce • McLean (VA)

On-site
USD 114,000 - 125,000
Site Reliability Operations Engineer
Site Reliability Operations Engineer

salesforce.com, inc. • Seattle (WA)

On-site
USD 94,000 - 142,000