Staff Engineer, Memory Systems Architecture

Conductor

San Jose (CA)

On-site

USD 163,000 - 253,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology firm is looking for a Staff Engineer in Memory Systems Architecture located in San Jose, California. This position involves mitigating DRAM failure rates, developing RAS algorithms, and collaborating with customers to enhance fault management. Ideal candidates should possess a Bachelor's degree with at least 10 years of experience in hardware fault management or related fields. The role emphasizes an innovative mindset and offers competitive compensation in the range of $163,000 to $253,000 USD per year.

Qualifications

  • 10+ years of relevant industry experience in hardware fault management.
  • Understanding of DRAM and HBM failure modes.

Responsibilities

  • Recommend solutions to mitigate field DRAM failure rate.
  • Establish value add of in-field fault management architecture.
  • Develop platform RAS algorithms for memory fault management.

Skills

Hardware fault management
Data center fleet management
Reliability Availability Serviceability (RAS)
ECC design and verification
Collaborative mindset

Education

Bachelor's degree with 10+ years experience or Master's with 8+ years or PhD with 5+ years

Job description

Staff Engineer, Memory Systems Architecture

San Jose, California, United States

Samsung Semiconductor is hiring now for a Staff Engineer, Memory Systems Architecture. The conventional DRAM failure analysis was physical electrical FA and physical FA. But, in the era of Data center, it is easier to track the field failure information. With this data set, Fault management team’s role is finding DRAM failure mode, abnormality and failure rate projection.

You will be part of an incubation team working on in-field telemetry intended to transform the Customer Quality Experience for Samsung memory products. Fault Management is the future of quality to minimize system downtime within AI/ML hardware deployments and workloads of the future. We analyze trends and patterns from enormous memory fleet telemetry to bucketize failures and perform virtual root-cause analysis. Telemetry analysis helps us design solutions to proactively avoid system downtime. We conduct research and develop both in-house and collaboratively in the industry with the opportunity to publish our findings through whitepapers and conferences. We are looking for innovative and passionate thinkers who can work in a start‑up environment and are excited to shape the future of data centers around the world. Join us in our mission!

What You'll Do
  • Based on the knowledge of SOC controller and memory operation including RAS feature, find and recommend better solution to mitigate the field DRAM failure rate.
  • Communicate better ECC scheme to customers based on Samsung DRAM failure mode (DQ and burst).
  • Interface with customers to establish the value add of enabling in-field fault management architecture.
  • Contribute to the standardization of DRAM/HBM failure logging in the OCP.
  • Propose and develop platform RAS (Reliability Availability Serviceability) algorithms for memory fault management such as page offlining, hPPR and conduct POC with known failure DIMMs in the real server and application.

Location: Daily onsite presence at our San Jose headquarters in alignment with our Flexible Work Policy.

Job ID: 42886

What You Bring
  • Bachelor’s degree with 10+ years of relevant industry experience, or Master’s with 8+ years or PhD with 5+ years hardware fault management, reliability, data center fleet management experience or related technical field preferred. (Must)
  • Knowledge of platform memory subsystem, platform RAS (Reliability Availability Serviceability) such as ECC, page offlining, hPPR and hardware sparing.
  • ECC design and verification and reverse engineering experience.
  • Understanding of the address mapping between CPU and memory.
  • Memory controller register modification.
  • DRAM and HBM failure mode understanding.
  • Inclusive mindset, adapting style to situations and diverse global norms.
  • Avid learner, approaches challenges with curiosity and resilience, seeks data to help build understanding.
  • Collaborative, builds relationships, offers support and welcomes approaches.
  • Innovative and creative, proactively explores new ideas and adapts quickly to change.

Base Pay Range: $163,000 - $253,000 USD

Equal Opportunity Employment Policy

Samsung Semiconductor takes pride in being an equal opportunity workplace dedicated to fostering an environment where all individuals feel valued and empowered to excel, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status.

When selecting team members, we prioritize talent and qualities such as humility, kindness, and dedication. We extend comprehensive accommodations throughout our recruiting processes for candidates with disabilities, long-term conditions, neurodivergent individuals, or those requiring pregnancy‑related support. All candidates scheduled for an interview will receive guidance on requesting accommodations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, Memory Systems Architecture
Staff Engineer, Memory Systems Architecture

Samsungsemiconductor • San Jose (CA)

On-site
USD 163,000 - 253,000
4+ weeks paid time off
Medical/Dental/Vision/401k
Flexible Work Policy
+2
Staff Engineer, Memory Systems Architecture
Staff Engineer, Memory Systems Architecture

Samsung Semiconductor • San Jose (CA)

On-site
USD 163,000 - 253,000
Incentive opportunities
Medical/Dental/Vision
401(k)
+2
Staff Engineer, Memory Systems Architecture
Staff Engineer, Memory Systems Architecture

Samsung Semiconductor Inc. • San Jose (CA)

On-site
USD 163,000 - 253,000
Charitable giving match
Paid time off 4+ weeks
Fertility/adoption stipend
+3
Senior Engineer, DRAM Applications
Senior Engineer, DRAM Applications

Socket.dev • San Jose (CA)

On-site
USD 138,000 - 206,000
Medical, Dental, Vision benefits
401(k) plan
Paid time off and holidays
+1
Senior Engineer, DRAM Applications
Senior Engineer, DRAM Applications

Samsung Semiconductor Inc. • San Jose (CA)

On-site
USD 138,000 - 206,000
4+ weeks of paid time off
Inclusive rewards plan
Flexible work environment
+1
Technical Account Manager, DRAM Business Enablement
Technical Account Manager, DRAM Business Enablement

Conductor • San Jose (CA)

On-site
USD 163,000 - 253,000
4+ weeks of paid time off
Flexible working environment
Support for fertility care or adoption
+1
Technical Account Manager, DRAM Business Enablement
Technical Account Manager, DRAM Business Enablement

Samsungsemiconductor • San Jose (CA)

On-site
USD 163,000 - 253,000
4+ weeks Paid Time Off
Flexible Work Environment
Comprehensive benefits including 401(k)
Senior Engineer, DRAM Applications
Senior Engineer, DRAM Applications

Samsungsemiconductor • San Jose (CA)

On-site
USD 138,000 - 206,000
Paid time off
Health benefits
Flexible work environment
DRAM Applications Engineer - SI/PI & Validation Lead
DRAM Applications Engineer - SI/PI & Validation Lead

Conductor • San Jose (CA)

On-site
Senior Engineer, DRAM Applications
Senior Engineer, DRAM Applications

Samsung Semiconductor • San Jose (CA)

On-site
USD 138,000 - 206,000