Reliability, Availability and Serviceability Expert, Datacenter AI Products Development

1 month ago


Santa Clara, United States NVIDIA Full time
For two decades, we have pioneered visual computing, the art and science of computer graphics - with our invention of the GPUs, the engine of modern AI technologies, the field has expanded to encompass AI-powered video games, social networking and web search, IC & other product design, medical diagnosis, and scientific research. Today, visual computing is the critical computing engine for deep learning-based AI including ChatGPT, becoming increasingly central to how people entertain and interact, and there has never been a more exciting time to join us to enable visual computing and AI to the next chapter. We are looking for one product development engineer as a SME to drive key aspects of RAS/Resilience features from Chip to module to server for our next-generation products for AI Applications. We are expecting you to bring deep knowledge and experience in RAS/Resilience testing, characterization, analysis, benchmarking, and risk assessment of large AI training or HPC cluster systems with InfiniBand or enhanced Ethernet.

What you’ll be doing:

The focal point SME for manufacturing test requirements, test methodology, test plan and test flow for AI system RAS/Resilience features to ensure good test coverage and successful production ramp-ups.

Own the AI system RAS/Resilience models, Benchmarking and Risk assessment.

Own the troubleshooting and root-causing of AI system RAS/Resilience related failures at factory and in the field.

Drive the end-to-end RAS efforts of chip-board-system to reduce FIT rates.

Lead the data analysis of RAS/Resilience logs to refine, revise and overhaul test methodology and manufacturing flows; influence and drive software tools/infrastructure required for new product development, validation, and productization.

Opportunity to work closely and partner with architecture, hardware, software, and product engineering teams through the product development lifecycle.

Be ready to be challenged to assess new hardware features and architect manufacturing RAS tests, flows, methodologies.

You'll nurture a deep understanding of NVIDIA's AI hardware and software architecture.

What we need to see:

BS or higher in EE, CE, CS, Mathematics, or equivalent experience.

12+ years proven hands-on experiences in design, testing, benchmarking, and risk assessment of system RAS / Resiliency features of large Compute or AI or HPC systems.

Proficient in Compute System RAS/Resilience model theory and methodology.

Proficient in HPC or AI system architecture and Cluster Interconnect technologies.

Proficient in using test equipment, Linux commands and benchmark utilities to test and trouble-shoot compute system RAS & Resiliency features.

Strong problem-solving and trouble-shooting expertise; and institutionalizing root-cause analysis.

Self-initiative, strong interpersonal skills, and flexibility to adapt to new technologies.

Solid Knowledge and/or Experience in HPC or MLPerf benchmarking is a plus.

NVIDIA is widely considered to be one of the technology world’s most desirable employers We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you

The base salary range is 188,000 USD - 356,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

  • Santa Clara, United States NVIDIA Full time

    For two decades, we have pioneered visual computing, the art and science of computer graphics - with our invention of the GPUs, the engine of modern AI technologies, the field has expanded to encompass AI-powered video games, social networking and web search, IC & other product design, medical diagnosis, and scientific research. Today, visual computing is...


  • Santa Clara, United States NVIDIA Full time

    For two decades, we have pioneered visual computing, the art and science of computer graphics - with our invention of the GPUs, the engine of modern AI technologies, the field has expanded to encompass AI-powered video games, social networking and web search, IC & other product design, medical diagnosis, and scientific research. Today, visual computing is...


  • Santa Clara, United States NVIDIA Full time

    For two decades, we have pioneered visual computing, the art and science of computer graphics - with our invention of the GPUs, the engine of modern AI technologies, the field has expanded to encompass AI-powered video games, social networking and web search, IC & other product design, medical diagnosis, and scientific research. Today, visual computing is...

  • Product Manager

    4 days ago


    Santa Clara, United States Astera Labs Full time

    Astera Labs is a global leader in purpose-built connectivity solutions that unlock the full potential of cloud and AI infrastructure. Our Intelligent Connectivity Platform integrates PCIe, CXL and Ethernet semiconductor-based solutions based on a software-defined architecture that is both scalable and customizable. Inspired by trusted partnerships with...


  • Santa Clara, United States NVIDIA Full time

    We are now looking for a Sr. Product Development Engineer - DataCenter! NVIDIA Corporation is a world leader in visual computing technology. The GPU, which the company invented, serves as the visual cortex of modern computers and is at the heart of our products and services. NVIDIA has transformed into a specialized platform company that targets four large...


  • Santa Clara, United States Resemble AI Full time

    About The Company Resemble AI is at the forefront of voice AI technology, pioneering in the field of Custom Voices. Our advanced Deep Learning models enable us to produce highly realistic Speech Synthesis, marking a significant milestone in AI evolution. At Resemble AI, we are committed to pushing the boundaries of what's possible in AI-driven speech...

  • AI Developer

    2 weeks ago


    Santa Clara, United States US Main Full time

    Job Title: AI Developer Location: San Francisco Bay Area, USA AKube.ai (www.akube.ai) is a Silicon Valley based AI startup looking to revolutionize Software Engineering using AI beyond LLM’s. The company has been founded by seasoned tech entrepreneurs and has already built the first version of its product and is starting customer pilots. We are...


  • Santa Clara, United States Global AI Platform Corporation Full time

    Product Marketing ManagerLocation: SF Bay Area - Santa ClaraAbout GAP:Founded in July 2023, Global AI Platform Corporation emerges as a key player in the global AI arena. With headquarters in Santa Clara, California, and Pangyo, South Korea, our mission revolves around pioneering advancements in AI technology. We're developing a “Personal AI Assistant”...

  • AI Developer

    1 week ago


    Santa Clara, CA, United States US Main Full time

    Job Title: AI Developer Location: San Francisco Bay Area, USA AKube.ai ( is a Silicon Valley based AI startup looking to revolutionize Software Engineering using AI beyond LLM’s. The company has been founded by seasoned tech entrepreneurs and has already built the first version of its product and is starting customer pilots. We are seeking...

  • Datacenter Technician

    3 weeks ago


    Santa Clara, United States TCWGlobal Full time

    Datacenter Technician Santa Clara, CA 95050 (*local candidates- onsite)$65-75hr (Weekly pay + benefits)6 month contact (Excellent potential for extension or permanent)Full-time: M-F onsite Our client is a rapidly growing and innovative tech company. redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an...

  • Datacenter Technician

    3 weeks ago


    Santa Clara, United States TCWGlobal Full time

    Datacenter Technician Santa Clara, CA 95050 (*local candidates- onsite)$65-75hr (Weekly pay + benefits)6 month contact (Excellent potential for extension or permanent)Full-time: M-F onsite Our client is a rapidly growing and innovative tech company. redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an...

  • Datacenter Technician

    3 weeks ago


    Santa Clara, United States TCWGlobal Full time

    Datacenter Technician Santa Clara, CA 95050 (*local candidates- onsite)$65-75hr (Weekly pay + benefits)6 month contact (Excellent potential for extension or permanent)Full-time: M-F onsite Our client is a rapidly growing and innovative tech company. redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an...


  • Santa Clara, United States NVIDIA Full time

    Technical Product Marketing Manager – AI and Game Development page is loaded Technical Product Marketing Manager – AI and Game Development Apply locations US, CA, Santa Clara time type Full time posted on Posted 3 Days Ago job requisition id JR1976790 The ACE for Gaming Product Management team is looking for a world class technical marketing manager to...

  • Datacenter Technician

    3 weeks ago


    Santa Clara, United States TCWGlobal Full time

    Job DescriptionJob DescriptionSystem AdministratorTechnician- Datacenter Santa Clara, CA 95050 (*local candidates- onsite)$65-75hr (Weekly pay + benefits)6 month contact (Excellent potential for extension or permanent)Full-time: M-F onsiteOur client is a rapidly growing and innovative tech company. redefining computer graphics, PC gaming, and accelerated...

  • Datacenter Technician

    3 weeks ago


    Santa Clara, United States TCWGlobal Full time

    Job DescriptionJob DescriptionSystem AdministratorTechnician- Datacenter Santa Clara, CA 95050 (*local candidates- onsite)$65-75hr (Weekly pay + benefits)6 month contact (Excellent potential for extension or permanent)Full-time: M-F onsiteOur client is a rapidly growing and innovative tech company. redefining computer graphics, PC gaming, and accelerated...


  • Santa Clara, United States Beth Page tech Full time

    Job DescriptionJob DescriptionRole: AI Product Marketing Manager - Full TimeLocation: Santa Clara, CARole Overview: The Product Marketing Manager is responsible for developing and implementing marketingstrategies to promote an innovative AI-powered product. This role requires a deep understandingof market trends and technological advancements, as well as the...


  • Santa Clara, United States Amazon Development Center U.S., Inc. Full time

    PhD, or Master's degree and 4+ years of CS, CE, ML or related field experience - Experience in patents or publications at top-tier peer-reviewed conferences or journals - Experience programming in Java, C++, Python or related language - Experience with neural deep learning methods and machine learning Machine learning (ML) has been strategic to Amazon from...


  • Santa Clara, United States Zyxware Technologies Full time

    Role: Product Marketing Manager Location: Santa Clara, CA Duration: Full Time Role Overview: The Product Marketing Manager is responsible for developing and implementing marketing strategies to promote an innovative AI-powered product. This role requires a deep understanding of market trends and technological advancements, as well as the ability to create...


  • Santa Clara, United States Amazon Development Center U.S., Inc. - B02 Full time

    Our team's mission is to help customers manage their network security complexity with ease. We're looking for an experienced engineer to help us build a new, greenfield product that radically improves on this mission by leveraging GenAI to help customers meet their security needs. This project is a novel application of AI that goes beyond current industry...


  • Santa Clara, CA, United States Celestial AI Full time

    About Celestial AIAs the industry strives to meet the demands of the AI workloads, bottlenecks in data transfers between processors and memory have hindered progress. The Photonic Fabric based Memory Fabric provides an optically scalable solution to the ‘Memory Wall’ problem, enabling tens of Terabytes of memory capacity at full HBM bandwidths with low...