Senior AI-HPC Cluster Engineer

3 weeks ago


Austin, United States NVIDIA Full time

NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can tackle, and that matter to the world. This is our life’s work, to amplify human imagination and intelligence. Make the choice to join us today

As a member of the GPU AI/HPC Infrastructure team, you will provide leadership in the design and implementation of ground breaking GPU compute clusters that run demanding deep learning, high performance computing, and computationally intensive workloads. We seek an expert to identify architectural changes and/or completely new approaches for our GPU Compute Clusters. As an expert, you will help us with the strategic challenges we encounter including: compute, networking, and storage design for large scale, high performance workloads, effective resource utilization in a heterogeneous compute environment, evolving our private/public cloud strategy, capacity modeling, and growth planning across our global computing environment.

What you'll be doing:
  • Building and improving our ecosystem around GPU-accelerated computing including developing large scale automation solutions

  • Maintaining and building deep learning clusters at scale

  • Supporting our researchers to run their flows on our clusters including performance analysis and optimizations of deep learning workflows

  • Root cause analysis and suggest corrective action for problems large and small scales

  • Finding and fixing problems before they occur

What we need to see:
  • Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience.

  • Minimum 5 years of experience designing and operating large scale compute infrastructure.

  • Experience analyzing and tuning performance for a variety of AI/HPC workloads.

  • Working knowledge of cluster configuration managements tools such as Ansible, Puppet, Salt.

  • Experience with AI/HPC advanced job schedulers, and ideally familiarity with schedulers such as Slurm, K8s, RTDA or LSF

  • In depth understating of container technologies like Docker, Singularity, Shifter, Charliecloud

  • Proficient in Centos/RHEL and/or Ubuntu Linux distros including Python programming and bash scripting

  • Experience with AI/HPC workflows that use MPI

Ways to stand out from the crowd:
  • Experience with NVIDIA GPUs, Cuda Programming, NCCL and MLPerf benchmarking

  • Experience with Machine Learning and Deep Learning concepts, algorithms and models

  • Familiarity with InfiniBand with IBOP and RDMA

  • Understanding of fast, distributed storage systems like Lustre and GPFS for AI/HPC workloads

  • Familiarity with deep learning frameworks like PyTorch and TensorFlow

NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you.

The base salary range is 148,000 USD - 339,250 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

  • Austin, United States NVIDIA Full time

    NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by...


  • Austin, United States Advanced Micro Devices , Inc. Full time

    Overview: WHAT YOU DO AT AMD CHANGES EVERYTHING We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences the building blocks for the data center, artificial intelligence, PCs, gaming and embedded....

  • HPC System Admin

    2 weeks ago


    Austin, United States NR Consulting Full time

    Job Title: HPC System Admin Work Location: Austin, TX Position Type: Contract with possible extension Duration: 12 + Months Job Description: Project Details: Responsible for architecting and implementing Linux High Performance Computing (HPC) clusters. Performs system architecture duties on a Linux High performance computing (HPC) cluster including cluster...

  • HPC Engineer

    2 days ago


    Austin, United States Optiver Full time

    Optiver is seeking a Research Infrastructure Engineer to contribute significantly to the development and management of our research infrastructure across both on-premises and cloud platforms. This role involves hands-on work in scaling and supporting high-performance computing (HPC) and storage systems, which are critical for our growing demand in research...


  • Austin, United States Advanced Micro Devices , Inc. Full time

    WHAT YOU DO AT AMD CHANGES EVERYTHING We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences - the building blocks for the data center, artificial intelligence, PCs, gaming and embedded. Underpinning our...


  • Austin, United States Advanced Micro Devices , Inc. Full time

    WHAT YOU DO AT AMD CHANGES EVERYTHING We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences - the building blocks for the data center, artificial intelligence, PCs, gaming and embedded. Underpinning our...


  • Austin, United States Sustainable Talent Full time

    Job DescriptionJob DescriptionAre you ready to make your mark in the forefront of technological innovation? As an HPC Cluster Engineer, you'll play a pivotal role in shaping the future of AI, deep learning, and machine learning initiatives. Join us and leverage Nvidia's cutting-edge GPU technology to drive groundbreaking discoveries and revolutionize...


  • Austin, United States ShiftCode Analytics Full time

    Visa: USC, GC or GC-EAD Duration: 9 months with potential extension Location: Onsite in Austin, TX They'll give preference to someone who is currently local to Austin and then will consider people willing to relocate. Requirements: -Experience with HPC Systems environments and Infrastructures technologies and workloads -B.S. in CS/CE/EE, or at...


  • Austin, United States AMD Full time

    WHAT YOU DO AT AMD CHANGES EVERYTHING We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences – the building blocks for the data center, artificial intelligence, PCs, gaming and embedded. Underpinning our...


  • Austin, United States SailPoint Technologies Full time

    JOB DUTIES: Maintain production build tools to automate the build, test and release process for Continuous Integration / Continuous Delivery as well as development, maintenance and support of the continuous build infrastructure. Configure and maintain production build files for deployment, deploy and maintain scalable infrastructure for our SaaS Identity...

  • HPC Network Engineer

    1 month ago


    Austin, United States Algo Capital Group Full time

    HPC Network EngineerStanding at the forefront of industry excellence, this Algorithmic Trading firm is renowned for its cutting-edge technology and innovative approach to financial markets. Leveraging sophisticated algorithms and high-performance computing systems, they execute trades with precision and efficiency; pioneering new strategies and techniques to...


  • Austin, United States Algo Capital Group Full time

    HPC Network EngineerStanding at the forefront of industry excellence, this Algorithmic Trading firm is renowned for its cutting-edge technology and innovative approach to financial markets. Leveraging sophisticated algorithms and high-performance computing systems, they execute trades with precision and efficiency; pioneering new strategies and techniques to...

  • HPC Network Engineer

    1 month ago


    Austin, United States Algo Capital Group Full time

    HPC Network EngineerStanding at the forefront of industry excellence, this Algorithmic Trading firm is renowned for its cutting-edge technology and innovative approach to financial markets. Leveraging sophisticated algorithms and high-performance computing systems, they execute trades with precision and efficiency; pioneering new strategies and techniques to...


  • Austin, United States Advanced Micro Devices , Inc. Full time

    Overview: WHAT YOU DO AT AMD CHANGES EVERYTHING We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences the building blocks for the data center, artificial intelligence, PCs, gaming and embedded....


  • Austin, United States NXP Semiconductors Full time

    HPC DevOps EngineerAustin, US (Hybrid)This is what you will do as HPC DevOps engineer at NXPYou are expected to work very closely with your global colleagues within R&D IT and help deliver the HPC services (High Performance Computing and Virtual Desktop Infrastructure) to our engineering and R&D customers. Your AMEC team has operational responsibility for...


  • Austin, United States NXP Semiconductors Full time

    HPC DevOps EngineerAustin, US (Hybrid)This is what you will do as HPC DevOps engineer at NXPYou are expected to work very closely with your global colleagues within R&D IT and help deliver the HPC services (High Performance Computing and Virtual Desktop Infrastructure) to our engineering and R&D customers. Your AMEC team has operational responsibility for...


  • Austin, United States NOVO Full time

    Are you an experienced software engineer looking for your next challenge? Want to join a team of accomplished A-players building technology to transform how law is practiced? Are you excited by the explosion of AI and want to dive head-first into the frontier? At Novo, our mission is to make legal representation more accessible for everyone. That’s why...


  • Austin, United States Cluster Full time

    Company: Building autonomous unmanned aerial vehiclesOpportunity: Our client is looking for a Senior Aerospace Engineer to lead their team of hardware and aerospace engineers in the design, development and testing of hardware systems. Requirements: You have 8+ years of experience working preferably in the aerospace industry or relevant...


  • Austin, United States Cluster Full time

    Company: Building autonomous unmanned aerial vehicles Opportunity: Our client is looking for a Senior Aerospace Engineer to lead their team of hardware and aerospace engineers in the design, development and testing of hardware systems. Requirements: You have 8+ years of experience working preferably in the aerospace industry or relevant environment....


  • Austin, United States NVIDIA Full time

    We are looking for experienced Systems SW Compiler Engineers for an exciting role in our PTX (Parallel Thread Execution) Compiler Development team. Join the PTX Compiler team and help drive the PTX compiler evolution. PTX enables all GPU Computing applications including HPC, Deep Learning and Autonomous Driving. PTX provides a stable programming model and...