We have other current jobs related to this field that you can find below


  • Santa Clara, California, United States ServiceNow Full time

    Company OverviewAt ServiceNow, we harness technology to create a better world for everyone, driven by our talented workforce. We prioritize speed and innovation to meet the demands of our customers and communities.Joining ServiceNow means becoming part of a dynamic team of innovators who possess a relentless curiosity and a commitment to creativity.We...


  • Santa Clara, California, United States ServiceNow Full time

    Company OverviewAt ServiceNow, we harness technology to enhance global operations, and our dedicated workforce makes it all possible. We operate swiftly because the world demands it, innovating uniquely for our clients and communities.By becoming part of ServiceNow, you join a dynamic team of innovators who possess a relentless curiosity and a passion for...


  • Santa Clara, United States Nvidia Full time

    Senior Site Reliability Engineer - StoragelocationsUS, CA, Santa Claratime typeFull timejob requisition idJR1979072NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and...


  • Santa Clara, United States Veear Full time

    Position: Site Reliability Engineer Location: Remote role Duration: 12+ Months Contract with possible extension Job Description: We seek development-heavy Site Reliability Engineers to design, build, maintain, and scale production services and server farms within our FedRAMP SASE product portfolio. We want passionate engineers who bring new ideas to all...


  • Santa Clara, California, United States Nvidia Full time

    NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables unique creativity and discovery, and powers what were...


  • Santa Clara, United States VeeAR Projects Inc. Full time

    Position: Site Reliability EngineerLocation: Remote roleDuration: 12+ Months Contract with possible extensionJob Description: We seek development-heavy Site Reliability Engineers to design, build, maintain, and scale production services and server farms within our FedRAMP SASE product portfolio. We want passionate engineers who bring new ideas to all facets...


  • Santa Clara, United States VeeAR Projects Inc. Full time

    Position: Site Reliability EngineerLocation: Remote roleDuration: 12+ Months Contract with possible extensionJob Description: We seek development-heavy Site Reliability Engineers to design, build, maintain, and scale production services and server farms within our FedRAMP SASE product portfolio. We want passionate engineers who bring new ideas to all facets...


  • Santa Clara, United States NVIDIA Full time

    NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and outstanding people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers,...


  • Santa Clara, United States NVIDIA Full time

    Senior Site Reliability Engineer, Data Science and ML Platforms Are you passionate about building and maintaining large-scale production systems that support advanced data science and machine learning applications? Do you want to join a team at the heart of NVIDIA's data-driven decision-making culture? If so, we have a great opportunity for you! NVIDIA is...


  • Santa Clara, United States Centrify Corporation Full time

    Our software runs on public clouds with 99.9% or better uptime and is mission critical for our customers. Our cloud operations team is where the rubber meets the road and needs innovative Site Reliability Engineers. Join a professional team of smart and hard-working professionals building enterprise-class cloud-based services in the rapidly growing market of...


  • Santa Clara, United States Sustainable Talent Full time

    Job DescriptionJob DescriptionJoin the Sustainable Talent team, supporting NVIDIA as a Senior Site Reliability Engineer supporting the Infrastructure, Planning, and Process organization. This is a W-2 full-time contract based in Santa Clara, CA, with Hybrid work options. We offer competitive pay $75 - $90/hr based on factors like experience, education,...


  • Santa Clara, United States NVIDIA Full time

    Senior Silicon Reliability Engineer page is loaded Senior Silicon Reliability Engineer Apply locations US, CA, Santa Clara time type Full time posted on Posted 3 Days Ago job requisition id JR1981353 NVIDIA has continuously reinvented itself over three decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern...


  • Santa Clara, United States NVIDIA Full time

    NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables unique creativity and discovery, and powers what were...


  • Santa Clara, United States DG Heating and Air Conditioning Inc Full time

    Are you passionate about building and maintaining large-scale production systems that support advanced data science and machine learning applications? Do you want to join a team at the heart of NVIDIAs data-driven decision-making culture? If so, we have a great opportunity for you! NVIDIA is seeking a Senior Site Reliability Engineer (SRE) for the Data...


  • Santa Clara, United States Kofi Group Full time

    To Apply for this Job Click HerePrincipal Site Reliability EngineerSan Francisco Bay Area, CAWe are partnering with a late-stage Cloud Security company that is looking for a Principal Level SRE The ideal candidate will have:Strong sense of architecture and design for fault tolerance, scale-out approaches, and stability Deep experience in building tools...


  • Santa Clara, California, United States Promote Project Full time

    About Promote Project: Promote Project is a leader in innovative technology solutions, dedicated to pushing the boundaries of what is possible in the realm of artificial intelligence and cloud computing. Our commitment to excellence is reflected in our talented workforce and our pursuit of groundbreaking advancements.Position Overview: We are seeking a...


  • Santa Clara, California, United States Promote Project Full time

    About the Company: Promote Project is at the forefront of innovation, leveraging cutting-edge technology to redefine the landscape of AI and computing. Our mission is to harness the power of advanced computing to create transformative solutions that impact various industries.Position Overview: We are seeking a Manager of Site Reliability Engineering to...


  • Santa Clara, United States NVIDIA Full time

    Senior System Reliability Engineer page is loaded Senior System Reliability Engineer Apply locations US, CA, Santa Clara time type Full time posted on Posted 2 Days Ago job requisition id JR1980220 NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer...


  • Santa Clara, United States NVIDIA Full time

    Senior System Reliability Engineer Locations: US, CA, Santa Clara Time Type: Full time Posted on: Posted 2 Days Ago Job Requisition ID: JR1980220 NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing —...


  • Santa Clara, California, United States Anello Full time

    About Anello Photonics:ANELLO Photonics is a leading-edge technology company based in Santa Clara, CA. The company has developed integrated photonic system-on-chip technology for next generation navigation. ANELLO's SIPHOGTM gyroscope is based on its patented photonic integrated circuit technology. The result is a product that is higher performance, much...

Senior Site Reliability Engineer

2 months ago


Santa Clara, United States Sustainable Talent Full time

Join the Sustainable Talent team, supporting NVIDIA as a Senior Site Reliability Engineer supporting the Infrastructure, Planning, and Process organization. This is a W-2 full-time contract based in Santa Clara, CA. We offer competitive pay based on factors like experience, education, location, etc. and provide full benefits, PTO, and amazing company culture

As a Senior Site Reliability Engineer, you will be part of a fast-paced crew that develops and maintains NVIDIA's sophisticated internal cloud provisioning product for GPUs and Tegra systems. The team works with various other business units within NVIDIA Software such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Driverless Cars to cater to their infrastructure & system's needs. You'll also be working in conjunction with various teams such as software engineering to deploy these new products and manage our infrastructure, associated processes and systems. Keen attention to detail, problem-solving abilities, and a solid knowledge base are essential.

What you'll be doing:

  • Working on systems deployed in NVIDIA's internal cloud making them available and reliable for our end users.
  • Monitor system performance and troubleshoot issues related to CPU, memory, disk, and network utilization.
  • Providing high quality of user support.
  • Monitoring KPIs and making sure that team's SLAs are met.
  • Managing and maintaining production Kubernetes clusters.
  • Drive automation of monitoring to gain more insight into applications and system health.
  • Craft and develop tools needed for automating workflows.
  • Develop, Improve and Maintain our infrastructure codebase.
  • Craft and implement critical metrics using various analytics methods and dashboards.
  • Take part in prototyping, crafting, and developing cloud infrastructure for NVIDIA.
  • Reuse AI techniques to extract useful signals about machines and jobs from the data generated.
What we need to see:
  • Experience of maintaining cloud infrastructure and highly-available production environment.
  • Experience managing systems installed data centers. Proficient with BMC (Redfish), KVM, and IPMI tools.
  • Working knowledge of OpenStack.
  • Background in Databases like SQL (MySQL) and timeseries DBs like Prometheus.
  • Strong knowledge of networking principles and protocols, including TCP/IP, DNS, DHCP, and VLANs.
  • Experience with data analytics/visualization tools like Kibana, Grafana, Splunk etc.
  • Strong Ansible skills. Experience with Ansible AWX.
  • Strong background with Jenkins and/or other CI/CD systems.
  • Proficient with Kubernetes, dockers & virtualization.
  • Proficient using source code management and binary repository systems like GitLab, GitHub, Artifactory, Perforce etc.
  • Knowledge of monitoring systems such as Zabbix, Prometheus, PagerDuty and/or similar systems.
  • Advanced knowledge of standard methodologies related to security.
  • 5+ years of proven experience.
  • Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent experience.
Ways to stand out from the crowd:
  • Previous experience with SRE teams managing on-prem infrastructure.
  • Experience managing NVIDIA hardware like GPUs and Tegras.
  • Thrives in a multi-tasking environment with constantly evolving priorities.
  • Prior experience with large scale operations team.
  • Experience with Windows server infrastructure.
  • Outstanding interpersonal skills and communication with all levels of management.
  • Experience with using and improving data centers.
  • Ability to analyze sophisticated problems into simple sub problems and then reuse available solutions to implement most of those.
  • Ability to design simple systems that can work efficiently without needing much support.


Sustainable Talent is a M/F+, disabled, and veteran equal employment opportunity and affirmative action employer.