Senior / Staff Site Reliability Engineer (SRE)

2 weeks ago

San Francisco, United States DevOps projects Full time

2025-10-25 Senior / Staff Site Reliability Engineer (SRE) Fluidstack is building GPU supercomputers for top AI labs, governments, and enterprises. Our customers include Mistral, Poolside, Black Forest Labs, Meta, and more. Our team is small, highly motivated, and focused on providing a world class supercomputing experience. We put out customers first in everything we do, working hard to not just win the sale, but to win repeated business and customer referrals. We hold ourselves and each other to high standards. We expect you to care deeply about the work you do, the products you build, and the experience our customers have in every interaction with us. You must work hard, take ownership from inception to delivery, and approach every problem with an open mind and a positive attitude. We value effectiveness, competence, and a growth mindset. About the Role Senior / Staff SREs at Fluidstack sit at the core of our infrastructure, working across software, hardware, and operations to ensure the reliability and performance of our global GPU cloud. They partner closely with teams including networking, platform engineering, and data center operations to build systems that scale with the demands of AI workloads. SREs are hands‑on and possess deep systems knowledge and strong communication skills. You’ll be responsible for tackling complex production issues, deploying resilient infrastructure, and continuously improving the stability and observability of our platform as we grow. A typical day may involve: Deploying clusters of 1,000+ GPUs using custom written playbooks; modifying these tools as necessary to provide the perfect solution for a customer. Validating correctness and performance of underlying compute, storage, and networking infrastructure, and working with providers to optimize these subsystems. Migrating petabytes of data from public cloud platforms to local storage, as quickly and cost effectively as possible. Debugging issues anywhere in the stack, from “this server’s fan is blocked by a plastic bag” to “optimizing S3 dataloaders from buckets in different regions”. Building internal tooling to decrease deployment time and increase cluster reliability, including automation where the customer benefits clearly outweigh the implementation overhead. This role will involve being part of an on‑call rotation up to one week per month. Focus A customer‑centric attitude, an accountability mindset, and a bias to action. A track record of shipping clean, well‑documented code in complex environments. An ability to create structure from chaos, navigate ambiguity, and adapt to the dynamic nature of the AI ecosystem. Strong technical and interpersonal communication skills, a low ego, and a positive mental attitude. Requirements An ideal candidate meets at least the following requirements: 2+ years of SRE, DevOps, Sysadmin, and/or HPC engineering experience. Great verbal and written communication skills in English. Experience deploying and operating Kubernetes and/or SLURM clusters. Experience in writing Go, Python, Bash. Experience using Ansible, Terraform, and other automation or IAC tools. Strong engineering background, preferably in Computer Science, Software Engineering, Math, Computer Engineering, or similar fields. Exceptional candidates have one or more of the following experiences: You have built and operated an AI workload at 1000+ GPU scale. You have built multi‑tenant, hyperscale Kubernetes based services. You have physically deployed infrastructure in a datacenter, managed bare metal hardware via MaaS or Netbox, etc. You have deployed and managed multi‑tenant InfiniBand or RoCE networks. You have deployed and managed petabyte scale all‑flash storage systems, including DDN, VAST, and/or Weka; or Ceph, LUSTRE, or similar open source tools. Benefits Retirement or pension plan, in line with local norms. Health, dental, and vision insurance. Generous PTO policy, in line with local norms. #J-18808-Ljbffr

Staff Site Reliability Engineer

3 weeks ago

San Francisco, United States Heartflow Full time

Heartflow is a medical technology company advancing the diagnosis and management of coronary artery disease, the #1 cause of death worldwide, using cutting-edge technology. The flagship product—an AI-driven, non-invasive cardiac test supported by the ACC/AHA Chest Pain Guidelines called the Heartflow FFRCT Analysis—provides a color‑coded, 3D model of a...
Staff Site Reliability Engineer

3 weeks ago

San Francisco, United States Heartflow Full time

Heartflow is a medical technology company advancing the diagnosis and management of coronary artery disease, the #1 cause of death worldwide, using cutting-edge technology. The flagship product—an AI-driven, non-invasive cardiac test supported by the ACC/AHA Chest Pain Guidelines called the Heartflow FFRCT Analysis—provides a color-coded, 3D model of a...
Senior Staff Site Reliability Engineer

1 week ago

San Francisco, United States WEX Full time

About the Team & RoleWe are looking for a highly motivated and high-potential Senior Staff Site Reliability Engineer (SRE) to join our team as a senior technical leader, driving transformational change and delivering significant business impact across WEX’s platform ecosystem.This is a truly exciting moment to be part of the SRE organization at WEX. Our...
Staff Site Reliability Engineer

3 weeks ago

San Francisco, CA, United States Heartflow Full time

Heartflow is a medical technology company advancing the diagnosis and management of coronary artery disease, the #1 cause of death worldwide, using cutting-edge technology. The flagship product—an AI-driven, non-invasive cardiac test supported by the ACC/AHA Chest Pain Guidelines called the Heartflow FFRCT Analysis—provides a color‑coded, 3D model of a...
Site Reliability Engineer

2 weeks ago

San Francisco, United States Air Apps Full time

Join to apply for the Site Reliability Engineer (SRE) role at Air AppsJoin to apply for the Site Reliability Engineer (SRE) role at Air AppsGet AI-powered advice on this job and more exclusive features.About Air AppsAt Air Apps, we believe in thinking bigger—and moving faster. We’re a family-founded company on a mission to create the world’s first...
Staff Site Reliability Engineer, User Protection SRE

3 weeks ago

San Francisco, United States Google Full time

Staff Site Reliability Engineer, User Protection SRE Join to apply for the Staff Site Reliability Engineer, User Protection SRE role at Google Applicants in San Francisco: Qualified applications with arrest or conviction records will be considered for employment in accordance with the San Francisco Fair Chance Ordinance for Employers and the California Fair...
Site Reliability Engineer

4 weeks ago

San Francisco, CA, United States Air Apps Full time

Join to apply for the Site Reliability Engineer (SRE) role at Air Apps Join to apply for the Site Reliability Engineer (SRE) role at Air Apps Get AI-powered advice on this job and more exclusive features. About Air Apps Are you ready to apply Make sure you understand all the responsibilities and tasks associated with this role before proceeding. At Air Apps,...
Staff Site Reliability Engineer

3 weeks ago

San Francisco, United States HeartFlow Full time

Heartflow is a medical technology company advancing the diagnosis and management of coronary artery disease, the #1 cause of death worldwide, using cutting-edge technology. The flagship product—an AI-driven, non-invasive cardiac test supported by the ACC/AHA Chest Pain Guidelines called the Heartflow FFR CT Analysis—provides a color-coded, 3D model of a...
CloudDevs: Senior Web site Reliability Engineer

2 weeks ago

San Francisco, United States The10minutecareersolution Full time

CloudDevs: Senior Web site Reliability Engineer (SRE) CloudDevs works with fast-moving, venture-backed startups throughout the US. We’re constructing a pool of world-class Web site Reliability Engineers for present roles and for upcoming alternatives. You’ll both be positioned straight into considered one of our associate startups or added to our vetted...
Staff Site Reliability Engineer, User Protection SRE

4 weeks ago

San Francisco, CA, United States Google Full time

Staff Site Reliability Engineer, User Protection SRE Increase your chances of an interview by reading the following overview of this role before making an application. Join to apply for the Staff Site Reliability Engineer, User Protection SRE role at Google Applicants in San Francisco: Qualified applications with arrest or conviction records will be...

Americas

Europe

Asia / Oceania

Africa

Senior / Staff Site Reliability Engineer (SRE)