Rack Systems Architect, Technical Lead

2 days ago

Phoenix, AZ, United States Ursus Inc Full-time
Rack Systems Architect, Technical Lead

Location:
New York, NY; Phoenix, AZ; or Austin, TX Duration: FTE Direct Hire Pay Range: $208K – $263K; offers equity

Qualifications:
10+ years of experience in rack-scale compute systems, high-performance computing (HPC), AI infrastructure, hyperscale data centers, server hardware, or integrated compute platform development. Demonstrated ownership of rack-level system architecture, including electrical, mechanical, thermal, networking, and integration requirements. Experience architecting and deploying high-density rack-scale computing environments supporting large accelerator deployments, AI training infrastructure, GPU clusters, or HPC systems. Experience developing and managing system-level power budgets, thermal budgets, cooling requirements, rack weight constraints, interconnect architectures, and infrastructure interfaces. Experience owning Interface Control Documents (ICDs) and defining boundaries between rack systems, facility systems, networking infrastructure, power systems, and cooling systems. Strong understanding of liquid-cooled rack architecture including CDUs, direct-to-chip cooling, thermal distribution systems, facility cooling integration, and serviceability considerations. Experience leading multidisciplinary technical teams including Electrical Engineering, Mechanical Engineering, Thermal Engineering, Systems Engineering, and Integration Engineering functions. Demonstrated experience making architecture trade-off decisions involving cost, performance, reliability, manufacturability, scalability, cooling, power density, transportability, and serviceability. Experience driving design reviews, technical readiness reviews, integration reviews, and engineering decision-making throughout product development lifecycles. Experience working directly with accelerator vendors, server vendors, ODMs, hyperscalers, or AI infrastructure providers.

Education:
Bachelor's degree in Electrical Engineering, Mechanical Engineering, Systems Engineering, Computer Engineering, or related technical discipline. Preferred

Qualifications:
Experience leading NVL-class, HGX-class, GB200-class, B100/B200-class, MI300-class, or equivalent rack programs. Deep familiarity with NVIDIA ecosystem architecture, accelerator integration, and GPU cluster deployments. Experience supporting hyperscale AI infrastructure environments at organizations such as NVIDIA, Meta, Microsoft, AWS, Google, Oracle Cloud, CoreWeave, Crusoe, Lambda, xAI, Dell, HPE, Supermicro, Quanta, Wistron, Celestica, or Foxconn. Experience with rack-level liquid cooling systems including direct-to-chip cooling, immersion cooling, rear-door heat exchangers, facility water systems, and CDU integration. Strong understanding of high-speed networking architectures, InfiniBand, Ethernet, optical interconnects, copper interconnects, scale-up fabrics, and scale-out architectures. Experience supporting transport qualification, seismic qualification, environmental qualification, or regulatory certification activities. Experience developing product roadmaps and aligning infrastructure designs to future accelerator generations. Experience supporting manufacturing, integration, rack deployment, system validation, or product lifecycle management activities. Advanced degree preferred in Electrical Engineering, Mechanical Engineering, Systems Engineering, Thermal Engineering, or Computer Engineering. How We Operate: Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role. Insane urgency. We drive everything forward as fast as possible. Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins. Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward. Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world. About The Team: The team owns the rack as a product: the unit of compute that designs, integrates, and deploys into gigawatt-scale AI data centers. Examples of key problems the team is working on: Architect 150kW-class liquid-cooled racks where power, thermal, weight, and interconnect budgets all bind at once and every tradeoff shows up somewhere else. Define and hold the interfaces between rack and facility so power, liquid, and network can each evolve without forcing a redesign of the other side. Keep the rack roadmap ahead of accelerator vendor roadmaps, so the next generation of dense compute lands in facilities designed before that hardware existed. Run a small engineering team with the review rigor of a large one, catching design faults before they reach integration. Position Summary: Own the rack as a product for 150kW-class liquid-cooled deployments: architecture, power budget, cooling, interconnect topology, structure, and