GPU Performance Engineer
2 weeks ago
We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Join us in shaping the future of AI and pushing the boundaries of what's possible in video generation.
We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.
The Role
You'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.
Key Responsibilities
- Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation
- Write high-performance CUDA and Triton kernels for critical model operations
- Optimize cold start latency from seconds to milliseconds for our serving infrastructure
- Tune memory access patterns, kernel fusion, and GPU utilization
- Collaborate with ML engineers to optimize model implementations
- Debug performance issues across the full stack from application to hardware
- Implement custom memory pooling and allocation strategies
- Share optimization techniques and build performance culture across teams
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field
- 5+ years systems programming experience with 3+ years focused on GPU optimization
- Expert proficiency with GPU profiling tools (Nsight Systems, nvprof)
- Strong CUDA programming skills with production kernel development
- Deep understanding of GPU architecture (memory hierarchy, SMs, warps)
- Track record of achieving significant performance improvements (5-10x)
- Experience with Python and C++ in production environments
- Experience with Triton kernel development
- Knowledge of CUTLASS or similar high-performance libraries
- Background in ML-specific optimizations (attention, transformers)
- RDMA/InfiniBand optimization experience
- Contributions to GPU libraries or frameworks
- Low-level debugging skills (PTX/SASS reading)
Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish.
-
Performance Engineer, GPU
2 weeks ago
San Francisco, CA, United States Anthropic Full timeAbout Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role:...
-
GPU Performance Engineer
2 weeks ago
San Diego, CA, United States Qualcomm Full timeCompany: Qualcomm Technologies, Inc. Job Area: Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: Job Description At Qualcomm, we believe in the power of technology. For decades, our innovations have transformed entire industries, improved billions of lives, and addressed many of society's biggest challenges. With the world...
-
GPU Performance Engineer
2 weeks ago
San Diego, CA, United States Qualcomm Full timeCompany: Qualcomm Technologies, Inc. Job Area: Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: Job Description At Qualcomm, we believe in the power of technology. For decades, our innovations have transformed entire industries, improved billions of lives, and addressed many of society's biggest challenges. With the world...
-
GPU Performance Analysis Engineer
2 weeks ago
San Diego, CA, United States Apple Full timeRole Number: 200572814-3543 Summary Do you love creating elegant solutions to highly complex challenges? As part of our Silicon Engineering Group, you’ll help design and manufacture our next-generation, high-performance, power-efficient GPU! You’ll ensure Apple products and services can seamlessly and efficiently handle the tasks that make them beloved...
-
GPU Performance Analysis Engineer
7 days ago
San Diego, CA, United States Apple Full timeRole Number: 200572814-3543 Summary Do you love creating elegant solutions to highly complex challenges? As part of our Silicon Engineering Group, you’ll help design and manufacture our next-generation, high-performance, power-efficient GPU! You’ll ensure Apple products and services can seamlessly and efficiently handle the tasks that make them beloved...
-
GPU Performance Analysis Engineer
2 weeks ago
San Diego, CA, United States Qualcomm Full timeCompany: Qualcomm Technologies, Inc. Job Area: Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: Qualcomm is one of the largest fabless design companies in the world providing hardware, software and related services to nearly every mobile device maker and operator in the global wireless marketplace. The Qualcomm Graphics System...
-
GPU Performance Modeling Engineer
2 weeks ago
San Diego, CA, United States Apple Full timeRole Number: 200631413-3543 Summary Do you love creating elegant solutions to highly complex challenges? Do you intrinsically see the importance in every detail? As part of our Silicon Technologies group, you’ll help design and manufacture our next-generation, high-performance, power-efficient GPU! You’ll ensure Apple products and services can seamlessly...
-
GPU Performance Modeling Driver Engineer
2 weeks ago
San Diego, CA, United States San Diego Staffing Full timeApple Gpu Modeling Team Engineer Imagine what you could do here! At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish! Dynamic, hard-working people and inspiring, innovative technologies are the norm here....
-
GPU Kernel Engineer
1 week ago
San Francisco, CA, United States Sciforium Full timeSciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time...
-
GPU Performance Software Development Engineer
2 weeks ago
San Jose, CA, United States Advanced Micro Devices , Inc. Full timeWHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create...