Sr. System Development Engineer, High-Performance Accelerator Servers for AI/ML

4 days ago


Seattle, Washington, United States Amazon Full time $136,100 - $235,200 per year
DESCRIPTION

Do you want to shape the future of Generative AI at AWS? Join the team building the foundation of the world's most advanced cloud for AI training and inference — where multi-billion-parameter models come to life at scale. Here, you'll design, deliver, and operate next-generation infrastructure that powers breakthrough innovation in AI/ML and HPC workloads. If you're passionate about pushing the limits of performance, efficiency, and scalability in the cloud, this is your opportunity to build the systems that define what's next for AWS — and for the entire AI industry.

You'll join a diverse AWS Hardware Engineering team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You'll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

The ideal candidate for this role will be an innovative self-starter. You are knowledgeable of the full technical stack - vertically from baremetal server hardware up to the software in userland, and everything in the middle. You have tremendous interest in cloud scale and curious how systems and software decisions impact the user. You insist on highest-standards and are able to develop tactical solutions/tools to diagnose and fix issues. You are an excellent systems debugger - finding interaction issues between components on server systems. You are a leader with strong organizational, planning, and communication skills. You are a builder

Key job responsibilities

You will be a technical leader solving complex problems. You will decompose big difficult server system testability, reliability and diagnosis problems into straightforward tasks, components or features that you will lead to deliver yourself and through others in parallel. You will use combination of hardware, software, system designs, x86 architecture, processes, diagnosis and operations knowledge.

A day in the life

Working with a variety of job roles (SDEs, SDETs, Hardware Engineers, TPMs, Managers, Principals) and groups (AWS Hardware Engineering, EC2, other AWS services) through server conception, design, test, launch, and operations. Driving high quality and reliability into future/new designs for AWS Accelerated server solutions for AWS Cloud.

About the team

*Why AWS*

Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

*Diverse Experiences*

Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying.

*Work/Life Balance*

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud.

*Inclusive Team Culture*

Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness.

*Mentorship and Career Growth*

We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

BASIC QUALIFICATIONS
  • 4+ years of non-internship professional software development experience
  • 4+ years of deploying and operating in a Linux/Unix environment experience
  • 4+ years of systems development in an IT or data center environment experience
  • 3+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby experience
  • 2+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • 2+ years of systems design, software development, operations, automation, and process improvement experience
  • Experience leading the design, build and deployment of complex and performant (reliable and scalable) software solutions in production
PREFERRED QUALIFICATIONS
  • 3+ years of development/programming/scripting language (Python/Java/Bash/Perl) experience
  • Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
  • Experience taking a leading role in building complex software or computing infrastructure that has been successfully delivered to customers
  • Experience debugging, integrating, and validating complex AI/ML and Cloud Computing servers.

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner.

Our compensation reflects the cost of labor across several US geographic markets. The base pay for this position ranges from $136,100/year in our lowest geographic market up to $235,200/year in our highest geographic market. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience. Amazon is a total compensation company. Dependent on the position offered, equity, sign-on payments, and other forms of compensation may be provided as part of a total compensation package, in addition to a full range of medical, financial, and/or other benefits. For more information, please visit This position will remain posted until filled. Applicants should apply via our internal or external career site.



  • Seattle, Washington, United States Amazon Web Services (AWS) Full time $136,100 - $235,200 per year

    DescriptionDo you want to build the backbone of Generative AI cloud at AWS? Do you want to build the future of the cloud for AI training and inference? Want to do industry leading work delivering continuous price performance improvements in the cloud for AI model training for multi billion variable LLMs? Come Join us in designing, delivering and operating...


  • Seattle, Washington, United States Amazon Web Services (AWS) Full time $151,300 - $261,500 per year

    DescriptionAnnapurna Labs builds custom Machine Learning accelerators that are at the forefront of AWS innovation and one of several AWS tools used for building Generative AI on AWS. The Neuron Compiler Engineering team is searching for a Senior Software Development Engineer to support the development infrastructure of a compiler to enable the world's...


  • Seattle, Washington, United States Scale AI Full time $179,400 - $310,500

    As a Software Engineer on the ML Infrastructure team, you will design and build the next generation of foundational systems that power all ML Infrastructure compute at Scale - from model training and evaluation to large-scale inference and experimentation.Our platform is responsible for orchestrating workloads across heterogeneous compute environments (GPU,...

  • AI/ML Engineer

    4 days ago


    Seattle, Washington, United States Optum Full time

    Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers,...


  • Seattle, Washington, United States Google Full time $197,000 - $291,000 per year

    Minimum qualifications:Bachelor's degree or equivalent practical experience.8 years of experience in software development.5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.5 years of experience building and developing large-scale infrastructure, distributed systems or networks, or...


  • Seattle, Washington, United States myGwork - LGBTQ+ Business Community Full time $136,100 - $235,200 per year

    This job is with Amazon, an inclusive employer and a member of myGwork – the largest global platform for the LGBTQ+ business community. Please do not contact the recruiter directly.DescriptionDo you want to build the backbone of Generative AI cloud at AWS? Do you want to build the future of the cloud for AI training and inference? Want to do industry...


  • Seattle, Washington, United States Sobek Full time $170,000 - $230,000 per year

    About Sobek AI and the RoleSobek AI is building the secure nervous system for life-sciences innovation networks and intergovernmental emergency response. Backed by $10M+ in grants and funding, we partner with globally significant organizations to accelerate mission-critical, distributed workflows with AI—while protecting IP and sensitive public-health...

  • AI Engineer

    4 days ago


    Seattle, Washington, United States Sign AI Full time $100,000 - $190,000 per year

    Well-capitalized startup seeks extremely talented AI Engineers to help us pioneer the future of Sign Language translation. Our vision is to create a human-level AI Sign Language Interpreter available on any device, anywhere. You must have experience architecting ML pipelines, training and evaluating multimodal models, and leveraging AI to do your job faster...


  • Seattle, Washington, United States Apple Full time $171,600 - $302,200 per year

    Do you want to shape the platform that enables the next generation of intelligent experiences on Apple products & services? In Apple's Machine Learning Platform Technology & Infra team we have built the platform that Apple uses for developing machine learning, artificial intelligence, and computer vision applications. As a team, we have a variety of...


  • Seattle, Washington, United States Apple Full time $171,600 - $302,200 per year

    We're building the foundation for intelligent, adaptive AI systems from multi-agent platforms and RAG pipelines to advanced evaluation and reasoning frameworks. We're looking for a Senior Applied ML Engineer to design, build, and scale machine learning systems that power next-generation AI applications. In this role, you'll work at the intersection of...