Sr. Software Engineer- AI/ML, AWS Neuron Apps

2 weeks ago


Seattle, United States Amazon Full time

Shape the Future of AI Accelerators at AWS Neuron Join the elite team behind AWS Neuron-the software stack powering AWS's next-generation AI accelerators Inferentia and Trainium. As a Senior Software Engineer in our Machine Learning Applications team, you'll be at the forefront of deploying and optimizing some of the world's most sophisticated AI models at unprecedented scale. What You'll Impact: • Pioneer distributed inference solutions for industry-leading LLMs such as GPT, Llama, Qwen • Optimize breakthrough language and vision generative AI models • Collaborate directly with silicon architects and compiler teams to push the boundaries of AI acceleration • Drive performance benchmarking and tuning that directly impacts millions of inference calls globally Key job responsibilities You will drive the Evolution of Distributed AI at AWS Neuron As a Technical Leader at the forefront of AWS's AI Accelerator, you'll architect the bridge between ML frameworks including PyTorch, JAX and AI hardware. This isn't just about just optimization-it's about revolutionizing how AI models run at scale. Technical Impact You'll Drive: • Spearhead distributed inference architecture for PyTorch and JAX using XLA • Engineer breakthrough performance optimizations for AWS Trainium and Inferentia • Develop ML tools to enhance LLM accuracy and efficiency • Transform complex tensor operations into highly optimized hardware implementations • Pioneer benchmarking methodologies that shape next-gen AI accelerator design What Makes This Role Unique: • Direct influence on AWS's AI infrastructure used by thousands of ML applications • Full-stack optimization from high-level frameworks to hardware-specific primitives • Creation of tools and frameworks that define industry standards for ML deployment • Collaboration with both open-source ML communities and hardware architecture teams Your Technical Arsenal Should Include: • Deep expertise in Python and ML framework internals • Strong understanding of distributed systems and ML optimization • Passion for performance tuning and system architecture A day in the life Work/Life Balance Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. About the team At AWS Neuron, we're revolutionizing how the world's most sophisticated AI models run at scale through Amazon's next-generation AI accelerators. Operating at the unique intersection of ML frameworks and custom silicon, our team drives innovation from silicon architecture to production software deployment. We pioneer distributed inference solutions for PyTorch and JAX using XLA, optimize industry-leading LLMs like GPT and Llama, and collaborate directly with silicon architects to influence the future of AI hardware. Our systems handle millions of inference calls daily, while our optimizations directly impact thousands of AWS customers running critical AI workloads. We're focused on pushing the boundaries of large language model optimization, distributed inference architecture, and hardware-specific performance tuning. Our deep technical experts transform complex ML challenges into elegant, scalable solutions that define how AI workloads run in production. BASIC QUALIFICATIONS - 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - 5+ years of programming experience using Python or C++ and PyTorch. - Experience with AI acceleration via quantization, parallelism, model compression, batching, KV caching, vllm serving - Experience with accuracy debugging & tooling, performance benchmarking of AI accelerators - Fundamentals of Machine learning and deep learning models, their architecture, training and inference lifecycles along with work experience on optimizations for improving the model execution. PREFERRED QUALIFICATIONS - Master's degree in computer science or equivalent - Master's degree in machine learning or equivalent - Experience with accuracy debugging & tooling, performance benchmarking of AI accelerators - Experience in developing CUDA kernels, HPC and inference optimization, tensors operations Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. Our compensation reflects the cost of labor across several US geographic markets. The base pay for this position ranges from $151,300/year in our lowest geographic market up to $261,500/year in our highest geographic market. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience. Amazon is a total compensation company. Dependent on the position offered, equity, sign-on payments, and other forms of compensation may be provided as part of a total compensation package, in addition to a full range of medical, financial, and/or other benefits. For more information, please visit This position will remain posted until filled. Applicants should apply via our internal or external career site.



  • Seattle, United States Amazon Web Services (AWS) Full time

    Sr. Software Engineer- AI/ML, AWS Neuron AppsJob ID: 2921620 | Amazon.com Services LLC - A57AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trn1 and Inf1 servers that use them. This role is for a senior software engineer in the Machine Learning Applications (ML Apps) team for AWS...


  • Seattle, United States Amazon Web Services (AWS) Full time

    Join to apply for the Software Engineer- AI/ML, AWS Neuron role at Amazon Web Services (AWS)3 days ago Be among the first 25 applicantsJoin to apply for the Software Engineer- AI/ML, AWS Neuron role at Amazon Web Services (AWS)DescriptionAWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machineDescriptionAWS Neuron is...


  • Seattle, United States Amazon Web Services (AWS) Full time

    Sr. Software Engineer- AI/ML, AWS Neuron Distributed TrainingAnnapurna Labs designs silicon and software that accelerates innovation. Customers choose us to create cloud solutions that solve challenges that were unimaginable a short time ago—even yesterday. Our custom chips, accelerators, and software stacks enable us to take on technical challenges that...


  • Seattle, United States Amazon Web Services (AWS) Full time

    Software Engineer- AI/ML, AWS Neuron Distributed TrainingJoin to apply for the Software Engineer- AI/ML, AWS Neuron Distributed Training role at Amazon Web Services (AWS)DESCRIPTIONAnnapurna Labs designs silicon and software that accelerates innovation. Customers choose us to create cloud solutions that solve challenges that were unimaginable a short time...


  • Seattle, United States Amazon Web Services (AWS) Full time

    Software Engineer-AI/ML, AWS Neuron InferenceJoin to apply for the Software Engineer-AI/ML, AWS Neuron Inference role at Amazon Web Services (AWS).AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud‑scale machine learning accelerators. This role is for a senior software engineer in the Machine Learning Inference Applications...


  • Seattle, United States Annapurna Labs Full time

    Shape the Future of AI Accelerators at AWS Neuron Join the elite team behind AWS Neuron—the software stack powering AWS's next-generation AI accelerators Inferentia and Trainium. As a Senior Software Engineer in our Machine Learning Applications team, you'll be at the forefront of deploying and optimizing some of the world's most sophisticated AI models at...


  • Seattle, United States Annapurna Labs Full time

    Shape the Future of AI Accelerators at AWS Neuron Join the elite team behind AWS Neuron—the software stack powering AWS's next-generation AI accelerators Inferentia and Trainium. As a Senior Software Engineer in our Machine Learning Applications team, you'll be at the forefront of deploying and optimizing some of the world's most sophisticated AI models at...


  • Seattle, United States Annapurna Labs (U.S.) Inc. Full time

    Shape the Future of AI Accelerators at AWS NeuronJoin the elite team behind AWS Neuron—the software stack powering AWS's next-generation AI accelerators Inferentia and Trainium. As a Senior Software Engineer in our Machine Learning Applications team, you'll be at the forefront of deploying and optimizing some of the world's most sophisticated AI models at...


  • Seattle, WA, United States Amazon Full time

    Shape the Future of AI Accelerators at AWS Neuron Join the elite team behind AWS Neuron-the software stack powering AWS's next-generation AI accelerators Inferentia and Trainium. As a Senior Software Engineer in our Machine Learning Applications team, you'll be at the forefront of deploying and optimizing some of the world's most sophisticated AI models at...


  • Seattle, United States Amazon Web Services (AWS) Full time

    Join to apply for the Sr. Machine Learning Compiler Engineer, AWS Neuron, Annapurna Labs role at Amazon Web Services (AWS)1 day ago Be among the first 25 applicantsDescriptionThe Product: AWS Machine Learning accelerators are at the forefront of AWS innovation and are used for building Generative AI on AWS. The Inferentia chip delivers top ML inference...