Sr. Infrastructure Reliability Engineer, Infrastructure Reliability

4 weeks ago


Chantilly VA United States Amazon Data Services, Inc. Full time
As an Infrastructure Reliability Engineer you will be proactively driving the reliability risk identification, assessment and mitigation for datacenter infrastructure equipment (Example: Air Handling Units, LV Generator, MV Transformers, LV SWGR, Breakers, UPS, Chillers etc.). You will also be responsible for root cause analysis of critical equipment failures and drive the continuous improvements to improve datacenter availability for AWS customers. You will work closely with both internal and outside partners including suppliers to drive key aspects of product specification, risk identification plan and execution. You must be ownership minded, independent, action and results oriented to succeed in an open collaborative environment.
The candidate should have experience in using Physics-of-Failure based approach to develop and implement both analytical and empirical approaches for product quality/reliability risk identification and assessment during product design, manufacture as well as deployment stages. The individual should be able to drive AWS application-specific requirements in carrying out both lifecycle environmental and operational stress driven risk analysis, including thermal, electrical, chemical and mechanical stresses so to identify overstress and fatigue-related product weaknesses. Candidate should be capable of evaluating not only product design quality/reliability risks, but also have the skills and experiences in assessing electronics manufacture process related quality/reliability issues. Knowledge of statistical techniques and models is required to analyze test as well as field data.

At the component level, the individual will drive critical component identification and the associated vendor selection and qualification requirements. The candidate will be expected to use knowledge of process capability for electronic component production as well as system-level performance requirements to establish critical to quality and reliability metrics.
At the system level, the individual will develop datacenter system level reliability model and related reliability quantification and risk analysis for datacenter configuration optimization. The candidate will be expected to be familiar with system reliability engineering tools, such as reliability block diagram, statistical modeling and data analytics.

During sustaining stage, candidate will be responsible for monitoring product performance in the field and will be responsible to drive root cause analysis of any critical failures and the associated corrective and preventive actions. The individual should also be able to drive effective vendor auditing and quarterly review process to drive the continuous improvements of datacenter availability.

The successful candidate should be considered as an expert in the reliability engineering field and have a proven track record of success in not only product reliability leadership, as well as business negotiations and program management. Strong skill-set in problem analysis and solving as well as communication and vendor management are necessary. Candidates should also be able to travel within US and internationally.

AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we’re looking for talented people who want to help.

You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

Key job responsibilities
Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness.
We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

Amazon Web Services (AWS) is committed to a diverse and inclusive workplace to deliver the best results for our customers. Amazon is an equal opportunity employer and does not discriminate on the basis of race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status; we celebrate the diverse ways we work. For individuals with disabilities who would like to request an accommodation, please let us know and we will connect you to our accommodation team. You may also reach them directly by visiting please https://www.amazon.jobs/en/disability/us.

We are open to hiring candidates to work out of one of the following locations:

Herndon, VA, USA
BASIC QUALIFICATIONS - Bachelor's or Master’s degree in Reliability Engineering, Physics, Electrical, Mechanical or Materials Engineering or related field
- 10+ years of Reliability Engineering work experience in high reliability industry
- 5+ years of experience with failure analysis activities and root cause analysis
- 5+ years of experience with accelerated life testing, stress analysis and finite element analysis

PREFERRED QUALIFICATIONS - Ph.D. in Reliability Engineering, Physics, Electrical, Mechanical or Materials Engineering or a related field.
- 10+ years of work experience in reliability risk identification and assessment from component to system level applying analytical, experimental and statistical approaches to evaluate product design and manufacture quality/reliability levels
- Experience with proactive and effective reliability approaches in a cost-effective manner throughout product design, manufacture and deployment stages
- Proven experience in working with external design and manufacturing supply chain partners.
- Familiarity with major data center infrastructure equipment reliability performance
- Ability to influence development teams, procurement and external partners
- Ability in managing multiple qualification activities and development schedules
- Excellent verbal and written communication skills
- Meets/exceeds Amazon’s functional/technical depth and complexity for this role
- Meets/exceeds Amazon’s leadership principles requirements for this role

Amazon is committed to a diverse and inclusive workplace. Amazon is an equal opportunity employer and does not discriminate on the basis of race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status. For individuals with disabilities who would like to request an accommodation, please visit https://www.amazon.jobs/en/disability/us.

  • New York, NY, United States Fourier Ltd Full time

    Joining a growing team to support, maintain and improve their automated trading systems. You'll be working in a fast paced and agile trading environment In the worlds most successful hedge fund to support maintain and improve their trading systems. This company looks for the most talented engineers on the market and rewards them accordingly. Build and...


  • Chantilly, United States ANALYTIC SOLUTIONS GROUP, LLC Full time

    As a Cloud Infrastructure Engineer, you’ll play a pivotal role in shaping our client’s cloud-based computing environments, ensuring they run seamlessly and securely. Responsibilities Design and Build: Collaborate with cross-functional teams to design, build, and maintain the hardware and software systems that underpin our cloud infrastructure. Your...


  • Chantilly, Virginia, United States Analytic Solutions Group Full time

    As a Cloud Infrastructure Engineer, you’ll play a pivotal role in shaping our client’s cloud-based computing environments, ensuring they run seamlessly and securely. Responsibilities Design and Build: Collaborate with cross-functional teams to design, build, and maintain the hardware and software systems that underpin our cloud infrastructure. Your...


  • Chantilly, United States Ampcus Full time

    Job Details: Role: Infrastructure Engineer Location : Chantilly, VA (Onsite) Duration: Fulltime Description: Strong AWS infrastructure experience Puppet Integration, Chef, Ansible, playbook Strong CI/CD pipeline experience, Experience with monitoring and observability, preferably with Splunk, Grafana, Prometheus Test automation experience,...


  • United, United States Forhyre Full time

    Job DescriptionJob DescriptionDo you enjoy solving technical issues, empathize with customer user experiences and want to keep up with the latest tech? We are looking for a Cloud Infrastructure Engineer that will work with talented software engineering and support teams to deploy, maintain and ensure reliability of our applications in a fast paced...


  • Chantilly, United States IDEMIA Full time

    Overview IDEMIA is the global leader in identity and security. Our mission is to create a safe and simple future where identity verification is indisputable, and only you can assert your identity. We are a distributed company leveraging the latest technologies to deliver world-class products in the private and public sectors of finance, telecom, identity,...


  • Reston, VA, United States ALTA IT Services Full time

    Site Reliability Engineering (SRE) Lead 100% Remote US Citizenship required per government contract  As a Site Reliability Engineering (SRE) Lead, you'll deliver mission-critical services that empower end users. As the ideal candidate, you'll use your extensive experience designing and implementing end-to-end continuous delivery pipelines and...


  • Reston, VA, United States ALTA IT Services Full time

    Site Reliability Engineering (SRE) Lead 100% Remote W2 ONLY US Citizenship required per government contract Must be able to obtain a DHS Public Trust clearance As a Site Reliability Engineering (SRE) Lead, you'll deliver mission-critical services that empower end users. As the ideal candidate, you'll use your extensive experience designing and implementing...


  • New York, NY, United States M-Logic Full time

    Role Summary: Our client is looking for a highly skilled Cloud Engineer to join a talented Infrastructure Team. As a Cloud Engineer, you will be responsible for designing, deploying, and maintaining our cloud infrastructure, with a particular focus on Kubernetes & Docker. You will be part of a team responsible for building and maintaining the backbone of our...


  • Los Gatos, California, United States Netflix Full time

    "At Netflix, we strive to bring joy to people across the world through amazing stories. As we grow internationally, we are continually enhancing our cloud-based infrastructure to improve our performance, scalability, and reliability.The SRE team's goal is to ensure customer joy by successfully managing risk and minimizing impact across Netflix. We do this...


  • Chicago, IL, United States CME Group Full time

    Description Position Overview: Data System Reliability Engineer (dSRE) CME Group: Where Futures Are Made CME Group is the world's leading and most diverse derivatives marketplace. But who we are goes deeper than that, here you can impact markets worldwide, transform industries and build a career shaping tomorrow. We invest in your success and you own...


  • Chicago, IL, United States CME Group Full time

    Description This role is hybrid requires to be 2 days on site in our Chicago office. This role does not allow to work outside of Illinois state. Position Overview: Data System Reliability Engineer (dSRE) CME Group: Where Futures Are Made CME Group is the world's leading and most diverse derivatives marketplace. But who we are goes deeper than...


  • Chantilly, United States TekStream Solutions Full time

    The job duties of the Linux Infrastructure Engineer are as follows:Engineer and support the rapidly evolving Red Hat Enterprise Linux (RHEL) common service requirements for mission critical network management systems and applicationsPatching and maintaining server buildsCoordinating and managing patching repositories in offline environments to ensure timely...


  • Chantilly, United States TekStream Solutions Full time

    The job duties of the Linux Infrastructure Engineer are as follows:Engineer and support the rapidly evolving Red Hat Enterprise Linux (RHEL) common service requirements for mission critical network management systems and applicationsPatching and maintaining server buildsCoordinating and managing patching repositories in offline environments to ensure timely...


  • Chantilly, Virginia, United States SAIC Career Site Full time

    Description SAIC is seeking a highly motivated Software Infrastructure SETA to provide SETA support to SAIC's Program, Landmark AOS in Chantilly, VA. Landmark AOS is a large SETA program, supporting the NRO's Ground Enterprise Directorate (GED), responsible for the acquisition of systems over the complete end-to-end life cycle. The Software Infrastructure...

  • Reliability Engineer

    4 weeks ago


    Salina, KS, United States Schwan's Full time

    Who we are! At Schwan’s Company, the opportunities are real, and the sky is the limit; this isn’t just a job, it’s a seat at the table. Around here, every job matters, every voice counts, and every person contributes in a big way. As part of our front lines, we look to you to execute business, build relationships, and take pride in your work because at...


  • CHANTILLY, United States SAIC Career Site Full time

    Description SAIC is seeking a highly motivated Software Infrastructure SETA to provide SETA support to SAIC’s Program, Landmark AOS in Chantilly, VA. Landmark AOS is a large SETA program, supporting the NRO’s Ground Enterprise Directorate (GED), responsible for the acquisition of systems over the complete end-to-end life cycle.  The Software...


  • Chantilly, United States Building Infrastructure Group, Inc. Full time

    Job DescriptionJob DescriptionPosition ResponsibilitiesResponsible for preparing detailed estimates in support of multiple small to large projects based on construction drawings and specifications including, but not limited to:Designing and delivering proposal documents to prospective clients and generating RFP’s based on applying established pricing...


  • Washington, VA, United States Leidos Full time

    Leidos has an opening for an Infrastructure Operations Manager to support the ESA V program. ESA V is an IT Services program supporting several customers within the Department of Justice and the rest of the US Federal Government. The program provides a range of IT services, including help desk, deskside support, Windows engineering and maintenance, managed...


  • Seattle, WA, United States The Calyx Institute Full time

    The Calyx Institute is seeking applications for Senior Systems Developer positions. A contributor to the Calyx engineering team, Senior Systems Developers will be primarily responsible for building and organizing components of our infrastructure. Ideal candidates should have experience with the architecture of systems used to support the development of...