Production Systems Engineer, Fleet AI Systems

3 weeks ago


Menlo Park, United States META Full time

Production Systems Engineer, Fleet AI Systems (NetZero)

Apply to this job

Location pin icon

Menlo Park, CA

Apply to this job

Meta is seeking a Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services. The RTP team is responsible for the Hardware Lifecycle of all Meta servers including pre-production hands-on system and hardware debugging and stress testing, enabling production-ready system monitoring, automated provisioning and automated remediation of issues. RTP Engineers work closely with hardware designers, system manufacturers, component vendors, capacity engineering, production engineering, Facebook services, and data center operations teams to test systems before release to our production data centers, and to track the health and lifecycle of servers in production.

Production Systems Engineer, Fleet AI Systems (NetZero) Responsibilities

  • Interface with external vendors and internal hardware, mechanical, power, thermal, manufacturing and software engineers to understand system architecture to develop and execute the test suites for various architectures
  • Proactively create experiments and tooling to detect and diagnose hardware/firmware/software health issues
  • Develop test framework for large-scale test automation inside fleet during product development and after mass production
  • Implement remediations across software and hardware stack according to plan, while keeping a thorough procedural record and data log
  • Develop and publish updates on resolutions and communicate findings internally. Troubleshoot, diagnose and root cause of system failures and isolate the components/failure scenarios while working with internal & external stakeholders
  • Develop visibility through data visualization and implement systemic solutions to hardware health issues
  • Drive necessary discussion with external and internal teams on test specification and methodologies to improve test quality continuously
  • Contribute to Meta's 2030 Net Zero targets by evaluating sustainability, carbon footprint of new hardware design and infrastructure design. Partner with Net Zero teams to implement strategies across infrastructure for reuse, recycling, energy aware computing and quality practices for deployed and decommissioned hardware
Minimum Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • 4+ years of experience in hardware system support, knowledge of server architecture and components
  • Experience with Energy Aware Computing and/or Sustainable Infrastructure Design
  • Experience with Linux and scripting. Experience in changing system configurations and measuring change impact
  • Experience working in a matrix organization. Engineering for different server system/data center products
Preferred Qualifications
  • 4+ years experience in Production support at scale
  • 4+ years experience in full system technologies, full system lifecycle
  • Experience supporting AI/HPC systems and/or related components at scale. Experience in post-production hyperscale post-production environments, solutions


For those who live in or expect to work from California if hired for this position, please click here for additional information.

Start preparing
Learn about how to prepare for your interview with our interview guide, tips, and interactive experiences.
Visit interview prep

Locations

Use ctrl + scroll to zoom the map

Zoom in

Zoom out

Recenter

Data Center

About Meta

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today-beyond the constraints of screens, the limits of distance, and even the rules of physics.

$132,000/year to $191,000/year + bonus + equity + benefits

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

Equal Employment Opportunity and Affirmative Action

Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or other applicable legally protected characteristics. You may view our Equal Employment Opportunity notice here .

Meta is committed to providing reasonable support (called accommodations) in our recruiting processes for candidates with disabilities, long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support. If you need support, please reach out to accommodations-ext@fb.com .

  • Menlo Park, United States META Full time

    Production Systems Engineer, Fleet AI Systems (NetZero) Apply to this job Location pin icon Menlo Park, CA Apply to this job Meta is seeking a Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our...


  • Menlo Park, California, United States META Full time

    Job Summary:Meta is seeking a highly skilled AI/HPC Systems Performance Engineer to join our team. As a key member of our infrastructure team, you will be responsible for designing, deploying, and operating high-performance networks to support our rapidly growing AI workloads.This is an exciting opportunity to work on cutting-edge technologies and contribute...


  • Menlo Park, California, United States Diffuse Bio Full time

    Key Responsibilities:Design and develop software and APIs to enable internal and external access to our AI systems.Build tools to automate and maintain computing clusters and data parsing pipelines.Collaborate with our team of researchers to develop cutting-edge AI solutions.Requirements:Bachelor's or Master's degree in Computer Science or a related...

  • Systems Engineer

    2 weeks ago


    Lexington Park, MD, United States BAE Systems Full time

    Job Description:BAE Systems is seeking an experienced Senior Engineer to implement and manage the VH92-A Mission Communications System Digital Ecosystem strategy. The selected candidate will manage a portfolio of Digital Ecosystem System Engineering (DESE), Data Architecture and Analytics, Product Lifecycle Management (PLM) Capability Development, Software...


  • Menlo Park, United States META Full time

    Summary: In this role, you will be a member of the Network AI Software team and part of the bigger DC networking organization. The team develops and owns the software stack around collective communication libraries around Meta.At the high level, the team aims to enable Meta-wide ML products and innovations to leverage our large-scale training and inference...


  • Menlo Park, California, United States META Full time

    Meta AI Systems Machine Learning Research Scientist InternMeta is seeking a Research Scientist Intern to join its Fundamental AI Research (FAIR) team. As a Research Scientist Intern, you will contribute to advancing the field of artificial intelligence by making fundamental advances in technologies to help interact with and understand our...


  • Menlo Park, California, United States META Full time

    About the Role:We are seeking a highly skilled Product Manager to lead the development of our next-generation AI infrastructure. The ideal candidate will have a strong background in product management, with a focus on AI and hardware.Key Responsibilities:Establish a shared vision and strategy for a portfolio of products that enable efficient and reliable...


  • Menlo Park, United States META Full time

    Summary: In this role, you will be a member of the MTIA (Meta Training & Inference Accelerator) Software team and part of the bigger industry-leading PyTorch AI framework organization. MTIA Software Team has been developing a comprehensive AI Compiler strategy that delivers a highly flexible platform to train & serve new DL/ML model architectures, combined...


  • Menlo Park, United States OSI Engineering Full time

    Job Overview:We are looking for an experienced Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure as well as senior product managers to create innovative...


  • Menlo Park, United States OSI Engineering Full time

    Job Overview:We are looking for an experienced Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure as well as senior product managers to create innovative...


  • Menlo Park, United States OSI Engineering Full time

    Job Overview:We are looking for an experienced Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure as well as senior product managers to create innovative...


  • Menlo Park, United States OSI Engineering Full time

    Job Overview:We are looking for an experienced Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure as well as senior product managers to create innovative...


  • Menlo Park, United States Raya Full time

    Our client, is a public company and a leading Enterprise AI software provider for accelerating digital transformation.Their Platform provides comprehensive services to build enterprise-scale AI applications more efficiently and cost-effectively than alternative approaches. The company’s Platform supports the value chain in any industry with prebuilt,...


  • Menlo Park, United States Raya Full time

    Our client, is a public company and a leading Enterprise AI software provider for accelerating digital transformation. Their Platform provides comprehensive services to build enterprise-scale AI applications more efficiently and cost-effectively than alternative approaches. The company’s Platform supports the value chain in any industry with prebuilt,...


  • Menlo Park, United States META Full time

    Summary: In this role, you will be a member of the Network AI Software team and part of the bigger DC networking organization. The team develops and owns the software stack around collective communication libraries around Meta.At the high level, the team aims to enable Meta-wide ML products and innovations to leverage our large-scale training and inference...


  • Allen Park, United States Apex Systems Full time

    Job#: 2052298 Job Description: Responsibilities: Design and develop AI systems for software quality and warranty. Lead the development and maintenance of AI and machine learning algorithms. Ensure the delivery of stable, high-quality infotainment products. Collaborate with various teams to meet product and engineering excellence. Thrive in a fast-paced,...


  • Menlo, Georgia, United States OSI Engineering Full time

    Job Overview:We are seeking a highly skilled Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure and senior product managers to create innovative customer...


  • menlo, United States OSI Engineering Full time

    Job Overview:We are looking for an experienced Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure as well as senior product managers to create innovative...


  • menlo, United States OSI Engineering Full time

    Job Overview:We are looking for an experienced Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure as well as senior product managers to create innovative...


  • menlo, United States OSI Engineering Full time

    Job Overview:We are looking for an experienced Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure as well as senior product managers to create innovative...