Production Systems Engineer, Fleet AI Systems

2 weeks ago


Menlo Park, United States META Full time

Production Systems Engineer, Fleet AI Systems (NetZero)

Apply to this job

Location pin icon

Menlo Park, CA

Apply to this job

Meta is seeking a Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services. The RTP team is responsible for the Hardware Lifecycle of all Meta servers including pre-production hands-on system and hardware debugging and stress testing, enabling production-ready system monitoring, automated provisioning and automated remediation of issues. RTP Engineers work closely with hardware designers, system manufacturers, component vendors, capacity engineering, production engineering, Facebook services, and data center operations teams to test systems before release to our production data centers, and to track the health and lifecycle of servers in production.

Production Systems Engineer, Fleet AI Systems (NetZero) Responsibilities

  • Interface with external vendors and internal hardware, mechanical, power, thermal, manufacturing and software engineers to understand system architecture to develop and execute the test suites for various architectures
  • Proactively create experiments and tooling to detect and diagnose hardware/firmware/software health issues
  • Develop test framework for large-scale test automation inside fleet during product development and after mass production
  • Implement remediations across software and hardware stack according to plan, while keeping a thorough procedural record and data log
  • Develop and publish updates on resolutions and communicate findings internally. Troubleshoot, diagnose and root cause of system failures and isolate the components/failure scenarios while working with internal & external stakeholders
  • Develop visibility through data visualization and implement systemic solutions to hardware health issues
  • Drive necessary discussion with external and internal teams on test specification and methodologies to improve test quality continuously
  • Contribute to Meta's 2030 Net Zero targets by evaluating sustainability, carbon footprint of new hardware design and infrastructure design. Partner with Net Zero teams to implement strategies across infrastructure for reuse, recycling, energy aware computing and quality practices for deployed and decommissioned hardware
Minimum Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • 4+ years of experience in hardware system support, knowledge of server architecture and components
  • Experience with Energy Aware Computing and/or Sustainable Infrastructure Design
  • Experience with Linux and scripting. Experience in changing system configurations and measuring change impact
  • Experience working in a matrix organization. Engineering for different server system/data center products
Preferred Qualifications
  • 4+ years experience in Production support at scale
  • 4+ years experience in full system technologies, full system lifecycle
  • Experience supporting AI/HPC systems and/or related components at scale. Experience in post-production hyperscale post-production environments, solutions


For those who live in or expect to work from California if hired for this position, please click here for additional information.

Start preparing
Learn about how to prepare for your interview with our interview guide, tips, and interactive experiences.
Visit interview prep

Locations

Use ctrl + scroll to zoom the map

Zoom in

Zoom out

Recenter

Data Center

About Meta

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today-beyond the constraints of screens, the limits of distance, and even the rules of physics.

Meta is committed to providing reasonable support (called accommodations) in our recruiting processes for candidates with disabilities, long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support. If you need support, please reach out to accommodations-ext@fb.com .

$124,000/year to $191,000/year + bonus + equity + benefits

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

  • Menlo Park, California, United States META Full time

    Production Systems Engineer - Fleet AI SystemsMeta is seeking a highly skilled Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services.ResponsibilitiesInterface with external vendors and...


  • Menlo Park, California, United States META Full time

    Production Systems Engineer, Fleet AI SystemsMeta is seeking a highly skilled Production Systems Engineer to join our Release to Production (RTP) team. As a key member of our team, you will be responsible for the Hardware Lifecycle of all Meta servers, including pre-production hands-on system and hardware debugging and stress testing, enabling...


  • Menlo Park, California, United States META Full time

    Job Title: Production Systems Engineer - Fleet AI SystemsMeta is seeking a highly skilled Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services.Responsibilities:Interface with external...


  • Menlo Park, California, United States META Full time

    Job Title: Production Systems Engineer, Fleet AI SystemsMeta is seeking a highly skilled Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services.Responsibilities:Interface with external...


  • Menlo Park, United States META Full time

    Meta is seeking a Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services. The RTP team is responsible for the Hardware Lifecycle of all Meta servers including pre-production hands-on system...


  • Menlo Park, California, United States META Full time

    Job SummaryMeta's AI Training and Inference Infrastructure is growing exponentially to support ever-increasing use cases of AI. This results in a dramatic scaling challenge that our engineers have to deal with on a daily basis. We need to build and evolve our network infrastructure that connects myriads of training accelerators like GPUs together.Key...


  • Menlo Park, California, United States META Full time

    Job Summary:Meta is seeking a highly skilled AI/HPC Systems Performance Engineer to join our team. As a key member of our infrastructure team, you will be responsible for designing, deploying, and operating high-performance networks to support our rapidly growing AI workloads.This is an exciting opportunity to work on cutting-edge technologies and contribute...


  • Tinley Park, Illinois, United States HNM Systems Full time

    Job Title: Senior Software Engineer - Generative AIHNM Systems is a leading provider of Communication and Information Technology staffing and consulting services. We are currently seeking a highly skilled Senior Software Engineer to join our team and contribute to the development of our Generative AI solutions.Job Summary:We are looking for a talented Senior...


  • Menlo Park, California, United States META Full time

    Meta Hardware Systems EngineerMeta is seeking a skilled Hardware Systems Engineer to join our Release to Production (RTP) team. As a key member of this team, you will be responsible for the end-to-end Hardware Lifecycle of all Meta servers, including prototyping of experimental HW, pre-production hands-on system and hardware debugging and stress testing,...


  • Menlo Park, California, United States Cyngn Full time

    About CyngnCyngn is a leading autonomous vehicle company based in Menlo Park, CA. We're a collaborative and diverse team that's passionate about innovation and continuous learning.Our self-driving technology can be deployed in various commercial domains across different vehicle form factors. We're seeking experienced leaders to join our team and help move...


  • Menlo Park, California, United States META Full time

    Job SummaryMeta is seeking a highly skilled Hardware Systems Engineer to join our Release to Production (RTP) team. As a key member of this team, you will be responsible for the end-to-end Hardware Lifecycle of all Meta servers, including prototyping, debugging, and stress testing.The RTP team is responsible for ensuring the efficient operation of our...

  • Systems Engineer

    4 weeks ago


    Lexington Park, United States BAE Systems Full time

    Job Description:BAE Systems is seeking an experienced Senior Engineer to implement and manage the VH92-A Mission Communications System Digital Ecosystem strategy. The selected candidate will manage a portfolio of Digital Ecosystem System Engineering (DESE), Data Architecture and Analytics, Product Lifecycle Management (PLM) Capability Development, Software...

  • Systems Engineer

    4 weeks ago


    Lexington Park, United States BAE Systems Full time

    Job Description:BAE Systems is seeking an experienced Senior Engineer to implement and manage the VH92-A Mission Communications System Digital Ecosystem strategy. The selected candidate will manage a portfolio of Digital Ecosystem System Engineering (DESE), Data Architecture and Analytics, Product Lifecycle Management (PLM) Capability Development, Software...


  • Menlo Park, California, United States Cyngn Full time

    About CyngnCyngn is a publicly traded autonomous vehicle company based in Menlo Park, CA. We have a culture of collaboration, diversity, and continuous learning. Our self-driving technology can be deployed in various commercial domains across various vehicle form factors.About the RoleWe are seeking a skilled Full Stack Engineer to contribute to the...


  • Menlo Park, California, United States META Full time

    Job Title: Production EngineerMeta is seeking a highly skilled Production Engineer to join our team. As a Production Engineer, you will be responsible for designing, developing, and maintaining software services to ensure optimal performance and capacity for growth.Key Responsibilities:Develop and maintain back-end data warehouse services, front-end...


  • Lexington Park, Maryland, United States BAE Systems Full time

    Job Title: Senior Systems EngineerWe are seeking an experienced Senior Systems Engineer to join our team at BAE Systems. As a key member of our Digital Ecosystem team, you will be responsible for implementing and managing the VH92-A Mission Communications System Digital Ecosystem strategy.Key Responsibilities:Manage a portfolio of Digital Ecosystem System...


  • Menlo Park, California, United States Diffuse Bio Full time

    Key Responsibilities:Design and develop software and APIs to enable internal and external access to our AI systems.Build tools to automate and maintain computing clusters and data parsing pipelines.Collaborate with our team of researchers to develop cutting-edge AI solutions.Requirements:Bachelor's or Master's degree in Computer Science or a related...


  • Lexington Park, United States BAE Systems USA Full time

    About the RoleBAE Systems USA is seeking an experienced Senior Systems Engineer to join our team in St. Mary's County, Maryland. As a Senior Systems Engineer, you will be responsible for implementing and managing the VH92-A Mission Communications System Digital Ecosystem strategy.Key ResponsibilitiesManage a portfolio of Digital Ecosystem System Engineering...


  • Lexington Park, Maryland, United States BAE Systems USA Full time

    Job DescriptionBAE Systems is seeking an experienced Senior Engineer to implement and manage the VH92-A Mission Communications System Digital Ecosystem strategy.Key ResponsibilitiesManage a portfolio of Digital Ecosystem System Engineering (DESE), Data Architecture and Analytics, Product Lifecycle Management (PLM) Capability Development, Software...


  • Menlo Park, United States OSI Engineering Full time

    Job Overview: We are looking for an experienced Staff/Principal Engineer to lead the development of AI capabilities. As the technical lead, you will focus on architecting and building high-quality front-end solutions while collaborating closely with platform engineers working on the AI infrastructure as well as senior product managers to create innovative...