Site Reliability Engineer
4 weeks ago
Application Deadline 1/6/25
Who We Are
At Cisco, we are a global leader in networking and IT, driving innovation and redefining how people connect, communicate, and collaborate. Our mission is to shape the future of the internet by creating unprecedented value and opportunity for our customers, employees, investors, and ecosystem partners. We are committed to encouraging a diverse and partnership environment where everyone can thrive and encourage our collective success.
Who You Are
We are seeking a highly skilled and experienced Senior Engineer to join our team, focusing on the design and development of AI services and capabilities tailored to IT’s GPU-based AI Clusters observability. This role involves reshaping how we lead alerts, metrics, and logs by introducing deep learning and GenAI to enhance reliability services. The ideal candidate will have a strong background in artificial intelligence, machine learning, and GPU-based AI infrastructure, with a consistent record of delivering innovative solutions that enhance system monitoring, performance, and reliability.
Key Responsibilities:
-
Design, build, and maintain observability systems for leading NVIDIA DGX clusters, ensuring flawless monitoring of AI workloads, hardware utilization (GPUs), and system health.
-
Develop monitoring tools and dashboards that supervise key metrics such as GPU utilization, memory, temperature, latency, network bandwidth, model performance, and system availability.
-
Build custom alerting systems for AI/ML workflows, enabling proactive issue detection (e.g., GPU failures, hardware bottlenecks, system crashes).
-
Collaborate with IT and MLOps teams to design efficient, scalable solutions for deploying, monitoring, and leading machine learning models on DGX systems.
-
Optimize DGX infrastructure by implementing standard processes for observability, ensuring high performance and reducing operational costs.
-
Supervise system-level metrics such as hardware temperature, power consumption, and GPU/CPU health, preventing hardware degradation or failure.
-
Develop solutions for supervising AI/ML model performance across DGX clusters, integrating logging and supervising for model training, inference, and deployment processes.
-
Integrate observability tools (e.g., Prometheus, Grafana, Splunk) with NVIDIA-specific tools (e.g., DCGM, NVIDIA GPU Cloud) for real-time monitoring and alerting.
-
Work closely with data scientists and machine learning engineers to ensure effective resource utilization and model observability, including the identification of performance bottlenecks and tuning for optimal GPU usage.
-
Drive solving and root cause analysis for failures and anomalies in both the DGX hardware and AI/ML models running on the infrastructure.
-
Ensure compliance with ethical AI standards by monitoring fairness, model drift, and performance consistency.
-
Document standard methodologies and processes for managing, deploying, and monitoring AI workloads on DGX clusters.
Minimum Qualifications:
-
Bachelor’s degree in Computer Science, Software Engineering, Data Science, or related fields.
-
7+ years of experience software engineering, systems engineering, or DevOps roles.
-
3+ years of experience in high-performance computing (HPC) or AI/ML environments.
Preferred Qualifications:
-
Strong experience leading NVIDIA DGX systems or similar GPU-based computing clusters.
-
Proficiency in GPU monitoring tools such as NVIDIA Data Center GPU Manager (DCGM) and related NVIDIA libraries/APIs.
-
Experience with AI/ML model deployment and monitoring on large-scale infrastructure, including model performance metrics (latency, throughput, accuracy).
-
Hands-on experience with observability tools such as Prometheus, Grafana, Splunk or similar, especially in high-performance computing environments.
-
Proficiency in scripting/programming languages (e.g., Python, Bash, Go) for automating cluster management and monitoring tasks.
-
Experience with container orchestration technologies (e.g., Docker, Kubernetes), including NVIDIA’s GPU operator for Kubernetes.
-
Familiarity with AI/ML lifecycle management tools such as ML flow, Kubeflow, or similar.
-
Strong understanding of HPC environments, including distributed computing, storage, and networking for AI/ML workloads.
-
Experience with infrastructure monitoring and solving at both hardware (GPU, CPU, memory) and software (AI/ML models, applications) levels.
-
Strong analytical and problem-solving skills, with the ability to interpret complex data and develop actionable insights.
-
Excellent verbal and written communication skills, with the ability to convey technical concepts to non-technical partners.
-
Ability to work effectively in a collaborative team environment and lead multiple projects simultaneously.
-
Experience with NVIDIA NGC (NVIDIA GPU Cloud) and DGX OS software stack for large-scale AI workloads.
-
Understanding of AI workload orchestration with frameworks such as Slurm or Kubernetes in GPU-based clusters.
-
Knowledge of NVIDIA Deep Learning frameworks (TensorFlow, PyTorch) and their performance optimization on DGX infrastructure.
-
Experience with AIOps tools for automated anomaly detection and solving of large-scale AI infrastructure.
-
Certification or experience with cloud platforms that offer GPU instances (AWS, GCP, Azure).
-
Familiarity with network performance tuning in HPC environments and large-scale AI workloads.
-
Familiarity with DevOps practices and tools, including CI/CD pipelines and infrastructure as code. Knowledge of Graphs, Graph DB's and Graph Theory. Familiarity with Terraform, Helm Chart, Ansible, or similar tools.
Why Cisco
#WeAreCisco, where each person is unique, but we bring our talents to work as a team and make a difference powering an expansive future for all.
We adopt digital, and help our customers implement change in their digital businesses. Some may think we’re "old" (36 years strong) and only about hardware, but we’re also a software company. And a security company. We even invented an intuitive network that adapts, predicts, learns and protects. No other company can do what we do - you can’t put us in a box
But "Digital Transformation" is an empty buzz phrase without a culture that allows for innovation, creativity, and yes, even failure (if you learn from it.)
Day to day, we focus on the give and take. We give our best, give our egos a break, and give of ourselves (because giving back is built into our DNA.) We take accountability, bold steps, and take difference to heart. Because without diversity of thought and a dedication to equality for all, there is no moving forward.
So, you have colorful hair? Don’t care. Tattoos? Show off your ink. Like polka dots? That’s cool. Pop culture geek? Many of us are. Passion for technology and world changing? Be you, with us
Cisco is an Affirmative Action and Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, national origin, genetic information, age, disability, veteran status, or any other legally protected basis.
Cisco will consider for employment, on a case by case basis, qualified applicants with arrest and conviction records.
-
Site Reliability Engineering Lead
3 weeks ago
Durham, North Carolina, United States Pearson Full timeDrive Operational Excellence as a Site Reliability Engineering LeadCultivate high-performing SRE teams at Pearson, fostering innovation, learning, and collaboration. This role involves supervising Junior, Titled, and Senior SREs, championing operational excellence, and orchestrating engineering goals with business objectives.Main Responsibilities:Champion...
-
Principal, Site Reliability Engineering
4 weeks ago
Durham, United States Pearson Full timeRole Overview: The Principal Site Reliability Engineering is a pivotal leadership role accountable for guiding Pearson's SRE teams towards increased operational excellence, system reliability, and strategic alignment with organizational objectives. This role is responsible for supervising Junior, Titled, and Senior SREs, fostering an environment of...
-
Senior Site Reliability Engineer
4 weeks ago
Durham, United States DataVisor Full timeDataVisor is a next generation security company that utilizes industry leading unsupervised machine learning to detect fraudulent activity for financial transactions, mobile user acquisition, social networks, commerce and money laundering. Our solution is used by some of the largest internet properties in the world, including Pinterest, FedEx, AirAsia,...
-
Sr. Reliability Engineer I
6 days ago
Durham, United States Biogen Full timeCompany Description Job Description About This Role The Sr. Reliability Engineer I applies Reliability Engineering methodologies to optimize design requirements and performance of critical assets across the site. Originates and develops analysis methods for determining reliability of components, equipment and processes. Acquires data and analyzes the...
-
Sr. Devops
4 weeks ago
Durham, United States CapB InfoteK Full timeCapB is a global leader on IT Solutions and Managed Services. Our R&D is focused on providing cutting edge products and solutions across Digital Transformations from Cloud, AI/ML, IOT, Blockchain to MDM/PIM, Supply chain, ERP, CRM, HRMS and Integration solutions. For our growing needs we need consultants who can work with us on salaried or contract basis. We...
-
Site Reliability Architect
3 weeks ago
Durham, North Carolina, United States Cisco Full timeJob SummaryWe are seeking a highly skilled and experienced Senior Engineer to join our team as an AI Engineering Lead, focusing on the design and development of AI services and capabilities tailored to IT's GPU-based AI Clusters observability. This role involves reshaping how we lead alerts, metrics, and logs by introducing deep learning and GenAI to enhance...
-
Site Reliability Engineer
4 weeks ago
Durham, United States Smart IT Frame LLC Full timeSRE Engineer ContractDurham NC Should have experience between 8 to 10 yearsExtensive experience working with linux flavors like rhel/centos os, shells, filesystems and utilitiesKnowledge of distributed computing and experience working with container orchestration frameworks including on-prem and rancher kubernetes and good knowledge on kubernetes...
-
Reliability Manager
6 days ago
Durham, United States Biogen Full timeJob Description As a Reliability Program Manager at Biogen, you'll be an integral contributor within the Facilities Engineering Technical Authority, responsible for spearheading a comprehensive reliability program. This pivotal role involves ensuring reliability, availability, and serviceability of our facilities, equipment, and systems throughout...
-
Process Reliability Engineer
3 days ago
Durham, North Carolina, United States Corning Incorporated Full time**Job Summary**Corning Incorporated is hiring a Process Reliability Engineer to support our manufacturing operations. The ideal candidate will have experience in technical manufacturing support and be able to provide on-shift technical support for manufacturing process focused on process stability and equipment reliability.Key Responsibilities:Provide...
-
Reliable Infrastructure Engineer
4 weeks ago
Durham, North Carolina, United States YO HR CONSULTANCY Full timeJob Title: Reliable Infrastructure EngineerThe role of a Reliable Infrastructure Engineer at YO HR CONSULTANCY is to design, implement, and maintain scalable and highly available infrastructure solutions. With a strong background in Kubernetes and Linux, the ideal candidate will have experience working with container orchestration frameworks, distributed...
-
Platform Reliability Engineering Architect
5 days ago
Durham, United States NetApp Full timeAbout NetAppNetApp is the intelligent data infrastructure company, turning a world of disruption into opportunity for every customer. No matter the data type, workload or environment, we help our customers identify and realize new business possibilities. And it all starts with our people.If this sounds like something you want to be part of, NetApp is the...
-
Platform Reliability Engineering Architect
6 days ago
Durham, United States NetApp Full timeAbout NetAppNetApp is the intelligent data infrastructure company, turning a world of disruption into opportunity for every customer. No matter the data type, workload or environment, we help our customers identify and realize new business possibilities. And it all starts with our people.If this sounds like something you want to be part of, NetApp is the...
-
Platform Reliability Engineering Architect
3 days ago
Durham, United States NetApp Full timeAbout NetApp NetApp is the intelligent data infrastructure company, turning a world of disruption into opportunity for every customer. No matter the data type, workload or environment, we help our customers identify and realize new business possibilities. And it all starts with our people. If this sounds like something you want to be part of, NetApp is...
-
Reliability Systems Architect
5 days ago
Durham, North Carolina, United States DataVisor Full timeAbout Us:DataVisor is a leading security company that leverages industry-leading unsupervised machine learning to detect and prevent fraudulent activities. Our innovative solution is used by some of the world's largest internet properties to protect them from ever-increasing risks of fraud. Our award-winning software is powered by a team of world-class...
-
Smart IT Frame LLC | Site Reliability Engineer
4 weeks ago
durham, United States Smart IT Frame LLC Full timeSRE Engineer ContractDurham NC Should have experience between 8 to 10 yearsExtensive experience working with linux flavors like rhel/centos os, shells, filesystems and utilitiesKnowledge of distributed computing and experience working with container orchestration frameworks including on-prem and rancher kubernetes and good knowledge on kubernetes...
-
Reliability Automation Expert
2 weeks ago
Durham, North Carolina, United States DataVisor Full timeDataVisor's cutting-edge security solutions utilize advanced machine learning techniques to prevent financial transaction fraud.We seek an experienced Senior Site Reliability Engineer to join our dynamic team. The ideal candidate will have a solid grasp of distributed systems, expertise in automation, and a drive to build robust systems.This role involves...
-
Site Reliability Engineering Lead
4 days ago
Durham, North Carolina, United States NetApp, Inc. Full timeJob RequirementsIn this role, we require at least 8 years of experience. You should have experience in writing, troubleshooting, and bug fixing product code. Additionally, you should be proficient in scripting and infrastructure automation using tools like Ansible, Python, Go, Perl, or Ruby. A deep understanding of Containers, Kubernetes, and Serverless...
-
Project Manager/ Project Engineer
3 weeks ago
Durham, United States Thomas & Hutton Full timeProject Manager / Project Engineer - Civil Site DevelopmentThomas & Hutton is a growing, well-established professional services firm providing consulting services throughout the southeast. We are an award-winning company that has been recognized as one of the best places to work in Georgia and South Carolina. Some of our many services include Civil,...
-
Cloud Engineering Lead
2 days ago
Durham, North Carolina, United States NetApp Full timeAbout NetAppAt NetApp, we're a team of innovative thinkers who help our customers turn challenges into business opportunities. We're passionate about harnessing the power of data to drive growth and innovation.We're looking for a skilled Sr. Site Reliability Engineer to join our team. As a Sr. Site Reliability Engineer, you'll play a key role in ensuring the...
-
Cloud Operations Engineer
5 days ago
Durham, North Carolina, United States DataVisor Full timeAbout the Role:We're seeking a talented Senior Site Reliability Engineer to join our growing team at DataVisor. As a key member of our engineering team, you'll play a vital role in ensuring the reliability, scalability, and performance of our infrastructure. With a passion for building scalable systems and automation, you'll collaborate closely with our team...