Reliability Engineer
18 hours ago
Austin, Texas, United States
Apple
Full-time
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
By continuing, you agree to our Terms & Privacy Policy.
At Apple, we focus deeply on our customers’ experience. Apple Ads brings this same approach to advertising, helping people find exactly what they’re looking for and helping advertisers grow their businesses.\\n\\nOur technology powers ads and sponsorships across Apple Services, including the App Store, Apple News, MLS Season Pass and Apple Maps. Everything we do is designed for trust, connection, and impact: We respect user privacy, integrate advertising thoughtfully into the experience, and deliver value for advertisers of all sizes—from small app developers to big, global brands. Because when advertising is done right, it benefits everyone.
As a Data Site Reliability Engineer in Apple Ads, you will focus on designing, building, and managing cloud-based systems and infrastructure that stores, process, transforms and analyzes large datasets. You will ensure data and corresponding infrastructure is accessible, reliable, and secure within the cloud environment, working with various Apple Internal and 3rd party cloud providers like AWS.
Design, build and manage large scale distributed data systems on AWS using services such as EMR, EKS, MSK, Iceberg, etc.\\nDesign and develop internal tooling and automation frameworks to improve data infrastructure reliability, cost-efficiency, and observability.\\nCollaborate with cross-functional engineering teams like advertising, machine learning and data platforms to define infrastructure architecture, troubleshoot complex issues, and drive production excellence.\\nLead initiatives to eliminate toil via automation, improve efficiency and scalable deployment patterns.\\nContribute to strategic decision making related to adoption of a multi-cloud echo system to maintain operational excellence on the current stack while still enabling the path for future innovation.
2+ years of experience in Infrastructure Engineering with a strong focus on building, scaling and operating cloud based distributed data systems. \\nProven expertise running cloud infrastructure built on managed services offered by cloud providers like AWS\\nStrong technical grasp and experience working on Open Source technologies designed for large scale data processing, like Apache Spark, Flink, Kafka, Iceberg or other similar technologies\\nStrong expertise and experience with container orchestration systems like Kubernetes\\nExperience leveraging ML and GenAI capabilities to improve infrastructure and Operational efficiency
Experience designing, building, and operating large-scale distributed data infrastructure (e.g., Spark, Flink, Kafka) supporting petabyte-scale batch and streaming pipelines with high availability and low latency.\\nTrack record owning reliability and performance of cloud-native data platforms (AWS/Kubernetes), including capacity planning, cost optimization, and incident response for production-critical systems.\\nAbility to partner with data science, ML, and product teams to define infrastructure standards, automate deployment/observability tooling, and mentor engineers on distributed systems best practices.\\nTrack record of driving automation, cost optimization and performance tuning at scale for data systems.\\nExperience designing, analyzing and troubleshooting large-scale distributed systems.\\nExperience driving adoption of new technologies and influencing designing of systems. \\nStrong programming skills on object-oriented, functional or procedural programming languages, preferably Java / Scala / Kotlin / Python\\nExperience designing and managing Infrastructure as Code with Helm and CRD, ensuring repeatable, secure, and scalable deployments.
As a Data Site Reliability Engineer in Apple Ads, you will focus on designing, building, and managing cloud-based systems and infrastructure that stores, process, transforms and analyzes large datasets. You will ensure data and corresponding infrastructure is accessible, reliable, and secure within the cloud environment, working with various Apple Internal and 3rd party cloud providers like AWS.
Design, build and manage large scale distributed data systems on AWS using services such as EMR, EKS, MSK, Iceberg, etc.\\nDesign and develop internal tooling and automation frameworks to improve data infrastructure reliability, cost-efficiency, and observability.\\nCollaborate with cross-functional engineering teams like advertising, machine learning and data platforms to define infrastructure architecture, troubleshoot complex issues, and drive production excellence.\\nLead initiatives to eliminate toil via automation, improve efficiency and scalable deployment patterns.\\nContribute to strategic decision making related to adoption of a multi-cloud echo system to maintain operational excellence on the current stack while still enabling the path for future innovation.
2+ years of experience in Infrastructure Engineering with a strong focus on building, scaling and operating cloud based distributed data systems. \\nProven expertise running cloud infrastructure built on managed services offered by cloud providers like AWS\\nStrong technical grasp and experience working on Open Source technologies designed for large scale data processing, like Apache Spark, Flink, Kafka, Iceberg or other similar technologies\\nStrong expertise and experience with container orchestration systems like Kubernetes\\nExperience leveraging ML and GenAI capabilities to improve infrastructure and Operational efficiency
Experience designing, building, and operating large-scale distributed data infrastructure (e.g., Spark, Flink, Kafka) supporting petabyte-scale batch and streaming pipelines with high availability and low latency.\\nTrack record owning reliability and performance of cloud-native data platforms (AWS/Kubernetes), including capacity planning, cost optimization, and incident response for production-critical systems.\\nAbility to partner with data science, ML, and product teams to define infrastructure standards, automate deployment/observability tooling, and mentor engineers on distributed systems best practices.\\nTrack record of driving automation, cost optimization and performance tuning at scale for data systems.\\nExperience designing, analyzing and troubleshooting large-scale distributed systems.\\nExperience driving adoption of new technologies and influencing designing of systems. \\nStrong programming skills on object-oriented, functional or procedural programming languages, preferably Java / Scala / Kotlin / Python\\nExperience designing and managing Infrastructure as Code with Helm and CRD, ensuring repeatable, secure, and scalable deployments.