Staff Site Reliability Engineer
4 weeks ago
We are seeking a highly skilled Staff Site Reliability Engineer to join our Data Engineering team. As a key member of our team, you will be responsible for maintaining and enhancing the reliability of our data infrastructure.
Your work will directly impact the availability and performance of our data services, enabling the organization to make better decisions.
You will collaborate closely with data engineers and software engineers to develop and drive 100% automation, best practices for deep monitoring and alerting.
This role will report to our Director of Data Engineering.
About You- Bachelor's degree in Computer Science, Information Technology, or a related field.
- 12+ years of experience in site reliability engineering, database operations, or a related role with a focus on data platforms, data stores, data operations.
- Extensive experience with AWS cloud platform and their data-related services.
- Proficiency in monitoring tools (e.g., Datadog, CloudWatch, DevOps Guru, DB Performance Insights).
- Proficiency in one or more programming languages (e.g. Python, Java).
- Proficiency in automation frameworks (e.g., Terraform, Cloud Formation).
- Strong understanding of various performance metrics both at a high level and at a low level like Disk/IO saturation.
- Experience in identifying and eliminating the bottlenecks in the system.
- Strong understanding of database internals like types of indexes, schemas, query plans.
- Strong understanding of database systems (e.g., SQL, NoSQL) and experience in managing large-scale data infrastructures.
- Strong understanding and hands-on implementation of CI/CD pipelines and DataOps practices.
- Experience with data governance, compliance, and lifecycle management.
- Ability to own and execute projects while effectively collaborating with the team to influence and shape the vision of the data engineering organization.
We value:
- Courage. We believe that when we overcome fear, we enable our best selves.
- Curiosity. We are curious, which is the gateway to empathy, inclusion, and understanding.
- Service. We serve our community with humility, enabling joy and belonging for others.
- Kaizen. We have a growth mindset committed to constant forward progress.
We are an equal opportunity employer and value diversity at Crunchyroll. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
-
Staff Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Crunchyroll Full timeAbout CrunchyrollWe're a global entertainment company dedicated to delivering the art and culture of anime to a passionate community. Our mission is to help everyone belong, and we're looking for talented individuals to join our team.The RoleWe're seeking a Staff Site Reliability Engineer to maintain and enhance the reliability of our data infrastructure. As...
-
Cloud Platform Staff Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Zilliz Full timeJob Title: Cloud Platform Staff Site Reliability EngineerWe are seeking a highly skilled Cloud Platform Staff Site Reliability Engineer to join our team at Zilliz. As a key member of our SRE team, you will be responsible for ensuring the reliability, availability, and performance of our distributed database systems.Key Responsibilities:Design and build tools...
-
Senior Staff Site Reliability Engineer
3 weeks ago
San Francisco, California, United States WEX Full timeAbout the RoleThe WEX Site Reliability Engineering team is seeking a technical leader to drive the design and implementation of complex systems at scale. As a Senior Staff SRE, you will work closely with engineering teams to ensure that our systems are reliable, performant, and secure.Key ResponsibilitiesProvide technical guidance and mentorship to other...
-
Senior Staff Site Reliability Engineer
3 weeks ago
San Francisco, California, United States WEX Full timeThe WEX Site Reliability Engineering team is seeking a Senior Staff SRE who is passionate about developing software and solutions focused on observability, incident response, reliability, and performance.The team will be part of the Benefits Reliability organization which supports our internal stakeholders and our Benefits Platform teams.As part of the...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Unreal Gigs Full timeJob Title: Site Reliability EngineerAt Unreal Gigs, we're seeking a highly skilled Site Reliability Engineer to join our team. As a Site Reliability Engineer, you will be responsible for ensuring the high availability, scalability, and performance of our complex distributed systems.Key Responsibilities:Design and implement monitoring, logging, and alerting...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Unreal Gigs Full timeJob Title: Site Reliability EngineerAt Unreal Gigs, we're seeking a highly skilled Site Reliability Engineer to join our team. As a Site Reliability Engineer, you will be responsible for ensuring the high availability, scalability, and performance of our complex distributed systems.Key Responsibilities:Design and implement monitoring, logging, and alerting...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States DaVita Full timeAbout the RoleThe WEX Site Reliability Engineering team is seeking a skilled Site Reliability Engineer to join our Platform Reliability organization. As a key member of our team, you will be responsible for developing software and solutions focused on observability, incident response, reliability, and performance.You will collaborate with our engineering...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Instabase Full timeAbout InstabaseAt Instabase, we're passionate about harnessing the power of AI innovation to democratize access to cutting-edge technology and empower organizations to solve complex unstructured data problems. With a strong presence in the market and a talented team, we're committed to delivering top-tier solutions that drive business success.Job...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Instabase Full timeAbout InstabaseInstabase is a global company with offices in San Francisco, New York, London, and Bengaluru. We're a people-first organization that values experimentation, curiosity, and customer obsession.Job SummaryWe're seeking a Site Reliability Engineer to join our Site Reliability and Platform Engineering team. As a key member of our team, you'll be...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Withorb Full timeAbout UsOrb is a cutting-edge technology company on a mission to revolutionize the way businesses approach revenue growth. Our team is passionate about building a robust infrastructure that enables our customers to unlock their full potential.Job DescriptionWe are seeking a highly skilled Site Reliability Engineer to join our team. As a key member of our...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Outdefine Full timeAbout the JobWe are seeking a highly skilled Site Reliability Engineer to join our team at Outdefine. As a key member of our engineering team, you will be responsible for ensuring the reliability, scalability, and performance of our ecommerce platform.Key ResponsibilitiesDesign and implement scalable and highly available cloud infrastructure using Kubernetes...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Roman Health Pharmacy LLC Full timeAbout the RoleWe are seeking a highly skilled Site Reliability Engineer to join our team at Xero. As a key member of our Reliability Enablement team, you will play a critical role in ensuring the reliability and performance of our systems.Key ResponsibilitiesInvestigate operational surprises and support teams in post-incident activitiesConduct in-depth...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States YO HR CONSULTANCY Full timeJob Title: Site Reliability EngineerJob Description:At YO HR CONSULTANCY, we are seeking a highly skilled Site Reliability Engineer to join our team.Key Responsibilities:* Extensive experience working with Linux flavors like RHEL/CentOS OS, shells, filesystems, and utilities* Knowledge of distributed computing and experience working with container...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States SpeedCast Full timeJob Title: Site Reliability EngineerAt Speedcast, we're seeking a highly skilled Site Reliability Engineer to join our team. As a Site Reliability Engineer, you will play a critical role in ensuring the reliability, scalability, and performance of our communication products.Key Responsibilities:Analyze and design continuous integration/continuous delivery...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Orb Full timeAbout the RoleOrb is seeking a skilled Site Reliability Engineer to join our team. As a key member of our engineering organization, you will play a critical role in maintaining and scaling our robust infrastructure, ensuring stability, scalability, and performance.You will be responsible for tackling complex engineering challenges, from scaling our data...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Swish Analytics Full time{"h1": "Site Reliability Engineer at Swish Analytics"} Swish Analytics is a sports analytics and betting startup that's revolutionizing the industry with cutting-edge predictive data products. We're on a mission to make oddsmaking a challenge rooted in engineering, mathematics, and sports betting expertise, not intuition. We're looking for a team-oriented...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States GRNET Full timeGRNET is seeking a highly skilled Site Reliability Engineer to join its team. As an SRE, you will be responsible for designing and implementing fault-tolerant, scalable, and distributed services. You will work closely with the team to bring your technical opinion and vision to the table, handle problems that require under-the-hood investigation, and lead...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Perplexity Full timeSite Reliability EngineerPerplexity is seeking a highly skilled Site Reliability Engineer to join our team in revolutionizing the way people interact with the internet. As a key member of our infrastructure team, you will be responsible for designing, implementing, and scaling the systems that support our web and mobile products.Key ResponsibilitiesDesign...
-
Site Reliability Engineer
4 weeks ago
San Francisco, California, United States Xai Full timeAbout xAIxAI is a cutting-edge technology company that specializes in developing large-scale, highly-reliable distributed systems. Our team of software engineers is passionate about building high-quality software and tackling complex technical challenges.The RoleWe are seeking an experienced Site Reliability Engineer to join our dynamic team in London. As a...
-
Site Reliability Engineer
3 weeks ago
San Francisco, California, United States WEX Full timeJob SummaryThe WEX Site Reliability Engineering team is seeking a highly motivated and quick-learning individual to join our team as a Site Reliability Engineer Level 1. As a key member of our team, you will be responsible for ensuring the reliability, performance, and security of our systems.Key Responsibilities:Actively participate in training and...