Senior Engineer

13 hours ago

Los Angeles, CA, United States United States Digital Space LLC Full-time

the company helps people stay better connected with the things they love. the company’s edge cloud platform enables customers to create great digital experiences quickly, securely, and reliably by processing, serving, and securing our customers’ applications as close to their end-users as possible — at the edge of the Internet. The platform is designed to take advantage of the modern internet, to be programmable, and to support agile software development. the company’s customers include many of the world’s most prominent companies, including GitHub, Yelp, Paramount, and JetBlue.

We're building a more trustworthy Internet. Come join us.

**Posting Open Date: Sept. 25, 2026Anticipated Posting Close Date*: Oct. 23, 2026Job posting may close early due to the volume of applicants.*

Senior Engineer - Production Cloud and Container Services

Fleet Operations and Production Engineering at the company is looking for a Senior Engineer to join our Cloud and Container Services team. This role is focused on helping to scale and manage the company’s Kubernetes based platform for control plane services. This platform is built on top of multiple public cloud services and contains many Kubernetes ecosystem components. We’re working to scale out our platform to support growth and at the same time evolve to address new business priorities. A successful candidate will help expand our platform feature set, support existing users and onboard new services, and drive efficiency while maintaining a secure platform.

What You'll Do:

  • Design, build, and operate infrastructure (cloud, the company datacenter) to enable reliable and rapid deployment. This includes effective monitoring and resilient operations in a large-scale multi-cloud and hybrid cloud / on-premise environment. The majority of workloads are containerized but some are using native cloud services such as compute and storage.
  • Diagnose and resolve performance and reliability issues across the stack: application, operating system, network, 3rd party services and APIs, including cross-application dependencies.
  • Deploy and support complex 3rd party and internally developed applications.
  • Write tools to automate maintenance and deployment of servers, services, and applications.
  • Collaborate with internal users and continually evolve the platform and its operations using solid engineering practices.
  • Drive projects sometimes independently and sometimes collaboratively across time zones.
  • Configure access and manage operations within a multi-cloud environment.

What We're Looking For:

  • Experience running high availability systems and supporting distributed infrastructure. Most Senior level Engineers at the company have more than 5 years of related experience.
  • You have designed services with fault tolerance and across geographies.
  • You have deployed and managed multi-tiered services.
  • Understanding of Linux systems, high and low level. You have used tcpdump and tracing tools.
  • Experience building and operating production-grade kubernetes clusters in multiple regions, clouds or data centers. You have experience with tooling in the CNCF space, such as: prometheus, flux, helm, etc.
  • Experience with programming languages such as Go and Python. You can read code and reason about what it does. You can write code within an existing large code base such as adding features. You can create medium-sized programs from scratch such as custom kubernetes controllers, custom prometheus exporters, and building tooling to help manage infrastructure.
  • Experience with infrastructure and configuration management tooling. You have used Terraform to manage infrastructure.
  • Experience provisioning and managing users and resources with cloud providers such as AWS and GCP. You have provisioned users and accounts using both graphical user interfaces and infrastructure as code frameworks. You have experience provisioning components on public cloud and understand how they work together in creating a multi-tiered service. You have deployed and managed services built on top of public cloud components such as EC2, S3, and GKE.
  • Experience with CI/CD and GitOps tooling. You can iterate infrastructure via pipelines through code changes. You have experience using Github including creating and reviewing pull requests.
  • Experience with monitoring tools such as Prometheus, Datadog and Grafana.
  • Experience working on a distributed team. You have experience collaborating across timezones. You can articulate challenges of distributed teams and how you mitigate them.
  • Experience leveraging AI tooling to accelerate your work and enable new capabilities. You