DevOps Engineer

5 days ago

United States Switch Intelligence Inc. Full-time
About Switch   Switch Intelligence, Inc. is an early-stage company building the platform that runs a company's commercial motion across marketing, sales, and customer operations. We are industry agnostic, and our buyers are companies between $250 million and $3 billion in revenue in manufacturing, healthcare, business services, distribution, and technology alike. Switch is based in NYC and founded by an operator who spent the last 20 years leading commercial organizations across the world's largest technology companies.  We are looking for a strong DevOps Engineer to build the cloud, deployment, and operational foundation of the Switch Intelligence Platform — a multi-tenant, source-agnostic data streaming and intelligence system . You will stand up infrastructure as code on Google Cloud Platform , build CI/CD for application and data services, and establish secrets management, monitoring, and cost controls that a small engineering team can operate without a dedicated platform function. The platform runs on Docker and Kubernetes (GKE) and includes PostgreSQL , a Neo4j knowledge graph , event streaming, and a lakehouse layer. Engagement Structure   This is a phased contract engagement. Cloud account and project structure, CI/CD for the application layer, environment separation, secrets management, and an observability baseline.  Infrastructure as code for the data platform, orchestration deployment, pipeline deployment, and cost controls.  Phase 3 — ongoing retainer, 8 to 16 hours per month. Design and implement cloud infrastructure as code on Google Cloud Platform using Terraform, Pulumi, or equivalent, with reproducible environments rather than hand-configured resources.  Build and maintain CI/CD pipelines for application services, data services, and infrastructure changes.  Containerize services and deploy them to Kubernetes (GKE) , including workload configuration, autoscaling, and resource limits.  Implement secrets management and access controls scoped to least privilege across environments.  Establish environment separation across development, staging, and production, including data isolation between environments.  Deploy and operate infrastructure for PostgreSQL , Neo4j , event streaming (Kafka or Google Pub/Sub), Redis , and object storage.  Deploy and operate orchestration infrastructure for data pipelines using Airflow, Dagster, Prefect, or similar.  Implement monitoring, logging, alerting, and tracing that surfaces production issues and suppresses noise.  Implement cost controls, budget alerting, and resource tagging across GCP projects.  Design and document backup, restore, and disaster recovery procedures for stateful services.  Build network configuration including VPC design, private service access, ingress, and TLS termination.  Produce runbooks and architecture documentation as a deliverable of each phase, and conduct a recorded handover walkthrough with the engineering team.  Collaborate closely with the Data Engineering and Backend functions on deployment requirements and environment needs.  Troubleshoot complex production infrastructure and deployment issues.  5+ years of professional DevOps, Platform, or Site Reliability Engineering experience .  ~ Demonstrated experience building production infrastructure from zero at an early-stage company, rather than maintaining an existing platform.  ~ Hands-on experience with Google Cloud Platform , or equivalent depth in AWS or Azure with demonstrated ability to transfer to GCP.  ~ Experience building CI/CD pipelines with GitHub Actions, Jenkins, GitLab CI, Bitbucket Pipelines, or similar.  ~ Experience deploying and operating stateful data infrastructure , including relational databases and message queues or streaming platforms.  ~ Experience with secrets management tooling such as Google Secret Manager, HashiCorp Vault, or similar.  ~ Experience with monitoring and observability tooling such as Prometheus, Grafana, Datadog, Google Cloud Operations, or similar.  ~ Working knowledge of Linux systems administration and networking fundamentals.  ~ Strong scripting ability in Python, Bash, or Go .  ~ Strong understanding of Git and software development workflows.  ~ Demonstrated ability to write clear technical documentation in English.  ~ Experience deploying and operating Neo4j or another graph database in production.  Experience operating Kafka , Google Pub/Sub , or similar event streaming infrastructure at production scale.  Experience deploying data orchestration frameworks such as Airflow, Dagster, or Prefect.  Experience with lakehouse or warehouse infrastructure such as BigQuery, Snowflake, or Databricks.  Experience with Ansible, SonarQube, Artifactory , or similar DevOps tooling.  Experience with FinOps practices and cloud cost optimization.  You should be comfortable taking an early-stage codebase with no deployment story and deliv