Observability Engineer

7 days ago

Phoenix, AZ, United States SN Cloud Solutions Full-time

Job Description

We are looking for an experienced Observability Engineer to design, implement, and maintain enterprise observability solutions across cloud-native and distributed environments. The ideal candidate should have strong hands-on experience with Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and modern observability practices.

Key Responsibilities

  • Design, implement, and maintain enterprise-wide observability and monitoring solutions.
  • Configure and administer Dynatrace for application performance monitoring, infrastructure monitoring, distributed tracing, and real-user monitoring.
  • Develop and manage Splunk dashboards, alerts, searches, and monitoring solutions.
  • Work with OpenSearch/Elasticsearch for centralized logging, log analytics, indexing, and visualization.
  • Monitor Kubernetes clusters, containers, microservices, and cloud-native applications.
  • Implement monitoring for application metrics, logs, traces, and events.
  • Develop and maintain dashboards, alerts, SLOs, SLIs, and operational metrics.
  • Troubleshoot application, infrastructure, and performance issues using observability tools.
  • Configure APM, distributed tracing, log aggregation, and infrastructure monitoring.
  • Integrate observability platforms with CI/CD pipelines and cloud-native environments.
  • Work with development and DevOps teams to identify performance bottlenecks and production issues.
  • Automate monitoring and observability tasks using Python, Shell scripting, or similar technologies.
  • Support incident management, root-cause analysis, and production troubleshooting.
  • Implement observability best practices across Linux, Kubernetes, microservices, APIs, and cloud platforms.
  • Ensure monitoring solutions provide actionable insights for application availability, performance, and reliability.

Required Skills

  • Strong experience as an Observability Engineer / SRE / Monitoring Engineer / APM Engineer.
  • Hands-on experience with Dynatrace.
  • Strong experience with Splunk.
  • Experience with OpenSearch and/or Elasticsearch.
  • Strong knowledge of Kubernetes and containerized environments.
  • Good Linux administration and troubleshooting skills.
  • Experience with cloud-native observability.
  • Strong understanding of logs, metrics, traces, APM, and distributed tracing.
  • Experience creating monitoring dashboards, alerts, and performance reports.
  • Knowledge of REST APIs and scripting/automation.
  • Experience with microservices and distributed applications.

Preferred Skills

  • Experience with AWS, Azure, or Google Cloud Platform.
  • Knowledge of Prometheus and Grafana.
  • Experience with OpenTelemetry (OTel).
  • Knowledge of Docker and CI/CD pipelines.
  • Experience with Python, Shell, or PowerShell scripting.
  • Understanding of SRE practices, SLI/SLO/SLA, and error budgets.
  • Experience troubleshooting high-volume production environments.