Observability Engineer
7 days ago
Phoenix, AZ, United States
SN Cloud Solutions
Full-time
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
By continuing, you agree to our Terms & Privacy Policy.
Job Description
We are looking for an experienced Observability Engineer to design, implement, and maintain enterprise observability solutions across cloud-native and distributed environments. The ideal candidate should have strong hands-on experience with Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and modern observability practices.
Key Responsibilities
- Design, implement, and maintain enterprise-wide observability and monitoring solutions.
- Configure and administer Dynatrace for application performance monitoring, infrastructure monitoring, distributed tracing, and real-user monitoring.
- Develop and manage Splunk dashboards, alerts, searches, and monitoring solutions.
- Work with OpenSearch/Elasticsearch for centralized logging, log analytics, indexing, and visualization.
- Monitor Kubernetes clusters, containers, microservices, and cloud-native applications.
- Implement monitoring for application metrics, logs, traces, and events.
- Develop and maintain dashboards, alerts, SLOs, SLIs, and operational metrics.
- Troubleshoot application, infrastructure, and performance issues using observability tools.
- Configure APM, distributed tracing, log aggregation, and infrastructure monitoring.
- Integrate observability platforms with CI/CD pipelines and cloud-native environments.
- Work with development and DevOps teams to identify performance bottlenecks and production issues.
- Automate monitoring and observability tasks using Python, Shell scripting, or similar technologies.
- Support incident management, root-cause analysis, and production troubleshooting.
- Implement observability best practices across Linux, Kubernetes, microservices, APIs, and cloud platforms.
- Ensure monitoring solutions provide actionable insights for application availability, performance, and reliability.
Required Skills
- Strong experience as an Observability Engineer / SRE / Monitoring Engineer / APM Engineer.
- Hands-on experience with Dynatrace.
- Strong experience with Splunk.
- Experience with OpenSearch and/or Elasticsearch.
- Strong knowledge of Kubernetes and containerized environments.
- Good Linux administration and troubleshooting skills.
- Experience with cloud-native observability.
- Strong understanding of logs, metrics, traces, APM, and distributed tracing.
- Experience creating monitoring dashboards, alerts, and performance reports.
- Knowledge of REST APIs and scripting/automation.
- Experience with microservices and distributed applications.
Preferred Skills
- Experience with AWS, Azure, or Google Cloud Platform.
- Knowledge of Prometheus and Grafana.
- Experience with OpenTelemetry (OTel).
- Knowledge of Docker and CI/CD pipelines.
- Experience with Python, Shell, or PowerShell scripting.
- Understanding of SRE practices, SLI/SLO/SLA, and error budgets.
- Experience troubleshooting high-volume production environments.