SRE Operations Lead

18 hours ago

Charlotte, North Carolina, United States Tcs Usa Full-time $110,000 - $120,000 Contract

SRE Operations Lead - AWS, GitLab & AIOps

Must Have Technical/Functional Skills

• SRE / Application Reliability Engineering (ARE) and 24x7 production operations

• Major Incident Management (P1/P2), ServiceNow, and Incident / Problem / Change Management

• SLI, SLO, SLA governance; MTTR reduction; service reliability KPIs

• AWS: CloudWatch, Route 53, S3, CloudFront, Lambda, ECR, and Bedrock

• Observability: Dynatrace, Grafana, Splunk, and CloudWatch

• Control-M and enterprise batch operations

• GitLab CI/CD, REST APIs, Personal Access Tokens, security scanning, and DevSecOps

• Python GitLab library and API-based automation

• Automation, AIOps, event correlation, self-healing, and stakeholder management

Roles & Responsibilities

• Lead 24x7 SRE operations and coordinate P1/P2 major incident response through service restoration and follow-up.

• Own reliability measures including SLIs, SLOs, SLAs, MTTR improvement, and service reliability KPIs.

• Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, and Splunk.

• Lead automation, AIOps, self-healing, event correlation, and AI-driven operations initiatives.

• Oversee AWS platform operations, batch processing, and Control-M environments.

• Integrate Claude AI on AWS Bedrock with GitLab using APIs, PATs, and custom workflows.

• Develop AI-driven analysis of GitLab project data, vulnerabilities, pipelines, and security findings.

• Design GitLab API automation, custom workflows, DevSecOps controls, and CI/CD pipeline improvements.

• Build and maintain Python-based GitLab integrations and REST API solutions.

• Support vulnerability remediation and onboarding/configuration of security scanning tools.

• Manage AWS Lambda, ECR, and Bedrock for deployment and automation; optimize Lambda configuration, concurrency, and scaling.

• Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR.

Generic Managerial Skills

• Lead geographically distributed operations teams and coordinate effectively during critical incidents.

• Communicate reliability risks, service performance, and remediation plans to technical and business stakeholders.

• Drive governance, prioritization, continuous improvement, and cross-team collaboration.

• Mentor engineers and promote automation-first, blameless, and reliability-focused ways of working.