Senior AI Ops
2 days ago
Atlanta, GA, United States
NTT DATA Services
Full-time
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
By continuing, you agree to our Terms & Privacy Policy.
Req ID:
390771
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a
Senior AI Ops / DevOps Engineer (FTE / Hybrid)
to join our team in
Atlanta
,
Georgia (US-GA)
,
United States (US)
. Day to Day Job Duties The Senior AI Ops / DevOps Engineer will architect, build, and manage next-generation AI-driven CI/CD and cloud operations ecosystems. This role will go beyond traditional DevOps automation by integrating LLM agents, Model Context Protocol servers, intelligent observability, and secure AI-assisted workflows into the software delivery lifecycle. Architect, build, and manage AI-enabled CI/CD pipelines that improve developer productivity, code quality, release reliability, and deployment speed. Design and deploy production-grade Model Context Protocol clients and servers to securely connect enterprise LLMs with engineering tools, repositories, cloud infrastructure, and observability platforms. Develop custom MCP servers using Python, TypeScript, Node.js, or JavaScript to expose logs, infrastructure metrics, deployment data, and internal tools to authorized AI agents. Integrate LLM agents into developer workflows to support automated code review, vulnerability detection, test generation, release validation, and infrastructure recommendations. Build and maintain robust CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, ArgoCD, Jenkins, or similar tools. Implement ChatOps 2.0 capabilities that allow engineers to interact with deployment pipelines, cloud environments, logs, and operational workflows using secure conversational interfaces. Create safe autonomous remediation workflows for log analysis, incident triage, root-cause analysis, and infrastructure issue resolution. Build guardrails that allow AI agents to generate, inspect, and safely execute Infrastructure as Code using Terraform, OpenTofu, Terragrunt, Pulumi, Crossplane, or similar tools. Manage containerized workloads using Docker and Kubernetes platforms such as AWS EKS, Azure AKS, or Google GKE. Integrate AI-driven observability workflows with platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK. Implement AI safety controls including role-based access control, least-privilege execution, human-in-the-loop approvals, audit logging, rollback mechanisms, and secure tool access. Partner with software engineering, DevOps, SRE, security, platform, and data/AI teams to identify opportunities for intelligent automation. Create reusable automation frameworks, runbooks, dashboards, documentation, and enablement materials for engineering teams. Drive an “automate everything” culture by reducing manual toil and improving operational efficiency across cloud and software delivery processes. Basic Qualifications Minimum 7+ years of experience in DevOps, Cloud Engineering, SRE, Platform Engineering, or Infrastructure Automation. Minimum 4+ years of hands-on experience designing and managing CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD, or similar platforms. Minimum 3+ years of experience managing scalable cloud environments in AWS, Azure, or GCP, with strong preference for AWS. Strong hands-on experience with Kubernetes, Docker, and production container orchestration platforms such as EKS, AKS, or GKE. Advanced proficiency with Infrastructure as Code tools such as Terraform, OpenTofu, Terragrunt, Pulumi, CloudFormation, or Crossplane. Strong programming and scripting experience using Python, TypeScript, JavaScript, Bash, or Go. Practical experience working with LLM APIs such as OpenAI, Anthropic, or similar enterprise AI platforms. Experience with AI orchestration or agentic frameworks such as LangChain, CrewAI, LlamaIndex, or similar tools. Strong understanding of the Model Context Protocol ecosystem and experience designing or integrating MCP clients and servers. Experience integrating DevSecOps controls into CI/CD pipelines, including SAST, DAST, dependency scanning, container scanning, secrets scanning, and vulnerability management. Strong knowledge of secret management and security tooling such as HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, or similar platforms. Experience with observability, monitoring, logging, and alerting platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK. Familiarity with security and compliance frameworks such as SOC2, ISO27001, or enterprise audit control environments. Ability to troubleshoot complex pipeline, infrastructure, deployment, and production issues across cloud-native environments. Preferred / Nice to Have Experience building AI-assis
390771
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a
Senior AI Ops / DevOps Engineer (FTE / Hybrid)
to join our team in
Atlanta
,
Georgia (US-GA)
,
United States (US)
. Day to Day Job Duties The Senior AI Ops / DevOps Engineer will architect, build, and manage next-generation AI-driven CI/CD and cloud operations ecosystems. This role will go beyond traditional DevOps automation by integrating LLM agents, Model Context Protocol servers, intelligent observability, and secure AI-assisted workflows into the software delivery lifecycle. Architect, build, and manage AI-enabled CI/CD pipelines that improve developer productivity, code quality, release reliability, and deployment speed. Design and deploy production-grade Model Context Protocol clients and servers to securely connect enterprise LLMs with engineering tools, repositories, cloud infrastructure, and observability platforms. Develop custom MCP servers using Python, TypeScript, Node.js, or JavaScript to expose logs, infrastructure metrics, deployment data, and internal tools to authorized AI agents. Integrate LLM agents into developer workflows to support automated code review, vulnerability detection, test generation, release validation, and infrastructure recommendations. Build and maintain robust CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, ArgoCD, Jenkins, or similar tools. Implement ChatOps 2.0 capabilities that allow engineers to interact with deployment pipelines, cloud environments, logs, and operational workflows using secure conversational interfaces. Create safe autonomous remediation workflows for log analysis, incident triage, root-cause analysis, and infrastructure issue resolution. Build guardrails that allow AI agents to generate, inspect, and safely execute Infrastructure as Code using Terraform, OpenTofu, Terragrunt, Pulumi, Crossplane, or similar tools. Manage containerized workloads using Docker and Kubernetes platforms such as AWS EKS, Azure AKS, or Google GKE. Integrate AI-driven observability workflows with platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK. Implement AI safety controls including role-based access control, least-privilege execution, human-in-the-loop approvals, audit logging, rollback mechanisms, and secure tool access. Partner with software engineering, DevOps, SRE, security, platform, and data/AI teams to identify opportunities for intelligent automation. Create reusable automation frameworks, runbooks, dashboards, documentation, and enablement materials for engineering teams. Drive an “automate everything” culture by reducing manual toil and improving operational efficiency across cloud and software delivery processes. Basic Qualifications Minimum 7+ years of experience in DevOps, Cloud Engineering, SRE, Platform Engineering, or Infrastructure Automation. Minimum 4+ years of hands-on experience designing and managing CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD, or similar platforms. Minimum 3+ years of experience managing scalable cloud environments in AWS, Azure, or GCP, with strong preference for AWS. Strong hands-on experience with Kubernetes, Docker, and production container orchestration platforms such as EKS, AKS, or GKE. Advanced proficiency with Infrastructure as Code tools such as Terraform, OpenTofu, Terragrunt, Pulumi, CloudFormation, or Crossplane. Strong programming and scripting experience using Python, TypeScript, JavaScript, Bash, or Go. Practical experience working with LLM APIs such as OpenAI, Anthropic, or similar enterprise AI platforms. Experience with AI orchestration or agentic frameworks such as LangChain, CrewAI, LlamaIndex, or similar tools. Strong understanding of the Model Context Protocol ecosystem and experience designing or integrating MCP clients and servers. Experience integrating DevSecOps controls into CI/CD pipelines, including SAST, DAST, dependency scanning, container scanning, secrets scanning, and vulnerability management. Strong knowledge of secret management and security tooling such as HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, or similar platforms. Experience with observability, monitoring, logging, and alerting platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK. Familiarity with security and compliance frameworks such as SOC2, ISO27001, or enterprise audit control environments. Ability to troubleshoot complex pipeline, infrastructure, deployment, and production issues across cloud-native environments. Preferred / Nice to Have Experience building AI-assis