Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure

Apply Now ↗
📍 Chennai, India

About this role

Job Summary

Microsoft Azure, Red Hat OpenShift, Red Hat Enterprise Linux, Ansible, Red Hat Cluster, Kubernetes Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Key Responsibilities

Microsoft Azure, Red Hat OpenShift, Red Hat Enterprise Linux, Ansible, Red Hat Cluster, Kubernetes Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Skill Requirements

Microsoft Azure, Red Hat OpenShift, Red Hat Enterprise Linux, Ansible, Red Hat Cluster, Kubernetes Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Other Requirements

Microsoft Azure, Red Hat OpenShift, Red Hat Enterprise Linux, Ansible, Red Hat Cluster, Kubernetes Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Frequently Asked Questions

Is the salary disclosed for the Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure position at HCLTech?
The salary for this Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure role at HCLTech is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure position at HCLTech located?
This Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure role at HCLTech is based in Chennai, India. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
How do I apply for the Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure position at HCLTech?
Click the "Apply Now" button on this page. You will be redirected to HCLTech's official application portal hosted on successfactors where you can submit your application directly.
When was the Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure job at HCLTech posted?
This Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure position at HCLTech was posted on Sep 23, 2026. Apply as soon as possible — early applications are often reviewed first.
Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure
HCLTech
Apply for this role ↗

You'll be redirected to HCLTech's official application page on successfactors.