Track Lead (Support & Operations)
About this role
Job Summary
Site Reliability Engineer (SRE) Lot 3 - HCBU | L2 - Senior SRE Engineer Experience Operating Model Focus Reporting 5-8 Years 24x7 Reliability Engineering Operations and Transformation SRE Lead / Reliability Architect ROLE PURPOSE Drive service reliability, resilience, observability, and automation across cloud, infrastructure, platform, and application services. Apply SRE practices to improve availability, reduce operational toil, accelerate recovery, and support the transition toward proactive and autonomous operations.
Key Responsibilities
KEY RESPONSIBILITIES Reliability and SLO Management Define and maintain SLIs, SLOs, error budgets, golden signals, and service-health measures for critical services. Monitor SLO performance and error-budget consumption, identify reliability risks, and coordinate corrective actions with service owners. Contribute to service classification, availability targets, operational readiness, and reliability improvement roadmaps. Observability and AIOps Implement and enhance full-stack observability across applications, APIs, cloud, infrastructure, Kubernetes, databases, and network services. Build dashboards, service maps, alerts, synthetic checks, log analytics, tracing, and dependency views using enterprise monitoring platforms. Improve signal quality through event correlation, alert tuning, noise reduction, anomaly detection, and actionable routing. Automation and Toil Reduction Identify repetitive operational activities and develop automated runbooks, health checks, self-service workflows, and self-healing solutions. Create reliable automation using Python, PowerShell, Shell, Terraform, Ansible, CI/CD, and platform-native capabilities. Ensure automation includes testing, approvals, audit evidence, exception handling, rollback, and human oversight where required. Incident, Problem, and Resilience Engineering Support P1/P2 incident response, technical triage, service restoration, and cross-tower coordination. Lead or contribute to RCA, post-incident reviews, problem investigations, corrective actions, and recurrence-prevention measures. Conduct capacity reviews, failover validation, DR exercises, resilience testing, performance analysis, and single-point-of-failure assessments. Governance and Continuous Improvement Prepare reliability reports covering SLO compliance, MTTR, incidents, recurring failures, telemetry gaps, automation benefits, and technical debt. Participate in service reliability reviews and maintain evidence for agreed reliability and operational controls. Collaborate with Infrastructure, Cloud, DevOps, Database, Network, Security, Application, and ITSM teams to embed reliability by design.
Skill Requirements
REQUIRED SKILLS AND EXPERIENCE Hands-on experience in SRE, production engineering, platform operations, cloud operations, or infrastructure reliability. Practical knowledge of SLI/SLO definition, error budgets, availability, latency, throughput, saturation, and service-health measurement. Experience with Datadog, Splunk, Grafana, Prometheus, AppDynamics, Dynatrace, Azure Monitor, GCP Operations Suite, or similar tools. Operational experience with Azure and/or GCP, Kubernetes/container platforms, Linux, and Windows environments. Automation and scripting skills in Python, PowerShell, Bash, or Shell, with exposure to APIs and Git-based version control. Knowledge of incident, problem, change, availability, capacity, and service-continuity processes using ServiceNow and ITIL practices. Strong troubleshooting, analytical, documentation, stakeholder communication, and cross-functional collaboration skills. PREFERRED SKILLS Terraform and Ansible AIOps, event correlation, and automated remediation Chaos engineering and resilience testing Azure DevOps, GitHub Actions, GitLab CI/CD, or Jenkins Distributed tracing, real-user monitoring, and synthetic monitoring Cloud-native architecture, FinOps awareness, and security-by-design practices KEY DELIVERABLES AND SUCCESS MEASURES Approved SLI/SLO definitions, service-health dashboards, and error-budget reporting for assigned services. Reduced alert noise, detection time, restoration time, repeat incidents, and manual operational effort. Increased adoption of automated runbooks, self-healing, proactive detection, and preventive remediation. Improved service availability, resilience, capacity readiness, observability coverage, and SLO compliance. Complete RCA, post-incident actions, operational documentation, and auditable evidence for reliability controls.
Other Requirements
PREFERRED CERTIFICATIONS SRE Foundation | Certified Kubernetes Administrator (CKA) | Microsoft Azure Administrator / Architect | Google Professional Cloud DevOps Engineer | Datadog Certification | ITIL v4
Frequently Asked Questions
Is the salary disclosed for the Track Lead (Support & Operations) position at HCLTech?
Where is the Track Lead (Support & Operations) position at HCLTech located?
How do I apply for the Track Lead (Support & Operations) position at HCLTech?
When was the Track Lead (Support & Operations) job at HCLTech posted?
You'll be redirected to HCLTech's official application page on successfactors.