SME - Kubernetes, Terraform

Apply Now โ†—
๐Ÿ“ Mississauga, Canada

About this role

Job Summary

AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.

Key Responsibilities

AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.

Skill Requirements

AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.

Other Requirements

AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.

Frequently Asked Questions

Is the salary disclosed for the SME - Kubernetes, Terraform position at HCLTech?
The salary for this SME - Kubernetes, Terraform role at HCLTech is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the SME - Kubernetes, Terraform position at HCLTech located?
This SME - Kubernetes, Terraform role at HCLTech is based in Mississauga, Canada. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
How do I apply for the SME - Kubernetes, Terraform position at HCLTech?
Click the "Apply Now" button on this page. You will be redirected to HCLTech's official application portal hosted on successfactors where you can submit your application directly.
When was the SME - Kubernetes, Terraform job at HCLTech posted?
This SME - Kubernetes, Terraform position at HCLTech was posted on Sep 17, 2026. Apply as soon as possible โ€” early applications are often reviewed first.
SME - Kubernetes, Terraform
HCLTech
Apply for this role โ†—

You'll be redirected to HCLTech's official application page on successfactors.