Senior Technical Support Specialist

Apply Now โ†—
๐Ÿ“ Noida, India

About this role

Job Summary

GenAI & ML Operations Engineer

Role Summary

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.


Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.

Continuous Improvement & Enhancements

  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Root Cause Analysis
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Unsupported image type.

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • GCP workloads, GCP Deployments, Vertex AI, Gemini Enterprise Agent Development Kit, BigQuery, BigTable, Redis, Postgres
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.

Key Responsibilities

GenAI & ML Operations Engineer

Role Summary

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.


Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.

Continuous Improvement & Enhancements

  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Root Cause Analysis
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Unsupported image type.

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • GCP workloads, GCP Deployments, Vertex AI, Gemini Enterprise Agent Development Kit, BigQuery, BigTable, Redis, Postgres
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.

Skill Requirements

GenAI & ML Operations Engineer

Role Summary

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.


Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.

Continuous Improvement & Enhancements

  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Root Cause Analysis
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Unsupported image type.

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • GCP workloads, GCP Deployments, Vertex AI, Gemini Enterprise Agent Development Kit, BigQuery, BigTable, Redis, Postgres
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.

Other Requirements

None

Frequently Asked Questions

Is the salary disclosed for the Senior Technical Support Specialist position at HCLTech?
The salary for this Senior Technical Support Specialist role at HCLTech is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Senior Technical Support Specialist position at HCLTech located?
This Senior Technical Support Specialist role at HCLTech is based in Noida, India. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
How do I apply for the Senior Technical Support Specialist position at HCLTech?
Click the "Apply Now" button on this page. You will be redirected to HCLTech's official application portal hosted on successfactors where you can submit your application directly.
When was the Senior Technical Support Specialist job at HCLTech posted?
This Senior Technical Support Specialist position at HCLTech was posted on Oct 8, 2026. Apply as soon as possible โ€” early applications are often reviewed first.
Senior Technical Support Specialist
HCLTech
Apply for this role โ†—

You'll be redirected to HCLTech's official application page on successfactors.