Our Staff TechOps & Support Engineer Ensures the stability, performance, and availability of production systems by managing infrastructure, monitoring environments, responding to incidents, automating operational tasks, supporting deployments, and maintaining security, backups, and documentation Β
Responsibilities
- Ensure the availability, performance, and resilience of production environments by proactively monitoring systems and resolving complex operational issues.
- Administer and optimize CI/CD pipelines, containerized platforms, IBM Cloud Pak Β Stacks, Confluent, Elasticsearch, and supporting infrastructure to improve operational efficiency.
- Lead incident response, perform root cause analysis, and implement preventive actions to reduce recurring issues and improve service reliability.
- Develop and enhance monitoring, logging, alerting, automation, and operational processes to improve system observability and reduce manual effort.
- Provide technical leadership and mentorship to junior engineers while supporting deployments, release management, and production readiness activities.
- Maintain system security, backup, disaster recovery, and compliance standards while ensuring accurate operational documentation and runbooks.
- Collaborate with Development, Platform Engineering, QUALITY, Network, and Security teams to resolve complex technical challenges and continuously improve production operations.
- Bachelor's degree or Diploma in Computer Science, Engineering, or a related field.
- 4β6 years of experience in Technical Operations, Production Support, Site Reliability Engineering (SRE), DevOps, or System Administration.
- Strong experience with Linux/Windows administration, Docker, Kubernetes, OpenShift, CI/CD pipelines, automation tools, scripting (Bash, PowerShell, Python), and configuration management.
- Hands-on experience with IBM Cloud Pak Stacks (CP4BA, CP4I, CP4D), Confluent, Elasticsearch, Databases (SQL, DB2, MongoDB, etc.), JVM troubleshooting, monitoring, and observability platforms.
- Strong understanding of networking, security best practices, IAM, production support, backup, disaster recovery, scalability, performance tuning, and service reliability principles.
- Proven experience in incident management, troubleshooting, root cause analysis, and implementing operational improvements within enterprise production environments.
- Excellent technical leadership, analytical, communication, mentoring, and collaboration skills with the ability to manage multiple priorities effectively.