Site Reliability Engineer

Apply Now โ†—
๐Ÿ“ Barrington, United StatesContract
IT Services

About this role

We are seeking an experienced Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, secure, and reliable production systems. The ideal candidate will have strong software engineering skills combined with hands-on experience in cloud infrastructure, Kubernetes, automation, monitoring, and observability.

The SRE will work closely with development, infrastructure, and operations teams to improve system reliability, automate operational processes, and resolve complex production issues.



Requirements

  • Design, deploy, and maintain highly reliable and scalable production systems.
  • Develop automation and reliability tooling using Go, Python, Java, or Rust.
  • Manage and support cloud environments across AWS, Azure, or GCP.
  • Deploy and manage containerized applications using Docker and Kubernetes.
  • Implement and maintain monitoring, logging, metrics, and distributed tracing solutions.
  • Build and enhance observability using OpenTelemetry (OTel) and related technologies.
  • Troubleshoot complex infrastructure, application, and production issues.
  • Participate in incident response, root-cause analysis, and post-incident reviews.
  • Automate repetitive operational tasks and improve engineering efficiency.
  • Implement Infrastructure as Code using Terraform and/or Ansible.
  • Develop and maintain CI/CD pipelines for reliable and automated deployments.
  • Monitor system performance, availability, capacity, and overall reliability.
  • Identify reliability risks and implement proactive solutions.
  • Collaborate with software developers to improve application reliability and performance.
  • Establish and improve SRE best practices, operational procedures, and reliability standards.

Required Skills & Experience

  • Proven experience as a Site Reliability Engineer, Production Engineer, DevOps Engineer, or similar role.
  • Strong programming experience with at least one of:
    • Go/Golang
    • Python
    • Java
    • Rust
  • Hands-on experience with AWS, Azure, or GCP.
  • Strong experience with Kubernetes and Docker.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with OpenTelemetry and observability.
  • Knowledge of monitoring and visualization tools such as Prometheus and Grafana.
  • Experience with Terraform, Ansible, or similar Infrastructure-as-Code tools.
  • Strong understanding of CI/CD pipelines and DevOps practices.
  • Experience with production incident management and Root Cause Analysis (RCA).
  • Strong knowledge of automation, scripting, networking, and distributed systems.



Frequently Asked Questions

Is the salary disclosed for the Site Reliability Engineer position at workiy?
The salary for this Site Reliability Engineer role at workiy is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Site Reliability Engineer position at workiy located?
This Site Reliability Engineer role at workiy is based in Barrington, United States. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Site Reliability Engineer role at workiy full-time or part-time?
This is listed as a Contract position. It is posted as a Site Reliability Engineer role at workiy.
How do I apply for the Site Reliability Engineer position at workiy?
Click the "Apply Now" button on this page. You will be redirected to workiy's official application portal hosted on zohorecruit where you can submit your application directly.
When was the Site Reliability Engineer job at workiy posted?
This Site Reliability Engineer position at workiy was posted on Aug 27, 2026. Apply as soon as possible โ€” early applications are often reviewed first.
Site Reliability Engineer
workiy
Apply for this role โ†—

You'll be redirected to workiy's official application page on zohorecruit.