System Reliability Engineer

veeamsoftware· Corp Tech, Enterprise Applications 1009142
Apply Now ↗
📍 San Jose, Costa Rica

About this role

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

We are looking for a Site Reliability Engineer (SRE) to join our team and ensure the continuous, reliable operation of company services. This role involves proactive monitoring, incident response, and building resilient observability and escalation practices across our infrastructure.

What You'll Do

  • Ensure monitoring and uninterrupted operation of company services
  • Write and maintain alerting rules and runbooks
  • Perform triage of incoming incidents and initial diagnosis of issues
  • Build and maintain escalation chains for incident response
  • Perform technical incident resolution activities according to runbooks
  • Participate in on-call rotations and post-incident reviews (RCA/postmortems)
  • Continuously improve observability coverage and reduce alert noise/false positives
  • Collaborate with development and infrastructure teams to identify reliability risks and implement preventive measures

What You'll Bring

  • Experience with observability tools (Grafana, ELK, VictoriaMetrics)
  • Experience working with Linux
  • Experience working with Kubernetes (k8s)
  • Experience with AWS and Azure cloud platforms
  • Ability to analyze incidents, identify root causes, and propose remediation steps

Bonus Skills

  • Experience with Infrastructure as Code (Terraform, Ansible, or similar)
  • Scripting skills (Python, Bash) for automation of operational tasks
  • Understanding of DevOps and CI/CD principles
  • Experience with incident management tools (PagerDuty, Opsgenie, etc.)
  • Effective communication skills and a collaborative approach to teamwork

What You'll Get

  • Comprehensive Health Coverage – Fully employer-paid medical, dental, and vision insurance for employees and eligible dependents, including virtual care services
  • Wellbeing & Mental Health Support – Access to an Employee Assistance Program (EAP) that includes confidential therapy sessions, plus legal and financial counseling services
  • Financial Protection Benefits – Company-provided life and disability insurance to help support employees and their families
  • Paid Time Off & Global Recharge Days – Vacation time, statutory holidays, and additional company-wide VeeaMe Days dedicated to rest, wellbeing, and self-care
  • Family-Friendly Leave Programs – Competitive maternity, paternity, adoption, and other leave benefits that support employees through important life moments
  • Give Back to Your Community – Employees receive paid volunteer time each year through the Veeam Cares program

Please note: The position is based in San Jose, Costa Rica. If the applicant is permanently located outside of Costa Rica, Veeam reserves the right to decline the application. All applications must be submitted in English.

#LI-FT
#LI-REMOTE

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Frequently Asked Questions

Is the salary disclosed for the System Reliability Engineer position at veeamsoftware?
The salary for this System Reliability Engineer role at veeamsoftware is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the System Reliability Engineer position at veeamsoftware located?
This System Reliability Engineer role at veeamsoftware is based in San Jose, Costa Rica. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Which team or department does the System Reliability Engineer at veeamsoftware belong to?
This System Reliability Engineer position is part of the Corp Tech, Enterprise Applications 1009142 department at veeamsoftware. See the full job description for more information about the team structure and responsibilities.
How do I apply for the System Reliability Engineer position at veeamsoftware?
Click the "Apply Now" button on this page. You will be redirected to veeamsoftware's official application portal hosted on greenhouse where you can submit your application directly.
When was the System Reliability Engineer job at veeamsoftware posted?
This System Reliability Engineer position at veeamsoftware was posted on Jul 10, 2026. Apply as soon as possible — early applications are often reviewed first.
System Reliability Engineer
veeamsoftware
Apply for this role ↗

You'll be redirected to veeamsoftware's official application page on Greenhouse.