Lead Engineer, Issue Management & Triage

diligentrobotics· 900-Central Ops
Apply Now ↗
📍 Austin, Texas, United States

About this role

What we’re doing isn’t easy, but nothing worth doing ever is. 

Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. As a Fleet Engineer, you'll own the reliability and continuous improvement of our deployed robotic fleet — leading hands-on investigations into how and why robots fail in the field, across the mobile base, charging/docking, motion and power, connectivity (modem), and sensor hardware. You'll combine remote data analysis with bench/lab failure analysis at our Austin HQ, turning field-technician reports and fleet data into clear problem statements, validated root causes, and corrective actions driven to closure with engineering, operations, manufacturing, and vendors.

 We are hiring a Lead Engineer, Issue Management & Triage to lead the systems, tooling, and team at the intersection of our Customers, Remote Operations Center (ROC), and Engineering. This is a highly technical, hands-on role focused on building the infrastructure that powers how we detect, triage, diagnose, and resolve issues across a deployed robotic fleet. You will work deeply with Engineering teams to design classification frameworks, build internal tools, and develop automation pipelines that improve reliability at scale. 

Location: Austin preferred, Remote possible (U.S.)
Travel: if remote up to ~50% travel to Austin, TX (especially in the your first 90 days)

What You’ll Do:

Own Issue Management & Triage Systems

  • Design and own end-to-end systems for issue intake, triage, and escalation.
  • Define severity frameworks, SLAs, and ensure issues are consistently structured for engineering prioritization.

Build Tools & Automation (Hands-On)

  • Develop automation and pipelines to ingest, process, and classify operational data, reducing manual triage effort.
  • Contribute directly to codebases (Python, backend services) and partner with Engineering on system integrations (logs, telemetry, alerts).

Bridge Operations & Engineering

  • Act as the primary technical interface between the Remote Operations Center (ROC) and Engineering.
  • Translate real-world issues into prioritized, categorized technical problems for resolution alignment.

Performance Measurement & Classification Frameworks

  • Develop systems and taxonomies to systematically measure and classify robot performance, failure modes, and degradation across the fleet.
  • Build dashboards and reporting systems to track trends, severity, and impact.

Root Cause Analysis & Continuous Improvement

  • Establish best practices for Root Cause Analysis (RCA) and identify systemic issues.
  • Drive long-term fixes and create feedback loops to influence improvements in hardware, software, and autonomy.

What We’re Looking For:

  • 7+ years in relevant technical or program management roles (e.g., engineering, incident management)
  • 3+ years of people management
  • Experience with complex, real-world systems (robotics, autonomous/distributed systems, or hardware-software products)
  • Proven track record building operational tools, systems, or infrastructure for workflows

Technical Skills

  • Strong programming experience (Python preferred; backend or data systems experience a plus)
  • Experience with:
    • Data pipelines and telemetry systems
    • Monitoring, alerting, and logging infrastructure
    • Internal tools and automation systems
  • Ability to design scalable systems for classification, prioritization, and workflow automation
  • Familiarity with platforms like Jira, Zendesk, SQL, Looker, Foxglove, or similar

Systems & Product Thinking

  • Strong systems thinker, translating ambiguous operational problems into structured technical solutions
  • Experience defining metrics, taxonomies, and performance frameworks
  • Data-driven approach to prioritization and decision-making

Mindset

  • Hands-on and willing to dive into technical problems when needed
  • Strong ownership and bias toward action
  • Comfortable operating in a fast-paced, scaling environment
  • Passion for improving real-world system performance and reliability

Frequently Asked Questions

Is the salary disclosed for the Lead Engineer, Issue Management & Triage position at diligentrobotics?
The salary for this Lead Engineer, Issue Management & Triage role at diligentrobotics is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Lead Engineer, Issue Management & Triage position at diligentrobotics located?
This Lead Engineer, Issue Management & Triage role at diligentrobotics is based in Austin, Texas, United States. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Which team or department does the Lead Engineer, Issue Management & Triage at diligentrobotics belong to?
This Lead Engineer, Issue Management & Triage position is part of the 900-Central Ops department at diligentrobotics. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Lead Engineer, Issue Management & Triage position at diligentrobotics?
Click the "Apply Now" button on this page. You will be redirected to diligentrobotics's official application portal hosted on greenhouse where you can submit your application directly.
When was the Lead Engineer, Issue Management & Triage job at diligentrobotics posted?
This Lead Engineer, Issue Management & Triage position at diligentrobotics was posted on Sep 11, 2026. Apply as soon as possible — early applications are often reviewed first.
Lead Engineer, Issue Management & Triage
diligentrobotics
Apply for this role ↗

You'll be redirected to diligentrobotics's official application page on Greenhouse.