Senior AI Infrastructure Engineer

heidihealth.com.au· Engineering
Apply Now ↗
📍 MelbourneFullTime

About this role

We’re Heidi.

 

We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. In just 18 months, our clinical AI products have absorbed the administrative chaos of 73 million patient visits. Today, we support over 2.5 million patient sessions a week across 190+ countries.

 

Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s healthcare capacity and strengthen the human connection at its heart.

 

We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again.

 

If you don’t choose easy and you want to build something way bigger than yourself then, choose the challenge, choose Heidi.

 

The role

 

This role sits in the model team, the researchers and engineers who train, deploy, and own the AI models behind every Heidi product. You’ll build and operate the infrastructure that makes those models fast, reliable, and cost-effective at scale.

Your work will span production model serving, GPU cluster management, and the infrastructure supporting training and evaluation. You’ll decide how workloads share compute, diagnose performance bottlenecks, and build the deployment and observability tools that help the team move quickly with confidence.

We’re looking for a hands-on engineer who has deployed and operated models in production, understands the demands of GPU workloads, and can take a system from initial design through rollout, incidents, and ongoing improvement. You’ll partner closely with researchers and our platform engineers, with ownership of the systems you build.

 
 

What you’ll do

 
  • Build and own model-serving infrastructure. Take models from checkpoint to production, with repeatable deployment pipelines, request routing, autoscaling, fallback paths, and controlled rollouts and rollbacks across regions.

  • Manage GPU clusters and workload scheduling. Improve resource allocation across inference, training, and evaluation. Build scheduling policies around workload priority, quotas, hardware topology, and recovery requirements so online services stay responsive while other workloads make productive use of capacity.

  • Improve inference performance. Profile real workloads and improve latency, throughput, and memory efficiency. Evaluate batching, KV-cache management, quantization, speculative decoding, and parallelism strategies against production traffic and quality requirements.

  • Support distributed training and model iteration. Give researchers reliable ways to launch fine-tuning and training jobs, manage model artifacts, save and restore checkpoints, and move validated models into serving. Reduce time lost to failed jobs, slow data loading, and manual setup.

  • Make deployments observable and incidents traceable. Connect application requests and sessions to the exact model, deployment configuration, and worker that served them. Build dashboards and alerts covering model latency, queueing, errors, GPU health, memory pressure, and workload performance.

  • Own production reliability. Define service objectives, investigate incidents across the application, inference engine, GPU, and network layers, and build recovery procedures that work. Turn recurring failures into fixes, automated checks, and useful runbooks.

  • Make compute costs actionable. Track GPU usage, idle capacity, and inference cost by model and workload. Use capacity forecasts and measured performance to guide deployment choices and improve cost per successful request without sacrificing quality or reliability.

  • Build a platform the model team can use independently. Automate provisioning, configuration, benchmarking, and releases. Partner with the engineers behind ASR, note generation, Evidence, and Dictate so new models can be deployed and evaluated through consistent, well-supported workflows.

What you'll need

 
  • A strong engineering foundation. Hands-on AI infrastructure experience. At least 1 year building and operating infrastructure for large language models, including model deployment, inference serving, or distributed training. You can design the system, write the code, and own it in production.

  • Production model deployment experience. You’ve deployed and maintained LLMs or other demanding ML workloads using engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or comparable systems. You understand the work between loading a model and running a reliable service.

  • GPU cluster and orchestration experience. You’ve managed GPU workloads using Kubernetes, Slurm, or an equivalent platform, with practical experience in scheduling, resource allocation, capacity planning, and failure recovery.

  • Performance debugging skills. You can use traces, metrics, and profiling tools to distinguish compute, memory, communication, and scheduling bottlenecks. You understand how batch size, context length, precision, and multi-GPU execution affect performance and cost.

  • Strong software and systems skills. You’re proficient in Python and comfortable with backend or systems development in Go, C++, Rust, or a comparable language. You have practical experience with Linux, containers, deployment automation, and distributed services.

  • Operational ownership. You’ve owned production incidents, built useful monitoring, and made releases recoverable. You can explain the trade-offs behind a design and work effectively with researchers, product engineers, and infrastructure partners.

 

Nice to have

 
  • Experience with distributed training frameworks such as PyTorch FSDP or Megatron, or infrastructure for reinforcement learning and rollout generation.

  • Experience tuning inference engines, serving MoE models, or implementing quantization, speculative decoding, and prefill/decode disaggregation.

  • Familiarity with GPU interconnects, NCCL, RDMA, topology-aware scheduling, or diagnosing multi-node communication problems.

  • CUDA or Triton kernel development, contributions to AI infrastructure projects, or experience building cluster operators and scheduling integrations.

  • Experience operating infrastructure across multiple regions or providers, particularly for healthcare or other sensitive production workloads.

 

How we show up

  • Build for the next decade, not next quarter. Our targets are outrageous on purpose. The world's health doesn't have the luxury of incrementalism.

  • Lead, don't wait. We treat tomorrow's problems today. Sometimes we build what's needed before it's wanted, and we're fine with that.

  • Follow the evidence. Trust the patient. We pursue truth relentlessly. But when the subjective and objective disagree, we treat the patient, not the numbers. Ego is a comorbidity we can't afford.

  • Own the outcome. Everyone here carries the company. Raise problems with solutions, solve them end-to-end, and never be a bystander.

  • Ship, measure, go again. A button today, a workflow tomorrow. More iterations beat better planning. We're precise at pace, not reckless.

  • Live in clinicians' reality. Not the ideal workflow, the twenty-patients-before-lunch actual one. We build for exhausted humans, and we'd better be decent ones while we do it.

 

Why Heidi?

 

You’ll join a team focused on real-world impact over imaginary valuations and glossy PR. We live and breathe the challenges of modern health systems, and are laser-focused on exacting the change we’d like to see. We’re medicos, engineers, builders, and designers who’ve felt the moral and practical toll of what non-care feels like. True A-players progress extremely fast here.

 

The nature of the scale-up game is demanding, but we value sustainable performance and mental health. You're trusted to perform, and you set your schedule. We operate on outcomes > inputs, not process theatre. We all take the bins out, metaphorically and literally.

 

Building what we’re building isn’t always easy. But we didn’t choose easy, we chose to build something that actually matters. We hold ourselves to a higher standard because healthcare demands it. If you join Heidi, you recognise that the deeper question isn’t whether AI can solve the global healthcare crisis, but whose hands will shape it. The work is hard, but you will trust and admire the people you work beside, and rest easy knowing you’re doing the defining work of your career.

 

We take care of you.

 

We offer a $1,000 annual learning and development budget, a $150/month health and wellness allowance, a $500 home office budget, 26 weeks paid primary parental leave and 18 weeks paid secondary parental leave, fertility support up to $10,000, four weeks of work from anywhere per year, and serious equity.

Frequently Asked Questions

Is the salary disclosed for the Senior AI Infrastructure Engineer position at heidihealth.com.au?
The salary for this Senior AI Infrastructure Engineer role at heidihealth.com.au is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Senior AI Infrastructure Engineer position at heidihealth.com.au located?
This Senior AI Infrastructure Engineer role at heidihealth.com.au is based in Melbourne. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Senior AI Infrastructure Engineer role at heidihealth.com.au full-time or part-time?
This is listed as a FullTime position. It is posted as a Senior AI Infrastructure Engineer role in the Engineering department at heidihealth.com.au.
Which team or department does the Senior AI Infrastructure Engineer at heidihealth.com.au belong to?
This Senior AI Infrastructure Engineer position is part of the Engineering department at heidihealth.com.au. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Senior AI Infrastructure Engineer position at heidihealth.com.au?
Click the "Apply Now" button on this page. You will be redirected to heidihealth.com.au's official application portal hosted on ashby where you can submit your application directly.
When was the Senior AI Infrastructure Engineer job at heidihealth.com.au posted?
This Senior AI Infrastructure Engineer position at heidihealth.com.au was posted on Sep 14, 2026. Apply as soon as possible — early applications are often reviewed first.
Senior AI Infrastructure Engineer
heidihealth.com.au
Apply for this role ↗

You'll be redirected to heidihealth.com.au's official application page on Ashby ATS.