About this role
Job Title: MLOps / Serving Engineer Experience : 5+ years Location : Hyderabad OR Pune Notice Period: 0-30 days Work mode - Hybrid We are seeking an experienced MLOps / Serving Engineer who can design and operate the production serving infrastructure for fine-tuned LLMs on AWS — optimised inference engines, shadow-mode and staged rollout pipelines, monitoring dashboards, and the path from experimental model to full production traffic. Key Responsibilities: Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic Execute staged rollout: canary (5%) → gradual ramp (25% → 50% → 100%) with automated rollback on quality degradation Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request Design auto-scaling; implement high-availability (2× instances); automated rollback triggers on end-to-end quality metrics Requirements 5+ years MLOps or ML infrastructure engineering Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback Experience onto Continuous batching, INT8 quantization, KV-cache management Expertise on Docker, Kubernetes (EKS) for ML workloads Worked on CloudWatch, Prometheus, Grafana Benefits Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind. Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones. Retirement Benefits: PF and Gratuity provided as per standard government regulations. Flexible Work Options: Enjoy hybrid work arrangements & flexible working hours Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays. Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
Frequently Asked Questions
Is the salary disclosed for the MLOps / Serving Engineer position at dataeconomy?
The salary for this MLOps / Serving Engineer role at dataeconomy is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the MLOps / Serving Engineer position at dataeconomy located?
This MLOps / Serving Engineer role at dataeconomy is based in Hyderabad, Telangana, India. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the MLOps / Serving Engineer role at dataeconomy full-time or part-time?
This is listed as a Full time position. It is posted as a MLOps / Serving Engineer role at dataeconomy.
How do I apply for the MLOps / Serving Engineer position at dataeconomy?
Click the "Apply Now" button on this page. You will be redirected to dataeconomy's official application portal hosted on zohorecruit where you can submit your application directly.
When was the MLOps / Serving Engineer job at dataeconomy posted?
This MLOps / Serving Engineer position at dataeconomy was posted on Sep 8, 2026. Apply as soon as possible — early applications are often reviewed first.