Data Engineer — ML Training Data Pipeline
About this role
Job Title:Data Engineer - ML Training Data Pipeline Notice period: 0-30 Days Experience : 5+ Years Location: Hyderabad OR Pune We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS. What We Expect: Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch) Convert between chat-completion formats (e.g., OpenAI → Llama 3.1 tool-calling format) Implement smart deduplication and sampling to balance training distribution Design identity-aware train/test splits that measure true generalization Build data validation gates to detect schema drift and format anomalies Create a continuous pipeline that auto-processes new production traces for retraining Requirements Experience: 6+ years data engineering focused on ML data pipelines Python: Strong — pandas, pyarrow, JSONL processing at scale ML Data Libraries: HuggingFace Datasets, Arrow-based storage Data Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting Deduplication: Content hashing, identity-based grouping strategies AWS: S3, EC2, batch processing workflows Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright. Benefits Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind. Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones. Retirement Benefits: PF and Gratuity provided as per standard government regulations. Flexible Work Options: Enjoy hybrid work arrangements & flexible working hours. Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays. Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
Frequently Asked Questions
Is the salary disclosed for the Data Engineer — ML Training Data Pipeline position at dataeconomy?
The salary for this Data Engineer — ML Training Data Pipeline role at dataeconomy is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Data Engineer — ML Training Data Pipeline position at dataeconomy located?
This Data Engineer — ML Training Data Pipeline role at dataeconomy is based in Hyderabad, Telangana, India. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Data Engineer — ML Training Data Pipeline role at dataeconomy full-time or part-time?
This is listed as a Full time position. It is posted as a Data Engineer — ML Training Data Pipeline role at dataeconomy.
How do I apply for the Data Engineer — ML Training Data Pipeline position at dataeconomy?
Click the "Apply Now" button on this page. You will be redirected to dataeconomy's official application portal hosted on zohorecruit where you can submit your application directly.
When was the Data Engineer — ML Training Data Pipeline job at dataeconomy posted?
This Data Engineer — ML Training Data Pipeline position at dataeconomy was posted on Sep 8, 2026. Apply as soon as possible — early applications are often reviewed first.
Data Engineer — ML Training Data Pipeline
dataeconomy
You'll be redirected to dataeconomy's official application page on zohorecruit.