Data Crawling - Consultant
rubick· Technology - R&D / Engineering
About this role
Data Crawling Consultant
Role Overview
You will be responsible for building and maintaining Rubick's data acquisition engine by extracting, validating, and structuring product information from eCommerce websites and marketplaces. This role focuses on developing scalable web crawling solutions that power Product Discovery, Search, Catalog Intelligence, and Market Intelligence platforms.
The role involves three areas: Part 1: The Fundamentals | Part 2: AI-Driven Data Crawling Excellence | Part 3: Innovation & Improvements
Part 1: The Fundamentals
- Develop, maintain, and optimize web crawlers to extract product data from eCommerce websites and marketplaces.
- Collect structured and unstructured product information, including product details, pricing, images, specifications, and availability.
- Clean, validate, and organize extracted datasets to ensure accuracy and consistency.
- Monitor crawler performance, identify failures, and resolve data extraction issues.
- Collaborate with Product, Engineering, and Data teams to ensure reliable and timely data collection.
- Maintain documentation for crawling processes, extraction rules, and data quality standards.
Part 2: AI-Driven Data Crawling Excellence
- Leverage AI-powered extraction techniques to improve data accuracy and extraction efficiency.
- Build intelligent crawling workflows using Python and modern web automation frameworks.
- Develop scalable solutions for handling dynamic websites, JavaScript-rendered content, and anti-bot mechanisms.
- Utilize browser automation, APIs, proxies, and scheduling tools to maximize crawl success rates.
- Implement automated data validation and monitoring systems to maintain high-quality datasets.
- Collaborate with AI and Data Engineering teams to support Machine Learning models and Product Intelligence systems.
Part 3: Innovation & Improvements
- Continuously optimize crawling performance for speed, scalability, and reliability.
- Identify opportunities to automate repetitive extraction and validation workflows.
- Improve data collection strategies by adopting new crawling technologies and AI-assisted solutions.
- Build reusable crawling frameworks and standardized extraction pipelines.
- Stay updated with the latest web scraping libraries, browser automation tools, and industry best practices.
- Contribute to knowledge repositories, documentation, and process improvements across the data acquisition function.
To Have
- 1+ year of experience in Web Crawling, Data Scraping, Product Matching, Data Extraction, or similar roles.
- Strong proficiency in Python for web scraping and automation.
- Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks.
- Good understanding of HTML, CSS, XPath, JSON, and DOM structures.
- Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping.
- Strong analytical and problem-solving skills.
- Ability to work with large datasets while maintaining high data quality.
- Experience in eCommerce, Retail, Marketplace Operations, or Catalog Management is preferred.
- Exposure to AI-assisted data extraction, Product Intelligence, or Machine Learning data pipelines is a strong advantage.
About Rubick : Company Overview
Rubick OS is a leading platform for brands and marketplaces, enabling seamless management of catalog automation, pricing and competitive intelligence, product listing intelligence, CAST systems, and brand and assortment intelligence.
Our solutions are used across five countries by major brands and marketplaces, including Amazon, Myntra, Reliance Group, Nykaa, Flipkart, Rare Rabbit, Decathlon, Celio, The Bay, Kiabi, Myer, Jumbo, and many more.
Work Model
Work from office – Monday to Saturday
Team Size
250+ Employees
Investors
Invested by Innospark US, Betatron Hong Kong, MJV India, 1Crowd India, and other investors.
Rubick.ai is a place to work if you want to build something foundational rather than incremental—it’s aiming to become the AI operating layer for ecommerce, solving complex, real-world problems across cataloging, pricing, and marketplace operations at scale.
You get high ownership, direct impact, and exposure to various parts of the business, making you an entrepreneur in-house. Rubick operates in the AI segment, which means you will be at the forefront of the technological revolution.
Rubick is profitable and growing and is looking to work with 10,000 brands in the next 2–3 years.
Join us in this exciting journey.
Visit Us
Frequently Asked Questions
Is the salary disclosed for the Data Crawling - Consultant position at rubick?
The salary for this Data Crawling - Consultant role at rubick is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Is the Data Crawling - Consultant job at rubick remote?
Yes, this Data Crawling - Consultant position at rubick is remote, with team members based in Bengaluru, KA, India, Remote, Oregon, United States. You can work from home or anywhere in the supported regions.
Is the Data Crawling - Consultant role at rubick full-time or part-time?
This is listed as a FULL TIME position. It is posted as a Data Crawling - Consultant role in the Technology - R&D / Engineering department at rubick.
Which team or department does the Data Crawling - Consultant at rubick belong to?
This Data Crawling - Consultant position is part of the Technology - R&D / Engineering department at rubick. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Data Crawling - Consultant position at rubick?
Click the "Apply Now" button on this page. You will be redirected to rubick's official application portal hosted on keka where you can submit your application directly.
When was the Data Crawling - Consultant job at rubick posted?
This Data Crawling - Consultant position at rubick was posted on Jul 29, 2026. Apply as soon as possible — early applications are often reviewed first.
Data Crawling - Consultant
rubick
You'll be redirected to rubick's official application page on keka.