Sponsored
Verified Job Work from home

Seeking Senior Freelance Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/D

Kolkata, West Bengal
Work from home
#794030
Remote / WFH
Outthinc Global Communications

Job Description

Build From Scratch | Freelance / Long-Term Collaboration | Profit-Sharing Opportunity for the Right Partner

We are looking for a senior-level expert who can architect and develop a complete web data acquisition platform from scratch for our upcoming AI-Powered Market Intelligence Platform.

This is not a basic web-scraping project. We need someone who can build the underlying infrastructure for AI-directed, targeted and scalable data collection.

Core Architecture

The intended workflow is:

Business Objective → LLM/AI Planning → Data Acquisition Platform → Search/Google Advanced Search/Dork Discovery → URL/Source Discovery → Proxy Management → Scrapers/APIs → Targeted Data Extraction → Validation → Deduplication → Knowledge Base → AI/ML

We already have separate LLM/AI experts. The selected developer will work closely with them to convert AI-generated research/data requirements into executable, reliable scraping jobs.

Key Responsibilities

The expert must be capable of developing:
• Complete scraping platform from scratch
• Scrapy / Playwright / Selenium / HTTP-based scraping infrastructure
• Proxy management layer with proxy pools, assignment, health monitoring, failover and multiple provider integration
• Scraping orchestrator for job queues, scheduling, concurrency, retries, workers and monitoring
• Google advanced-search / Dork-style discovery and search-result/URL extraction using appropriate and compliant mechanisms
• Source/domain discovery and filtering
• Modular connectors for eBay, Amazon, Shopify, Etsy, Reddit, YouTube review platforms and other sources
• Targeted field-level extraction rather than unnecessary bulk scraping
• Data validation, normalization and deduplication
• Source/URL/query/data lineage
• Structured database and data pipeline
• APIs through which our LLM team can create, monitor, modify and retrieve scraping jobs
• Monitoring, logging and error handling

Critical Requirement: AI-Directed Scraping

The system must allow our LLM layer to determine before scraping:
• What to collect
• Where to collect it
• Which queries/keywords to use
• Which sources/URLs are relevant
• Which fields are required
• Filters/timeframes
• Maximum/sufficient data volume
• When to stop

Therefore:

LLM Planning → Targeted Collection → Validation → AI Analysis

rather than:

Scrape Everything → Store Everything → AI Filters It

The objective is to reduce data volume, scraping/proxy costs, storage, processing and LLM token consumption while improving relevance and precision.

LLM Integration

Our LLM experts will handle:
• LLMs and reasoning
• Research intelligence
• RAG
• AI agents
• Market/product/review intelligence
• Recommendations

The selected developer will handle the data acquisition infrastructure and collaborate with the LLM team on:
• AI-to-scraper APIs
• Data schemas
• Collection specifications
• Intelligent stop conditions
• Feedback loops
• Additional targeted data requests

The architecture should support:

AI → Data → AI → Additional Data → AI

Required Experience

Strong practical experience in:
• Large-scale web scraping
• Python / Scrapy / Playwright / Selenium
• Proxy infrastructure and proxy management
• Search/query-driven data discovery
• Google advanced-search/Dork-based workflows
• Distributed scraping / job orchestration
• REST APIs
• PostgreSQL / Redis
• Docker / Linux / cloud infrastructure
• Data validation and deduplication
• Scalable data pipelines

Experience with marketplaces, review platforms, search engines, market intelligence or AI/LLM integration is strongly preferred.

Important

We are not looking for a freelancer who can only write individual scrapers.

We need someone capable of:

Architecting → Developing → Integrating → Scaling

the complete Search + Proxy + Scraping + Data Acquisition Infrastructure from scratch.

The system must be designed for legitimate and responsible collection of publicly accessible/authorised data and must not rely on bypassing authentication, security controls or other access restrictions.

MVP

The MVP will initially be validated using a controlled source/dataset such as eBay, together with search/discovery functionality, and then expanded to additional marketplaces, review sources and other data sources.

Application Requirements

Please provide:
• 2–3 relevant projects involving large-scale scraping/proxy infrastructure.
• Your proposed technical architecture for the above system.
• Recommended technology stack.
• Approach for AI/LLM → scraping integration.
• Approach for proxy management and cost optimisation.
• Estimated MVP timeline and development cost.
• Explanation of how you would scale from 10,000 → 100,000 → 1M+ records.

Position

Senior Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/Developer

Project: Build from Scratch | MVP → Long-Term Platform | Close collaboration with existing LLM/AI team
View more Seeking Jobs in Kolkata →
Sponsored

Similar Openings in Work from home

More jobs you might like

Ahmedabad, Gujarat Work from home 5000.00 ₹

Join Maple Grove Consulting as a Remote Editorial Assistant! This part-time with weekend availability remote role involves collaborating wit...

Posted 1h ago View Details
Ahmedabad, Gujarat Work from home 10000.00 ₹

Sunrise Digital Services is hiring a Remote Logistics Coordinator to join our distributed team on a part-time basis. In this role you will b...

Posted 1h ago View Details
Delhi, Delhi Work from home 15000.00 ₹

Join Bright Occasion Planning as a Freelance Wedding Planner (Remote)! This flexible hours remote role involves sourcing and negotiating wit...

Posted 1h ago View Details