Seeking Senior Freelance Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/D
Job Description
We are looking for a senior-level expert who can architect and develop a complete web data acquisition platform from scratch for our upcoming AI-Powered Market Intelligence Platform.
This is not a basic web-scraping project. We need someone who can build the underlying infrastructure for AI-directed, targeted and scalable data collection.
Core Architecture
The intended workflow is:
Business Objective → LLM/AI Planning → Data Acquisition Platform → Search/Google Advanced Search/Dork Discovery → URL/Source Discovery → Proxy Management → Scrapers/APIs → Targeted Data Extraction → Validation → Deduplication → Knowledge Base → AI/ML
We already have separate LLM/AI experts. The selected developer will work closely with them to convert AI-generated research/data requirements into executable, reliable scraping jobs.
Key Responsibilities
The expert must be capable of developing:
• Complete scraping platform from scratch
• Scrapy / Playwright / Selenium / HTTP-based scraping infrastructure
• Proxy management layer with proxy pools, assignment, health monitoring, failover and multiple provider integration
• Scraping orchestrator for job queues, scheduling, concurrency, retries, workers and monitoring
• Google advanced-search / Dork-style discovery and search-result/URL extraction using appropriate and compliant mechanisms
• Source/domain discovery and filtering
• Modular connectors for eBay, Amazon, Shopify, Etsy, Reddit, YouTube review platforms and other sources
• Targeted field-level extraction rather than unnecessary bulk scraping
• Data validation, normalization and deduplication
• Source/URL/query/data lineage
• Structured database and data pipeline
• APIs through which our LLM team can create, monitor, modify and retrieve scraping jobs
• Monitoring, logging and error handling
Critical Requirement: AI-Directed Scraping
The system must allow our LLM layer to determine before scraping:
• What to collect
• Where to collect it
• Which queries/keywords to use
• Which sources/URLs are relevant
• Which fields are required
• Filters/timeframes
• Maximum/sufficient data volume
• When to stop
Therefore:
LLM Planning → Targeted Collection → Validation → AI Analysis
rather than:
Scrape Everything → Store Everything → AI Filters It
The objective is to reduce data volume, scraping/proxy costs, storage, processing and LLM token consumption while improving relevance and precision.
LLM Integration
Our LLM experts will handle:
• LLMs and reasoning
• Research intelligence
• RAG
• AI agents
• Market/product/review intelligence
• Recommendations
The selected developer will handle the data acquisition infrastructure and collaborate with the LLM team on:
• AI-to-scraper APIs
• Data schemas
• Collection specifications
• Intelligent stop conditions
• Feedback loops
• Additional targeted data requests
The architecture should support:
AI → Data → AI → Additional Data → AI
Required Experience
Strong practical experience in:
• Large-scale web scraping
• Python / Scrapy / Playwright / Selenium
• Proxy infrastructure and proxy management
• Search/query-driven data discovery
• Google advanced-search/Dork-based workflows
• Distributed scraping / job orchestration
• REST APIs
• PostgreSQL / Redis
• Docker / Linux / cloud infrastructure
• Data validation and deduplication
• Scalable data pipelines
Experience with marketplaces, review platforms, search engines, market intelligence or AI/LLM integration is strongly preferred.
Important
We are not looking for a freelancer who can only write individual scrapers.
We need someone capable of:
Architecting → Developing → Integrating → Scaling
the complete Search + Proxy + Scraping + Data Acquisition Infrastructure from scratch.
The system must be designed for legitimate and responsible collection of publicly accessible/authorised data and must not rely on bypassing authentication, security controls or other access restrictions.
MVP
The MVP will initially be validated using a controlled source/dataset such as eBay, together with search/discovery functionality, and then expanded to additional marketplaces, review sources and other data sources.
Application Requirements
Please provide:
• 2–3 relevant projects involving large-scale scraping/proxy infrastructure.
• Your proposed technical architecture for the above system.
• Recommended technology stack.
• Approach for AI/LLM → scraping integration.
• Approach for proxy management and cost optimisation.
• Estimated MVP timeline and development cost.
• Explanation of how you would scale from 10,000 → 100,000 → 1M+ records.
Position
Senior Web Scraping, Proxy & AI-Directed Data Acquisition Platform Architect/Developer
Project: Build from Scratch | MVP → Long-Term Platform | Close collaboration with existing LLM/AI team
