01.
Robots Exclusion Protocol (robots.txt & Meta Directives)
Our crawlers evaluate target site governance directives prior to initiating extraction sessions. In accordance with RFC 9309:
- Disallow Path Processing: URLs containing paths specifically prohibited by
robots.txt for all user-agents or our specific bot identifiers are eliminated from our discovery queue before HTTP dispatch.
- Crawl-Delay Adherence: When a target domain declares a
Crawl-delay: <seconds> directive, our scheduler automatically spaces out successive requests to ensure compliance.
- HTML Meta Directives: Pages containing
<meta name="robots" content="noindex, nofollow"> or HTTP response header X-Robots-Tag: noindex are stripped from extraction pipelines.
02.
Politeness Throttling, Concurrency & Load Balancing
To eliminate server saturation risks, our platform applies hard architectural limits to every crawl job:
Per-Host Concurrency Caps
No target web host receives more than 1 to 5 simultaneous worker threads. Free trial accounts are hardcoded to a single concurrent thread.
Smart Request Pacing (RPM)
The platform enforces minimum delay intervals (typically 1,000msβ2,500ms) between page fetches on the same target domain to emulate human reading speeds.
Strict Depth Ceilings
Crawls are strictly bounded by configured Depth Limits (Level 2 for trials; maximum Level 8 for Enterprise), preventing runaway loop traversal.
03.
User-Agent Transparency & Webmaster Identification
Our crawlers identify themselves truthfully and provide webmasters with direct contact links:
User-Agent: Mozilla/5.0 (compatible; WebExtractoPro-Bot/2.4; +https://webextracto.com/ethical-scraping)
Webmasters reviewing access logs can immediately trace our crawler traffic back to this documentation, review our operational standards, and initiate automated opt-out suppression.
04.
Dynamic Backoff & Error Rate Circuit Breakers
If target servers show signs of load, Web Extracto Pro adopts immediate defensive backoff behaviors:
- HTTP 429 (Too Many Requests): The crawler immediately quadruples the delay interval and backs off for a minimum of 60 seconds before retrying.
- HTTP 503 (Service Unavailable): Requests to the host domain are temporarily halted to permit server recovery.
- Latency Spike Detection: If server response time exceeds 2,500ms, the system downshifts concurrency to 1 thread automatically.
- Persistent Error Abort: If consecutive errors exceed 15% of total scanned pages, the active session is terminated gracefully.
05.
Universal Domain Blacklist & Webmaster Opt-Out
We maintain a centralized Universal Domain Blacklist. If you own or manage a domain and do not wish our platform to index publicly visible contact info:
Submit Domain Removal & Permanent Blacklist Request
No legal threats or complex forms required. Verified domain owners are blacklisted within 24 hours.
Request Domain Removal →