A production-ready Go REST API that proxies requests through Scrape.do, extracts a specific CSS selector from the HTML response, and returns only that HTML fragment.
If the finalData parameter is provided, the service retries through up to 5 different Scrape.do modes until that value is found in the extracted content.
- Architecture
- Project Structure
- Getting Started
- Configuration
- API Reference
- Retry Mechanism
- Cache Behavior
- Security
- Logging
- Monitoring
- Error Codes
- Dependencies
The service follows a layered architecture. Each layer has a single responsibility and is wired to the layer below it via dependency injection.
HTTP Request
│
▼
┌─────────────────────────────────┐
│ Middleware Stack │
│ Metrics → RateLimiter → Mux │
└────────────────┬────────────────┘
│
▼
┌─────────────────────────────────┐
│ Controller │ ← Input validation, cache read/write,
│ (controller.go) │ response formatting, logging
└────────────────┬────────────────┘
│
▼
┌─────────────────────────────────┐
│ Service │ ← Retry logic, business flow,
│ (service.go) │ finalData check
└────────────────┬────────────────┘
│
┌────────┴────────┐
▼ ▼
┌──────────────┐ ┌─────────────┐
│ HttpClient │ │ Parser │ ← Scrape.do HTTP request │ CSS selector extract
│ (client.go) │ │ (parser.go) │
└──────────────┘ └─────────────┘
| Layer | File | Responsibility |
|---|---|---|
| Controller | controller.go |
Input validation, cache management, HTTP response |
| Service | service.go |
5-step retry flow, finalData check |
| HttpClient | client.go |
Scrape.do requests, SSRF-safe dialer, connection pooling |
| Parser | parser.go |
CSS selector extraction from HTML |
| Security | ssrf.go, sanitizer.go |
SSRF protection, HTML/selector sanitization |
| Cache | cache.go |
In-memory TTL cache |
| Middleware | ratelimit.go, metrics.go |
Rate limiting, Prometheus metrics |
scrape-service/
├── cmd/
│ └── server/
│ └── main.go # Entry point, graceful shutdown
├── internal/
│ ├── config/
│ │ └── config.go # Environment-based configuration
│ ├── controller/
│ │ └── controller.go # HTTP handler, validation, cache
│ ├── service/
│ │ └── service.go # 5-step retry logic
│ ├── httpclient/
│ │ └── client.go # SSRF-safe HTTP client, connection pooling
│ ├── parser/
│ │ └── parser.go # CSS selector extraction via goquery
│ ├── security/
│ │ ├── ssrf.go # SSRF protection, custom dialer
│ │ └── sanitizer.go # HTML sanitization, selector validation
│ ├── cache/
│ │ └── cache.go # In-memory cache wrapper
│ ├── errors/
│ │ └── errors.go # Custom error types
│ └── middleware/
│ ├── ratelimit.go # Per-IP token bucket rate limiting
│ └── metrics.go # Prometheus histogram & counter
├── .env.example # Example configuration file
├── go.mod
└── README.md
- Go 1.21+
- A valid Scrape.do API token
# 1. Clone the repository
git clone <repo-url>
cd scrape-service
# 2. Create the configuration file
cp .env.example .env
# Edit .env if needed — defaults work out of the box
# 3. Download dependencies
go mod download
# 4. Run
go run ./cmd/serverOutput:
{"time":"...","level":"INFO","msg":"loaded config","server":{"port":"8080"},"timeouts":{"request_timeout":"30s"},"cache":{"ttl":"1m0s"},"limits":{"max_url_length":2048,"max_selector_length":256},"rate_limit":{"rps":10,"burst":20}}
{"time":"...","level":"INFO","msg":"server listening","addr":"http://localhost:8080"}go build -o scrape-service ./cmd/server
./scrape-servicecurl -s -X POST http://localhost:8080/scrape \
-H "Content-Type: application/json" \
-d '{
"url": "https://news.ycombinator.com",
"selector": ".titleline",
"scrapedoToken": "YOUR_TOKEN_HERE"
}'All configuration is provided via environment variables. You can also create a .env file in the project root.
| Variable | Default | Description |
|---|---|---|
PORT |
8080 |
HTTP server listen port |
REQUEST_TIMEOUT_MS |
30000 |
Maximum wait time per Scrape.do step (ms) |
CACHE_TTL_SECONDS |
60 |
How long successful responses are cached (seconds) |
MAX_URL_LENGTH |
2048 |
Maximum allowed URL length |
MAX_SELECTOR_LENGTH |
256 |
Maximum allowed CSS selector length |
RATE_LIMIT_RPS |
10 |
Maximum requests per second per IP |
RATE_BURST |
20 |
Token bucket burst capacity |
# Server
PORT=8080
# Timeouts
REQUEST_TIMEOUT_MS=30000
# Cache
CACHE_TTL_SECONDS=60
# Input limits
MAX_URL_LENGTH=2048
MAX_SELECTOR_LENGTH=256
# Rate limiting
RATE_LIMIT_RPS=10
RATE_BURST=20Scrapes the target URL, extracts the specified CSS selector, and returns it.
{
"url": "https://example.com/products/123",
"selector": ".product-price",
"finalData": "$29.99",
"scrapedoToken": "YOUR_SCRAPE_DO_TOKEN"
}| Field | Type | Required | Description |
|---|---|---|---|
url |
string | ✅ | Target URL to scrape. Must use http or https. |
selector |
string | ✅ | CSS selector of the content to return (e.g. .price, #title, h1) |
finalData |
string | ❌ | A string that must appear in the extracted content. When set, cache is bypassed. |
scrapedoToken |
string | ✅ | Your Scrape.do API token. Never logged. |
<span class="product-price">$29.99</span>Content-Type: text/html; charset=utf-8X-Cache: HITorX-Cache: MISSheader is included
Only the matched selector's HTML is returned. The full page HTML is never returned.
{
"code": "SelectorNotFoundError",
"message": "selector '.product-price' not found in HTML"
}Returns the service health status.
curl http://localhost:8080/health{"status": "ok"}Returns metrics in Prometheus exposition format. Monitoring systems should scrape this endpoint.
curl http://localhost:8080/metricsThe service escalates through 5 Scrape.do modes in order. At each step the following checks are performed:
- Did the HTTP request succeed?
- Is the selector present in the returned HTML?
- If
finalDatawas provided, does the extracted content contain it?
If all checks pass, the result is returned. If any check fails, the next step is tried.
Step 1: Simple Request
──────────────────────────────────────────────
GET http://api.scrape.do/?token=TOKEN&url=TARGET_URL
Failed → Step 2
Step 2: Super Mode (advanced anti-bot bypass)
──────────────────────────────────────────────
GET ...&super=true
Failed → Step 3
Step 3: Render Mode (JavaScript execution)
──────────────────────────────────────────────
GET ...&super=true&render=true
Failed → Step 4
Step 4: WaitSelector (wait for the selector to appear in the DOM)
──────────────────────────────────────────────
GET ...&super=true&render=true&blockResources=false&waitSelector=SELECTOR
Failed → Step 5
Step 5: PlayWithBrowser (full browser automation)
──────────────────────────────────────────────
GET ...&super=true&render=true&blockResources=false&playwithbrowser=ENCODED_JSON
All steps failed → TimeoutError (504)
In the final step, the following browser action sequence is sent to Scrape.do:
[
{ "Action": "Wait", "Timeout": 5000 },
[
{
"Action": "WaitSelector",
"WaitSelector": ".SELECTOR",
"Timeout": 10000
}
]
]This payload is URL-encoded and passed as the playwithbrowser query parameter.
Each step has its own 30-second timeout. The maximum total wait time is 5 steps × 30 seconds = 150 seconds.
| Condition | Cache Read? | Cache Write? |
|---|---|---|
finalData not set |
✅ Yes | ✅ Yes |
finalData set |
❌ No | ❌ No |
Why is cache skipped when finalData is set?
finalData requires a specific string to be present in the extracted content. This is typically used to confirm that the page has fully loaded and contains live data. Returning a stale cached result in this case would defeat the purpose — a fresh, up-to-date response is always required.
Cache key:
{url}|{selector}
A second request with the same URL and selector combination will return X-Cache: HIT until the CACHE_TTL_SECONDS window expires.
The service blocks requests targeting internal networks or loopback addresses. This protection is enforced at the TCP dial level — not just at DNS resolution time — which means DNS rebinding attacks cannot bypass it.
Blocked networks:
| CIDR | Description |
|---|---|
127.0.0.0/8 |
Loopback |
::1/128 |
IPv6 Loopback |
10.0.0.0/8 |
Private |
172.16.0.0/12 |
Private |
192.168.0.0/16 |
Private |
169.254.0.0/16 |
Link-local (AWS metadata, etc.) |
100.64.0.0/10 |
Shared address space |
198.18.0.0/15 |
Benchmark |
0.0.0.0/8 |
"This" network |
Before the extracted HTML is returned to the client, the following are stripped:
- All
<script>tags and their contents - All inline event handler attributes (
onclick,onload,onerror, etc.)
CSS selectors containing the characters {, }, ;, (, ), \ are rejected. These characters could be used to perform CSS injection or manipulate the parser.
The following headers are never forwarded to the client:
Set-Cookie— prevents session data leakageAuthorization— prevents credential leakage
The Scrape.do API token is never included in any log line. URLs are logged by hostname only (e.g. example.com).
The service produces structured JSON logs via log/slog.
When the service starts, all active configuration values are logged:
{
"level": "INFO",
"msg": "loaded config",
"server": { "port": "8080" },
"timeouts": { "request_timeout": "30s" },
"cache": { "ttl": "1m0s" },
"limits": { "max_url_length": 2048, "max_selector_length": 256 },
"rate_limit": { "rps": 10, "burst": 20 }
}When a request arrives, all configuration values that apply to that request are logged:
{
"level": "INFO",
"msg": "scrape request received",
"request": {
"target": "example.com",
"selector": ".product-price",
"final_data_set": true
},
"config": {
"cache_enabled": false,
"cache_ttl": "1m0s",
"request_timeout": "30s",
"max_url_length": 2048,
"max_selector_length": 256,
"rate_limit_rps": 10,
"rate_burst": 20
}
}{
"level": "INFO",
"msg": "scrape request completed",
"target": "example.com",
"selector": ".product-price",
"cache": "MISS",
"result_bytes": 312,
"duration_ms": 1843
}Each failed retry step is logged individually:
{
"level": "INFO",
"msg": "retry step failed",
"step": "simple",
"selector": ".product-price",
"target": "example.com",
"error": "[SelectorNotFoundError] selector '.product-price' not found in HTML"
}{
"level": "ERROR",
"msg": "scrape request failed",
"target": "example.com",
"selector": ".product-price",
"error": "[TimeoutError] request timed out after all retry steps",
"duration_ms": 5120
}The API token is never present in any log line. URLs are logged as hostname only.
The GET /metrics endpoint exposes the following metrics in Prometheus format:
| Metric | Type | Labels | Description |
|---|---|---|---|
scrape_request_duration_seconds |
Histogram | path, status |
Request duration distribution |
scrape_requests_total |
Counter | path, status |
Total number of requests |
# Average response time over the last 5 minutes
rate(scrape_request_duration_seconds_sum[5m]) / rate(scrape_request_duration_seconds_count[5m])
# Per-second error rate (5xx)
rate(scrape_requests_total{status=~"5.."}[1m])
# 95th percentile latency
histogram_quantile(0.95, rate(scrape_request_duration_seconds_bucket[5m]))
| Code | HTTP Status | When Returned |
|---|---|---|
InvalidSelectorError |
400 | Selector contains invalid characters |
ValidationError |
400 | Required field is missing or malformed |
SelectorNotFoundError |
404 | Selector was not found after all retry steps |
FinalDataNotFoundError |
422 | Selector was found but finalData string was not present |
TargetRequestFailedError |
502 | Scrape.do could not reach the target URL |
TimeoutError |
504 | All 5 retry steps failed |
RateLimitError |
429 | This IP has sent too many requests |
MethodNotAllowed |
405 | An HTTP method other than POST was used |
BadRequest |
400 | Request body is not valid JSON |
| Package | Version | Purpose |
|---|---|---|
| PuerkitoBio/goquery | v1.9.1 | HTML parsing and CSS selector extraction |
| patrickmn/go-cache | v2.1.0 | In-memory TTL cache |
| prometheus/client_golang | v1.19.0 | Prometheus metrics |
| golang.org/x/time | v0.5.0 | Token bucket rate limiter |
All other dependencies (cascadia, protobuf, etc.) are indirect dependencies of the above.
The token bucket algorithm is used. A separate limiter is created per IP address, stored in a sync.Map for thread safety.
RATE_LIMIT_RPS=10→ Each IP can send a maximum of 10 requests per secondRATE_BURST=20→ Bursts of up to 20 requests are allowed to handle sudden spikes
When the limit is exceeded:
HTTP 429 Too Many Requests
{"code": "RateLimitError", "message": "too many requests"}When the service receives a SIGINT or SIGTERM signal:
- It stops accepting new connections
- It waits up to 30 seconds for active requests to complete
- Requests that finish within that window receive normal responses
- After 30 seconds, the server is forcefully shut down
{"level":"INFO","msg":"shutting down gracefully..."}
{"level":"INFO","msg":"server stopped"}