Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

scrape-service

A production-ready Go REST API that proxies requests through Scrape.do, extracts a specific CSS selector from the HTML response, and returns only that HTML fragment.

If the finalData parameter is provided, the service retries through up to 5 different Scrape.do modes until that value is found in the extracted content.


Table of Contents


Architecture

The service follows a layered architecture. Each layer has a single responsibility and is wired to the layer below it via dependency injection.

HTTP Request
     │
     ▼
┌─────────────────────────────────┐
│         Middleware Stack         │
│  Metrics → RateLimiter → Mux    │
└────────────────┬────────────────┘
                 │
                 ▼
┌─────────────────────────────────┐
│           Controller            │  ← Input validation, cache read/write,
│       (controller.go)           │    response formatting, logging
└────────────────┬────────────────┘
                 │
                 ▼
┌─────────────────────────────────┐
│            Service              │  ← Retry logic, business flow,
│          (service.go)           │    finalData check
└────────────────┬────────────────┘
                 │
        ┌────────┴────────┐
        ▼                 ▼
┌──────────────┐   ┌─────────────┐
│  HttpClient  │   │   Parser    │  ← Scrape.do HTTP request  │ CSS selector extract
│ (client.go)  │   │ (parser.go) │
└──────────────┘   └─────────────┘

Layer Responsibilities

Layer File Responsibility
Controller controller.go Input validation, cache management, HTTP response
Service service.go 5-step retry flow, finalData check
HttpClient client.go Scrape.do requests, SSRF-safe dialer, connection pooling
Parser parser.go CSS selector extraction from HTML
Security ssrf.go, sanitizer.go SSRF protection, HTML/selector sanitization
Cache cache.go In-memory TTL cache
Middleware ratelimit.go, metrics.go Rate limiting, Prometheus metrics

Project Structure

scrape-service/
├── cmd/
│   └── server/
│       └── main.go              # Entry point, graceful shutdown
├── internal/
│   ├── config/
│   │   └── config.go            # Environment-based configuration
│   ├── controller/
│   │   └── controller.go        # HTTP handler, validation, cache
│   ├── service/
│   │   └── service.go           # 5-step retry logic
│   ├── httpclient/
│   │   └── client.go            # SSRF-safe HTTP client, connection pooling
│   ├── parser/
│   │   └── parser.go            # CSS selector extraction via goquery
│   ├── security/
│   │   ├── ssrf.go              # SSRF protection, custom dialer
│   │   └── sanitizer.go         # HTML sanitization, selector validation
│   ├── cache/
│   │   └── cache.go             # In-memory cache wrapper
│   ├── errors/
│   │   └── errors.go            # Custom error types
│   └── middleware/
│       ├── ratelimit.go         # Per-IP token bucket rate limiting
│       └── metrics.go           # Prometheus histogram & counter
├── .env.example                 # Example configuration file
├── go.mod
└── README.md

Getting Started

Requirements

Steps

# 1. Clone the repository
git clone <repo-url>
cd scrape-service

# 2. Create the configuration file
cp .env.example .env
# Edit .env if needed — defaults work out of the box

# 3. Download dependencies
go mod download

# 4. Run
go run ./cmd/server

Output:

{"time":"...","level":"INFO","msg":"loaded config","server":{"port":"8080"},"timeouts":{"request_timeout":"30s"},"cache":{"ttl":"1m0s"},"limits":{"max_url_length":2048,"max_selector_length":256},"rate_limit":{"rps":10,"burst":20}}
{"time":"...","level":"INFO","msg":"server listening","addr":"http://localhost:8080"}

Build a Binary

go build -o scrape-service ./cmd/server
./scrape-service

Quick Test

curl -s -X POST http://localhost:8080/scrape \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://news.ycombinator.com",
    "selector": ".titleline",
    "scrapedoToken": "YOUR_TOKEN_HERE"
  }'

Configuration

All configuration is provided via environment variables. You can also create a .env file in the project root.

Variable Default Description
PORT 8080 HTTP server listen port
REQUEST_TIMEOUT_MS 30000 Maximum wait time per Scrape.do step (ms)
CACHE_TTL_SECONDS 60 How long successful responses are cached (seconds)
MAX_URL_LENGTH 2048 Maximum allowed URL length
MAX_SELECTOR_LENGTH 256 Maximum allowed CSS selector length
RATE_LIMIT_RPS 10 Maximum requests per second per IP
RATE_BURST 20 Token bucket burst capacity

.env.example

# Server
PORT=8080

# Timeouts
REQUEST_TIMEOUT_MS=30000

# Cache
CACHE_TTL_SECONDS=60

# Input limits
MAX_URL_LENGTH=2048
MAX_SELECTOR_LENGTH=256

# Rate limiting
RATE_LIMIT_RPS=10
RATE_BURST=20

API Reference

POST /scrape

Scrapes the target URL, extracts the specified CSS selector, and returns it.

Request Body

{
  "url": "https://example.com/products/123",
  "selector": ".product-price",
  "finalData": "$29.99",
  "scrapedoToken": "YOUR_SCRAPE_DO_TOKEN"
}
Field Type Required Description
url string ✅ Target URL to scrape. Must use http or https.
selector string ✅ CSS selector of the content to return (e.g. .price, #title, h1)
finalData string ❌ A string that must appear in the extracted content. When set, cache is bypassed.
scrapedoToken string ✅ Your Scrape.do API token. Never logged.

Success Response 200 OK

<span class="product-price">$29.99</span>
  • Content-Type: text/html; charset=utf-8
  • X-Cache: HIT or X-Cache: MISS header is included

Only the matched selector's HTML is returned. The full page HTML is never returned.

Error Response

{
  "code": "SelectorNotFoundError",
  "message": "selector '.product-price' not found in HTML"
}

GET /health

Returns the service health status.

curl http://localhost:8080/health
{"status": "ok"}

GET /metrics

Returns metrics in Prometheus exposition format. Monitoring systems should scrape this endpoint.

curl http://localhost:8080/metrics

Retry Mechanism

The service escalates through 5 Scrape.do modes in order. At each step the following checks are performed:

  1. Did the HTTP request succeed?
  2. Is the selector present in the returned HTML?
  3. If finalData was provided, does the extracted content contain it?

If all checks pass, the result is returned. If any check fails, the next step is tried.

Step 1: Simple Request
──────────────────────────────────────────────
GET http://api.scrape.do/?token=TOKEN&url=TARGET_URL

Failed → Step 2

Step 2: Super Mode (advanced anti-bot bypass)
──────────────────────────────────────────────
GET ...&super=true

Failed → Step 3

Step 3: Render Mode (JavaScript execution)
──────────────────────────────────────────────
GET ...&super=true&render=true

Failed → Step 4

Step 4: WaitSelector (wait for the selector to appear in the DOM)
──────────────────────────────────────────────
GET ...&super=true&render=true&blockResources=false&waitSelector=SELECTOR

Failed → Step 5

Step 5: PlayWithBrowser (full browser automation)
──────────────────────────────────────────────
GET ...&super=true&render=true&blockResources=false&playwithbrowser=ENCODED_JSON

All steps failed → TimeoutError (504)

PlayWithBrowser Payload

In the final step, the following browser action sequence is sent to Scrape.do:

[
  { "Action": "Wait", "Timeout": 5000 },
  [
    {
      "Action": "WaitSelector",
      "WaitSelector": ".SELECTOR",
      "Timeout": 10000
    }
  ]
]

This payload is URL-encoded and passed as the playwithbrowser query parameter.

Each step has its own 30-second timeout. The maximum total wait time is 5 steps × 30 seconds = 150 seconds.


Cache Behavior

Condition Cache Read? Cache Write?
finalData not set ✅ Yes ✅ Yes
finalData set ❌ No ❌ No

Why is cache skipped when finalData is set?

finalData requires a specific string to be present in the extracted content. This is typically used to confirm that the page has fully loaded and contains live data. Returning a stale cached result in this case would defeat the purpose — a fresh, up-to-date response is always required.

Cache key:

{url}|{selector}

A second request with the same URL and selector combination will return X-Cache: HIT until the CACHE_TTL_SECONDS window expires.


Security

SSRF Protection

The service blocks requests targeting internal networks or loopback addresses. This protection is enforced at the TCP dial level — not just at DNS resolution time — which means DNS rebinding attacks cannot bypass it.

Blocked networks:

CIDR Description
127.0.0.0/8 Loopback
::1/128 IPv6 Loopback
10.0.0.0/8 Private
172.16.0.0/12 Private
192.168.0.0/16 Private
169.254.0.0/16 Link-local (AWS metadata, etc.)
100.64.0.0/10 Shared address space
198.18.0.0/15 Benchmark
0.0.0.0/8 "This" network

XSS Protection

Before the extracted HTML is returned to the client, the following are stripped:

  • All <script> tags and their contents
  • All inline event handler attributes (onclick, onload, onerror, etc.)

Selector Injection Protection

CSS selectors containing the characters {, }, ;, (, ), \ are rejected. These characters could be used to perform CSS injection or manipulate the parser.

Header Security

The following headers are never forwarded to the client:

  • Set-Cookie — prevents session data leakage
  • Authorization — prevents credential leakage

Token Security

The Scrape.do API token is never included in any log line. URLs are logged by hostname only (e.g. example.com).


Logging

The service produces structured JSON logs via log/slog.

Startup — Config Log

When the service starts, all active configuration values are logged:

{
  "level": "INFO",
  "msg": "loaded config",
  "server": { "port": "8080" },
  "timeouts": { "request_timeout": "30s" },
  "cache": { "ttl": "1m0s" },
  "limits": { "max_url_length": 2048, "max_selector_length": 256 },
  "rate_limit": { "rps": 10, "burst": 20 }
}

Per Request — Incoming Log

When a request arrives, all configuration values that apply to that request are logged:

{
  "level": "INFO",
  "msg": "scrape request received",
  "request": {
    "target": "example.com",
    "selector": ".product-price",
    "final_data_set": true
  },
  "config": {
    "cache_enabled": false,
    "cache_ttl": "1m0s",
    "request_timeout": "30s",
    "max_url_length": 2048,
    "max_selector_length": 256,
    "rate_limit_rps": 10,
    "rate_burst": 20
  }
}

Per Request — Completion Log

{
  "level": "INFO",
  "msg": "scrape request completed",
  "target": "example.com",
  "selector": ".product-price",
  "cache": "MISS",
  "result_bytes": 312,
  "duration_ms": 1843
}

Retry Step Log

Each failed retry step is logged individually:

{
  "level": "INFO",
  "msg": "retry step failed",
  "step": "simple",
  "selector": ".product-price",
  "target": "example.com",
  "error": "[SelectorNotFoundError] selector '.product-price' not found in HTML"
}

Error Log

{
  "level": "ERROR",
  "msg": "scrape request failed",
  "target": "example.com",
  "selector": ".product-price",
  "error": "[TimeoutError] request timed out after all retry steps",
  "duration_ms": 5120
}

The API token is never present in any log line. URLs are logged as hostname only.


Monitoring

Prometheus Metrics

The GET /metrics endpoint exposes the following metrics in Prometheus format:

Metric Type Labels Description
scrape_request_duration_seconds Histogram path, status Request duration distribution
scrape_requests_total Counter path, status Total number of requests

Example PromQL Queries

# Average response time over the last 5 minutes
rate(scrape_request_duration_seconds_sum[5m]) / rate(scrape_request_duration_seconds_count[5m])

# Per-second error rate (5xx)
rate(scrape_requests_total{status=~"5.."}[1m])

# 95th percentile latency
histogram_quantile(0.95, rate(scrape_request_duration_seconds_bucket[5m]))

Error Codes

Code HTTP Status When Returned
InvalidSelectorError 400 Selector contains invalid characters
ValidationError 400 Required field is missing or malformed
SelectorNotFoundError 404 Selector was not found after all retry steps
FinalDataNotFoundError 422 Selector was found but finalData string was not present
TargetRequestFailedError 502 Scrape.do could not reach the target URL
TimeoutError 504 All 5 retry steps failed
RateLimitError 429 This IP has sent too many requests
MethodNotAllowed 405 An HTTP method other than POST was used
BadRequest 400 Request body is not valid JSON

Dependencies

Package Version Purpose
PuerkitoBio/goquery v1.9.1 HTML parsing and CSS selector extraction
patrickmn/go-cache v2.1.0 In-memory TTL cache
prometheus/client_golang v1.19.0 Prometheus metrics
golang.org/x/time v0.5.0 Token bucket rate limiter

All other dependencies (cascadia, protobuf, etc.) are indirect dependencies of the above.


How Rate Limiting Works

The token bucket algorithm is used. A separate limiter is created per IP address, stored in a sync.Map for thread safety.

  • RATE_LIMIT_RPS=10 → Each IP can send a maximum of 10 requests per second
  • RATE_BURST=20 → Bursts of up to 20 requests are allowed to handle sudden spikes

When the limit is exceeded:

HTTP 429 Too Many Requests
{"code": "RateLimitError", "message": "too many requests"}

Graceful Shutdown

When the service receives a SIGINT or SIGTERM signal:

  1. It stops accepting new connections
  2. It waits up to 30 seconds for active requests to complete
  3. Requests that finish within that window receive normal responses
  4. After 30 seconds, the server is forcefully shut down
{"level":"INFO","msg":"shutting down gracefully..."}
{"level":"INFO","msg":"server stopped"}

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors