Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 65 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,6 +112,71 @@ public version. Entries below describe the Hub as it stands at that release.

### Fixed

- **Instagram scraped public posts only, and said nothing about it.** Since
2026-07 the scraper ran fully anonymously, because attaching cookies then
made yt-dlp take Instagram's authenticated web API, which 404'd on every
post. That path works again, and anonymous-only had quietly become the
binding constraint: a queue drained 10 of 75 posts per run while the other
65 failed every attempt, 50 of them on Instagram's own ruling ("This
content isn't available to everyone: It can't be seen by certain
audiences", which matched no classifier rule and churned as a retryable
`unknown`) and 15 on yt-dlp's "empty media response". Sampled the day of
the fix, 13 of 13 such posts failed anonymously and 13 of 13 extracted with
the operator's cookies. The scraper now goes anonymous first and retries
once with the session cookies for a post hidden from logged-out viewers —
immediately, instead of burning three anonymous attempts on a wall it
cannot pass — and the media leg follows the metadata leg's auth mode. The
ruling classifies as a login wall; with cookies attached an empty media
response means throttling again, as it did before. Instagram's health check
reports the cookies' real state instead of "anonymous access".

- **A YouTube queue of dead videos could never drain.** Every id in the queue
had already failed, so retries had distilled it down to videos that were
gone or blocked. Three things then interlocked: the metadata leg runs with
`ignore_no_formats_error`, so a refused video still returned an info dict
and was saved as an empty placeholder row (no author, -1 plays, created
2000-01-01 — 224 such rows in the last 25 scrape files); the media leg then
answered a bare "Video unavailable", which reads as a removal but is also
what a throttled session returns, so the verdict is deliberately
distrusted; and 15 of those in a row tripped the permanent-storm guard,
which aborts the batch — and an aborted batch charges no retry budget. The
result was 190 items queued, 45 attempted, 0 drained, run after run, with
the queue file untouched for two days. Fixed by giving the metadata leg the
tv player client, the only one that states *why* YouTube will not play a
video, and capturing the reason the flag otherwise swallows. A video the
platform has no record of is now a failure rather than a placeholder row,
and both it and a video whose record is intact but which names a region
whitelist or a rights claim are treated as *corroborated*: verdicts backed
by per-item evidence, which neither feed the storm guards nor are demoted
by them. A bare "Video unavailable" with the record intact is still
distrusted, so the protection added after 2026-09-18 stands. Replayed
against the live queue, the next run prunes 187 of 190 and downloads the
one video that plays.

- **YouTube read two spellings of the same failure oppositely.** "Video
unavailable" classified as a permanent removal while "This video is
unavailable" matched no rule and churned as a retryable `unknown` — every
one of the nine seen was gone for good. Takedown and block phrasings the tv
client reports are classified too: copyright claim blocks as their own
`blocked` category (kept distinct from `geo_blocked` and `removed` so a run
from another vantage point can single them out), copyright *takedowns* as
removals, and a captcha challenge as a bot check.

- **One item's success reset every other item's retry strikes.** A batch that
pruned anything deleted the whole per-platform strike sidecar, so a queue
that trickled forward could carry a tail that failed every run
indefinitely — no strikes ever accumulated against it. Progress now clears
only the strikes of the ids it actually pruned, and every id that leaves
the queue drops its media-retry strikes even when the batch was aborted.

- **Instagram and YouTube scraping is refused on Cloud Run.** Neither works
from a datacenter IP whatever cookies or PO tokens are attached, and a run
there would burn the queue and trip guards whose state, shared through the
bucket, then holds off the local install that can actually drain it. Both
scrapers are marked `residential_ip_only`: the queue worker declines to
start and the enrichment supervisor leaves the queue alone without charging
the plan a stall. TikTok is unaffected.

- **Viability floors counted rows that are not viewing.** A donation was
admitted on 10+ rows of *any* kind, so an export could become a collection
with no viewing activity in it at all. TikTok was the clearest case:
Expand Down
5 changes: 5 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,11 @@ python web_interface/run_queue_annotator.py
python web_interface/run_queue_scraper.py --platform tiktok
```

The Instagram and YouTube scrapers have to be run this way, from a
residential connection: both platforms wall off Cloud Run's datacenter IPs
whatever cookies are attached, so the deployed services decline to scrape
them and leave those queues to a local install.

## Verification

Every change should pass the gate before merging:
Expand Down
7 changes: 6 additions & 1 deletion docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -163,7 +163,12 @@ identical *permanent* classifications (a flagged session mis-reporting live
items as removed) or identical *transient* ones (a bot wall failing every
item retryably) abort the batch, stop self-chaining, and raise a persistent
per-platform scraper alert; the failed-scrapes record stores each item's
failure category so storms are diagnosable after the fact. On the analysis
failure category so storms are diagnosable after the fact. Because those
guards read a homogeneous run as a broken session, and a queue of nothing but
retries is homogeneous by construction, a scraper can mark a verdict
**corroborated** by per-item evidence — it then neither extends nor resets a
storm run and is pruned even when the guard trips (see
[pipeline.md](pipeline.md)). On the analysis
side, every study refresh writes a
**methods/provenance note** (`{study}_methods.json`) summarising filters,
counts, and the contract/model versions behind the data — see
Expand Down
3 changes: 2 additions & 1 deletion docs/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,8 @@ Slack and email integrations silently no-op when unconfigured. Everything stores
| **Python 3.12** | everything | Matches production (`python:3.12-slim` on Cloud Run); `ruff` and CI target 3.12 too. `brew install python@3.12` / `apt install python3.12` |
| `ffmpeg` | YouTube HD media only | yt-dlp needs it to merge DASH video+audio. TikTok/Instagram downloads and photo-slideshow assembly work without it (bundled `imageio-ffmpeg`). `brew install ffmpeg` / `apt install ffmpeg` |
| `node` *or* `deno` | YouTube media from datacenter IPs | Runs yt-dlp's JS challenge solver. Usually unnecessary on a home (residential) connection. |
| Google Chrome, logged in | authenticated scraping | Cookies are read from the local Chrome profile — **macOS only** (approve the Keychain prompt on first use). Instagram scraping effectively requires this; TikTok/YouTube degrade to public-content access. On Linux, provide a Netscape cookies file instead: `YTDLP_COOKIE_FILE_TIKTOK` / `_INSTAGRAM` / `_YOUTUBE` for one platform, or `YTDLP_COOKIE_FILE` for all of them (the per-platform form wins; either is ignored unless the file exists). |
| Google Chrome, logged in | authenticated scraping | Cookies are read from the local Chrome profile — **macOS only** (approve the Keychain prompt on first use). YouTube scrapes signed in; Instagram needs the cookies for posts it hides from logged-out viewers (public posts scrape anonymously); TikTok degrades to public-content access. On Linux, provide a Netscape cookies file instead: `YTDLP_COOKIE_FILE_TIKTOK` / `_INSTAGRAM` / `_YOUTUBE` for one platform, or `YTDLP_COOKIE_FILE` for all of them (the per-platform form wins; either is ignored unless the file exists). |
| A residential connection | Instagram and YouTube scraping | Both wall off datacenter IPs whatever cookies or PO tokens are attached, so their scrape queues are drained by a local install and the Cloud Run deployment declines to run them (`BaseScraper.residential_ip_only`). TikTok scrapes fine from Cloud Run. |

## Install

Expand Down
62 changes: 62 additions & 0 deletions docs/pipeline.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,6 +114,68 @@ plus optional overrides (throttle limits, health check, slideshow hooks).
All three current scrapers (TikTok, Instagram, YouTube) are yt-dlp-based;
cookies are managed per-platform by `scraper_cookies.py`.

**Where each scraper runs.** Instagram and YouTube do not scrape from
datacenter IPs in practice — both wall Cloud Run off whatever cookies or
PO tokens are attached — so their queues are drained by a local install on a
residential IP. That is a property of the scraper
(`BaseScraper.residential_ip_only`): on Cloud Run the queue worker refuses to
start and the enrichment supervisor leaves those queues alone, instead of
burning them against the wall and tripping guards whose state, shared through
the bucket, would then hold the local install off too.

**Authentication.** TikTok and YouTube scrape signed in (locally, from the
operator's own Chrome profile). Instagram goes anonymous first and retries
with the session cookies only for a post it hides from logged-out viewers,
which keeps the account's footprint to the posts that need it and survives
either path breaking — both have (2026-07: attaching cookies broke every
extraction; 2026-09: anonymous alone left 65 of 75 queued posts unfetchable).

**Permanent vs transient.** Each scraper classifies a failure into its own
taxonomy, and `classify_error` maps it to `permanent:<reason>` (pruned from
the queue, recorded in the failed-scrapes ledger) or `transient:<reason>`
(kept for a later run). One broken session can make every item read as
permanently gone, so three guards sit on top — all of them abort the batch,
and the storm guards also raise a scraper alert for a human:

| Guard | Trips on | The items |
|---|---|---|
| circuit breaker | 15 consecutive throttle verdicts | stay queued |
| permanent-storm guard | 15 consecutive identical *permanent* verdicts | demoted to transient, stay queued |
| transient-storm guard | 25 consecutive identical *transient* verdicts | already transient; chaining stops |

**Corroborated verdicts.** Those guards assume a healthy queue produces
heterogeneous outcomes — which a queue of nothing but retries never does,
since retrying only the failures distils it down to items that fail. A
scraper may therefore mark a permanent verdict *corroborated*
(`attrs['verdict_corroborated']`, see `BaseScraper.fetch`) when it rests on
per-item evidence independent of the error text: such a verdict neither
extends nor resets a storm run, and is pruned even when the guard trips.
YouTube corroborates two cases, both read from the metadata leg, which adds
the tv player client — the only one that states *why* a video will not play —
and captures the reason yt-dlp otherwise swallows under
`ignore_no_formats_error`:

* the platform has no record of the video (no channel, no view count, no
duration) **and** the stated reason is itself a removal. A video with no
record is a failure, never the empty placeholder row it used to be saved as;
* the record is intact but YouTube refuses to play it here, naming a region
whitelist or a rights claim. Its metadata is scraped, the media leg is
skipped, and the id leaves the queue with its metadata-only row standing.

A bare "Video unavailable" with the record intact is never corroborated: that
is exactly what a throttled session returns.

**Retry budgets** bound everything that stays queued, in per-platform sidecars
(`scrape_queues.py`). An item transiently failing through
`MAX_ZERO_PROGRESS_STRIKES` runs in which the queue as a whole made no
progress is given up on and recorded as failed; an item whose metadata scraped
but whose media did not is retried for `MAX_MEDIA_RETRY_STRIKES` healthy runs
and then pruned with its metadata-only row standing. An aborted batch never
charges either budget — the verdicts implicate the session, not the items —
but ids that left the queue always drop their strikes, and a batch that makes
progress clears only the strikes of the ids it pruned (one item's success is
no evidence for another's).

The canonical cross-platform scrape schema lives in
`config/scrape_contract.toml`: base fields every platform emits (including
generic popularity counts `fave_count`/`comment_count`/... and per-K
Expand Down
Loading
Loading