Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -166,6 +166,26 @@ public version. Entries below describe the Hub as it stands at that release.
in other collections (e.g. left by a data migration). The run now reports
the two apart, per collection, and the panel explains each in plain words.

- **Instagram logged the scraper's session out, and the scraper kept using
it.** The first run with the cookie fallback scraped 14 gated posts in
about a minute — two logged-in API calls each, two threads, no pacing —
and Instagram then logged the session out, answering the logged-in API
with its login page. yt-dlp falls back to logged-out extraction only when
that redirect arrives as an HTTP error, so it surfaced as "Failed to parse
JSON", a retryable `unknown`, and the scraper sent 87 more attempts into
the dead session before the transient-storm guard stopped it. Three
changes: logged-in requests are now spaced across threads
(`[misc] scraper_instagram_auth_interval`, default 20 s); the media leg
downloads from the info dict the metadata leg already extracted, so a gated
post costs one logged-in call instead of two (both legs now select the same
format — reprocessing a default DASH selection in a second yt-dlp instance
drew HTTP 403 from the CDN every time); and the login page in place of
JSON is a new `session_expired` verdict. It stops every further logged-in
request that run, keeps the items queued without charging either retry
budget, stops the run, stays out of the storm guards, raises a scraper
alert asking for a fresh login in Chrome, and holds the enrichment
supervisor off until a run completes without it.

- **Instagram scraped public posts only, and said nothing about it.** Since
2026-07 the scraper ran fully anonymously, because attaching cookies then
made yt-dlp take Instagram's authenticated web API, which 404'd on every
Expand Down
10 changes: 10 additions & 0 deletions docs/pipeline.md
Original file line number Diff line number Diff line change
Expand Up @@ -137,6 +137,16 @@ with the session cookies only for a post it hides from logged-out viewers,
which keeps the account's footprint to the posts that need it and survives
either path breaking — both have (2026-07: attaching cookies broke every
extraction; 2026-09: anonymous alone left 65 of 75 queued posts unfetchable).
The logged-in session is spent sparingly, because Instagram logs it out when
it is not (2026-09-23, after ~14 gated posts in a minute): logged-in requests
are spaced `[misc] scraper_instagram_auth_interval` seconds apart across
threads (default 20), and the media leg downloads from the info dict the
metadata leg already extracted, so a gated post costs one logged-in call.
A logged-out session shows up as Instagram's login page where the API's JSON
should be; the scraper then classifies `session_expired`, makes no further
logged-in request that run, and the orchestrator stops the run without
charging any retry budget and raises a scraper alert asking for a fresh
login in Chrome.

**Permanent vs transient.** Each scraper classifies a failure into its own
taxonomy, and `classify_error` maps it to `permanent:<reason>` (pruned from
Expand Down
123 changes: 116 additions & 7 deletions fyp/scrape/instagram_dl.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,15 @@
account") stays permanently ``private``; the donated enrichment seed still
surfaces its caption/author.

The logged-in session is spent sparingly, because Instagram logs it out when
it is not (2026-09-23: after ~14 gated posts in ~60 s). Every logged-in request
is spaced ``[misc] scraper_instagram_auth_interval`` seconds apart across
threads (default 20); the media leg downloads from the info dict the metadata
leg already extracted, so a gated post costs one logged-in call, not two; and
Instagram's login page in place of the API's JSON — how a logged-out session
shows up — is ``session_expired``: no further logged-in request is made that
run, the run stops, and a scraper alert asks for a fresh login.

Image-only posts (single photos and carousels): extraction uses yt-dlp's
``ignore_no_formats_error`` so an image post returns a full info dict; the
image URLs come from its thumbnails (single post) or its playlist entries'
Expand All @@ -45,7 +54,7 @@
from json import loads as json_loads
from os import remove
from os.path import exists, join
from time import sleep
from time import monotonic, sleep

import pandas as pd
import yt_dlp
Expand All @@ -54,6 +63,7 @@

from fyp.scrape import scraper_cookies
from fyp.scrape.platform_scraper import (
SESSION_EXPIRED,
SLIDESHOW_SECONDS_PER_IMAGE,
BaseScraper,
cleanup_temp_files,
Expand Down Expand Up @@ -89,6 +99,67 @@ def _cf():
_META_MAX_RETRIES = 3
_DL_MAX_RETRIES = 2

# One format for both legs, so the info dict the metadata leg returns can be
# downloaded as-is (see _download_media). Processed with yt-dlp's default
# selection (DASH video+audio), an info dict carries that selection's residue,
# and a second YoutubeDL instance re-processing it for this format got HTTP 403
# from the CDN every time (measured 2026-09-23); with one format it downloads.
_FORMAT = 'best[ext=mp4]/best'

# The logged-in session. On 2026-09-23 Instagram logged it out after ~14
# authenticated posts in ~60 s — two logged-in API calls each, two threads, no
# pacing — and the scraper then sent 87 more attempts into the dead session.
# Logged-in requests are now spaced across threads, and once the session is
# seen logged out no further logged-in request is made in this process (one
# process is one scraper run).
_AUTH_LOCK = threading.Lock()
_auth_next_at = [0.0]
_SESSION_DEAD = threading.Event()


def _auth_interval() -> float:
"""Minimum seconds between logged-in requests (``[misc] scraper_instagram_auth_interval``)."""
try:
return max(0.0, float(_cf()["misc"].get("scraper_instagram_auth_interval", 20)))
except Exception:
return 20.0


def _pace_authenticated() -> None:
"""Block until the next logged-in request may go, across all threads."""
with _AUTH_LOCK:
wait = _auth_next_at[0] - monotonic()
if wait > 0:
sleep(wait)
_auth_next_at[0] = monotonic() + _auth_interval()


def _session_logged_out(detail: str | None) -> bool:
"""True when a logged-in request got Instagram's login page instead of JSON.

Instagram answers a logged-out session's API call with a 200 redirect to
``/accounts/login/``. yt-dlp falls back to logged-out extraction only when
that redirect arrives as an HTTP error, so this one surfaces as "Failed to
parse JSON" — a retryable ``unknown`` until 2026-09-23, retried three times
per post.
"""
return 'failed to parse json' in (detail or '').lower()


def _mark_session_dead(item_id: str) -> None:
"""Record that the session was logged out; say so once, loudly."""
if not _SESSION_DEAD.is_set():
_SESSION_DEAD.set()
logger.warning("Scrape %s: Instagram answered a logged-in request with its login page — "
"the session was logged out. No further logged-in requests this run; "
"log in to Instagram in Chrome again before the next one.", item_id)


def _reset_session_state() -> None:
"""Forget the pacing clock and a dead session (tests; a run is one process)."""
_SESSION_DEAD.clear()
_auth_next_at[0] = 0.0


def _classify_error(exc: Exception) -> tuple[str, str]:
"""Classify a yt-dlp error into (category, detail) for retry decisions.
Expand Down Expand Up @@ -251,6 +322,10 @@ def _extract_metadata(url: str, item_id: str, verbose: bool = False):
cookies = scraper_cookies.cookie_opts("instagram")
if not cookies:
return info, fail
if _SESSION_DEAD.is_set():
return None, _empty_fail(SESSION_EXPIRED,
"hidden from logged-out viewers, and the session was logged "
"out earlier in this run — not retried with the cookies")
logger.info("Scrape %s: hidden from logged-out viewers — retrying with the session cookies",
item_id)
info, fail = _extract_metadata_as(url, item_id, cookies, verbose=verbose)
Expand All @@ -265,20 +340,29 @@ def _extract_metadata_as(url: str, item_id: str, cookies: dict, verbose: bool =
A login wall ends the pass at once: repeating the same request cannot get
past it (the anonymous pass used to burn three attempts per gated post).
With the cookies attached, an empty media response is throttling again
and retries with backoff like any rate limit.
and retries with backoff like any rate limit; every attempt is paced
(:func:`_pace_authenticated`), and Instagram's login page in place of the
API's JSON ends the pass as ``session_expired`` and stops all further
logged-in requests this run.
"""
ydl_opts: dict = {
'quiet': True,
'no_warnings': not verbose,
**cookies,
'skip_download': True,
'no_color': True,
'format': _FORMAT,
'ignore_no_formats_error': True,
'extractor_retries': 3,
'socket_timeout': 30,
}

for attempt in range(_META_MAX_RETRIES):
if cookies:
if _SESSION_DEAD.is_set():
return None, _empty_fail(SESSION_EXPIRED, "the session was logged out earlier "
"in this run — not retried")
_pace_authenticated()
try:
with yt_dlp.YoutubeDL(ydl_opts) as ydl:
return ydl.extract_info(url, download=False), None
Expand All @@ -287,6 +371,9 @@ def _extract_metadata_as(url: str, item_id: str, cookies: dict, verbose: bool =
logger.warning("Scrape %s metadata attempt %d/%d failed%s: [%s] %s",
item_id, attempt + 1, _META_MAX_RETRIES,
" (with cookies)" if cookies else "", category, detail)
if cookies and _session_logged_out(detail):
_mark_session_dead(item_id)
return None, _empty_fail(SESSION_EXPIRED, detail)
if category == "login_required" or (not cookies and _login_gated(category, detail)):
return None, _empty_fail(category, detail)
if category in _RETRYABLE and attempt < _META_MAX_RETRIES - 1:
Expand Down Expand Up @@ -334,13 +421,21 @@ def _download_media(
stream_to_bucket=None,
verbose: bool = False,
authenticated: bool = False,
info: dict | None = None,
) -> tuple[bool, str | None, str, float | None]:
"""Download the post's video to temp and move/upload it.

The first attempt downloads from ``info`` — the dict the metadata leg
already extracted — so it makes no Instagram request at all: a gated post
costs one logged-in API call, not two (2026-09-23: the re-extraction
doubled the logged-in traffic that got the session logged out). Only a
retry re-extracts the post from ``url``, paced like any logged-in request.

Args:
authenticated: Attach the session cookies — set when the metadata leg
needed them (the download re-extracts the post, so a post hidden
from logged-out viewers fails anonymously here too).
needed them, because a re-extraction of a post hidden from
logged-out viewers fails anonymously too.
info: The metadata leg's info dict, extracted with :data:`_FORMAT`.

Returns:
``(ok, error_category, error_detail, duration)`` — category/detail are
Expand All @@ -358,16 +453,27 @@ def _download_media(
'outtmpl': out_template,
'no_color': True,
'overwrites': True,
'format': 'best[ext=mp4]/best',
'format': _FORMAT,
'merge_output_format': 'mp4',
'retries': 3,
'socket_timeout': 30,
}

for attempt in range(_DL_MAX_RETRIES):
reuse = info is not None and attempt == 0
if authenticated and not reuse:
if _SESSION_DEAD.is_set():
return (False, SESSION_EXPIRED, "the session was logged out earlier in this "
"run — not re-extracted with the cookies", None)
_pace_authenticated()
try:
with yt_dlp.YoutubeDL(dl_opts) as ydl:
ydl.download([url])
if reuse:
ydl.process_ie_result(
{k: v for k, v in info.items() if not k.startswith('_fyp_')},
download=True)
else:
ydl.download([url])

downloaded = join(temp_dir, f"{item_id}.mp4")
if not exists(downloaded):
Expand Down Expand Up @@ -404,6 +510,9 @@ def _download_media(
logger.warning("Scrape %s download attempt %d/%d failed: [%s] %s",
item_id, attempt + 1, _DL_MAX_RETRIES, category, detail)
_cleanup_temp_files(temp_dir, item_id)
if authenticated and not reuse and _session_logged_out(detail):
_mark_session_dead(item_id)
return False, SESSION_EXPIRED, detail, None
if category in _RETRYABLE and attempt < _DL_MAX_RETRIES - 1:
backoff = 3 * (3 ** attempt)
logger.info("Retrying download %s in %ds...", item_id, backoff)
Expand Down Expand Up @@ -970,7 +1079,7 @@ def fetch(
ok, media_category, media_detail, media_duration = _download_media(
url, item_id, save_path,
stream_to_bucket=stream_to_bucket, verbose=verbose,
authenticated=bool(info.get('_fyp_authenticated')))
authenticated=bool(info.get('_fyp_authenticated')), info=info)
if ok:
data_row.loc[0, 'video_downloaded'] = True
# Backfill the duration metadata extraction no longer returns
Expand Down
9 changes: 9 additions & 0 deletions fyp/scrape/platform_scraper.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,15 @@
logger = logging.getLogger(__name__)


# Error category for "the platform logged the scraper's OWN session out
# mid-run" (2026-09-23: Instagram answered the logged-in API with its login
# page after ~14 authenticated posts). It says nothing about the item: it is
# transient (the item stays queued), a batch that saw it charges no retry
# budget, it feeds neither storm guard, and the orchestrator stops the run and
# raises a scraper alert — only an operator logging in again can clear it.
SESSION_EXPIRED = "session_expired"


def empty_fail(error_type: str = "unknown", error_detail: str = "", *,
corroborated: bool = False) -> pd.DataFrame:
"""Return an empty DataFrame tagged with error classification metadata.
Expand Down
39 changes: 38 additions & 1 deletion fyp/scrape/scrape.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@
from fyp.scrape import scrape_contract as sc
from fyp.scrape import scrape_queues, scrape_versioning, scraper_alerts
from fyp.scrape.platform_scraper import (
SESSION_EXPIRED,
SLIDESHOW_SECONDS_PER_IMAGE,
THROTTLE_CATEGORIES,
ThrottleController,
Expand Down Expand Up @@ -1093,6 +1094,11 @@ def download_video_threads(
# chaining; the items are transient so they stay queued as-is.
t_storm_state = {"classification": None, "consecutive": 0, "tripped": False}
t_storm_threshold = _transient_storm_threshold()
# The platform logged the scraper's own session out (SESSION_EXPIRED).
# The scraper stops its logged-in requests itself and the batch carries on
# with what needs no login; this flag stops the run after it, keeps the
# batch from charging retry budget, and raises the session alert.
session_state = {"expired": False}
abort_event = threading.Event()
# Memory safety valve: set once the container's memory cgroup crosses
# MEMORY_STOP_FRACTION. Workers that have not started downloading yet defer
Expand All @@ -1113,6 +1119,12 @@ def _breaker_track(category, corroborated: bool = False) -> None:
f"aborting batch; remaining items stay queued.")
else:
breaker_state["consecutive"] = 0
if category == SESSION_EXPIRED:
# A known cause, not a mystery: kept out of both storm runs
# (neither extended nor reset) so the session alert is the one
# an operator sees.
session_state["expired"] = True
return
if storm_state["tripped"] or t_storm_state["tripped"]:
# Frozen once tripped: post-abort "batch_aborted" results are
# transient and would otherwise wipe the storm classification.
Expand Down Expand Up @@ -1422,7 +1434,21 @@ def _mem_watch():
# results again with no storm. Best-effort on both sides (never blocks
# scraping), and skipped in dry runs.
if not dry_run:
if storm_state["tripped"]:
if session_state["expired"]:
# First: it is the one failure only a person can fix, and a
# platform keeps a single alert.
scraper_alerts.raise_alert(
platform=scraper.platform,
kind=scraper_alerts.KIND_SESSION_EXPIRED,
category=f"transient:{SESSION_EXPIRED}",
message=(
f"{scraper.platform.capitalize()} logged the scraper's session out during "
f"scraping. Log in to {scraper.platform.capitalize()} again in Chrome (as the "
f"research account), then re-run the scraper. Items that need the login "
f"stay queued and were not charged any retry budget."
),
)
elif storm_state["tripped"]:
scraper_alerts.raise_alert(
platform=scraper.platform,
kind=scraper_alerts.KIND_PERMANENT_STORM,
Expand Down Expand Up @@ -1486,6 +1512,7 @@ def _mem_watch():
empty_results.attrs['transient_storm_tripped'] = t_storm_state["tripped"]
empty_results.attrs['transient_storm_category'] = t_storm_state["classification"]
empty_results.attrs['memory_stop'] = mem_stop_event.is_set()
empty_results.attrs['session_expired'] = session_state["expired"]
empty_results.attrs['batch_deadline_hit'] = deadline_hit
empty_results.attrs['media_retry_ids'] = list(media_retry_ids)
return empty_results, permanent_failed_ids, transient_failed_ids
Expand Down Expand Up @@ -1513,6 +1540,7 @@ def _mem_watch():
results.attrs['transient_storm_tripped'] = t_storm_state["tripped"]
results.attrs['transient_storm_category'] = t_storm_state["classification"]
results.attrs['memory_stop'] = mem_stop_event.is_set()
results.attrs['session_expired'] = session_state["expired"]
results.attrs['batch_deadline_hit'] = deadline_hit
# Ids saved metadata-only (media failed): callers charge the media-retry
# budget with these — they are also in transient_failed_ids so the queue
Expand Down Expand Up @@ -1620,6 +1648,15 @@ def _on_threads_change(n):
logger.warning(" Batch deadline hit — the completed rows are saved; the "
"unfinished items stay in the queue for the next run.")

if results_from_scraper.attrs.get('session_expired'):
logger.warning(" The platform logged the scraper's session out — stopping the batch "
"loop. Items that need the login stay in the queue, uncharged; log in "
"again in Chrome, then re-run the scraper.")
aborted = True
if reporter is not None:
reporter.emit_data({"session_expired": True})
break

if results_from_scraper.attrs.get('circuit_breaker_tripped'):
logger.warning(" Rate-limit circuit breaker tripped — stopping the batch loop. "
"Unfinished items stay in the queue; re-run the scraper later.")
Expand Down
4 changes: 4 additions & 0 deletions fyp/scrape/scraper_alerts.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,10 @@
# A run of consecutive identical *transient* verdicts (e.g. a new bot wall
# breaking the extractor for every item — 2026-08-10 TikTok incident).
KIND_TRANSIENT_STORM = "transient_storm"
# The platform logged the scraper's own login session out mid-run
# (platform_scraper.SESSION_EXPIRED; 2026-09-23 Instagram). Needs a human to
# log in again, so it outranks the storm alerts.
KIND_SESSION_EXPIRED = "session_expired"



Expand Down
Loading
Loading