Skip to content

Spend Instagram's logged-in session sparingly, and stop using it once Instagram logs it out - #6

Merged
pwikstrom merged 1 commit into
mainfrom
claude/scraper-failures-instagram-youtube-f90476
Sep 22, 2026
Merged

pwikstrom merged 1 commit into
mainfrom
claude/scraper-failures-instagram-youtube-f90476

Conversation

@pwikstrom

Copy link
Copy Markdown
Owner

The first run with the cookie fallback scraped 14 gated posts in about a
minute -- two logged-in API calls each, two threads, no pacing -- and
Instagram then logged the session out. It answers a logged-out session's API
call with a 200 redirect to its login page; yt-dlp falls back to logged-out
extraction only when that redirect arrives as an HTTP error, so it surfaced
as "Failed to parse JSON", a retryable unknown, and the scraper sent 87 more
attempts into the dead session before the transient-storm guard stopped it.

  • A logged-out session is a new session_expired verdict. It ends every
    further logged-in request for the run (later gated posts skip the cookies),
    keeps the items queued, charges neither retry budget, stays out of both
    storm guards, stops the run, raises a scraper alert asking for a fresh
    login -- ahead of the storm alerts, since a platform keeps one alert -- and
    holds the enrichment supervisor off until a run completes without it.

  • Logged-in requests are spaced across threads by
    [misc] scraper_instagram_auth_interval (default 20 s). Public posts keep
    anonymous speed.

  • The media leg downloads from the info dict the metadata leg already
    extracted, so a gated post costs one logged-in call instead of two; only a
    retry re-extracts, paced. Both legs now select the same format: an info
    dict processed with yt-dlp's default DASH selection drew HTTP 403 from the
    CDN when a second instance reprocessed it (cookies and impersonation were
    ruled out). Measured live: one extraction per post, where it was two.

… Instagram logs it out

The first run with the cookie fallback scraped 14 gated posts in about a
minute -- two logged-in API calls each, two threads, no pacing -- and
Instagram then logged the session out. It answers a logged-out session's API
call with a 200 redirect to its login page; yt-dlp falls back to logged-out
extraction only when that redirect arrives as an HTTP error, so it surfaced
as "Failed to parse JSON", a retryable unknown, and the scraper sent 87 more
attempts into the dead session before the transient-storm guard stopped it.

- A logged-out session is a new session_expired verdict. It ends every
  further logged-in request for the run (later gated posts skip the cookies),
  keeps the items queued, charges neither retry budget, stays out of both
  storm guards, stops the run, raises a scraper alert asking for a fresh
  login -- ahead of the storm alerts, since a platform keeps one alert -- and
  holds the enrichment supervisor off until a run completes without it.

- Logged-in requests are spaced across threads by
  [misc] scraper_instagram_auth_interval (default 20 s). Public posts keep
  anonymous speed.

- The media leg downloads from the info dict the metadata leg already
  extracted, so a gated post costs one logged-in call instead of two; only a
  retry re-extracts, paced. Both legs now select the same format: an info
  dict processed with yt-dlp's default DASH selection drew HTTP 403 from the
  CDN when a second instance reprocessed it (cookies and impersonation were
  ruled out). Measured live: one extraction per post, where it was two.
@pwikstrom
pwikstrom merged commit c63f143 into main Sep 22, 2026
1 check passed
@pwikstrom
pwikstrom deleted the claude/scraper-failures-instagram-youtube-f90476 branch September 28, 2026 07:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant