Spend Instagram's logged-in session sparingly, and stop using it once Instagram logs it out - #6
Merged
Conversation
… Instagram logs it out The first run with the cookie fallback scraped 14 gated posts in about a minute -- two logged-in API calls each, two threads, no pacing -- and Instagram then logged the session out. It answers a logged-out session's API call with a 200 redirect to its login page; yt-dlp falls back to logged-out extraction only when that redirect arrives as an HTTP error, so it surfaced as "Failed to parse JSON", a retryable unknown, and the scraper sent 87 more attempts into the dead session before the transient-storm guard stopped it. - A logged-out session is a new session_expired verdict. It ends every further logged-in request for the run (later gated posts skip the cookies), keeps the items queued, charges neither retry budget, stays out of both storm guards, stops the run, raises a scraper alert asking for a fresh login -- ahead of the storm alerts, since a platform keeps one alert -- and holds the enrichment supervisor off until a run completes without it. - Logged-in requests are spaced across threads by [misc] scraper_instagram_auth_interval (default 20 s). Public posts keep anonymous speed. - The media leg downloads from the info dict the metadata leg already extracted, so a gated post costs one logged-in call instead of two; only a retry re-extracts, paced. Both legs now select the same format: an info dict processed with yt-dlp's default DASH selection drew HTTP 403 from the CDN when a second instance reprocessed it (cookies and impersonation were ruled out). Measured live: one extraction per post, where it was two.
pwikstrom
deleted the
claude/scraper-failures-instagram-youtube-f90476
branch
September 28, 2026 07:30
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The first run with the cookie fallback scraped 14 gated posts in about a
minute -- two logged-in API calls each, two threads, no pacing -- and
Instagram then logged the session out. It answers a logged-out session's API
call with a 200 redirect to its login page; yt-dlp falls back to logged-out
extraction only when that redirect arrives as an HTTP error, so it surfaced
as "Failed to parse JSON", a retryable unknown, and the scraper sent 87 more
attempts into the dead session before the transient-storm guard stopped it.
A logged-out session is a new session_expired verdict. It ends every
further logged-in request for the run (later gated posts skip the cookies),
keeps the items queued, charges neither retry budget, stays out of both
storm guards, stops the run, raises a scraper alert asking for a fresh
login -- ahead of the storm alerts, since a platform keeps one alert -- and
holds the enrichment supervisor off until a run completes without it.
Logged-in requests are spaced across threads by
[misc] scraper_instagram_auth_interval (default 20 s). Public posts keep
anonymous speed.
The media leg downloads from the info dict the metadata leg already
extracted, so a gated post costs one logged-in call instead of two; only a
retry re-extracts, paced. Both legs now select the same format: an info
dict processed with yt-dlp's default DASH selection drew HTTP 403 from the
CDN when a second instance reprocessed it (cookies and impersonation were
ruled out). Measured live: one extraction per post, where it was two.