|
All checks were successful
CI Trade-In / changes (pull_request) Successful in 8s
CI / changes (pull_request) Successful in 8s
CI Trade-In / frontend-checks (pull_request) Has been skipped
CI / backend-tests (pull_request) Has been skipped
CI / openapi-codegen-check (pull_request) Has been skipped
CI Trade-In / backend-tests (pull_request) Successful in 53s
CI / frontend-tests (pull_request) Has been skipped
Ports the cookie-injection + organic SERP-origin session->cookies->fetch_detail wiring (PR #2430, PR #2433), previously only reachable via the manual debug endpoint, into a scheduled orchestrator mirroring avito_detail_backfill.py's shape: snapshot-once-then-loop, budget guard, SIGTERM-drain, jittered delay, heartbeat every 25 attempts. DomClickBlockedError (QRATOR challenge / browser-fetch failure) increments a consecutive-block counter and aborts via mark_done (not mark_failed) past max_consecutive_blocks -- a block is an expected operational outcome for a QRATOR reputation-based anti-bot, not a task failure. DomClickParseError (__SSR_STATE__ schema drift) is a plain failed++, neutral to the breaker. Unlike Avito there is no curl fallback (one BrowserFetcher per run) and no IP-rotation recovery (single dedicated residential proxy). Migration 175 seeds the schedule with enabled=false -- activation is a deliberate manual step after smoke-testing against a live cookie session. default_params are more conservative than Avito's (batch_size=200, request_delay_sec=12, max_consecutive_blocks=3) since a burned QRATOR reputation/session is costlier to recover than an Avito IP-block. Also fixes a stale docstring in scraper_kit's domclick/detail.py that incorrectly named DataDome as the anti-bot mechanism -- it's QRATOR (exit-IP reputation-based), per the extensive live verification in #2000. |
||
|---|---|---|
| .. | ||
| app | ||
| data/sql | ||
| scripts | ||
| tests | ||
| .dockerignore | ||
| Dockerfile | ||
| pyproject.toml | ||