6 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5015109f31 |
refactor(tradein/tasks): newbuilding_enrich_backfill на kit cian.newbuilding + config= (#2397 Part D3)
All checks were successful
CI Trade-In / changes (pull_request) Successful in 8s
CI / changes (pull_request) Successful in 7s
CI Trade-In / frontend-checks (pull_request) Has been skipped
CI / backend-tests (pull_request) Has been skipped
CI / frontend-tests (pull_request) Has been skipped
CI / openapi-codegen-check (pull_request) Has been skipped
CI Trade-In / backend-tests (pull_request) Successful in 1m2s
|
||
|
|
85059aeb1b |
refactor(tradein/scheduler): удалить legacy scheduler_loop + scraper-scheduling, kit единственный путь (#2397 Part C)
Топология подтверждена перед удалением (docker-compose.prod.yml): tradein-backend (uvicorn app.main:app) — SCHEDULER_ENABLE=false; tradein-scraper (python -m app.scheduler_main) — SCHEDULER_ENABLE=true + USE_KIT_SCHEDULER=true. Kit-путь (_run_kit_scheduler → scraper_kit.orchestration.scheduler + product_handlers) самодостаточен: не импортирует ничего из app.services.scheduler.scheduler_loop или app.services.scrape_pipeline. Все НЕ-sweep джобы, которые kit-scheduler диспетчерит через build_product_handlers, идут напрямую в app.tasks.*/ app.services.* (либо lazy-импортят import_rosreestr_dkp/_execute_cian_backfill из scheduler.py) — мимо удаляемой legacy-машинерии. app/services/scheduler.py: 2098 → 418 строк. Удалено: scheduler_loop, get_due_schedules, reap_zombies, _claim_run, _defer_next_run_at, _spawn_tracked/ _drain_inflight/_inflight_tasks, все 27 trigger_*_run-функций, импорт app.services.scrape_pipeline, константы SCHEDULER_TICK_SEC/ZOMBIE_THRESHOLD_HOURS (достижимы были только через удалённый scheduler_loop-путь). Оставлено (живые импортёры вне удалённого): compute_next_run_at + has_running_run (admin.py), import_rosreestr_dkp + _execute_cian_backfill (lazy-импорты в product_handlers.py — job-тела kit-handler'ов). main.py: убран `from app.services.scheduler import scheduler_loop` + lifespan-блок запуска (`if settings.scheduler_enable: asyncio.create_task(scheduler_loop())`); прод-backend всегда шёл с SCHEDULER_ENABLE=false, так что это был мёртвый код. scheduler_main.py: убрана ship-dark развилка #2192 (USE_KIT_SCHEDULER=false → legacy scheduler_loop fallback) — _run_kit_scheduler() теперь безусловный путь. Поле settings.use_kit_scheduler оставлено в конфиге (Settings extra="ignore" защищает от startup-краха на leftover env var), но на ветвление не влияет. app.services.scrape_pipeline: 0 runtime-импортёров в app/+scripts/+packages/ после этого PR (только тесты, которые Part E удалит вместе с самим файлом) — подтверждено grep. scrape_pipeline.py не тронут (Part E). Тесты: удалены test_house_imv_backfill_scheduler.py (100% legacy-триггер, backfill_house_imv сервис покрыт в test_house_imv_backfill_browser_flag.py / test_backfill_wave2.py) и test_kit_registry_completeness.py (parity-инвариант против удалённого dispatch, дублирует test_scraper_kit_scheduler_parity.py). Точечно вырезаны "Scheduler wiring" секции (trigger_fn_exists/dispatch_branch_ wired/runs_in_executor) из ~10 файлов, тестирующих сами task-функции — сами task-тесты (SQL-shape, миграции, fake-db поведение) оставлены нетронутыми. test_scheduler.py: 825 → ~90 строк (остались только compute_next_run_at-тесты). test_scraper_kit_scheduler_parity.py: убрана golden-parity секция против удалённого scheduler_loop (SOURCE_TO_OLD_TRIGGER/_drive_old_one_tick/ test_routing_parity_per_source), остальное (claim/reap_zombies/dispatch/ registry-shape тесты kit-модуля) сохранено — источник этих инвариантов не app.services.scheduler, а сам scraper_kit.orchestration.scheduler. test_scheduler_main.py: 2 теста, патчившие app.services.scheduler.scheduler_loop, переведены на монкипатч sm._run_kit_scheduler (единственный путь после этого PR). test_sweep_imv_phase.py:171-371 (6 прямых импортов run_avito_city_sweep из scrape_pipeline) намеренно НЕ тронуты — Part E. Verify: полный pytest 3179 passed / 6 skipped / 1 known-unrelated fail (test_search_cache_hit, #2208, не связан с этим PR); ruff 0.7.4 чист на всех изменённых файлах; `python -c "import app.main; import app.scheduler_main"` OK. |
||
| 566b2f9617 |
fix(tradein/scraper): drain detached run-children on SIGTERM, not just coordinator (#1182 Phase 2)
Pre-push review: _await_scheduler ждал только scheduler_loop COORDINATOR, но вся scrape-работа крутится в detached asyncio.create_task детях (каждый trigger_* делал `task = create_task(_run())` без join). На SIGTERM coordinator выходил из while True и завершался → asyncio.run() teardown хард-кансельил ещё бегущих детей mid-await = ровно #1182 failure mode. Кооперативный checkpoint спасал ребёнка лишь когда его residual-время случайно перекрывало drain — вероятностно, не гарантированно. Fix — coordinator теперь дренажит детей перед выходом: - _spawn_tracked(coro): централизованный detached-spawn, кладёт задачу в module-level registry _inflight_tasks (strong-ref = RUF006 keep-alive) + done-callback ретривит exception и убирает из set'а. Заменил 26 одинаковых `task = create_task(_run()); task.add_done_callback(...)` сайтов. - _drain_inflight(): один asyncio.wait по живым детям с бюджетом _CHILD_DRAIN_TIMEOUT_S=80s (< scheduler_main 100s < docker grace 120s). Кооперативные дети (avito_detail_backfill, rosreestr-executor) дочекивают карточку/батч + mark_done и резолвятся; некооперативные упираются в timeout и падают на внешний hard-cancel. - scheduler_loop по выходу из tick-loop (только по SIGTERM) зовёт await _drain_inflight(). NB: raw asyncio.all_tasks()-minus-self здесь НЕЛЬЗЯ — в нашей топологии он захватывает _run parent-task (блокирован на wait_for(coordinator)) и shutdown_waiter → циклическое ожидание coordinator↔_run, всегда упирающееся в timeout. Точный registry это исключает. Tests: tests/test_scheduler.py — detached cooperative child drained-not-cancelled, non-cooperative child timeout→left for hard-cancel, no-op без детей, scheduler_loop→drain wiring. Обновил 3 source-inspection теста под новый _spawn_tracked паттерн. |
|||
| e1442d3b1d |
feat(tradein-scheduler): nightly newbuilding-enrichment schedule (#973)
All checks were successful
CI / changes (pull_request) Successful in 6s
CI / changes (push) Successful in 7s
CI / backend-tests (pull_request) Has been skipped
CI / frontend-tests (pull_request) Has been skipped
CI / backend-tests (push) Has been skipped
CI / frontend-tests (push) Has been skipped
Deploy Trade-In / test (push) Successful in 33s
Deploy Trade-In / build-backend (push) Successful in 1m3s
Deploy Trade-In / changes (push) Successful in 5s
Deploy Trade-In / build-frontend (push) Has been skipped
Deploy Trade-In / build-browser (push) Has been skipped
Deploy Trade-In / deploy (push) Successful in 40s
Регистрирует отдельный in-app scheduler-источник для cian_newbuilding enrichment-backfill (houses_price_dynamics / house_reliability_checks / house_reviews), который раньше бежал только инлайн в Cian full-sweep. - run_newbuilding_enrich() wrapper (heartbeat → backfill → mark_done/failed), делегирует backfill_newbuilding_enrichment (#972, уже в main). - trigger_newbuilding_enrich_run() + dispatch-ветка в scheduler_loop (зеркалит trigger_sber_index_pull_run): _claim_run guard, async task, zombie-reap наследуется; idempotency наследуется (skip уже обогащённых, per-house SAVEPOINT, bounded per-fire limit дренирует backlog 306 домов). - seed-миграция 103: source='newbuilding_enrich', окно 00:00-01:00 UTC (03:00-04:00 МСК, до 01:00+ UTC sweep-блока), limit=25, next_run_at на завтра. Засеяно enabled=FALSE (dormant): на prod весь external-HTTP scraping намеренно на паузе. Расписание+триггер готовы, но не стартуют до намеренного возобновления (UPDATE ... SET enabled=true). Не активирую одиночный источник при общей паузе. Closes #973 |
|||
| b0fe292a63 |
fix(tradein): working cian ЖК-url resolver via cat.php SERP (#972)
All checks were successful
CI / changes (push) Successful in 6s
CI / backend-tests (push) Has been skipped
CI / frontend-tests (push) Has been skipped
CI / changes (pull_request) Successful in 5s
Deploy Trade-In / test (push) Successful in 31s
Deploy Trade-In / build-backend (push) Successful in 1m20s
CI / backend-tests (pull_request) Has been skipped
CI / frontend-tests (pull_request) Has been skipped
Deploy Trade-In / changes (push) Successful in 5s
Deploy Trade-In / build-frontend (push) Has been skipped
Deploy Trade-In / build-browser (push) Has been skipped
Deploy Trade-In / deploy (push) Successful in 46s
The legacy resolve_cian_zhk_url hit /zhk/<id>/ which now 404s, leaving the 318 geo-matched cian houses (house_sources.ext_id set, cian_zhk_url NULL) unfetchable -> the newbuilding-enrichment backfill matched 0 houses. Add resolve_cian_zhk_url_via_search(nb_id): fetch the cat.php newbuilding SERP and extract the canonical zhk-<slug>.cian.ru url, ANCHORED on the <h1 data-name="Title"> header anchor (NOT naive first-match — the SERP carries promo/recommendation zhk-* links before the title that would otherwise resolve the wrong ЖК and silently corrupt enrichment). Validated against the real 2.81MB prod SERP + an adversarial poisoned-recommendation test. Wire into newbuilding_enrich_backfill: broaden selection to "has ext_id OR cian_zhk_url", resolve+persist the url under a SAVEPOINT before enriching, rate-limited + resumable + idempotent. Keep the old resolver (deprecated). Bounded prod proof (5 houses, direct): 5/5 urls resolved+persisted, 3/5 fully enriched (+18 price_dynamics, +3 reliability); the 2 misses were direct-mode anti-bot on the 2nd fetch. Full 318-run gated on the cian mobile proxy (mproxy.site) being restored. code-reviewer APPROVE (SQL/idempotency) + resolver hardened against wrong-ЖК. 19 tests green. Refs #972. |
|||
| bcb1285341 |
feat(tradein): backfill task for cian newbuilding enrichment (#972)
All checks were successful
Deploy Trade-In / changes (push) Successful in 7s
Deploy Trade-In / build-frontend (push) Has been skipped
Deploy Trade-In / build-browser (push) Has been skipped
Deploy Trade-In / deploy (push) Successful in 46s
Deploy Trade-In / test (push) Successful in 37s
Deploy Trade-In / build-backend (push) Successful in 45s
950-E1: app/tasks/newbuilding_enrich_backfill.py selects houses linked to
ext_source='cian_newbuilding' with a fetchable cian_zhk_url and runs the existing
fetch_newbuilding + save_newbuilding_enrichment over each, populating
houses_price_dynamics + house_reliability_checks (+ house_reviews when present).
- Idempotent + resumable: skips already-enriched houses; force= re-run dup-safe
(price_dynamics UPSERT via dim_key, reliability dedup guard with id-DESC
tiebreaker, reviews ON CONFLICT (source,ext_review_id)).
- Anti-bot aware: get_scraper_delay('cian') + jitter, SAVEPOINT-per-house
(one captcha/failure logs + continues, never aborts the batch).
- limit= for bounded proof runs; sizing counters (total/fetchable/pending).
Fix cian_newbuilding._extract_transport_rate: cian transportAccessibilityRate
drifted to a nested dict and broke save_newbuilding_enrichment's CAST bind on
LIVE data (cannot adapt type 'dict') — coerce dict/float/bad -> int|None
(bool excluded before int). Fixes the live enrichment crash, not just backfill.
house_reviews stays ~0 for now: cian reviews are in a separate XHR, not the ЖК
initialState — extending fetch_newbuilding for them is a documented follow-up.
Tested: 8 new unit tests; isolated-DB proof-run landed 6 price_dynamics + 1
reliability row, idempotency proven across re-runs. Refs #972.
|