gendesign/tradein-mvp/backend/app/services/house_dedup_merge.py
bot-backend 078dadc440
All checks were successful
CI / changes (pull_request) Successful in 7s
CI / backend-tests (pull_request) Has been skipped
CI / frontend-tests (pull_request) Has been skipped
CI / openapi-codegen-check (pull_request) Has been skipped
fix(tradein): cadastre ЖКХ-fallback + COALESCE merge-key + median-floors guard
Серия sale-share data-quality (по находкам юзера):
- house_dedup_merge: cluster_key → tradein_canon_addr(COALESCE(short_address,
  full_address, address)) — ловит дубли с район-noise в address (Вайнера 66).
- zhkh_flats_loader: cadastre-fallback (searchByAddress по cadastreNumber,
  guid-независимо) — резолвит дома где gar_house_guid ≠ ЖКХ houseGuid
  (Переходный 6: guid→0, cadastre→165). CLI --cadastre.
- мигр.153: плаузибилити-гард по медиане этажности листингов — denom >=
  GREATEST(total_floors, median(listings.total_floors), 8); отсекает gar-недосчёт
  башен без ЖКХ (gar=9 vs листинги 19-26 эт → NULL).

ruff clean, pytest 36 green.
2026-06-29 01:32:48 +03:00

600 lines
24 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""Recurring, idempotent, collision-safe house-deduplication merge (#1772).
WHY (recurrence, not a one-shot):
Migration 108_merge_duplicate_houses.sql did a ONE-TIME merge of duplicate `houses`
rows, but the duplicates RECUR: the matching pipeline keeps creating them
(Tier-3 geo-jitter splits one physical building across rows + per-source ext_id INSERTs).
A single migration cannot fix a steady-state inflow — we need a RECURRING merge that
re-runs on a schedule and is a no-op when there is nothing to merge.
WHAT this is:
A direct, set-based reimplementation of 108's proven
cluster → pick canonical → re-point children (UNIQUE-collision-safe) → delete losers
pipeline, run inside ONE transaction so a crash leaves the table untouched.
Cluster key: CANONICAL address via tradein_canon_addr() over the CLEAN address
COALESCE(short_address, full_address, address) (cadastral_number is 100% NULL on prod —
confirmed in migration 040 — so address is the real building key). The clean source matters:
`address` can carry район-noise the canon does not strip (e.g. «улица Вайнера, 66 · р-н Центр»
→ canon «вайнера66рнцентр»), while `short_address` holds the clean «улица Вайнера, 66»
(→ «вайнера66») — preferring the clean field lets such a row cluster with its twin. The canon
collapses spelling/район variants of the SAME building (ул→улица, strips город/район/«россия»/
«м-н»/«р-н» noise, removes spaces/punct, PRESERVES корпус) so e.g. «ул. Вайнера,66» and
«улица Вайнера, 66» share one cluster_key. Rows whose canon is NULL/blank are never clustered
(cluster_key NULL → ignored). Only canons shared by >1 house_id form a cluster.
GEO GUARD (anti over-merge): because the canon strips город/район, two different buildings
with the same street+number in different region-66 towns (e.g. «Ленина 5») would share a
canon. To avoid merging them, within a canon-cluster a house is a LOSER only if it is within
250 m of the keeper (ST_DistanceSphere on the WGS84 geom). Same-canon houses >250 m away — or
with NULL geom on either side — are left as separate rows (conservative; never over-merges).
KEEPER RULE (one canonical row per cluster), in priority order:
1. geom NOT NULL first (a geocoded row is the better survivor)
2. most linked listings (the row the corpus already points at)
3. most-populated fields (the richest metadata row)
4. min(id) (deterministic tie-break)
CHILD-COLLISION HANDLING — the migration-133 lesson:
Migration 133 broke on prod because a child UNIQUE constraint was missed. EVERY FK child
of houses(id) is handled below. The FK set was re-audited against the LIVE schema
(grep "REFERENCES houses" data/sql/*.sql) — it is the SAME 11 children 108 handled; no
new FK child was added after 108. For each UNIQUE-constrained child we DELETE the loser
rows that would collide with the keeper (deduping by the TARGET key so dup-vs-dup
collisions inside clusters >2 are also caught) BEFORE re-pointing the survivors.
IMPORTANT divergence from 108: houses_price_dynamics's UNIQUE was REPLACED by
migration 029 — it is now the 6-column
(house_id, source, room_count, prices_type, period, month_date)
NOT 108's stale 3-column (house_id, month_date, source). We dedup on the LIVE 6-column
key; deduping on the old key would still violate the live constraint on re-point.
Per-child handling (loser=L, keeper=K):
listings.house_id_fk no UNIQUE → plain re-point
house_placement_history.house_id no UNIQUE → plain re-point
house_reviews.house_id no blocking UNIQUE→ plain re-point
house_reliability_checks.house_id no blocking UNIQUE→ plain re-point
external_valuations.house_id no blocking UNIQUE→ plain re-point
house_sources UNIQUE(ext_source,ext_id) → delete colliding L, re-point rest
house_address_aliases UNIQUE(normalized_address)→ delete colliding L, re-point rest
houses_price_dynamics UNIQUE(6-col, migr 029) → dedup by target 6-col key, re-point
house_imv_evaluations UNIQUE(house_id) → dedup by target house_id, re-point
house_suggestions UNIQUE(house_id,ext_item_id)→ dedup by target key, re-point
address_mismatch_audit UNIQUE(house_id,audit_batch)→ dedup by target key, re-point
BACKFILL (reduces recurrence):
After deleting losers we backfill house_sources + house_address_aliases onto the keeper
(same as 108) so the matching pipeline's Tier-1/Tier-2 finds the keeper next scrape and
does not immediately re-split it.
IDEMPOTENCY:
Every UPDATE/DELETE keys off a temp mapping of (loser→keeper). On a clean table the
mapping is empty → every statement touches 0 rows → no-op. Re-running is safe.
DRY-RUN:
dry_run=True computes the same mapping and counts, then ROLLS BACK — no writes. Used by
the orchestrator to preview a merge before enabling the (DESTRUCTIVE) schedule.
psycopg v3: all SQL uses CAST(:x AS type), never the colon-colon bound-param cast form.
"""
from __future__ import annotations
import logging
import time
from dataclasses import dataclass, field
from sqlalchemy import text
from sqlalchemy.orm import Session
logger = logging.getLogger(__name__)
# Significant columns counted to score per-row completeness when picking the keeper.
# Mirrors 108's completeness expression (all confirmed present in the live houses schema).
_COMPLETENESS_EXPR = """
(h.address IS NOT NULL)::int +
(h.lat IS NOT NULL)::int +
(h.lon IS NOT NULL)::int +
(h.year_built IS NOT NULL)::int +
(h.house_type IS NOT NULL)::int +
(h.total_floors IS NOT NULL)::int +
(h.cadastral_number IS NOT NULL)::int +
(h.developer_name IS NOT NULL)::int +
(h.rating_score IS NOT NULL)::int +
(h.avito_id_hash IS NOT NULL)::int
"""
# Keeper ORDER BY, shared by the ROW_NUMBER() rank and the first_value() keeper pick so they
# agree row-for-row. Priority: geom present → most linked listings → most-populated → min id.
_KEEPER_ORDER = f"""
(h.geom IS NOT NULL) DESC,
listing_cnt DESC,
({_COMPLETENESS_EXPR}) DESC,
h.id ASC
"""
# Build the loser→keeper mapping as a TEMP table for this transaction. Only addresses shared
# by >1 house_id are clustered; the keeper is rn=1 per cluster, losers are rn>1.
_BUILD_MAPPING_SQL = text(
f"""
CREATE TEMP TABLE _1772_dup_mapping ON COMMIT DROP AS
WITH clustered AS (
SELECT
id,
CASE WHEN NULLIF(
tradein_canon_addr(COALESCE(short_address, full_address, address)), ''
) IS NOT NULL
AND tradein_canon_addr(COALESCE(short_address, full_address, address)) ~ '[0-9]'
THEN 'addr:'
|| tradein_canon_addr(COALESCE(short_address, full_address, address))
END AS cluster_key
FROM houses
),
clusters_with_count AS (
SELECT cluster_key, count(*) AS n
FROM clustered
WHERE cluster_key IS NOT NULL
GROUP BY cluster_key
HAVING count(*) > 1
),
dup_houses AS (
SELECT h.id, cl.cluster_key
FROM houses h
JOIN clustered cl ON cl.id = h.id
JOIN clusters_with_count cw ON cw.cluster_key = cl.cluster_key
),
listing_counts AS (
SELECT house_id_fk AS house_id, count(*) AS listing_cnt
FROM listings
WHERE house_id_fk IS NOT NULL
GROUP BY house_id_fk
),
ranked AS (
SELECT
dh.id,
dh.cluster_key,
lower(trim(COALESCE(h.short_address, h.full_address, h.address))) AS norm_address,
h.geom AS loser_geom,
ROW_NUMBER() OVER (
PARTITION BY dh.cluster_key
ORDER BY {_KEEPER_ORDER}
) AS rn,
first_value(dh.id) OVER (
PARTITION BY dh.cluster_key
ORDER BY {_KEEPER_ORDER}
) AS keeper_id,
first_value(h.geom) OVER (
PARTITION BY dh.cluster_key
ORDER BY {_KEEPER_ORDER}
) AS keeper_geom
FROM dup_houses dh
JOIN houses h ON h.id = dh.id
LEFT JOIN listing_counts lc ON lc.house_id = dh.id
)
-- Geo guard against cross-town over-merge: tradein_canon_addr strips город/район, so two
-- different buildings sharing a street+number canon (e.g. «Ленина 5» in different towns)
-- collapse to one cluster_key. A loser is only merged when it is geographically next to the
-- keeper (<=250 m — покрывает geocode-разброс одного здания, прод: дубли Мраморская 34к4 в
-- 222 m; города региона-66 в км+ друг от друга → 250 m безопасно от cross-town). Houses in
-- the same canon-cluster but >250 m away, or with NULL geom on either side, are left as
-- separate rows (conservative — never over-merges). Cluster_key также требует цифру в canon
-- (номер дома) — исключает вырожденные canon вида «екатеринбург»/«сооружение».
SELECT id AS loser_id, keeper_id, norm_address
FROM ranked
WHERE rn > 1
AND id <> keeper_id
AND keeper_geom IS NOT NULL
AND loser_geom IS NOT NULL
AND ST_DistanceSphere(loser_geom, keeper_geom) <= 250
"""
)
# Each step keys off _1772_dup_mapping → empty mapping ⇒ 0 rows touched ⇒ idempotent no-op.
_STEPS: list[tuple[str, str]] = [
# ── Plain re-point (no UNIQUE on the FK column) ───────────────────────────
(
"listings",
"""
UPDATE listings l
SET house_id_fk = m.keeper_id
FROM _1772_dup_mapping m
WHERE l.house_id_fk = m.loser_id
""",
),
(
"house_placement_history",
"""
UPDATE house_placement_history hph
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE hph.house_id = m.loser_id
""",
),
(
"house_reviews",
"""
UPDATE house_reviews hr
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE hr.house_id = m.loser_id
""",
),
(
"house_reliability_checks",
"""
UPDATE house_reliability_checks hrc
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE hrc.house_id = m.loser_id
""",
),
(
"external_valuations",
"""
UPDATE external_valuations ev
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE ev.house_id = m.loser_id
""",
),
# ── UNIQUE(ext_source, ext_id): delete colliding losers, re-point rest ─────
(
"house_sources(collision-delete)",
"""
DELETE FROM house_sources hs
USING _1772_dup_mapping m
WHERE hs.house_id = m.loser_id
AND EXISTS (
SELECT 1 FROM house_sources hs2
WHERE hs2.house_id = m.keeper_id
AND hs2.ext_source = hs.ext_source
AND hs2.ext_id = hs.ext_id
)
""",
),
(
"house_sources(re-point)",
"""
UPDATE house_sources hs
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE hs.house_id = m.loser_id
""",
),
# ── UNIQUE(normalized_address): delete colliding losers, re-point rest ─────
(
"house_address_aliases(collision-delete)",
"""
DELETE FROM house_address_aliases haa
USING _1772_dup_mapping m
WHERE haa.house_id = m.loser_id
AND EXISTS (
SELECT 1 FROM house_address_aliases haa2
WHERE haa2.house_id = m.keeper_id
AND haa2.normalized_address = haa.normalized_address
)
""",
),
(
"house_address_aliases(re-point)",
"""
UPDATE house_address_aliases haa
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE haa.house_id = m.loser_id
""",
),
# ── UNIQUE(house_id, source, room_count, prices_type, period, month_date) ──
# LIVE 6-col key (migration 029 replaced 108's stale 3-col). Dedup by the TARGET key
# (keeper survives, else min id) so dup-vs-dup collisions in clusters >2 are also caught.
(
"houses_price_dynamics(dedup)",
"""
DELETE FROM houses_price_dynamics t
USING (
SELECT t2.id,
ROW_NUMBER() OVER (
PARTITION BY COALESCE(m.keeper_id, t2.house_id),
t2.source, t2.room_count, t2.prices_type,
t2.period, t2.month_date
ORDER BY (m.keeper_id IS NULL) DESC, t2.id ASC
) AS rn
FROM houses_price_dynamics t2
LEFT JOIN _1772_dup_mapping m ON m.loser_id = t2.house_id
) d
WHERE t.id = d.id AND d.rn > 1
""",
),
(
"houses_price_dynamics(re-point)",
"""
UPDATE houses_price_dynamics hpd
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE hpd.house_id = m.loser_id
""",
),
# ── UNIQUE(house_id): one evaluation per keeper ───────────────────────────
(
"house_imv_evaluations(dedup)",
"""
DELETE FROM house_imv_evaluations t
USING (
SELECT t2.id,
ROW_NUMBER() OVER (
PARTITION BY COALESCE(m.keeper_id, t2.house_id)
ORDER BY (m.keeper_id IS NULL) DESC, t2.id ASC
) AS rn
FROM house_imv_evaluations t2
LEFT JOIN _1772_dup_mapping m ON m.loser_id = t2.house_id
) d
WHERE t.id = d.id AND d.rn > 1
""",
),
(
"house_imv_evaluations(re-point)",
"""
UPDATE house_imv_evaluations hie
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE hie.house_id = m.loser_id
""",
),
# ── UNIQUE(house_id, ext_item_id) ─────────────────────────────────────────
(
"house_suggestions(dedup)",
"""
DELETE FROM house_suggestions t
USING (
SELECT t2.id,
ROW_NUMBER() OVER (
PARTITION BY COALESCE(m.keeper_id, t2.house_id), t2.ext_item_id
ORDER BY (m.keeper_id IS NULL) DESC, t2.id ASC
) AS rn
FROM house_suggestions t2
LEFT JOIN _1772_dup_mapping m ON m.loser_id = t2.house_id
) d
WHERE t.id = d.id AND d.rn > 1
""",
),
(
"house_suggestions(re-point)",
"""
UPDATE house_suggestions hs
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE hs.house_id = m.loser_id
""",
),
# ── UNIQUE(house_id, audit_batch) ─────────────────────────────────────────
(
"address_mismatch_audit(dedup)",
"""
DELETE FROM address_mismatch_audit t
USING (
SELECT t2.id,
ROW_NUMBER() OVER (
PARTITION BY COALESCE(m.keeper_id, t2.house_id), t2.audit_batch
ORDER BY (m.keeper_id IS NULL) DESC, t2.id ASC
) AS rn
FROM address_mismatch_audit t2
LEFT JOIN _1772_dup_mapping m ON m.loser_id = t2.house_id
) d
WHERE t.id = d.id AND d.rn > 1
""",
),
(
"address_mismatch_audit(re-point)",
"""
UPDATE address_mismatch_audit ama
SET house_id = m.keeper_id
FROM _1772_dup_mapping m
WHERE ama.house_id = m.loser_id
""",
),
]
# Delete the loser houses — all FK children are re-pointed or CASCADE by now.
_DELETE_LOSERS_SQL = text(
"""
DELETE FROM houses h
USING _1772_dup_mapping m
WHERE h.id = m.loser_id
"""
)
# Backfill house_sources for keeper avito-rows so Tier-1 (source_exact) finds the keeper next
# scrape (reduces recurrence). Same shape as 108 Step 5.
_BACKFILL_SOURCES_SQL = text(
"""
INSERT INTO house_sources (house_id, ext_source, ext_id, confidence, matched_method, matched_at)
SELECT h.id, h.source, h.ext_house_id, 1.0, 'backfill_dedup_merge', NOW()
FROM houses h
WHERE h.source = 'avito'
AND h.ext_house_id IS NOT NULL
AND h.id IN (SELECT DISTINCT keeper_id FROM _1772_dup_mapping)
ON CONFLICT (ext_source, ext_id) DO NOTHING
"""
)
# Backfill normalized-address aliases onto the keeper so Tier-2 (address) finds it next scrape.
_BACKFILL_ALIASES_SQL = text(
"""
INSERT INTO house_address_aliases (house_id, normalized_address, fingerprint, source)
SELECT h.id, lower(trim(h.address)), NULL, 'backfill_dedup_merge'
FROM houses h
WHERE h.address IS NOT NULL
AND length(trim(h.address)) >= 5
AND h.id IN (SELECT DISTINCT keeper_id FROM _1772_dup_mapping)
ON CONFLICT (normalized_address) DO NOTHING
"""
)
@dataclass
class DedupMergeResult:
"""Counters from one merge pass (also the dry-run preview)."""
clusters_merged: int = 0 # distinct keepers that absorbed at least one loser
losers_deleted: int = 0 # duplicate house rows removed (0 on a clean table)
listings_repointed: int = 0 # listings.house_id_fk moved loser→keeper
children_deleted: int = 0 # collision/dedup deletions across all UNIQUE children
children_repointed: int = 0 # survivor child rows moved loser→keeper
dry_run: bool = False
duration_sec: float = field(default=0.0)
def to_counters(self) -> dict[str, int]:
return {
"clusters_merged": self.clusters_merged,
"losers_deleted": self.losers_deleted,
"listings_repointed": self.listings_repointed,
"children_deleted": self.children_deleted,
"children_repointed": self.children_repointed,
"dry_run": int(self.dry_run),
"duration_sec": int(self.duration_sec),
}
def merge_duplicate_houses(db: Session, *, dry_run: bool = False) -> dict[str, int]:
"""Cluster houses by canonical address and merge duplicates onto one canonical keeper.
Re-implements migration 108's proven collision-safe pipeline as a RECURRING job:
cluster (canon address + 250 m geo guard) → pick keeper →
re-point children (UNIQUE-collision-safe) → delete losers → backfill sources/aliases.
All in ONE transaction. dry_run=True computes counts then ROLLS BACK (no writes).
Idempotent: a clean table yields an empty mapping → every statement is a 0-row no-op.
Returns the counter dict (DedupMergeResult.to_counters()).
"""
start = time.monotonic()
result = DedupMergeResult(dry_run=dry_run)
try:
db.execute(_BUILD_MAPPING_SQL)
mapping = db.execute(
text(
"SELECT loser_id, keeper_id, norm_address FROM _1772_dup_mapping "
"ORDER BY keeper_id, loser_id"
)
).all()
result.losers_deleted = len(mapping)
result.clusters_merged = len({row.keeper_id for row in mapping})
if not mapping:
# Clean table — nothing to merge. Still roll back (we only opened a temp table).
db.rollback()
result.duration_sec = time.monotonic() - start
logger.info(
"merge_duplicate_houses: no duplicate-address clusters found (no-op) dry_run=%s",
dry_run,
)
return result.to_counters()
# Audit log: every loser→keeper move with its address, for traceability.
for row in mapping:
logger.info(
"merge_duplicate_houses: merge loser_id=%d → keeper_id=%d address=%r",
row.loser_id,
row.keeper_id,
row.norm_address,
)
for label, sql in _STEPS:
res = db.execute(text(sql))
rowcount = res.rowcount or 0
if label == "listings":
result.listings_repointed = rowcount
elif label.endswith("(collision-delete)") or label.endswith("(dedup)"):
result.children_deleted += rowcount
elif label.endswith("(re-point)") or label in (
"house_placement_history",
"house_reviews",
"house_reliability_checks",
"external_valuations",
):
result.children_repointed += rowcount
logger.debug("merge_duplicate_houses: step=%s rows=%d", label, rowcount)
db.execute(_DELETE_LOSERS_SQL)
db.execute(_BACKFILL_SOURCES_SQL)
db.execute(_BACKFILL_ALIASES_SQL)
if dry_run:
db.rollback()
logger.info(
"merge_duplicate_houses: DRY-RUN computed clusters=%d losers=%d "
"listings_repointed=%d children_deleted=%d children_repointed=%d — ROLLED BACK",
result.clusters_merged,
result.losers_deleted,
result.listings_repointed,
result.children_deleted,
result.children_repointed,
)
else:
db.commit()
logger.info(
"merge_duplicate_houses: COMMITTED clusters=%d losers=%d "
"listings_repointed=%d children_deleted=%d children_repointed=%d",
result.clusters_merged,
result.losers_deleted,
result.listings_repointed,
result.children_deleted,
result.children_repointed,
)
except Exception:
logger.exception("merge_duplicate_houses: FAILED — rolling back")
try:
db.rollback()
except Exception:
logger.exception("merge_duplicate_houses: rollback also failed")
raise
result.duration_sec = time.monotonic() - start
return result.to_counters()
# ── Run lifecycle wrapper (scheduler entrypoint) ─────────────────────────────
def run_house_dedup_merge(db: Session, *, run_id: int, params: dict) -> dict[str, int]:
"""Run-lifecycle wrapper for the recurring house-dedup merge (sync, DB-only).
Launched by the in-app scheduler (source='house_dedup_merge') via
trigger_house_dedup_merge_run, or manually. Mirrors run_cadastral_geo_match: pure
internal DB op, finalises scrape_runs (mark_done / mark_failed) with counters.
Params (default_params jsonb):
dry_run: compute + return counts then ROLLBACK without writing (default false).
The deploy seeds the schedule DISABLED; the orchestrator can flip a single
manual run to dry_run=true to preview before enabling.
"""
from app.services import scrape_runs as runs_mod
dry_run = bool(params.get("dry_run", False))
counters: dict[str, int] = {
"clusters_merged": 0,
"losers_deleted": 0,
"listings_repointed": 0,
"children_deleted": 0,
"children_repointed": 0,
"dry_run": int(dry_run),
}
try:
runs_mod.update_heartbeat(db, run_id, counters)
counters = merge_duplicate_houses(db, dry_run=dry_run)
runs_mod.mark_done(db, run_id, counters)
logger.info(
"run_house_dedup_merge: run_id=%d DONE clusters=%d losers=%d dry_run=%s",
run_id,
counters.get("clusters_merged", 0),
counters.get("losers_deleted", 0),
dry_run,
)
return counters
except Exception as exc:
logger.exception("run_house_dedup_merge: run_id=%d FAILED", run_id)
try:
db.rollback()
except Exception:
pass
runs_mod.mark_failed(db, run_id, str(exc)[:1000], counters)
raise