gendesign/tradein-mvp/backend/app/services/matching/conflict_resolution.py
lekss361 4470b1a764 fix(tradein): cross-source matching — address PR #470 BLOCK verdict (NOT NULL / geom type / abbreviation order / determinism)
CRITICAL #1: houses INSERT now populates source, ext_house_id, url (NOT NULL cols) with
  ON CONFLICT (source, ext_house_id) upsert; source_url param added with synthetic fallback.
CRITICAL #2: geom removed from INSERT col list — houses_set_geom_trg trigger auto-populates.
CRITICAL #3: Tier 3 ST_DWithin/ST_Distance now cast geom::geography on both sides to avoid
  mixed geometry/geography function lookup failure.
CRITICAL #4: normalize.py _PUNCT now keeps hyphens ([^\w\s\-]); abbreviation expansion for
  пр-кт/б-р/пр-д runs before hyphen collapse, eliminating leftover кт token.
HIGH #5: Tier 0 cadastral queries add ORDER BY id ASC for deterministic results (houses + listings).
HIGH #6: match_or_create_house docstring documents race-condition known limitation + Stage 8.x plan.
Medium #7: no-op ('литер', 'литер') entry removed from _ABBREV.
Medium #9: update_canonical_fields raises NotImplementedError instead of silent pass.
Medium #11: price_rub bigint truncation comment added to _upsert_listing_source.
Medium #14: upsert_listing_source exported from __init__.py __all__.
Tests: 29/29 pass; prkT test asserts no leftover кт; update_canonical_fields test updated.
2026-05-23 17:01:24 +03:00

62 lines
2.3 KiB
Python

"""Field priority rules for canonical record merging.
When multiple sources provide conflicting values for the same field,
these dicts define which source to prefer (highest-priority first).
Algorithm reference: decisions/Cross_Source_Matching_Strategy.md sec 4.5
"""
from __future__ import annotations
from sqlalchemy.orm import Session
# Higher-priority source listed first.
# For each column, prefer the first source in the list that has a non-NULL value.
HOUSE_FIELD_PRIORITY: dict[str, list[str]] = {
'year_built': ['cian', 'avito_houses', 'rosreestr'],
'address': ['avito', 'cian', 'rosreestr'],
'cadastral_number': ['rosreestr', 'cian', 'avito'],
'building_class': ['cian', 'avito_houses'],
'floors_count': ['cian', 'avito_houses'],
'series_name': ['cian'],
'entrances': ['cian'],
'flat_count': ['cian'],
}
LISTING_FIELD_PRIORITY: dict[str, list[str]] = {
# Price: always from latest-seen source wins among equal-priority sources.
'price_rub': ['avito', 'cian', 'yandex'],
'area_m2': ['rosreestr', 'cian', 'avito'],
'living_area_m2': ['cian', 'avito'],
'kitchen_area_m2': ['cian', 'avito'],
'floor': ['rosreestr', 'cian', 'avito'],
'rooms': ['rosreestr', 'cian', 'avito'],
'ceiling_height': ['cian', 'avito'],
'description': ['avito', 'cian'],
'phones': ['avito', 'cian'],
'photos': ['avito', 'cian'],
'cadastral_number': ['rosreestr', 'cian', 'avito'],
}
def update_canonical_fields(
db: Session,
listing_id: int,
ext_source: str,
lot_data: object,
) -> None:
"""Merge source fields into the canonical listings row.
Stage 8 v1: populate-NULL strategy only — write source value only when
the canonical column is currently NULL. Full priority arbitration with
conflict logging is planned for Stage 8.x when ≥2 sources are live.
Args:
db: Active SQLAlchemy session.
listing_id: Canonical listings.id.
ext_source: Source identifier (e.g. 'cian', 'avito').
lot_data: ScrapedLot or similar object with attribute access.
Unknown attributes are silently skipped.
"""
raise NotImplementedError("Stage 8.x — full priority arbitration not yet implemented")