Корень «−1.00 везде» (эпик #1953): compute_demand_supply_forecast брал
district-wide unit_velocity (847.5/мес, ВСЕ классы/комнаты) как спрос и
весь district-сток (~63k доступных) как предложение для КАЖДОЙ ячейки
what_to_build → один и тот же ratio во всех ячейках → все deficit_index
прижаты к −1.0. Плюс objective_lots — append-per-snapshot (~2.9× инфляция
строк), что симметрично раздувало обе базы → даже сегментация без дедупа
осталась бы вырожденной.
Фикс (blast radius — ТОЛЬКО forecast/deficit calc; platform-wide dedup = #1964):
- market_metrics.compute_market_metrics: +obj_class/+room_bucket (+cache key).
_STOCK_SQL и _SALES_WINDOW_SQL дедуплят до ПОСЛЕДНЕГО снапшота на физлот
(DISTINCT ON project_name,corpus_name,section,floor,lot_number ORDER BY …
snapshot_date DESC,id DESC), затем агрегируют. Class-фильтр (LOWER=LOWER,
class lowercase) + room-bucket (Source-B room_area-вокабуляр, зеркало
sales_series.room_area_bucket_of → what_to_build фильтрует без перевода).
ROLLUP/GROUPING сохранён; confidence считается на дедуплицированных counts.
- demand_supply_forecast: base_pace и open-сток теперь ПОСЕГМЕНТНЫЕ
(market_metrics(obj_class,room_bucket)). При заданном сегменте L2/L3
(hidden/future) ИСКЛЮЧЕНЫ из баланса — они класс/формат-агностичны, иначе
двоились бы по всем ячейкам. +_market_room_bucket VOCAB-мост (валидирующий
pass-through Source-B меток; неизвестное → None = без фильтра, не тихий 0-rows).
- what_to_build/_DEFAULT_CLASSES и recommendation Economy-маппинг: «эконом»→
«стандарт» (в objective_lots эконома НЕТ, стандарт=483k → раньше ячейка
матчила 0 строк и молча выпадала).
- report_assembler honesty-guard: если ВСЯ сетка прижата к ±1.0
(degenerate-fallback) — не эмитим «строить»/«избегать», показываем
«недостаточно гранулярных данных для посегментного вывода».
- data/sql/173_objective_lots_physflat_idx.sql: partial index под DISTINCT ON
(Index Only Scan + Unique, без Sort на 1.75M строк; idempotent, BEGIN/COMMIT).
Prod-verify (parcel 66:41:0205010:287, Железнодорожный, h=24): ячейки
ДИФФЕРЕНЦИРУЮТ (12 measured, 7 distinct) вместо all −1.0; MOI комфорт/студия
38.5 vs стандарт/студия 244.3 (точное совпадение с ожидаемым).
Тесты: регрессия «ячейки различаются (не all −1.0)» + vocab-translation +
honesty-guard + посегментное предложение. ruff clean; no :name::type.
- Add window_months > 0 guard to vel_by_room dict comprehension (mirrors _monthly_rate)
- Correct outer comment in recommendation.py: honest-zero for known-zero buckets,
fallback only when velocity_by_room=None or bucket absent from _FORECAST_TO_METRIC_BUCKETS
- Add comment in as_dict() noting velocity_by_room is intentionally not serialized
(internal pipeline attr consumed directly by recommendation.py)
Add `velocity_by_room: dict[str, float] | None` to `MarketMetrics` — per-bucket
unit velocity (ед./мес) derived from the existing `sold_by_room` ROLLUP data that
`_query_sales_window` already returns. No new SQL required.
Thread per-bucket velocity through `_demand_only_overlay` via the new
`_FORECAST_TO_METRIC_BUCKETS` constant that maps each forecast bucket to its
market_metrics room-bucket keys. "80+ м²" sums "4" + "5+" keys. Fallback to
aggregate `unit_velocity` when `velocity_by_room` is None (thin-data path).
Previously `base_pace` was identical for all 5 room-buckets, so §9.4 norm and §9.2
base_pace cancelled out in pace/max_pace and ranking was driven purely by §9.5
macro_coef (segment steepness proxy). Now §9.2 reflects real per-bucket observed
demand from objective_lots.contract_date data.
Callers of `compute_market_metrics` that don't use `velocity_by_room` are unaffected
(the new field is additive to the frozen dataclass). All existing callers verified —
none construct `MarketMetrics` directly except the one production site.
_price_sensitivity передавал сырое admin-имя ('Кировский') в _elasticity_coef,
который фильтрует objective_corpus_room_month.district по МИКРО-вокабуляру
(Втузгородок, ЖБИ, …) → регрессия получала 0 точек → всегда FALLBACK_ELASTICITY.
§9.2 district-level эластичность молча НЕ считалась в /analyze-пути (только
'Академический' совпадал в обоих вокабулярах случайно).
Fix: вызываем resolve_objective_districts() в _price_sensitivity и передаём
список микро через новый kwarg districts=[…] в _elasticity_coef. Резолвер
None ('не определён' / нет чистых алиасов) → пустой список → EKB-wide
регрессия. _elasticity_coef расширен с back-compat: districts=None →
legacy путь по district_name (другой caller в analytics_queries —
отдельный bug class, вне scope).
5 новых юнит-тестов TestPriceSensitivityDistrictResolution: admin→micros в
SQL bind, None→EKB-wide, regression preserved post-resolve, graceful.
76/76 market_metrics + 156/156 elasticity/sensitivity тестов зелёные.
ruff + psycopg v3 grep clean.
Closes#1211
_SALES_WINDOW_SQL делал GROUP BY ROLLUP (rooms_int), rooms_int nullable
(ETL пишет NULL для «неопределённого типа», sales_series.py:399 явно
обрабатывает None). Проданный лот с rooms_int IS NULL даёт ДВЕ строки
rooms_int IS NULL (NULL-группа + grand-total итог), неразличимые в
Python (оба if r["rooms_int"] is None).
MixedAggregate-план PG16 эмитит grand-total ПЕРВЫМ (среди hash-строк),
NULL-группа после → loop затирает units_total частичным счётом (живой
тест на PG16: 2000 → 200). Эффект: unit_velocity / absorption_rate
занижены, months_of_supply завышен → base_pace в demand_supply_forecast
неверный (recommendation.py:586) → reports/scoring врёт.
Patch:
- SQL: добавить GROUPING(rooms_int) AS is_total (=1 для grand-total).
- Python: ветвить по is_total, NULL-комнатную группу класть в
by_room['unknown'] (отдельный бакет), аккумулировать через +=
вместо assign (защита от будущих NULL-вариантов).
- Тесты: моки получили "is_total" поле (1 для grand-total, 0 иначе).
71/71 market_metrics тестов зелёные. ruff clean.
Closes#1214
Cold §22 forecast measured ~215-233s on prod: §9.x layers re-execute the same
horizon/segment-invariant DB loads with identical args hundreds of times per
report (profiled: get_competitors x69, market_metrics x124, get_monthly_macro
x290). Add a per-report ContextVar cache (forecast_cache(), opened once in the
orchestrator) + @cached(key_builder) on the expensive §9.x loaders so each
unique load runs ONCE and reuses the same frozen, read-only instance.
Output is byte-identical (memoized producers are frozen dataclasses / read-only
Pydantic, callers never mutate; cache is per-report, discarded on exit; no-op
outside the report build). No concurrency, no signature changes.
- forecast_request_cache.py: ContextVar cache + cached() decorator (no-op
outside context, reentrant, _MISS sentinel for cached None)
- @cached on competitors/future_supply/market_metrics/macro_series/
sales_series/macro_coefficient/demand_normalization/regression loaders
- orchestrator: wrap build_site_finder_report in forecast_cache()
- 58 tests: key discrimination (call-counting regression guard), no-op-outside,
per-report isolation, reentrancy, frozen-producer canary, amplification proof
(real get_monthly_macro xN->1)
code-reviewer APPROVE (keys correct, mutation-safe, output identical). 1265
forecast/cache tests green. No new deps. Refs #1129.
/analyze passes the official ЕКБ admin district (ekb_districts polygon, e.g.
'Кировский'), but objective_lots/corpus_room_month store informal micro-districts
('Втузгородок','ЖБИ') -> admin name matched 0 rows -> silent empty forecast.
Add resolve_objective_districts() (site_finder/district_resolver.py) mapping an
admin name to its clean micros via ekb_district_alias (note IS NULL), with
None -> EKB-wide fallback and raw-micro pass-through. Wire into the objective_lots
district filters of market_metrics (§9.2 stock+sales), supply_layers L1 (§9.3),
and sales_series Sources A+B (crm shares the micro vocab, prod-verified),
switching the scalar filter to psycopg3-safe = ANY(CAST(:districts AS text[])).
supply_layers L2/L3 keep the admin name (domrf_kn_objects.district_name is admin vocab).
Prod: Кировский/Ленинский/Орджоникидзевский obj_count 0 -> 32/64/31.
Tests mutation-verified non-vacuous. 192 module tests pass; ruff clean. Refs #969#949.
REOPENED. _SALES_WINDOW_SQL derived "sales in window" from objective_lots_history
snapshots, but history is only ~17 days deep — every currently-sold lot had a
sold-snapshot in the window, so window-sales collapsed into the entire cumulative
sold stock (Автовокзал 6mo: 33,245 vs real ~2,308). Inflated absorption_rate
(~235%/mo with confidence=high), months_of_supply, unit_velocity, liquidity,
demand_concentration → contaminated forecast #950/#952.
Count window sales directly from objective_lots by contract_date in the window
(the real sale date — present on 100% of sold lots: 41,091/41,091). Return
contract of _query_sales_window unchanged (units/area/by-room ROLLUP); downstream
formulas untouched. Removed the now-dead objective_lots_history JOIN/CTE.
Regression test: lots sold outside window (contract_date out of range) not counted
(41,091 cumulative vs 2,308 window → absorption 2.35→0.04). 288 tests green.
Verification = prod compute_market_metrics(Автовокзал) post-deploy. Refs #949