83% of tracker issues (7460 total) were pure noise drowning real signal:
- basic_auth 401 (3738 issues, 2019 distinct titles) — ops/glitchtip-auth-
forwarder sent EVERY 401 from bots scanning gendsgn.ru (GET /wp-admin/
install.php etc.) as an individual GlitchTip event, remote_ip baked into
message/tags inflated cardinality. Not an application error — expected
bot-scan traffic against a basic_auth-protected site.
- RetryError (2462 issues) — geocoder.py's three tenacity @retry-wrapped
Nominatim helpers (lookup/suggest/reverse) raised tenacity.RetryError on
exhaustion without reraise=True; RetryError.__str__() embeds a Future
repr() with a memory address that differs every call, so GlitchTip
grouped each exhausted retry as a distinct issue instead of one.
Fix at the source, not post-hoc issue cleanup:
- forwarder.py: before_send drops events tagged event_type in
{basic_auth_failed, basic_auth_storm}; forwarder's own capture_exception
(real script bugs) carries no such tag and passes through untouched.
- geocoder.py: reraise=True on all three @retry decorators — propagates
the real underlying exception (stable type + stacktrace) instead of the
unstable RetryError wrapper.
- sentry_scrub.stabilize_retry_error_fingerprint: belt-and-suspenders
before_send hook, composed into both app/main.py and scheduler_main.py
(geocoder runs in both processes — FastAPI request path and the
overnight geocode_missing_listings batch). Collapses any RetryError that
still slips through into one persistent issue per cause-exception type
name only — never IP/address/listing-id.
Content-ful categories (OperationalError, city-sweep, harvest_quarter,
cian/avito/yandex sweep failures, scrape_freshness_check — ~700 issues)
are untouched: filters key off event_type tag / exception type name only.