Conversation
Fixes #106: the chains filter on /v1/products/ was applied after the SQL LIMIT, so filtered searches returned fewer results than requested even when more matches existed. The filter is now pushed into the search query. Also rewrite the exact-search predicate onto the expression indexed by idx_chain_products_name_trgm (previously an unindexable name ILIKE that seq-scanned ~540k rows, ~160ms per search) and join products after the limit. Exact search is now diacritic-insensitive, matches brand, and returns results in relevance order.
Review follow-up: chain_products accumulates discontinued products (~28% of all rows, >50% for konzum/studenac), and the price-availability check in the response builder ran after the SQL limit, so searches could return far fewer than 'limit' results even when more matches existed. Both search queries now require a chain_prices row on the chain's effective date (latest load <= requested date, same rule as the price lookup), applied before grouping and limiting. The requested date is now also threaded into the search, so historical queries match against that date's availability. Measured on prod: +0-15ms vs the previous shape.
|
Follow-up commit addressing the P1 review finding: price/date filtering also ran after the LIMIT, so the PR as originally submitted still couldn't guarantee full result pages. Why this turned out to be material
Worse, the interaction with the chain filter introduced in this PR would have amplified it: within a single chain nearly all candidates tie at The fixBoth search queries now require a
Cost (EXPLAIN ANALYZE on prod)
*pre-aggregate-first shape; the EXISTS probes (~0.9 ms filtered, ~18 ms for 3.6k probes unfiltered) are roughly offset by joining Behavior change to be aware of
Not addressed here, filed mentally as follow-ups: purging/flagging the ~148k dead rows (pure perf/disk win now, no longer a correctness issue), and the suspiciously high churn ratios for Studenac/Brodokomerc that hint at unstable product codes in those chains' feeds. |
Fixes #106.
The bug
The
chainsfilter onGET /v1/products/was applied in Python after the SQLLIMIT:search_products()fetched the top-N matches globally, thenprepare_product_response()dropped everything not in the requested chains. So?q=mlijeko&limit=100&chains=konzumreturned the Konzum subset of the global top 100 (76 products) instead of up to 100 Konzum products — even though 225 Konzum products match "mlijeko" on prod today.Findings along the way
While analyzing the fix, EXPLAIN on prod showed the default (non-fuzzy) search couldn't use any index at all: its predicate was
cp.name ILIKE '%word%', but the only text index (idx_chain_products_name_trgm) is a GIN trigram index overlower(immutable_unaccent(name || ' ' || COALESCE(brand, '')))— a different expression. Every exact search was a parallel seq scan over ~540kchain_productsrows: ~160 ms and 3 backends per request, CPU-bound (warm cache didn't help), plus a hash join materializing the entireproductstable.Measured plans on prod (q=mlijeko):
(chain_id, code)unique indexproductsThe fix
Three coordinated changes to
search_products()/fuzzy_search_products(), no schema or index changes:chain_idsand filter withcp.chain_id = ANY($n)inside the query. The router resolves chain codes to IDs up front. Uses the existingUNIQUE (chain_id, code)index.lower(immutable_unaccent(name || ' ' || brand)) LIKE '%' || immutable_unaccent($n) || '%'with the word lowercased app-side — same convention the fuzzy path already uses. This makes the trigram GIN index usable, eliminating the seq scan.cp.product_id(equivalent to grouping byean, which is unique per product) and joinproductsonly for the returned rows. Also removes the second round-trip throughget_products_by_ean(), which didn't preserve ordering.Intentional behavior changes
Review notes
GET /v1/products/?q=mlijeko&limit=100&chains=konzumshould return 100 products (was 76), and default searches should drop from ~160 ms to ~10–55 ms of DB time.