Skip to content

Use the trigram index for conversation and sender searches when counts say it pays - #230

Open
MaxGhenis wants to merge 2 commits into
perf/v2-fts5-searchfrom
perf/v2-fts5-thread-search
Open

MaxGhenis wants to merge 2 commits into
perf/v2-fts5-searchfrom
perf/v2-fts5-thread-search

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #223 (base perf/v2-fts5-search, itself on #215); retarget to main after both merge. No migration.

#223 kept conversation- and sender-scoped message search on its LIKE over the scope's index range. That never regresses, but for a term rare in a long thread it reads the whole thread: 70–135 ms in a 127,000-message thread on a copy of the live store with ten times its history. This lets those searches use the trigram index when counts say it pays, without changing which rows any search returns, or their order.

What changed

SearchMessages sends a nonempty query narrowed by conversation or sender, with a literal run of three characters, to searchScope (internal/storage/sqlite/scoped_search.go). Other scoped queries (1–2 characters, empty) keep the old statement.

  • Conversation. It seeks the 500th-newest row of the conversation's range (a covering index walk) and runs the old LIKE at or above that row. If those rows hold limit matches, that is the answer. Otherwise it either tries the index (below) or runs the same LIKE below the cut; window matches followed by the rest's first limit − k are exactly the old walk. A range shorter than 500 rows runs the old statement. The statements run in one read transaction, so a write between them cannot move a message across the cut (returned twice, or not at all).
  • When the index is tried. From counts only, never timings, so a search always takes the same path. The rows the LIKE would still read are estimated from the window's hit rate, ceil((limit − k) × 501 / (k + 1)), and capped by a count of the rows below the cut. The index is tried when that is at least max(3000, messages/50) (messages = max(rowid)).
  • The index search is one statement: it materializes at most half that many candidates of an FTS5 MATCH that ANDs the query's trigrams, counts them, and, only when there are fewer than the bound, joins them to messages by rowid and applies the scope's filters and the LIKE. With too many candidates it returns the count alone and the LIKE continues. The postings are merged once either way.
  • trigramMatchQuery builds that MATCH: the query ends at its first NUL and is cut at %/_; each piece of ≥3 characters contributes every trigram as its own quoted term (a detail=none index refuses multi-trigram phrases). Unlike Answer v2 substring search from FTS5 trigram indexes #223's LIKE on the FTS column, a candidate from another conversation costs one rowid lookup, not a body fetch through the index: 1.6–2.7× cheaper per candidate in the probes.
  • Sender. In the same kind of read transaction, the sender's identities are resolved once and named in each statement (the filter's subquery scans identities, which has no index on canonical_value, every time a statement runs it). The sender's rows are counted, and the index is tried when there are at least twice the minimum; otherwise the LIKE runs with the identities named. For a sender with one identity that LIKE now reads in time order and stops at limit (the subquery form always read and sorted all of the sender's rows).
  • Plan detail. With a date range, the upper part leaves out since and the lower part until (the cut row lies inside the range, so they are implied). With both bounds SQLite seeks by the dates and only filters by the cut, reading index entries past it; plans are pinned.

Measured

Old statement vs SearchMessages, interleaved run by run (7 runs at 10x, 11 at 1x; medians), on copies of the live v2 store (fresh cp of store.sqlite3 + -wal, integrity-checked; the live store was never opened) and its 10x-deep copy (722k messages; scale-deep.sh). Limits 30 and 50. Term classes per scope: absent, absent here but rare/common elsewhere, rare/uncommon/frequent/common here, the scope's top words, four phrases, 1–2 characters, and every prefix of a word as typed. Load average 25–47 during these runs. 0 mismatches in 1,780 cases.

1x scope (rows) old → new, median ms total ms
busiest conversation (12,709) 6.04 → 1.79 536 → 216
same, last 365 days (all 12,709 in range) 6.03 → 1.90 580 → 232
conversation (7,694) 3.21 → 1.84 344 → 205
conversation (2,888) 1.07 → 1.52 103 → 146
conversation (467) 0.20 → 0.26 21 → 26
sender (4,450) 2.65 → 2.88 316 → 315
sender (3,713) 2.02 → 2.16 218 → 219
sender (3,390; two identities) 1.74 → 2.02 189 → 220
all 2,307 → 1,579 (−32%)
10x-deep scope (rows) old → new, median ms total ms
busiest conversation (127,090) 78.84 → 4.23 5,950 → 891
same, last 365 days (12,709 in range) 5.64 → 7.30 433 → 550
conversation (53,800) 52.22 → 4.27 3,245 → 509
conversation (20,620) 17.10 → 3.14 1,220 → 445
conversation (7,800) 4.39 → 5.37 288 → 355
conversation (2,980) 1.18 → 1.61 80 → 107
conversation (500) 0.31 → 0.44 24 → 33
sender (44,500) 46.62 → 6.89 4,578 → 1,028
sender (37,130) 40.02 → 6.90 3,898 → 943
sender (33,900; two identities) 39.70 → 11.11 4,176 → 2,172
all 23,892 → 7,033 (−71%)

By term class in the 127,090-message thread (10x), median ms: absent 99.8 → 2.5; typed prefixes 100.4 → 2.7; absent here but rare elsewhere 110.3 → 3.9, common elsewhere 108.1 → 7.9; uncommon here 103.4 → 5.0; phrases 36.7 → 14.6; frequent here 17.0 → 13.5; common here 0.81 → 1.12; top words 0.41 → 0.58; 1–2 characters 0.25 → 0.25 (same statement). Full tables: results/final/summary-{1x,deep10}.md.

What got slower, and why.

  • Ranges a little longer than the window but too short for the index. The store cannot know a thread's length without walking its index, so these pay for the boundary seek and the count that proves them too short: +0.45 ms on a 1.1 ms search (2,900 rows), +1.0 ms on 4.4 ms (7,800 rows at 10x), +1.7 ms on 5.6 ms (a 12,709-row date range at 10x, where the minimum is 14,473). In the plan benchmark the bare machinery, with the index never tried, costs 0.2–0.5 ms per conversation search at 1x.
  • Terms common in a conversation: +0.1–0.3 ms (the boundary seek and a second statement) on searches under 1 ms.
  • Senders at today's size: none holds twice the minimum, so each search pays for the count: medians +0.15–0.3 ms on 1.7–2.7 ms.
  • A sender with several identities, common terms, at 10x: 38–40 → 43–47 ms (the count, then an index search that gives up).

How the constants were chosen

results/plan-bench.md: 15 terms × 11 scopes, old statement and every plan interleaved.

  • Minimum rows 3,000, or 1/50 of the store. The index's merge of the query's postings is a cost no count sees: up to 1.3 ms at 1x and 10.8 ms at 10x for long words and phrases of common trigrams, whatever the number of candidates. 3,000 rows is what the LIKE reads in 1.3 ms at 1x; 1/50 is 14,473 rows at 10x. At 10x, 1/50 beat 1/30 and 1/100 (totals 2,318 / 2,401 / 2,671 ms against 3,940 old, at 3 rows per candidate).
  • Half as many candidates as rows (2 rows per candidate). Against 3 rows per candidate the totals were 1,990 vs 2,318 ms at 10x and 323 vs 316 ms at 1x.
  • A 500-message conversation window, not the 2,000 used across conversations: every conversation search seeks its boundary, and a smaller window leaves more of a mid-sized thread for the index. At 1x the 11-scope total was 371 ms with 2,000 and 324 with 500 (old 325); at 10x 500 cost 3% more (2,146 vs 2,076), and 250 never reached the minimum there.
  • Twice the minimum for senders. A sender has no window to answer its common terms first, so each pays for an index search that gives up.
  • One statement for count and search. Counting first and searching second merged the postings twice; fused, "let me know" in the 127k thread went 16.2 → 9.2 ms and "redevelopment" 9.9 → 5.3 ms.

Earlier variants (two statements; window 2,000; senders without resolved identities) and their runs are kept under results/final-*.

Invariants (all executed as tests)

  1. For every query (any bytes) and filter, SearchMessages returns exactly the plain LIKE statement's rows (whole rows, reflect.DeepEqual) in its order, through every path. TestSearchMessagesMatchesLikeProperty now randomizes the plan and adds 60 scoped searches per store; it fails unless all six paths are reached (each new path 65–700 times per run, 14 or more under -race).
  2. trigramMatchQuery's MATCH selects exactly the rows FTS5 reads as candidates for LIKE '%q%', for every pattern FTS5 narrows, and returns an expression exactly when likePatternUsesTrigrams does (TestTrigramMatchQuerySelectsFTS5LikeCandidates: 3,000 patterns with NULs, invalid UTF-8, quotes, wildcards and expressions of hundreds of trigrams; FTS5's candidates are made observable by rewriting the content table so the re-applied LIKE keeps every one).
  3. Cutting a conversation's range at any of its rows splits the LIKE walk: the first limit matches at or above the cut, then the first limit − k below, are the LIKE's rows, for every filter; the count below the cut is exact (TestConversationSplitMatchesLikeProperty).
  4. The index search is complete exactly when the index holds fewer candidates than its bound; then it returns the LIKE's rows (none included, as an empty slice), and otherwise the statement returns the count alone, never a message from a truncated list (the property, and TestScopeIndexSearchGivesUpAtItsBound).
  5. A sender search with identities named by ID returns the rows of the subquery form, in the index search and in the LIKE (the property); the sender count is the number of rows the sender's LIKE reads (TestSenderRangeCountCountsTheSendersMessages).
  6. A write between a search's statements cannot change the answer. TestConversationSearchReadsOneSnapshot moves a message across the cut mid-search and gets the pre-write LIKE rows, and shows the duplicate the same statements return outside a transaction; TestSenderSearchReadsOneSnapshot moves a message to a new identity of the same address after the identities are read, and shows it missing outside one.
  7. expectedLikeRows never grows as the window finds more matches; scopeIndexBound reads nothing when the estimate is below the minimum, and counts no further than it.
  8. Paths are pinned per filter (TestSearchMessagesChoosesItsPath), and plans (TestScopedSearchPlansReadTheirIndexes): the cut and the counts read only their covering index; each part seeks messages_conversation_time_idx from the cut with no sort, with and without dates; the index search reads the trigram postings first and joins by rowid; a sender's statements never read identities.

Mutation check (mutants.py, run on f896cf1; the two transaction mutants re-run on 1b1f50e): 38 planted bugs: 21 that change results (one of them only under a concurrent write) and 17 that change only a plan, a path, the speed or a test seam. 36 are killed. The two survivors leave every test passing and the pinned plans unchanged on SQLite 3.51.2: dropping the index search's outer ORDER BY (the rows already come back in the sub-query's order; it stays as the guarantee), and dropping INDEXED BY from the two parts (a store without statistics picks that index anyway; the hint holds if statistics ever exist). The first run left two more alive, a guard that let an over-full candidate list be searched and the missing transaction; both now have tests that kill them.

Not changed

Searches across conversations (#223's paths and statements), conversation-name search, short and empty queries, schema. MessageRepository.betweenSearchStatements is a test hook, nil in production.

Not done here, measured or noted: #223's cross-conversation trigram statement could use the same MATCH form (1.6–2.7× cheaper per candidate); a per-conversation message count kept by triggers would remove the cost of proving a thread too short, at the price of a migration and a write per message.

Evidence (harnesses, sweeps, the policy replay, plan benchmarks, raw results, mutation script and log): ~/reviews/openmessage-fts5-thread-search-2026-10-10/.

Review

Independent review requested from a GPT-6.1 Sol lane (subfleet run --task review --tier hard) against 1b1f50e; findings and responses will follow in a comment.

🤖 Generated with Claude Code

MaxGhenis and others added 2 commits October 10, 2026 09:42
…s say it pays

A v2 message search narrowed by conversation or sender ran its LIKE over the
scope's index range. That never regresses, but for a term rare in a long
thread it reads the whole thread: 70-135 ms in a 127,000-message thread on a
copy of the live store with ten times its history.

SearchMessages now sends such a search, when its query has a literal run of
three characters, to searchScope (scoped_search.go):

- A conversation's newest 500 messages are read first, with the same LIKE
  cut at a boundary key; a term common in the thread is answered there. The
  rest of the answer is either the same LIKE below the boundary or the
  trigram index's candidates in the conversation. Both statements run in
  one read transaction.
- The index is tried when the LIKE would still read at least
  max(3000, messages/50) rows: the window's hit rate applied to the matches
  still missing, capped by a count of the rows left. One statement then
  materializes at most half that many candidates of an FTS5 MATCH that ANDs
  the query's trigrams (trigramMatchQuery), counts them, and searches them
  only when there are fewer; otherwise the LIKE continues.
- A sender's identities are resolved once and named in each statement. Its
  LIKE is replaced by the index search when the sender holds at least twice
  that minimum.

The choice uses counts, never timings. Every path returns exactly the plain
LIKE statement's rows in its order: TestSearchMessagesMatchesLikeProperty
covers all six paths, TestTrigramMatchQuerySelectsFTS5LikeCandidates checks
the MATCH against FTS5's own candidates for the LIKE, and
TestConversationSplitMatchesLikeProperty the cut at every row.

Measured on copies of the live store (old statement vs SearchMessages,
interleaved, 1,780 cases, 0 mismatches): at ten times the history the median
search fell from 79 to 4.2 ms in the 127,000-message thread and from 40-47
to 7-11 ms for the busiest senders; at today's size from 6.0 to 1.8 ms in
the busiest thread. Ranges a little longer than the window but too short
for the index pay 0.45 ms (2,900 messages) to 1.0-1.7 ms (7,800-12,700
messages at ten times the history) for the boundary seek and the count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sender path resolves the address's identities in one statement and
searches their messages in another. Run it in the read transaction the
conversation path already uses, so a write between the two cannot move a
message to another identity of the same address and out of the answer.
TestSenderSearchReadsOneSnapshot shows both: the LIKE's rows at one
snapshot inside the transaction, the moved message missing outside it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant