Skip to content

Answer v2 substring search from FTS5 trigram indexes - #223

Draft
MaxGhenis wants to merge 2 commits into
perf/sqlite-plan-auditfrom
perf/v2-fts5-search
Draft

MaxGhenis wants to merge 2 commits into
perf/sqlite-plan-auditfrom
perf/v2-fts5-search

Conversation

@MaxGhenis

Copy link
Copy Markdown
Owner

Stacked on #215 (base perf/sqlite-plan-audit); retarget to main after #215 merges.

The v2 store's message search ran body LIKE '%q%' over every message (#215's audit: 38 ms on the live store, 585 ms at 30x), and conversation-name search scanned conversations, participants and identities on every /api/search keystroke (4.7 / 142 ms). This answers both from SQLite FTS5 trigram indexes without changing which rows any search returns, or their order.

What changed

  • Migration 0012 substring_search adds four FTS5 external-content trigram tables (messages_fts over messages.body, conversations_fts over title, conversation_participants_fts over display_name, identities_fts over display_name and canonical_value) with tokenize='trigram case_sensitive 0', detail=none, columnsize=0. Twelve AFTER INSERT/DELETE/UPDATE triggers keep each index equal to its table on every write, whatever statement or FK cascade makes it. The migration then builds the indexes from existing rows.
  • Same semantics. Each search puts its existing LIKE on the indexed column. SQLite hands that LIKE to FTS5, which returns rowids holding every trigram of the pattern's literal runs, and SQLite then re-applies the LIKE to each candidate: FTS5 does not mark LIKE constraints omit. Verified on the repo's modernc SQLite 3.51.2: wildcards, ASCII-only case folding (É ≠ é), no accent folding, NULs, invalid UTF-8 and 1–2 character queries all match the plain LIKE.
  • SearchMessages routing (internal/storage/sqlite/messages.go):
    • A search filtered by conversation or sender keeps today's LIKE statement, which is bounded by its index; the conversation one walks newest-first and stops at limit.
    • Any other nonempty query first reads the newest 2,000 messages (within the date range, if any) and stops at limit matches. If those 2,000 hold limit matches, or are the whole range, that is the answer: common terms and short queries finish in under 1 ms.
    • Otherwise, a query with a literal run of ≥3 characters reads messages_fts CROSS JOIN messages. CROSS JOIN pins FTS as the outer loop: with an account filter, the planner's own join order probed FTS once per message and took 7–34 s.
    • Shorter queries keep the old scan.
  • SearchConversationsByName unions one trigram arm per searched column. Each arm puts a superset pattern on the index (literal % becomes _, since SQLite only pushes an ESCAPE-free LIKE to a virtual table) and the exact escaped LIKE on the row.
  • storeDSN turns recursive_triggers on. Without it, a REPLACE deletes rows without firing delete triggers; the random-writes property caught that drift. No code uses REPLACE today; this makes it safe if any does.
  • Store.VerifySearchIndexes runs FTS5's integrity-check with rank 1. PRAGMA integrity_check does not compare an external-content index with its table. The legacy→v2 migration.Transform records the result as Validation.SearchIndexesValid, which the validation gate (now validationPassed, unit-tested gate by gate) requires.
  • Docs: CLAUDE.md, and a runbook section "What a v2 search matches" covering semantics, how to check and rebuild the indexes on a copy, and the upgrade-time build.

Measured

Copies of the live v2 store (fresh cp of store.sqlite3 + -wal, integrity- and FK-checked; the live store was never opened) and #215's 30x (2.17M messages) and 10x-deep (722k; one 127k-message thread) copies. Each cell is the median (max) over the term classes: absent, document frequency ~1/10/100/1000, the most common word and a common 3-letter word, 1- and 2-char, and every prefix of a word as typed. All cases are end-to-end SearchMessages vs the old LIKE statement on the same copy, with the machine load average at 20–90 throughout. 0 result mismatches in 413 cases.

filter (limit) live 1x: LIKE → new, ms 10x deep 30x
none (30, the UI) 29 (32) → 2.1 (7.0) 185 (304) → **2.2 (184)**¹ 528 (781) → 3.4 (17)
last 30 days (50) 23 (35) → 1.0 (1.3) 164 (363) → 0.9 (1.4) 509 (678) → 1.0 (1.2)
last 365 days (50) 27 (44) → 3.0 (6.6) 163 (205) → **1.8 (159)**¹ 505 (624) → **3.2 (481)**¹
account (30)² 25 (29) → 3.0 (40) 191 (329) → 2.4 (303) 564 (706) → 4.8 (852)
busiest conversation 6.6 → 6.6 (same SQL) 118 → 118 (same SQL) 6.4 → 6.2 (same SQL)
busiest sender 3.4 → 3.3 (same SQL) 53 → 53 (same SQL) 19 → 19 (same SQL)

¹ A rare 2-character query: no trigram, few recent hits, so it falls back to the old scan, unchanged.
² No production caller filters by account (v2read never sets AccountID). It shows the one case that can now be slower: a term found in many messages overall but in fewer than limit of the newest 2,000. The index then visits all its matches, measured at 1.34–1.6× the old scan for a word in 28% of messages. The same can happen unfiltered when a once-common term disappears from recent messages.

Conversation-name search (/api/search's second half), live 1x: 5–6 ms → 0.5–0.7 ms; the most common name token 6.3 → 3.8 ms. At 30x (41k conversations): 134–171 → 0.7–3 ms; the most common token 190 → 134 ms. 1–2 character queries are unchanged.

Migration 0012 (execution_ms in the ledger, including the migration runner's whole-database foreign_key_check): 1.7–1.8 s on the live copy at load ~30–40, up to 4.7 s at load ~90; 13–34 s at 10x deep; 47 s at 30x. It runs inside the daemon's startup migration (#215: read clients never migrate), before the v2 stack starts. Index size: messages_fts 7.1 MB against 6.8 MB of bodies (55 MB messages table); 66 MB at 10x, 190 MB at 30x. The three name indexes take 0.3 MB live.

Invariants (all executed as tests)

  1. For every query (any bytes) and filter, SearchMessages returns exactly the plain LIKE statement's rows (whole rows, reflect.DeepEqual) in its order, through every path. TestSearchMessagesMatchesLikeProperty (testing/quick, 30 random stores × 80 queries) draws text built to separate trigram folding from LIKE: non-ASCII case, the Kelvin sign, combining marks, ZWJ emoji, % _ \, NUL and invalid UTF-8. It randomizes the window size and fails unless every path is reached.
  2. When the recent window reports itself complete, it is that same answer (same property).
  3. SearchConversationsByName equals the escaped LIKE statement for every query (TestSearchConversationsByNameMatchesLikeProperty).
  4. After any sequence of writes, each index equals its table (FTS5 integrity-check, rank 1). TestSearchIndexesEqualTablesAfterRandomWritesProperty checks after every write: inserts, projection upserts that change or keep the body, edits, state changes, moves, REPLACE, deletes, and conversation deletes that cascade. Separately, the package's store test helpers verify all four indexes at teardown, so every existing repository test (projection, import, history insert, mutations, outbox repoint, repair, rebind) is also an index-maintenance test. TestSearchIndexesFollowRepositoryWrites walks the named write paths explicitly.
  5. The candidate pattern matches every value the escaped LIKE matches (TestTrigramCandidatePatternMatchesEveryLiteralMatch).
  6. likePatternUsesTrigrams agrees with FTS5's own decision. TestLikePatternUsesTrigramsMatchesFTS5 checks 3,000 random patterns, observed through a row present in the content table but absent from the index.
  7. VACUUM and VACUUM INTO (the backup path) keep rowids, so the indexes stay valid (TestSearchIndexesSurviveVacuumAndVacuumInto).
  8. The migration changes no rows, keeps earlier ledger rows, has a pinned checksum, and reopens idempotently, on both blank and populated pre-0012 stores.
  9. Plans are pinned: FTS outer loop under every filter; the window walks messages_time_idx with no sort; window counts are covering-index walks; conversation and sender searches stay on their indexes; each name arm reads its trigram index; and short queries keep Keep v2 reads off whole-table scans found by a SQL plan audit #215's pinned scan.

Mutation check (mutants.py in the evidence folder): 22 planted bugs, 16 that change results, let an index drift or weaken a check, and 6 that only change a plan or the speed. 21 are killed. The survivor drops the window's message_id tie-break. It is equivalent under the pinned plan, because INDEXED BY messages_time_idx already returns ties in message_id order; the tie-break stays so the window remains exact if that plan ever changes.

Migration numbering

#215 takes 0011; this is 0012. #196, #201, #166 and #210 also add a migration numbered 0011. Whichever lands first keeps its number and the rest renumber. Here the migration tests find substring_search by name; the lines that move are the version counts in internal/storage/sqlite/*_test.go, internal/migration/transform_test.go, and internal/migration/validate.go's version/checksum gate.

Not changed

Relevance ranking (none, as before); conversation- and sender-scoped message search (same statements). A follow-up chip covers using the index inside very large threads with a candidate-count estimate. Evidence (probes, harness, every run's raw output and summaries, mutation script): ~/reviews/openmessage-fts5-search-2026-10-09/.

Review

Independent review requested from a GPT-6.1 Sol lane (subfleet run --task review --tier hard) against commit 55401b8; findings and responses will follow in a comment.

🤖 Generated with Claude Code

The v2 store's message search ran `body LIKE '%q%'` over every message (38 ms
on the live store, 585 ms at 30x), and conversation-name search scanned
conversations, participants and identities on every /api/search keystroke.

Migration 0012 substring_search adds FTS5 external-content trigram indexes
over message bodies, conversation titles, participant names and identity
names/addresses (case_sensitive 0, detail=none, columnsize=0), kept equal to
their tables by AFTER INSERT/DELETE/UPDATE triggers and built from existing
rows. Searches put the same LIKE on the indexed column: SQLite reads
candidates from the trigram postings and re-applies the LIKE to each, so every
search returns exactly the rows and order it did before.

SearchMessages without a conversation or sender filter first reads the newest
2,000 messages (in the date range, if any) and stops at limit matches; that
answers common terms and short queries. Otherwise a query with a literal run
of three characters reads messages_fts (CROSS JOIN keeps it the outer loop;
the planner's own order took 7-34 s with an account filter), and shorter ones
keep the old scan. Conversation- and sender-scoped searches keep their bounded
LIKE. Conversation-name search unions one trigram arm per column, using a
superset pattern on the index and the escaped LIKE on the row.

Store connections turn recursive_triggers on so a REPLACE fires the delete
triggers. Store.VerifySearchIndexes runs FTS5's content integrity check; the
legacy-to-v2 migration requires it (Validation.SearchIndexesValid).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The race build exists to find data races, and modernc SQLite runs several
times slower under it: the three properties took about three minutes of the
package's race run. They still reach every search path over 8 random stores;
the plain test run keeps all 30.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant