Repository navigation
Conversation
The v2 store's message search ran `body LIKE '%q%'` over every message (38 ms on the live store, 585 ms at 30x), and conversation-name search scanned conversations, participants and identities on every /api/search keystroke. Migration 0012 substring_search adds FTS5 external-content trigram indexes over message bodies, conversation titles, participant names and identity names/addresses (case_sensitive 0, detail=none, columnsize=0), kept equal to their tables by AFTER INSERT/DELETE/UPDATE triggers and built from existing rows. Searches put the same LIKE on the indexed column: SQLite reads candidates from the trigram postings and re-applies the LIKE to each, so every search returns exactly the rows and order it did before. SearchMessages without a conversation or sender filter first reads the newest 2,000 messages (in the date range, if any) and stops at limit matches; that answers common terms and short queries. Otherwise a query with a literal run of three characters reads messages_fts (CROSS JOIN keeps it the outer loop; the planner's own order took 7-34 s with an account filter), and shorter ones keep the old scan. Conversation- and sender-scoped searches keep their bounded LIKE. Conversation-name search unions one trigram arm per column, using a superset pattern on the index and the escaped LIKE on the row. Store connections turn recursive_triggers on so a REPLACE fires the delete triggers. Store.VerifySearchIndexes runs FTS5's content integrity check; the legacy-to-v2 migration requires it (Validation.SearchIndexesValid). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 9, 2026
The race build exists to find data races, and modernc SQLite runs several times slower under it: the three properties took about three minutes of the package's race run. They still reach every search path over 8 random stores; the plain test run keeps all 30. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #215 (base
perf/sqlite-plan-audit); retarget tomainafter #215 merges.The v2 store's message search ran
body LIKE '%q%'over every message (#215's audit: 38 ms on the live store, 585 ms at 30x), and conversation-name search scanned conversations, participants and identities on every/api/searchkeystroke (4.7 / 142 ms). This answers both from SQLite FTS5 trigram indexes without changing which rows any search returns, or their order.What changed
substring_searchadds four FTS5 external-content trigram tables (messages_ftsovermessages.body,conversations_ftsovertitle,conversation_participants_ftsoverdisplay_name,identities_ftsoverdisplay_nameandcanonical_value) withtokenize='trigram case_sensitive 0',detail=none,columnsize=0. Twelve AFTER INSERT/DELETE/UPDATE triggers keep each index equal to its table on every write, whatever statement or FK cascade makes it. The migration then builds the indexes from existing rows.LIKEon the indexed column. SQLite hands that LIKE to FTS5, which returns rowids holding every trigram of the pattern's literal runs, and SQLite then re-applies the LIKE to each candidate: FTS5 does not mark LIKE constraintsomit. Verified on the repo's modernc SQLite 3.51.2: wildcards, ASCII-only case folding (É≠é), no accent folding, NULs, invalid UTF-8 and 1–2 character queries all match the plain LIKE.SearchMessagesrouting (internal/storage/sqlite/messages.go):limit.limitmatches. If those 2,000 holdlimitmatches, or are the whole range, that is the answer: common terms and short queries finish in under 1 ms.messages_fts CROSS JOIN messages. CROSS JOIN pins FTS as the outer loop: with an account filter, the planner's own join order probed FTS once per message and took 7–34 s.SearchConversationsByNameunions one trigram arm per searched column. Each arm puts a superset pattern on the index (literal%becomes_, since SQLite only pushes an ESCAPE-free LIKE to a virtual table) and the exact escaped LIKE on the row.storeDSNturnsrecursive_triggerson. Without it, aREPLACEdeletes rows without firing delete triggers; the random-writes property caught that drift. No code uses REPLACE today; this makes it safe if any does.Store.VerifySearchIndexesruns FTS5'sintegrity-checkwith rank 1.PRAGMA integrity_checkdoes not compare an external-content index with its table. The legacy→v2migration.Transformrecords the result asValidation.SearchIndexesValid, which the validation gate (nowvalidationPassed, unit-tested gate by gate) requires.Measured
Copies of the live v2 store (fresh
cpofstore.sqlite3+-wal, integrity- and FK-checked; the live store was never opened) and #215's 30x (2.17M messages) and 10x-deep (722k; one 127k-message thread) copies. Each cell is the median (max) over the term classes: absent, document frequency ~1/10/100/1000, the most common word and a common 3-letter word, 1- and 2-char, and every prefix of a word as typed. All cases are end-to-endSearchMessagesvs the old LIKE statement on the same copy, with the machine load average at 20–90 throughout. 0 result mismatches in 413 cases.¹ A rare 2-character query: no trigram, few recent hits, so it falls back to the old scan, unchanged.
² No production caller filters by account (
v2readnever setsAccountID). It shows the one case that can now be slower: a term found in many messages overall but in fewer thanlimitof the newest 2,000. The index then visits all its matches, measured at 1.34–1.6× the old scan for a word in 28% of messages. The same can happen unfiltered when a once-common term disappears from recent messages.Conversation-name search (
/api/search's second half), live 1x: 5–6 ms → 0.5–0.7 ms; the most common name token 6.3 → 3.8 ms. At 30x (41k conversations): 134–171 → 0.7–3 ms; the most common token 190 → 134 ms. 1–2 character queries are unchanged.Migration 0012 (
execution_msin the ledger, including the migration runner's whole-databaseforeign_key_check): 1.7–1.8 s on the live copy at load ~30–40, up to 4.7 s at load ~90; 13–34 s at 10x deep; 47 s at 30x. It runs inside the daemon's startup migration (#215: read clients never migrate), before the v2 stack starts. Index size:messages_fts7.1 MB against 6.8 MB of bodies (55 MBmessagestable); 66 MB at 10x, 190 MB at 30x. The three name indexes take 0.3 MB live.Invariants (all executed as tests)
SearchMessagesreturns exactly the plain LIKE statement's rows (whole rows,reflect.DeepEqual) in its order, through every path.TestSearchMessagesMatchesLikeProperty(testing/quick, 30 random stores × 80 queries) draws text built to separate trigram folding from LIKE: non-ASCII case, the Kelvin sign, combining marks, ZWJ emoji,%_\, NUL and invalid UTF-8. It randomizes the window size and fails unless every path is reached.SearchConversationsByNameequals the escaped LIKE statement for every query (TestSearchConversationsByNameMatchesLikeProperty).TestSearchIndexesEqualTablesAfterRandomWritesPropertychecks after every write: inserts, projection upserts that change or keep the body, edits, state changes, moves, REPLACE, deletes, and conversation deletes that cascade. Separately, the package's store test helpers verify all four indexes at teardown, so every existing repository test (projection, import, history insert, mutations, outbox repoint, repair, rebind) is also an index-maintenance test.TestSearchIndexesFollowRepositoryWriteswalks the named write paths explicitly.TestTrigramCandidatePatternMatchesEveryLiteralMatch).likePatternUsesTrigramsagrees with FTS5's own decision.TestLikePatternUsesTrigramsMatchesFTS5checks 3,000 random patterns, observed through a row present in the content table but absent from the index.TestSearchIndexesSurviveVacuumAndVacuumInto).messages_time_idxwith no sort; window counts are covering-index walks; conversation and sender searches stay on their indexes; each name arm reads its trigram index; and short queries keep Keep v2 reads off whole-table scans found by a SQL plan audit #215's pinned scan.Mutation check (
mutants.pyin the evidence folder): 22 planted bugs, 16 that change results, let an index drift or weaken a check, and 6 that only change a plan or the speed. 21 are killed. The survivor drops the window'smessage_idtie-break. It is equivalent under the pinned plan, becauseINDEXED BY messages_time_idxalready returns ties inmessage_idorder; the tie-break stays so the window remains exact if that plan ever changes.Migration numbering
#215 takes 0011; this is 0012. #196, #201, #166 and #210 also add a migration numbered 0011. Whichever lands first keeps its number and the rest renumber. Here the migration tests find
substring_searchby name; the lines that move are the version counts ininternal/storage/sqlite/*_test.go,internal/migration/transform_test.go, andinternal/migration/validate.go's version/checksum gate.Not changed
Relevance ranking (none, as before); conversation- and sender-scoped message search (same statements). A follow-up chip covers using the index inside very large threads with a candidate-count estimate. Evidence (probes, harness, every run's raw output and summaries, mutation script):
~/reviews/openmessage-fts5-search-2026-10-09/.Review
Independent review requested from a GPT-6.1 Sol lane (
subfleet run --task review --tier hard) against commit 55401b8; findings and responses will follow in a comment.🤖 Generated with Claude Code