- Phase 1: faithful, unsafe-friendly port. Same page-as-byte-buffer layout, same hand-computed offsets, same function boundaries as the C original. Goal is passing the existing test harness, not writing idiomatic Rust.
- Phase 2 (separate branch/commits): idiomatic refactor, once Phase 1 passes and I understand why the C original made each choice.
- Any "this could be more Rust-y" idea that occurs to me during Phase 1 goes in the Deferred Improvements list below, not into the code.
- NodeType enum instead of byte-flag branching
- Cursor<'a> borrowing from Table, vs. index-based page handles
- Page pool for the pager (single-threaded version of the retired/ready-to-hand-out design from the queue project)
- Concurrency of any kind — single-threaded throughout
- Crash recovery / durability beyond what the tutorial itself covers
- Anything past where the tutorial's C code actually stops Concurrency control (coarse RwLock, later crabbing if warranted) belongs to the follow-up DBMS project, not here.
- Part 1-2: REPL + SQL compiler skeleton
- Part 3-5: single-table storage, persistence
- Part 6-7: cursor abstraction
- Part 8-10: B-tree leaf nodes
- Part 11-13: B-tree internal nodes, splitting
- Part 14: duplicate keys, scanning
My next project will be self-directed: a real DBMS, built in phases, each one depending on the last.
- Newtype wrappers for
PageId,FrameId,TableId,Lsn(not bareusize/u32).LsnasNonZeroU64, with0reserved as "no LSN yet".- Highest-leverage item on this list: most of cstack_db's real bugs were two
different
usize-shaped concepts colliding (a key passed where an index was expected, a page number passed where a cell index was expected). A newtype per concept turns that into a compile error.
- Custom table schema with
String,int, andfloat.Stringstored as a length-prefixedVarInt. - Slotted page layout, fullness evaluated by free space instead of a fixed cell count. Byte-packed metadata (dirty bit, node type, etc).
- Sibling pointers on leaf nodes — bidirectional (prev and next) this time, to support reverse scans for roughly free.
- Reads pages off disk, deserializes into a proper
PageNodestruct (enum PageNode { Leaf(LeafData), Internal(InternalData) }), serializes back to bytes only at the swap boundary (eviction or shutdown).- Byte manipulation lives in exactly one place: the serialize/deserialize
pair. Everything above that boundary works with typed data, not raw
&mut [u8]— this is what makes the offset/node-type-confusion bug class from cstack_db structurally impossible here.
- Byte manipulation lives in exactly one place: the serialize/deserialize
pair. Everything above that boundary works with typed data, not raw
RwLock-guarded pages (RwLock<Option<Box<PageFrame>>>per frame).RwLockchosen deliberately for real cross-thread concurrency, not by default —RefCellgets the same&self-based win far cheaper if this stays single-threaded.
- Pin counting via RAII guard (increment on fetch, decrement on
Drop) — a bookkeeping mechanism, not the eviction policy itself. - Clock eviction policy: reference bit per frame, set on access, cleared as
the clock hand sweeps past without evicting. Only
pin_count == 0frames are eligible. - Dirty bit set on write-guard drop; only serialize-and-write on eviction/flush if dirty.
- One access path only (
fetch_page, always faults in on a miss) — no separate forcing/non-forcing variants. Half of cstack_db's session-restart bugs were exactly a call site using the read-only accessor that returnedNoneon a cache miss instead of loading from disk.
BTreestruct holding a reference to the BPM (similar toTableandPagerhere).- Delete, including underflow handling (merge with / borrow from a sibling).
Conspicuously absent from cstack_db — this is the mirror image of
split-on-insert, and at least as fiddly as the key-promotion bookkeeping in
internal_node_split_and_insertwas, just in reverse. - Free-list / space map for page allocation, replacing cstack_db's
get_unused_page_num(which only ever grows). Needed once delete exists so freed pages are reusable. - A minimal catalog/system table to durably store user-defined schema definitions themselves, since schemas are no longer fixed at compile time.
RwLock-guarded pages (built in Phase 2) used for real.- Latch crabbing (lock coupling) for B+Tree traversal: hold the parent's latch, acquire the child's, release the parent once the child is confirmed safe (won't split/merge). Standard technique for concurrent tree traversal without one thread pinning the root latch for an entire operation.
Hardest item on the list — sequence deliberately rather than attempting full ARIES in one pass:
- Redo-only logging + crash recovery first (log before touching the page, replay on restart).
- Undo + CLRs (compensation log records) for transaction abort, once redo-only recovery is solid.
- Invariant to never violate: the log record for a change is durable before the corresponding page write hits disk. Every page carries the LSN of its last modifying record.
- WAL records are generated at the BTree call site (logical redo/undo info —
"inserted key K at slot S in page P" — lives there, not in the buffer
pool).
WriteGuard::dropdoes not push to the WAL; it only marks the frame dirty. WAL-before-data is enforced at the other end: in the BPM's flush/eviction path, before writing a dirty page, force the log manager to durably flush up through that page's stamped LSN. - Build a crash-test harness as a real deliverable (kill the process mid-transaction, restart, assert recovery produces a consistent state) — the only way to know ARIES is actually correct rather than "looks right".
- Basic
TransactionManager: begin/commit/abort, transaction IDs, hooked into WAL. A single global lock serializing all transactions is a reasonable v1 concurrency model — get commit/abort/WAL integration correct before attempting real isolation levels (2PL/MVCC). - Basic
QueryEngine: thin dispatch layer once everything below it works, similar in spirit tovm.rshere.