Commits and Durability

Every commit of store.graph() (every traversal, every transaction(..), every import; see Transactions) is written to the store's write-ahead log as one record and made durable before it returns. A commit that changes nothing writes nothing.

#![allow(unused)]
fn main() {
use std::time::Duration;
use graphersal::persist::{Durability, Store, StoreOptions};

let store = Store::open("data", StoreOptions::new())?;                    // fsync per commit
let fast = Store::open(
    "other",
    StoreOptions::new().with_durability(Durability::Interval(Duration::from_millis(200))),
)?;
store.graph().write().traversal_mut().add_v("person").to_list()?;          // on disk when it returns
Ok::<(), Box<dyn std::error::Error>>(())
}

Durability

DurabilityA commit returns afterA crash of the processA crash of the machine (power)
EveryCommit (default)the record is written and fsyncedloses nothingloses nothing
Interval(d)the record is written; fsync at most every dloses nothingmay lose the commits of the last interval, never one in the middle
Osthe record is handed to the operating systemloses nothingmay lose what the OS had not written yet

The write-ahead guarantees

  • The journal runs last. The store's journal is a commit hook in its own slot after every other hook (your own hooks, schema checks, a veto), and the commit's sequence number is assigned before it writes. Once the record is written nothing can veto the commit any more: the WAL never holds a commit that did not happen.
  • A failed write or fsync vetoes the commit: it rolls back and the caller gets GraphError::CommitRejected with the reason. Memory and disk never diverge, and nobody gets "ok" for a commit that is not on disk.
  • Refused records are cut off. Before the store gives up, the refused record is cut off the WAL again (and synced), so a commit the caller was told failed is never replayed. If even that cut fails, its position is recorded in GRAPH, and the next open cuts there (every reader stops there meanwhile).
  • A panicking hook rolls back. A commit hook of your own that panics in before_commit rolls the unit back (the panic then continues to the caller); the journal runs after it and never writes that unit (Transactions). A panic inside the journal itself (a bug, or a store backend that panics while the WAL is written) cuts the record off again, poisons the journal and rolls the unit back too. Every case, with what a reopen gives, is in the table Failures while committing.
  • Poisoning. After a failed fsync the state of the file is unknown, so the store refuses every later commit and mark (PersistError::Poisoned) until it is reopened. Reopening runs the recovery and continues from the last durable commit.
  • After close() the graph refuses commits.

The durable end

store.durable_end() is the position after the last record whose sync completed (its WAL segment, byte offset and commit). It is readable without any lock: a backup copies the WAL only up to it, because a record after it may still be vetoed.

WAL segments

The WAL is a sequence of segment files wal/<first commit>.wal, 16 MiB each by default (StoreOptions::with_wal_segment_bytes; a record is never split, a segment closes after the record that passes the size). A segment is closed with an fsync before the next one starts, so only the last segment can end in an interrupted write. Starting a segment is recorded in GRAPH, so a lost, emptied or replaced newest segment is reported as damage instead of opening the store at an earlier commit (Damage).

Large records (from 4 KiB) are compressed (StoreOptions::with_codec).

The size of one commit

Every commit is one WAL record, and its body is at most 1 GiB before compression (format spec 9.3). The body holds every change of the unit with its before and after values, a dropped element with its full properties included, so a unit that drops and re-creates a few hundred thousand elements with long texts can reach it. Such a commit is refused (GraphError::CommitRejected from the hook journal, its source PersistError::CommitTooLarge): the unit rolls back, nothing is written, the store keeps working. Split the work into smaller units: in the playground and the dev server a whole script is one unit, so drop in a run of its own and create the data over several runs; in the CLI, Python and Rust every traversal is its own unit already. See Resource Limits.

Reading the committed changes (CDC)

#![allow(unused)]
fn main() {
use graphersal::persist::{Store, StoreOptions};
let store = Store::open("data", StoreOptions::new())?;
for change in store.changes_since(10)? {        // every commit after commit 10, in order
    let change = change?;
    println!("commit {} at {}: {} mutation(s)", change.commit_seq(), change.committed_at(),
             change.mutations().len());
}
Ok::<(), Box<dyn std::error::Error>>(())
}

changes_since(commit) streams the committed ChangeSets from the WAL, up to the durable end: every mutation with its before and after values, the commit time and metadata. It is the basis for change data capture, replication and audit. It fails when the WAL it needs was pruned. For a live process, a commit hook gets every ChangeSet as it commits; the WAL's byte format is public (format specification, section 9) for readers in other processes and languages.