Schemas

A graph can store a schema: which vertex and edge labels exist, which properties they carry, of which types and within which constraints, and how edges may connect labels. The schema is plain data in one format, the same in Rust, the script DSL, Python and the playground: our envelope outside, standard JSON Schema 2020-12 inside (a strict subset), so a model_json_schema() from Pydantic can be pasted as a label's schema.

The rule behind enforcement: the storage owns consistency. Every operation that changes the graph passes the schema inside one storage call, never in a step, and every traversal is one unit of work: it is applied completely or not at all.

A first schema

g.set_schema(`{
  "mode": "closed",
  "vertices": {
    "person": {"schema": {
      "type": "object",
      "properties": {
        "name": {"type": "string", "minLength": 1},
        "age":  {"type": ["integer", "null"], "minimum": 0}
      },
      "required": ["name"]
    }}
  },
  "edges": {
    "knows": {"connections": [{"from": "person", "to": "person"}],
              "schema": {"type": "object", "properties": {"weight": {"type": "number"}}}}
  }
}`)
g.add_v("person").property("name", "Ann").next()     // ok
g.add_v("person").property("age", -1).next()         // error: age >= 0 (and name is required)

g.get_schema() returns the stored schema as a map (() without one), g.infer_schema() the schema the data satisfies. Rendered as the final value of a session (the CLI, the playground) a schema shows as a table; --format mermaid|plantuml|json|jsonschema|markdown|tree picks another view.

Freeze the structure you have

A common way to start: build the graph without a schema, experiment until the shape is right, then freeze that shape so every later write has to follow it:

g.set_schema(g.infer_schema(), SchemaMode.closed)   // or SchemaMode.open / SchemaMode.none

set_schema(schema, mode) stores schema with its mode replaced by mode (an inferred schema has mode none, which enforces nothing). The mode is a SchemaMode token (SchemaMode.closed, SchemaMode::closed) or one of the strings "none", "open", "closed"; validate_schema(schema, mode) is the dry run. SchemaMode is a reserved name in the script scope.

BindingFreeze
DSLg.set_schema(g.infer_schema(), SchemaMode.closed) (also setSchema(.., "closed"))
Rustlet s = g.infer_schema()?.with_mode(SchemaMode::Closed); g_mut.set_schema(s)?
Pythongraph.set_schema(graph.infer_schema(), mode="closed")
PlaygroundSchema panel: Freeze as open / Freeze as closed (shown while the graph stores no enforcing schema)

What each mode freezes:

  • open: the labels and properties that exist get their types and required lists; new labels and new properties stay allowed (unchecked).
  • closed: additionally, no other label, no other property (at any depth) and no edge between other labels than the ones that exist.

It always succeeds on the data the schema was inferred from, in both modes: inference declares every kind seen at every position and requires only the keys every element has (see Inference). The one exception is closed with elements without a label: an unlabeled vertex or edge has no label to declare, and a closed schema admits only declared labels. Then the call fails, changes nothing, and the report names them:

Cannot apply the Closed schema: the existing data violates it (2 kinds of violations on 2 elements of 6 checked)
  1) 1 edge without a label (Closed admits only declared labels) [ids: "7"]
  2) 1 vertex without a label (Closed admits only declared labels) [ids: "3"]

Freeze such a graph in open mode (it admits unlabeled elements unchecked), or give those elements a label first with the label steps (add_label, drop_label, set_label) and freeze again:

g.v("3").add_label("thing").to_list();                                  // a vertex from the report
g.e().not(__.has_label()).set_label("link").to_list();   // every unlabeled edge
g.set_schema(g.infer_schema(), SchemaMode.closed)

(g.v().not(__.has_label()).add_label("thing") labels every unlabeled vertex; Rust: the same steps, or GraphStorage::add_vertex_label/set_edge_label.) (The generated large sample graph has unlabeled edges: it freezes open, not closed.)

Later changes go through patch_schema (add a property, relax a rule) like any other schema.

The format

Envelope

KeyTypeMeaning
$schemastring, optionalThe format: https://graphersal.dev/schema/v1 (any other value is an error).
metaobject, optionalFree-form data the library stores, round-trips and never interprets (a version, a status, server data). Changing it is a schema change.
mode"none" | "open" | "closed", default "none"What is enforced, see Modes.
verticesobjectVertex label → {"schema": <label schema>}.
edgesobjectEdge label → {"connections": [{"from": .., "to": ..}], "schema": <label schema>}.

A label schema is a JSON Schema object node describing the element's property map: {"type": "object", "properties": {..}, "required": [..], "additionalProperties": ..} plus annotations. It may carry $defs (for $ref), $schema and $id (accepted and not stored). A label without schema declares no properties. A connection's from/to is a vertex label or * (any label).

Unknown keys are errors everywhere, the envelope included: {"mdoe": "closed"} is rejected with its location (Invalid schema at /mdoe: unknown key 'mdoe'), never silently ignored.

The canonical form (what get_schema, to_json and every binding return) always writes $schema, mode, vertices and edges, keywords in a fixed order, integers as integers. The empty schema (mode none, no meta, no labels) is no schema: set_schema({}) removes the stored schema. The format is described by the meta-schema graphersal-schema-v1.json (graphersal::schema::META_SCHEMA_JSON).

Types

JSON SchemaEngine typeNote
"integer"int64A value outside i64 cannot be stored at all; 3.0 is not an integer.
"number"float64An int64 written to it is coerced when exactly representable; a stored int64 does not conform.
"string"string
"string" + "format": "uuid"uuidA canonical UUID string is coerced on write; a stored string does not conform.
"boolean", "null"boolean, null
"array" + itemsarray<T>items is a full schema (constraints and nullability apply to the items); without items any item.
"object" + propertiesobjectWith required and additionalProperties.
["integer", "null"]union<int64|null>A type array: any of the listed types.
anyOf: [..]union<..>At least one member accepts the value (Pydantic's Optional[X] is anyOf: [X, {"type": "null"}]).
{} (no type)anyAdmits every value.

Diagnostics, Display and GType keep the engine's names (int64, float64, uuid).

Keywords

KindKeywords
Structuretype, format (only uuid), items, properties, required, additionalProperties (true/false), anyOf
Constraintsminimum, maximum, exclusiveMinimum, exclusiveMaximum (numbers), minLength, maxLength, pattern (strings), minItems, maxItems (arrays), enum, const (any value)
Annotations (no effect on validation)title, description, $comment, examples, deprecated, default
References$ref: "#/$defs/Name" with $defs at the label schema's root; inlined when the schema is loaded

A constraint applies to the values of its own kind and ignores the others (JSON Schema semantics): in {"type": ["integer", "string"], "minimum": 10, "minLength": 3} the bound checks the integers and the length the strings. Next to a declared type, a keyword of another kind is an error (minLength on an integer is a mistake, not a no-op). enum/const compare JSON values (1 equals 1.0, a uuid compares as its canonical string).

Local references are inlined: a loaded schema has no $ref/$defs, and a reference cycle is an error (recursive schemas are not supported). Only annotations may sit next to a $ref; they override the definition's.

What is an error, and why

InputWhy it is rejected
An unknown key or keyword, anywhereA typo must never weaken a schema silently.
format other than uuid (date-time, date, email, ...)Dates come with a date type in 0.2.0. Accepting date-time as a plain string now would let 0.2.0 silently change what an existing schema accepts. Remove the format (the value is a string).
additionalProperties: {schema} (a map of T, Pydantic dict[str, T])No map type yet. Use {"type": "object", "additionalProperties": true}.
oneOf, allOf, not, if/then/elseNot implemented; anyOf covers oneOf when the members cannot both match.
prefixItems, contains, tuples (items: [..])An array has one item schema.
uniqueItems, multipleOfNot yet (additive later).
patternProperties, propertyNames, minProperties, dependent*, unevaluated*Keys are declared with properties, required, additionalProperties.
readOnly, writeOnly, content*Not among the accepted annotations.
A remote or non-$defs $ref, a reference cycleReferences are inlined at load.
nullable, variants, min, min_length, int64, ... (spellings of other schema dialects, such as OpenAPI's nullable)The hint names the JSON Schema spelling.

Every error is SchemaError::InvalidSchema { path, reason, hint }, path a JSON Pointer into the input (/vertices/person/schema/properties/age/minimum), and its help names the fix.

default is not applied

default is an annotation only, and it stays one:

  • it is never applied on write: an element created without the property does not get it;
  • it is never substituted on read: values("age"), valueMap(), the display and every export show nothing for an absent property;
  • changing a default changes no data (the diff classifies it as a compatible annotation change);
  • a required property with a default is still required: the default does not fill it.

Why, although it may look surprising: this is JSON Schema's own meaning of default ("may be used by a user interface", not by a validator); Pydantic writes default on every optional field ("default": null), so applying it would materialise values nobody asked for; and a default that re-valued data whenever the schema changes would be "magic" no one can audit. Clients and editors may use it (the playground's schema editor prefills a form with it). Should create-time defaults (SQL DEFAULT semantics: stored physically, only for new elements) ever be supported, they get a separate keyword, because giving default that meaning later would silently change existing schemas.

Required and nullable

required (on the object) and nullability (the type) are independent, as in JSON Schema:

null not admitted ("type": "string")null admitted ("type": ["string", "null"])
in requiredmust be present, non-nullmust be present, may be null
not in requiredmay be absent, never nullmay be absent or null

A missing required key is MissingRequiredProperty; a null where the type does not admit it is a TypeMismatch (expected string, got null). null stays a real stored value (property("k", ()) stores it, it never removes a property: use remove_property, see Removing Properties). Inside a JSON value {"a": null} differs from {}.

Required properties must be supplied when an element is created. The optimizer rule add_property_fold folds addV(..)/addE(..) and the property(k, v) steps that follow into one creation, so g.addV("Person").property("name", "Ann").property("age", 30) creates a valid element in one call. With the rule disabled, or for a property() that cannot be folded, the empty creation is rejected before anything is written. Removing a required property is rejected in open and closed mode.

Modes and additionalProperties

noneopenclosed
Coercion on write (canonical string to uuid, int64 to float64)nodeclared propertiesdeclared properties
Types, constraints, required of declared labels (any depth)noyesyes
Undeclared labelsallowedallowed, uncheckedrejected (also unlabeled elements)
Default additionalProperties of an object node-truefalse
Edge topology (connections)nonoyes

mode governs undeclared labels and topology, and it is the default additionalProperties of every node typed object that does not state its own. An explicit value wins, both ways: a free-form object in a closed graph is {"type": "object", "additionalProperties": true}, a strict object in an open graph says false. A node without type: "object" (an anyOf wrapper) puts no constraint on keys; its members do.

Multi-label vertices (see Multi-Label Vertices): every label checks the properties it declares and its own required list; a top-level property that no label declares is allowed when ANY label admits additional properties (the same "any label" rule as topology). Coercion uses the first label, in label order, that declares the key. closed rejects a vertex without a label or with an undeclared label. Adding or removing a label re-validates the vertex under the new label set and commits only on success.

Edge labels are single, but can be changed (set_label(..), GraphStorage::set_edge_label). The change checks the edge as stored under the new label (closed: a declared label and a declared connection; open and closed: the new label's required keys and declared types) and changes nothing on a violation. Stored values are not coerced by a label change (reads never coerce); later writes coerce by the new label's declarations.

none describes without enforcing; the GraphML import still restores the declared logical types from it (uuid, array, object; see UUID Values).

Edge topology (closed)

An edge label declares its allowed connections ({"from": "Person", "to": "Company"}, * matches every label). Because labels are a set, an edge is accepted if ANY label of the source and ANY label of the target matches a declared connection. An edge label without connections allows every pair. A label change that would make an incident edge violate the connections is rejected as a whole.

Nested values

The content of an object, array or anyOf value is validated at any depth: declared nested keys are type- and constraint-checked, required nested keys must be present, undeclared nested keys follow the object's additionalProperties. Errors name the canonical path ($.address.zip, $.items[1].n). A nested write (property(jpath("address.zip"), v)) coerces the leaf to the type declared at that location. See Path Keys (jpath).

Constraints

pattern is a regular expression with JSON Schema semantics: it is not anchored, so a value matches when the expression matches anywhere in it. Write ^...$ to match the whole value:

{ "type": "string", "pattern": "^[A-Z][a-z]+$" }
{ "type": "string", "pattern": "@" }
{ "type": "string", "pattern": "^\\d{4}-\\d{2}-\\d{2}$" }

The syntax is the ECMA-262-like subset of the Rust regex crate: literals, ., classes ([a-z], [^0-9], \d, \w, \s), anchors, groups, alternation | and the quantifiers *, +, ?, {n}, {n,}, {n,m}. Two differences to ECMA-262: look-around and back-references are not supported (matching stays linear in the input), and \d, \w, \s, \b are Unicode-aware (write [0-9] for ASCII digits). A JSON text doubles every backslash. The expression is compiled once, when the schema is loaded, and an invalid one rejects the schema.

minLength/maxLength count characters (code points). An empty array satisfies every items (it has no item of a wrong type); use minItems: 1 to require one.

The schema methods

Seven methods, the same everywhere. They are plain calls on the graph source, not traversal steps (like g.statistics()); each runs in the current unit (a transaction, a whole-script unit) or as a unit of its own.

MethodReturnsAuthorization
get_schema()the stored schema, or none (never inferred)Read on Schema
set_schema(schema), set_schema(schema, mode)- (replaces the stored schema; mode replaces the schema's mode)Update on Schema
patch_schema(patch)the new schema (JSON Merge Patch, RFC 7396, onto the stored one)Update on Schema
validate_schema(schema), validate_schema(schema, mode){violations, diff, compatible}: a dry run against the dataRead on Schema
validate_schema_patch(patch)the same for a patchRead on Schema
infer_schema()the schema the data satisfiesRead on Schema
diff_schema(schema){changes, compatible} from the stored schemaRead on Schema
InputOutput
Rust (GraphTraversalSource)GraphSchema (GraphSchema::from_json(text), from_value, or the builders), SchemaPatch (graphersal::schema)GraphSchema, SchemaValidation, SchemaDiff (to_value() gives the data shape)
DSL (both spellings: get_schema/getSchema, ...)a map or JSON texta map
Python (Graph.get_schema(), ...)a dict or JSON str (keyword policy=)a dict
Playground engine (wasm getSchema, ...)JSON textJSON

A map built in the DSL keeps its keys in alphabetical order (Rhai maps are sorted), so the property order of a schema that went through a map is alphabetical; JSON text and Python dicts keep it. The traversal step g.v()...infer_schema() infers over a stream, with the same inference.

In a DSL map literal, quote the keys that are not plain identifiers: $ref, $defs, $schema and $id start with $. default (a word Rhai reserves) is accepted as a plain key in the Graphersal DSL, quoted or not:

g.patch_schema(#{vertices: #{person: #{schema: #{type: "object", properties: #{
    age: #{type: ["integer", "null"], default: 7}
}}}}})

JSON text needs no quoting rules of its own: g.set_schema(text) takes the schema as written.

A schema change is a database change: it advances the commit sequence number, makes the run mutated, is recorded as Mutation::SetSchema { before, after } (the canonical JSON) in the same change set as the data, rolls back with its unit, and is reported in the result's changes: {data, schema} (Execution::changes(), the DSL's r.changes, the playground's response). Validation is immediate: the existing data is checked when the schema is set, not at commit.

Applying a schema to existing data

set_schema and patch_schema validate all existing data first, in one pass over the vertices and edges, when the mode is open or closed. If anything violates the schema, the graph and its current schema stay unchanged and the call fails with one aggregated error that lists every kind of violation, grouped by rule with a count and a few sample ids:

Cannot apply the Closed schema: the existing data violates it (5 kinds of violations on 22 elements of 22 checked)
  1) Person.age is declared mandatory but is missing on 12 vertices [ids: "p0", "p1", "p2", "p3", "p4" and 7 more]
  2) vertex label Alien is not declared (Closed): 5 vertices [ids: "x0", "x1", "x2", "x3", "x4"]
  3) Company.name is declared mandatory but is missing on 3 vertices [ids: "c0", "c1", "c2"]
  4) Person.age has the wrong type on 1 vertex (declared int64, found string) [ids: "t"]
  5) Robot.model is not declared (Closed) on 1 vertex [ids: "r"]
  • A group is (kind, label, field or path): missing required keys, undeclared labels or properties, type mismatches, constraint violations, edge topology violations (closed) and nested path violations (Person $.address.zip is declared mandatory but is missing on 1 vertex).
  • Every violation inside a value is listed, not only the first: all keys, all array items, at any depth, each under its own path ($.address.zip wrong type, $.address.street missing and $.address.city undeclared are three groups). Array indices are collapsed into [*]: a rule broken by any item of an array is ONE group ($.addresses[0].zip and $.addresses[2].zip missing are the group $.addresses[*].zip), which shows the first concrete path it met (Person $.addresses[*].zip is declared mandatory but is missing on 5 vertices (first at $.addresses[1].zip) [ids: ...]). Per node only the first failing constraint is named. An element counts once per group, however many of its items break the rule. A single write still fails on the first violation it meets.
  • Groups are ordered by count (largest first), then kind, element, label and field, so the same data gives the same report.
  • Each group shows its count and up to 5 sample ids (cut to 40 characters, escaped); at most 20 groups are printed, the structured report keeps up to 1000, further violations are only counted.
  • Stored data is checked exactly as it is: nothing is coerced and nothing is written.
  • validate_schema returns the same report as data: violations: {mode, elements_checked, violating_elements, total_violations, omitted_violations, groups: [{kind, element, label, field, example_field, detail, count, sample_ids, omitted_ids}]} (example_field is the first concrete path of a [*] field, null otherwise), with kind one of unknown_label, missing_required, type_mismatch, constraint_violation, undeclared_property, invalid_topology.

How to proceed: fix or migrate the offending data, relax the schema (make a key optional, admit null, declare the label or property), apply it with mode none, or start from infer_schema(), which is valid for the data it came from.

Making a key required: backfill first, in one transaction

A schema that requires a key the existing elements lack is rejected. Backfill first and tighten last, all in one transaction, so no reader ever sees the intermediate state:

let declare = SchemaPatch::from_json(
    r#"{"vertices": {"person": {"schema": {"properties": {"email": {"type": "string"}}}}}}"#)?;
let require = SchemaPatch::from_json(
    r#"{"vertices": {"person": {"schema": {"required": ["name", "age", "email"]}}}}"#)?;
graph.transaction(|g| -> TraverserResult<()> {
    // 1. In a closed schema, declare the key (optional) so it may be written.
    g.traversal_mut().patch_schema(&declare)?;
    // 2. Backfill.
    g.traversal_mut().v(None).has_label("person").property("email", "unknown").to_list()?;
    // 3. Make it required.
    g.traversal_mut().patch_schema(&require)?;
    Ok(())
})?;

(schema_api_tests::backfill_then_require_in_one_transaction runs this recipe.)

The same in a script run as one unit (atomic=True in Python, every playground run): patch to optional, backfill, patch to required. A schema default does not backfill anything.

Changing a schema: patch and diff

patch_schema applies a JSON Merge Patch (RFC 7396) to the canonical form of the stored schema ({} without one): an object merges member by member (property order is kept), null removes a member, any other value (an array too) replaces it. RFC 7396 cannot set a member to null; use set_schema for a "default": null. validate_schema_patch previews the result against the data.

diff_schema lists the changes (kind, element, label, path such as $.address.zip or $.tags[*], before, after) and whether each is backward compatible: data and clients written against the old schema keep working: the old data stays valid and every write that followed the old schema still succeeds. In short, a compatible change only adds (an optional key, a label, an allowed value) or relaxes (a wider type, a looser bound, a less strict mode); tightening a rule or removing a declaration something could depend on (a label, a key of a closed object) is incompatible. This is the usual schema-registry convention: adding an optional property is compatible, also to an open object, and so is removing a property declaration from an open object (the key becomes an undeclared one, which the object admits; the change is still listed as property_removed). A server's "compatible changes only" policy can veto the rest in a before_commit hook.

ChangeCompatible when
modeit gets less strict (closed → open → none); open → closed is incompatible
meta, annotations (title, description, default, ...)always
label addedthe old mode was closed (no client could write it), or its schema requires no key and admits undeclared keys (an open graph)
label removednever
connection added/removedthe label's connections still allow every old pair (an empty list allows all), or the new mode is not closed (topology binds only there)
property addedit is optional (not in required), also in an open object
property removedthe object admits undeclared keys (open: mode open/none or additionalProperties: true), whether the key was optional or required; from a closed object never
key made required / optionalnever / always
additionalPropertiesfalse → true
typethe new type admits every kind the old one did ("integer" → ["integer", "null"])
minimum, maxLength, minItems, ...removed or relaxed
pattern, constremoved
enumremoved, or a superset

A key the old schema listed in required without declaring it was written by every old client, with any value: declaring it later is compatible only if the declaration accepts everything. With a new mode none every change is compatible (nothing is enforced any more).

The diff is static: it compares the two schemas, never the data. Existing data that conflicts with a compatible change (an element that already holds a newly declared optional key with another type, say) is found by validate_schema / validate_schema_patch, and set_schema refuses to apply the schema until it is fixed.

Inference

infer_schema() (and the step) report what the data holds, at every position: the kinds seen (a type array, or anyOf when a string and a uuid share a position), items of arrays, properties of objects, and required only for keys every object at that position has, at any depth (also inside arrays and next to other kinds). Every label of a vertex is declared; connections use the endpoints' first labels. The result has mode none and is valid for the data it came from in open and in closed mode (an unlabeled element has no label to declare, so closed rejects it; see Freeze the structure you have). It never guesses from string content: a UUID-shaped string is a string. Inference statistics (counts) are not part of the format.

Renderers

FormatShows
table, markdownone row per property path (address.zip, roles[].name), type, required, nullable, constraints (0..150, length 1..80, pattern ^x$), notes (description, default marked as annotation)
treethe labels and their properties as a tree
jsonthe canonical format
jsonschemaone standalone JSON Schema per label ($defs: vertex:person, edge:knows), the mode's default additionalProperties written out
mermaidan erDiagram; nested objects are entities of their own (`person
plantumlthe label topology as entity diagrams, then a @startjson view of the condensed schema

The meta-graph (schema.to_graph() in the DSL, GraphSchema::to_graph in Rust) has one vertex per vertex label and one edge per connection, each with its schema node as the schema property.

The playground's Schema panel shows the same schema as editable tables, as JSON, as a graph of its labels (node size and count from the data) and as the Mermaid diagram, here for a small cloud inventory:

The label graph of a schema: vertex labels with their counts, edge labels as arrows

The Mermaid entity diagram of the same schema

Pasting a Pydantic model

Model.model_json_schema() loads as a label schema as it is ($defs/$ref, anyOf with null, title, description, default, enum, const, exclusive*, additionalProperties: false), except for two things to remove first:

  • datetime/date/time fields ("format": "date-time", ...): no date type before 0.2.0; drop the format (a plain string) or the field.
  • dict[str, T] fields ("additionalProperties": {..}): no map type; write {"type": "object", "additionalProperties": true}.

Tuples (prefixItems) and Decimal (anyOf with a string pattern for some versions) are not supported either. The top level of the model must be an object ("type": "object").

A failing query changes nothing

The schema is validated per storage call (one add_v with its folded properties, one single property(key, value), one drop, ...), and every traversal is one unit of work: when a call fails, the whole traversal is rolled back, including what its earlier steps wrote (see Transactions).

In the playground a refused write shows the failing step and how to fix it (here a closed schema that declares cpus as int64):

A schema violation in the Results panel with its help text

CaseTest
property(map) with several entriesschema_gap_tests::gap4_property_map_is_atomic
property_json (Rust API)schema_gap_tests::gap4_property_json_is_atomic
mergeV/mergeE option(Merge.onMatch, map) with several entriesschema_gap_tests::gap4_merge_v_on_match_is_atomic
addE over several start verticesschema_gap_tests::gap4_add_e_over_several_endpoints_is_atomic
any traversal with several stepstransaction_tests::a_schema_violation_rolls_back_the_whole_traversal

A GraphML import into an existing graph is one unit too: an invalid element rolls back the elements imported before it. GraphML carries no schema: set the schema first, then import. The CLI does both in that order with graphersal --graph g.graphml --schema schema.json [--schema-mode closed] (also for GraphSON, and with --server); a violation stops the start with exit code 1.