Study notes
Overview
OntoDB is a database management system for ontologies. Data is RDF. The query language is SPARQL. These notes are the map of each feature, and of how a later comparison becomes a writeup.
A query is a pipeline. A committed insert is that same pipeline plus a log. The memory path already answers a SPARQL subset. The index, the optimizer, and recovery are specified and still to build.
The pieces
Two paths
A read of data/tiny.ttl stays in memory:
- The shell collects a query until braces and quotes balance.
- The lexer and parser build an AST.
ameansrdf:type. - The binder resolves prefixes and interns constants.
- The planner builds a left-deep tree of triple scans, nested-loop joins, a filter, and a projection.
- Volcano
Init/Nextpulls rows of term ids fromMemStore. - The shell prints spellings looked up in
MemDictionary.
An insert, once the disk features exist, does not stop at the memory store:
- The dictionary turns three spellings into ids. The bytes live in a term heap.
- Those ids are packed into a 24-byte big-endian
TripleKey. - The indexed store inserts that key into each permutation the catalog has a root for.
- The B+ tree pins pages through guards. Frames live in the buffer pool. Eviction is LRU-K.
- Before a dirty page is written, the log is flushed through that page's pageLSN. That is the WAL rule.
- Commit forces a commit record. On the next open, ARIES runs analysis, redo, and undo.
What already runs
- Turtle, N-Triples, and RDF/XML through raptor2. OWL is stored as RDF. There is no reasoner.
PREFIX,SELECT(variables or*), basic graph patterns witha,;, and,, andFILTERwith= != < > <= >=and&&.DISTINCT,LIMIT,OFFSET, andORDER BYare parsed and placed in the plan. Their executors are not built yet, so executing them throws.INSERT DATAandDELETE DATA.\explainprints the plan and does not execute.\explain analyzeis plan 5.5.
OPTIONAL, UNION, property paths, aggregates, and DELETE/INSERT WHERE fail with a line and a column. That is a finished error, not a crash.
What this site is for
Read a feature before you build it. Each page says what the piece is, how OntoDB uses it, what works today, and the slice that comes next. Benchmarks and writeups is how a finished comparison is recorded. The code repository stays separate. PLAN.md there is the feature list.