OntoDB

Study notes

Overview

OntoDB is a database management system for ontologies. Data is RDF. The query language is SPARQL. These notes are the map of each feature, and of how a later comparison becomes a writeup.

A query is a pipeline. A committed insert is that same pipeline plus a log. The memory path already answers a SPARQL subset. The index, the optimizer, and recovery are specified and still to build.

The pieces

Two paths

A read of data/tiny.ttl stays in memory:

  1. The shell collects a query until braces and quotes balance.
  2. The lexer and parser build an AST. a means rdf:type.
  3. The binder resolves prefixes and interns constants.
  4. The planner builds a left-deep tree of triple scans, nested-loop joins, a filter, and a projection.
  5. Volcano Init / Next pulls rows of term ids from MemStore.
  6. The shell prints spellings looked up in MemDictionary.

An insert, once the disk features exist, does not stop at the memory store:

  1. The dictionary turns three spellings into ids. The bytes live in a term heap.
  2. Those ids are packed into a 24-byte big-endian TripleKey.
  3. The indexed store inserts that key into each permutation the catalog has a root for.
  4. The B+ tree pins pages through guards. Frames live in the buffer pool. Eviction is LRU-K.
  5. Before a dirty page is written, the log is flushed through that page's pageLSN. That is the WAL rule.
  6. Commit forces a commit record. On the next open, ARIES runs analysis, redo, and undo.

What already runs

OPTIONAL, UNION, property paths, aggregates, and DELETE/INSERT WHERE fail with a line and a column. That is a finished error, not a crash.

What this site is for

Read a feature before you build it. Each page says what the piece is, how OntoDB uses it, what works today, and the slice that comes next. Benchmarks and writeups is how a finished comparison is recorded. The code repository stays separate. PLAN.md there is the feature list.