OntoDB

Terms

Dictionary

The index never compares strings. It compares ids. The dictionary is the only place a spelling lives.

term_id_t is a uint64_t. Ids start at 1. INVALID_TERM_ID is the maximum uint64 and means unbound, not a term. Zero is not a valid id, so a zeroed key does not accidentally name the first term you interned.

One spelling

term_codec is the canonical form both the loader and the binder use.

Escapes are part of the spelling. Two sources that mean the same IRI must intern once. Insert of an existing spelling returns the original id.

Numeric datatypes (xsd:integer, xsd:decimal, xsd:double, xsd:float) compare by value in filters and in ORDER BY. Everything else compares by id for equality, and by the canonical spelling for ordered comparisons of the same kind.

Memory dictionary

MemDictionary is a mutex-guarded map in both directions. It is what the shell uses today. The mutex lets two sessions share the process. It is not isolation. Locks come later.

Term heap

On disk, the bytes of a spelling live in a slotted page, not in the B+ tree leaf. A leaf key is three ids, 24 bytes, so the fanout stays high. A leaf that stored the IRIs would be a leaf of strings, and a split would copy those strings.

TermHeap::Insert returns nullopt when the record does not fit on the page you asked. A record larger than an empty page throws StorageException. Get is valid only while the page is pinned. Delete frees the slot. A later insert may reuse it. FreeSpace includes the slot-directory entry, so the test captures free space before the delete and does not compare the page to itself.

The heap chains pages. Its iterator visits every live record once and holds no pin after operator* returns.

Disk dictionary

Both directions survive restart: spelling to id, and id to spelling. Ids stay dense from 1. A hash index from spelling to id, plus the heap from id to bytes, is enough. Do not intern a spelling twice after restart, or one IRI becomes two ids and every join on that IRI breaks.

The catalog remembers the dictionary root next to the permutation roots. Flush makes that metadata durable.