The default RDF dataset structure used by LinkedDataHub

Structure

The default LinkedDataHub dataset structure follows these conventions:

  1. The default graph is not used.
  2. Each document's description is stored in an RDF named graph whose name is the same as the document URI. The graph can be managed using the Graph Store Protocol.
  3. Documents form a parent/child hierarchy. There are two types of documents: containers, which can have other documents as children, and items, which cannot.

The default dataset is installed into a dataspace service during setup.

In TriG, the Northwind Beverages category document — one named graph, named by the document URI, holding the document's whole description:

<https://localhost:4443/categories/1/>
{
    <https://localhost:4443/categories/1/> a dh:Item ;
        dct:title "Beverages" ;
        sioc:has_container <https://localhost:4443/categories/> ;
        foaf:primaryTopic <https://localhost:4443/categories/1/#this> .

    <https://localhost:4443/categories/1/#this> a schema:ProductGroup ;
        schema:name "Beverages" ;
        schema:identifier "1" .
}

Hierarchical URIs

By default, the URIs of the document resources represent the same parent/child hierarchy: the root container's URI equals the dataspace's base URI, and descendant document URIs are relative to their parent container's.

https://localhost:4443/                    # root container = the base URI
https://localhost:4443/categories/         # its child container
https://localhost:4443/categories/1/       # an item in the container
https://localhost:4443/categories/1/#this  # the category itself — a fragment of its document

Uploaded files are an exception to this rule. They are content-addressed in the uploads/{sha1sum} namespace, where sha1sum is the SHA1 hash of the file content — the Beverages image, for example, lives at uploads/ec56b79671b30ecb98316d9b766fba198286b3b6 no matter which document references it.

Default datasets

The default datasets of administration and end-user dataspaces can be found in platform/datasets/admin.trig and platform/datasets/end-user.trig, respectively.

The end-user dataset contains the root container document and descriptions of the public DBpedia and Wikidata SPARQL services. The administration dataset contains the root container, the sign-up and ontology-clearing endpoint documents, the acl/ document tree — containers for agents, public keys, users, groups, authorization requests and authorizations, the default authorizations, and the owners, writers and readers groups — as well as the ontologies/ container.

RDF terms (classes, properties etc.) used in the default datasets come from the well-known FOAF and SIOC vocabularies (with additional LinkedDataHub-specific assertions) as well as system ontologies.

Storage

The dataset of each dataspace is stored in an RDF triplestore which is accessed as a SPARQL service. End-user dataspaces and admin dataspaces have separate datasets and are backed by two different services.

The default triplestore is Apache Jena Fuseki, and multiple dataspaces can share the same backend service. See the triplestores reference for compatibility requirements and configuration examples.

Read more about documents in the data model reference.