Read and write RDF data from and to LinkedDataHub dataspaces over HTTP

LinkedDataHub implements a uniform, generic RESTful Linked Data API as defined by the SPARQL 1.1 Graph Store HTTP Protocol. It adds a few conventions of its own, however — such as automatic document metadata — as well as some constraints.

Authentication

The LinkedDataHub UI supports two authentication methods:

See how those authentication methods can be configured or how to get an account on LinkedDataHub.

HTTP API access using the CLI or curl currently does not support the OIDC method.

Access control

All HTTP access to documents is subject to access control. Requesting a document with insufficient access rights will result in a 403 Forbidden response. That means either:

  • the user is not authenticated and public access to the document is not allowed
  • the user is authenticated but the associated agent does not have an authorization to perform the action on the requested document

Content negotiation

LinkedDataHub implements proactive content negotiation based on the request Accept header value. The following RDF media types are supported:

Supported formats and their media types
Format Media type Direction
Turtle text/turtle Request, response
N-Triples application/n-triples Request, response
RDF/XML application/rdf+xml Request, response
JSON-LD application/ld+json Request, response
RDF/POST application/x-www-form-urlencoded Request only

The same document, negotiated as N-Triples:

GET /products/1/ HTTP/1.1
Host: localhost:4443
Accept: application/n-triples

HTTP/1.1 200 OK
Content-Type: application/n-triples
ETag: "1e635b4-ntriples"

<https://localhost:4443/products/1/#this> <https://schema.org/name> "Chai" .
…

Error responses

LinkedDataHub provides machine-readable error responses in the requested RDF format. An example of 403 Forbidden:

@prefix xsd:  <http://www.w3.org/2001/XMLSchema#> .
@prefix http: <http://www.w3.org/2011/http#> .
@prefix sc:   <http://www.w3.org/2011/http-statusCodes#> .
@prefix dct:  <http://purl.org/dc/terms/> .

[ a http:Response ;
    dct:title "Access not authorized" ;
    http:reasonPhrase "Forbidden" ;
    http:sc sc:Forbidden ;
    http:statusCodeValue "403"^^xsd:long
] .

Managing documents

Every document is also a named graph in the dataspace's RDF dataset. LinkedDataHub supports the SPARQL Graph Store Protocol's direct graph identification as the HTTP CRUD protocol for managing document data.

GSP indirect graph identification is not supported starting with LinkedDataHub version 5.x.

The API also supports the PATCH HTTP method which is optional in GSP. It accepts graph-scoped SPARQL updates that will modify the requested document. Only the INSERT/WHERE and DELETE WHERE forms are supported; GRAPH patterns are not allowed.

Trailing slashes in document URIs are enforced using 308 Permanent Redirect responses.

HTTP methods and the responses they return
Method Description Success Failure Reason
GET Returns the data of a document 200 OK 404 Not Found Document with request URI not found
406 Not Acceptable Media type not supported
POST Appends data to a document 204 No Content 400 Bad Request RDF syntax error
404 Not Found Document with request URI not found
413 Payload Too Large Request body too large
415 Unsupported Media Type Media type not supported
422 Unprocessable Entity Constraint violation
PUT Upserts a document 200 OK
201 Created
308 Permanent Redirect
400 Bad Request RDF syntax error
Malformed document URI
413 Payload Too Large Request body too large
415 Unsupported Media Type Media type not supported
422 Unprocessable Entity Constraint violation
DELETE Removes the requested document 204 No Content 404 Not Found Document with request URI not found
405 Method Not Allowed Deleting root, owner, or secretary documents is not allowed
PATCH Modifies a document using SPARQL Update 204 No Content 422 Unprocessable Entity SPARQL update string violates syntax constraints

Any method may additionally answer 403 Forbidden — see Access control. Any method that writes may answer 428 Precondition Required when the request carries no precondition, or 412 Precondition Failed when the one it carries is out of date — see Conditional writes.

Document metadata

Since LinkedDataHub 5, the document hierarchy is managed automatically.

By default, LinkedDataHub treats an RDF document as an item by giving it the dh:Item type and attaching it to the parent container using sioc:has_container. If the client wants to create a container instead, it has to explicitly add the dh:Container type on the document resource; the new container will be attached to its parent using sioc:has_parent. In either case, the URI of the new document will be relative to its parent's.

LinkedDataHub will also manage additional document metadata, such as its owner and creation/modification timestamps.

For example, this HTTP request creating the Northwind regions container (Turtle syntax):

PUT /regions/ HTTP/1.1
Host: localhost:4443
Content-Type: text/turtle

@prefix dh:     <https://w3id.org/atomgraph/linkeddatahub/document-hierarchy#> .
@prefix dct:    <http://purl.org/dc/terms/> .

<> a dh:Container ;
    dct:title "Regions" .

will produce the following document triples:

@prefix dh:     <https://w3id.org/atomgraph/linkeddatahub/document-hierarchy#> .
@prefix dct:    <http://purl.org/dc/terms/> .
@prefix xsd:    <http://www.w3.org/2001/XMLSchema#> .
@prefix sioc:   <http://rdfs.org/sioc/ns#> .
@prefix acl:    <http://www.w3.org/ns/auth/acl#> .

<https://localhost:4443/regions/>
    a dh:Container ;
    dct:created "2025-03-31T21:46:21.984Z"^^xsd:dateTime ;
    dct:creator <https://admin.localhost:4443/acl/agents/865c2431-8436-4ae8-b300-2a531a013cd0/#this> ;
    dct:title "Regions" ;
    sioc:has_parent <https://localhost:4443/> ;
    acl:owner <https://admin.localhost:4443/acl/agents/865c2431-8436-4ae8-b300-2a531a013cd0/#this> .

The HTTP request to produce a new item can be empty:

PUT /regions/1/ HTTP/1.1
Host: localhost:4443
Content-Type: text/turtle

It will create an item document with the following triples:

@prefix dh:     <https://w3id.org/atomgraph/linkeddatahub/document-hierarchy#> .
@prefix dct:    <http://purl.org/dc/terms/> .
@prefix xsd:    <http://www.w3.org/2001/XMLSchema#> .
@prefix sioc:   <http://rdfs.org/sioc/ns#> .
@prefix acl:    <http://www.w3.org/ns/auth/acl#> .

<https://localhost:4443/regions/1/>
    a dh:Item ;
    dct:created "2025-03-31T20:45:42.802Z"^^xsd:dateTime ;
    dct:creator <https://admin.localhost:4443/acl/agents/865c2431-8436-4ae8-b300-2a531a013cd0/#this> ;
    sioc:has_container <https://localhost:4443/regions/> ;
    acl:owner <https://admin.localhost:4443/acl/agents/865c2431-8436-4ae8-b300-2a531a013cd0/#this> .

Graph-scoped updates

A PATCH applies a SPARQL update scoped to the requested document's graph. The request body must contain a single update operation, in one of two forms: DELETE/INSERT with a WHERE clause (either template can be omitted, giving plain INSERT { } WHERE { } or DELETE { } WHERE { }), or the DELETE WHERE shorthand. Every other form is rejected with 422 Unprocessable Entity, as is a body with more than one operation. Appending a description to the Chai product:

PATCH /products/1/ HTTP/1.1
Host: localhost:4443
Content-Type: application/sparql-update

PREFIX schema: <https://schema.org/>

INSERT
{
    <https://localhost:4443/products/1/#this> schema:description "A delicate black tea blend" .
}
WHERE {}

The update executes against that one graph, so GRAPH patterns are not allowed — this request is rejected with 422 Unprocessable Entity:

PATCH /products/1/ HTTP/1.1
Host: localhost:4443
Content-Type: application/sparql-update

PREFIX schema: <https://schema.org/>

INSERT
{
    GRAPH <https://localhost:4443/products/1/>    # GRAPH is implied by the request URI — not allowed here
    {
        <https://localhost:4443/products/1/#this> schema:description "A delicate black tea blend" .
    }
}
WHERE {}

Built-in constraints

LinkedDataHub has a few built-in constraints that are not found in the standard Graph Store Protocol:

  • It is not possible to delete the root document (returns 405 Method Not Allowed)
  • It is not possible to modify or delete the documents of the owner agent and the secretary agent (returns 405 Method Not Allowed)
  • A document can only be created with a URL relative to an existing container (i.e. resolving .. against the new document's URL must identify an existing container)

The built-in constraints are similar to, but separate from the ontology constraints.

Conditional writes

A write to a document that already exists must say which state it was written against. Every write reads the document's graph, changes it in memory and writes the whole graph back, so two writers that name no precondition overwrite each other with nothing to show that anything was lost.

Such a write is answered 428 Precondition Required. To make it conditional, send either:

  • If-Match with the document's current ETag, to modify what you read
  • If-None-Match: *, to create a document only if it does not exist yet

Creating a document needs no precondition — there is no prior state to have been written against — so a PUT to a URL that does not exist is accepted as it always was.

If the document changed since you read it, the write is refused 412 Precondition Failed, and that response carries the document's current ETag so you can re-read, reapply your change and retry without an extra request. If-None-Match: * against a document that already exists is likewise refused 412, which is how a client can create-or-append in two steps.

An ETag identifies a negotiated representation, not the graph alone: the same document answers a different tag as RDF/XML than as Turtle. Read the validator with the same Accept the write will send, or the write is refused 412.

HEAD /my/document/ HTTP/1.1
Accept: application/n-triples

HTTP/1.1 200 OK
ETag: "f731c06f5ac4fec6490dff1399bd5440"

PATCH /my/document/ HTTP/1.1
Accept: application/n-triples
Content-Type: application/sparql-update
If-Match: "f731c06f5ac4fec6490dff1399bd5440"

HTTP/1.1 204 No Content

HEAD is answered for any access mode the agent holds, not acl:Read alone, so an agent granted acl:Write or acl:Append can read the validator its writes have to quote without being able to read the document. GET still requires acl:Read.

The command line interface does this for you: every ldh command that writes reads the document's validator first and sends it.

Versioned documents

In dataspaces with versioning enabled, document URLs accept additional query parameters that expose the version history through the standard Memento (RFC 7089) protocol:

?version=<commit-sha>
A historical version of the document, served with a Memento-Datetime header and an immutable Cache-Control. Snapshots are read-only: write methods answer 405 Method Not Allowed
?timemap
The document's version history — a TimeMap described with PROV-O in RDF formats, or in application/link-format as RFC 7089 requires
?timegate
Datetime negotiation: a request with an Accept-Datetime header answers 302 Found with the closest version in Location

See the versioning reference for details.

Executing SPARQL

Every LinkedDataHub dataspace provides a SPARQL endpoint on sparql path (relative to the dataspace's base URI). It supports the SPARQL 1.1 Protocol and serves as a proxy for the backend endpoint of the dataspace.

The protocol's dataset parameters can be used to scope a query to a single document's graph: passing default-graph-uri=<document-uri> makes that document's named graph the query's default graph. The dataset specified this way overrides any FROM/FROM NAMED clauses in the query, and GRAPH patterns match nothing since no named graphs are part of the dataset. This is the read-side counterpart of the graph-scoped PATCH method. Reading the Chai product's name out of its document alone:

curl -k -G "https://localhost:4443/sparql" \
  -H "Accept: application/sparql-results+json" \
  --data-urlencode 'query=PREFIX schema: <https://schema.org/> SELECT ?name WHERE { ?product a schema:Product ; schema:name ?name }' \
  --data-urlencode 'default-graph-uri=https://localhost:4443/products/1/'

Linked Data proxy

LinkedDataHub works as a Linked Data proxy (from the end-user perspective, as a Linked Data browser) when a URL is provided using the uri query parameter. All HTTP methods are supported.

If the URL dereferences successfully as RDF, LinkedDataHub forwards its response body (re-serializing it to enable content negotiation). During a write request, the request body is forwarded to the provided URL.

If the URL returns an HTML document, LinkedDataHub extracts every embedded JSON-LD script (<script type="application/ld+json">) and parses it as RDF using JSON-LD 1.1. Pages annotated with schema.org markup can therefore be browsed as Linked Data. The schema.org JSON-LD context is bundled with LinkedDataHub and served locally, so no external network request is made to resolve it. A supplier's homepage carrying this markup:

<script type="application/ld+json">
{
    "@context": "https://schema.org/",
    "@type": "Corporation",
    "@id": "#this",
    "legalName": "Exotic Liquids",
    "telephone": "(171) 555-2222"
}
</script>

browses as these triples:

<https://supplier.example/#this> a schema:Corporation ;
    schema:legalName "Exotic Liquids" ;
    schema:telephone "(171) 555-2222" .

The proxy only accepts external (non-relative to the current dataspace's base URI) URLs; local URLs have to be dereferenced directly.

The proxy forwards the origin's ETag and Last-Modified validators and passes conditional request headers (If-Match, If-None-Match, If-Modified-Since, If-Unmodified-Since) through. Preconditioned writes against proxied documents are therefore evaluated at the origin. Error statuses from the origin reach the client unchanged.

Federation

The Linked Data proxy makes dataspaces interoperable across origins: one dataspace's client can browse, query, and write to another origin's dataspace, whether on the same LinkedDataHub instance or a different one.

Browsing a UNESCO Thesaurus concept from the Northwind dataspace, for example, the proxy forwards the remote dataspace's hypermedia along with the RDF:

GET /?uri=https%3A%2F%2Funesco-thesaurus.demo.linkeddatahub.com%2Fconcepts%2Fconcept3683%2F HTTP/1.1
Host: northwind-traders.demo.linkeddatahub.com
Accept: application/rdf+xml

HTTP/1.1 200 OK
Content-Type: application/rdf+xml
Link: <https://unesco-thesaurus.demo.linkeddatahub.com/sparql>; rel="http://www.w3.org/ns/sparql-service-description#endpoint"
  • The remote dataspace's SPARQL endpoint is discovered from the Link headers the proxy forwards, so views and charts over remote documents query the remote endpoint
  • Writes are graph-scoped SPARQL updates (PATCH) or appends (POST) sent through the proxy under the origin's If-Match precondition
  • Requests carry the delegated identity of the requesting agent, so the remote dataspace's access control arbitrates every operation

In the user interface this means you can browse a remote dataspace and add data to its documents the same way as to local ones, provided the remote access control allows it.

The example above uses the public demo hosts because they are reachable without any setup, but nothing about federation requires a second instance: two dataspaces on one LinkedDataHub are distinct origins, so the same exchange happens between northwind-traders.demo.localhost:4443 and unesco-thesaurus.demo.localhost:4443 once both are declared.

System endpoints

Endpoints that LinkedDataHub serves in addition to the document hierarchy. Paths are relative to the dataspace's base URI.

System endpoints by path
Path Description
Admin and end-user apps
ns In-memory namespace ontology as well as its SPARQL endpoint
sitemap.xml The dataspace's public documents, in the sitemaps protocol: generated at startup from the public read rules of the dataspace's own admin service, so an origin lists only its own documents. An end-user dataspace with nothing public has none, and neither does an admin dataspace
robots.txt Allows crawling and names the sitemap when the dataspace has one; disallows everything when it has none
settings The dataspace's own description in the dataspace configuration. PATCH with a SPARQL update edits it — this is how packages are installed and uninstalled. Editing is restricted to end-user dataspaces: a PATCH on an admin dataspace's settings answers 422 Unprocessable Entity
Admin app only
admin/access Access metadata (for the authenticated agent)
admin/access/request Access request (for the authenticated agent)
admin/sign%20up WebID-TLS agent signup
oauth2/login/google OpenID Connect with Google login
oauth2/login/orcid OpenID Connect with ORCID login
oauth2/authorize/google OpenID Connect with Google callback
oauth2/authorize/orcid OpenID Connect with ORCID callback
admin/clear Clears every cached ontology graph and imports closure from memory. With an optional uri form parameter, also purges that ontology's proxy caches and reloads its closure from the admin SPARQL endpoint

Caching

GET and HEAD RDF responses from the backend triplestores (not LinkedDataHub responses) are cached automatically by LinkedDataHub using Varnish as an HTTP proxy cache. You can check the age of the response by inspecting the Age response header (the value is in seconds).

LinkedDataHub sends ETag response headers derived from a digest of the document's URL and its RDF content. Every representation (HTML, RDF/XML, Turtle, etc.) gets a distinct ETag value, and language-negotiated HTML representations factor the accepted languages into the tag. Treat the value as opaque: it is a validator to quote back in a conditional request, and nothing about the document can be read from it.

HTML responses carry a Content-Language negotiated against the languages the UI ships with, and advertise Vary: Accept-Language so shared caches keep the renderings apart.

Caching of LinkedDataHub responses can be enabled on the nginx HTTP proxy server by uncommenting the add_header Cache-Control directives in the platform/nginx.conf.template file. Caching of /uploads/ and /static/ namespaces is enabled by default.