CSV imports and mapping to RDF using SPARQL CONSTRUCT

CSV is a plain-text format for tabular data.

Read the shared mapping query rules first — they apply to CSV and RDF imports alike.

A CSV import is a combination of multiple resources:

File
The CSV file to be mapped to RDF and imported (ldh:file)
Mapping query
A user-defined CONSTRUCT query that produces RDF (spin:query)
Delimiter
The single-character column separator (ldh:delimiter). The CLI defaults it to ,

A CSV import in LinkedDataHub has two steps:

  1. generic conversion creates an intermediary, generic CSV/RDF representation for each CSV row
  2. vocabulary conversion maps the CSV/RDF to the final RDF representation using the mapping query

Imports run in the background — see import execution.

The mapping is done one row at a time, with each row resulting in a newly created document, which must be attached to the document hierarchy. The documents have to be URI resources. The server will automatically assign URIs to the documents constructed in the default graph. Alternatively, it is possible to explicitly specify the document graph using a GRAPH block in the CONSTRUCT template (which is a Jena-specific extension of SPARQL 1.1).

The resulting RDF data is validated against constraints as it is created. Constraint violations, if any, are attached to the import item.

The running example in the following sections is the Northwind Traders category data, exactly as the tutorial and the demo app import it:

categoryID,categoryName,description,pictureHash
1,Beverages,Soft drinks coffees teas beers and ales,ec56b79671b30ecb98316d9b766fba198286b3b6
2,Condiments,Sweet and savory sauces relishes spreads and seasonings,7e9f535a3bf693ec856702856fffd3be5a3df4f1
3,Confections,Desserts candies and sweet breads,e781b6c8e83be3f23aa241416f92447a66dc35e4

An import itself is described in RDF, pairing the uploaded file with the mapping query:

<#categories-import> a ldh:CSVImport ;
    dct:title "Categories" ;
    spin:query <#categories-query> ;
    ldh:file <https://localhost:4443/uploads/58ee67ce53cfb3ad73368eda24a089314bb0bf41> ;
    ldh:delimiter "," .

Generic conversion

The data table is converted to a graph by treating rows as resources, columns as predicates, and cells as xsd:string literals. The approach is the same as CSV on the Web minimal mode.

@base <https://localhost:4443/> .

_:8228a149-8efe-448d-b15f-8abf92e7bd17
    <#categoryID> "1" ;
    <#categoryName> "Beverages" ;
    <#description> "Soft drinks coffees teas beers and ales" ;
    <#pictureHash> "ec56b79671b30ecb98316d9b766fba198286b3b6" .

_:ec59dcfc-872a-4144-822b-9ad5e2c6149c
    <#categoryID> "2" ;
    <#categoryName> "Condiments" ;
    <#description> "Sweet and savory sauces relishes spreads and seasonings" ;
    <#pictureHash> "7e9f535a3bf693ec856702856fffd3be5a3df4f1" .

_:e8f2e8e9-3d02-4bf5-b4f1-4794ba5b52c9
    <#categoryID> "3" ;
    <#categoryName> "Confections" ;
    <#description> "Desserts candies and sweet breads" ;
    <#pictureHash> "e781b6c8e83be3f23aa241416f92447a66dc35e4" .

Vocabulary conversion

This step provides a semantic lift for the generic RDF output of the previous step by mapping it to classes and properties from specific vocabularies. It also connects instances in the imported data to the documents in LinkedDataHub's dataset.

The mapping is a user-defined SPARQL CONSTRUCT that follows the shared mapping query rules.

Example

The demo's category mapping constructs a schema:ProductGroup per row and mints each row its own document with a GRAPH block:

PREFIX foaf:       <http://xmlns.com/foaf/0.1/>
PREFIX dct:        <http://purl.org/dc/terms/>
PREFIX schema:     <https://schema.org/>

CONSTRUCT
{
    GRAPH ?graph
    {
        ?graph dct:title ?categoryName ;
            foaf:primaryTopic ?category .

        ?category a schema:ProductGroup ;
            dct:title ?categoryName ;
            dct:description ?description ;
            schema:name ?categoryName ;
            schema:identifier ?categoryID ;
            schema:description ?description ;
            foaf:depiction ?picture .
    }
}
WHERE
{
    ?category_row <#categoryID> ?categoryID ;
        <#description> ?description ;
        <#categoryName> ?categoryName ;
        <#pictureHash> ?pictureHash .

    BIND(uri(concat(str($base), "categories/")) AS ?container)
    BIND(uri(concat(str(?container), encode_for_uri(?categoryID), "/")) AS ?graph)
    BIND(uri(concat(str(?graph), "#this")) AS ?category)
    BIND(uri(concat(str($base), "uploads/", encode_for_uri(?pictureHash))) AS ?picture)
}

(A mapping can also construct dh:Item resources in the default graph and let the server assign the document URIs, as described above — the GRAPH form gives the import full control over the URIs instead.)

When the import is complete, the imported documents appear as children of the {base}categories/ container. The Beverages row becomes:

PREFIX  dct:  <http://purl.org/dc/terms/>
PREFIX  foaf: <http://xmlns.com/foaf/0.1/>
PREFIX  schema: <https://schema.org/>

<https://localhost:4443/categories/1/>
    dct:title "Beverages" ;
    foaf:primaryTopic <https://localhost:4443/categories/1/#this> .

<https://localhost:4443/categories/1/#this> a schema:ProductGroup ;
    dct:title "Beverages" ;
    dct:description "Soft drinks coffees teas beers and ales" ;
    schema:name "Beverages" ;
    schema:identifier "1" ;
    schema:description "Soft drinks coffees teas beers and ales" ;
    foaf:depiction <https://localhost:4443/uploads/ec56b79671b30ecb98316d9b766fba198286b3b6> .

Management

CSV import actions and their CLI commands
Action CLI command
Create CSV import from an uploaded file and a query (--file, --query, --delimiter) ldh add csv-import
Import a local CSV file end-to-end (--csv-file, --query-file, --delimiter) ldh import csv

If you are ready to import some CSV, follow the user guide on creating a CSV import.