Read, write and compare RDF.

Feed a parser input as it arrives and handle statements as they close. Build terms under a type system that rules out the invalid ones. Decide whether two graphs denote the same thing.

STEP 01

Add the crates to an Alire application

Configure the Flyology organization index, then let Alire resolve the crates as ordinary dependencies. flyology_rdf depends only on flyology_iri; flyology_n3 and flyology_sparql depend on it for terms and for the shared scanner.

Flyology organization index
alr index --reset-community
alr index --add=git+https://github.com/flyology-ada/alire-index.git \
  --name=flyology --before=community
alr with flyology_rdf

Add flyology_n3 or flyology_sparql the same way when you need them.

alire.toml
[[depends-on]]
flyology_rdf = "*"
flyology_n3 = "*"        --  optional
flyology_sparql = "*"    --  optional

flyology_rdf has no runtime dependency and no tasking. Its only scheduler-related feature is an optional cooperative-yield callback, so it runs the same in a plain sequential program and inside a lightweight-task runtime.

STEP 02

How the parts fit together

The pipeline has four parts, and only the middle two concern you.

One scanner
The scanner decides token boundaries in one place for all three grammars, with a dialect switch. Deciding them twice produces a family of conformance bugs: a scanner that recognizes PREFIX by looking ahead misreads a prefixed name whose prefix is prefix.
A parser you feed
The parser holds at most one partial token and the statement it is assembling. It does not hold the document, and it does not hold what it has already given you.
A sink you write
Events arrive as they complete: a graph declaration, a statement, a diagnostic. You decide what happens to them: collect, stream onward, or count and discard.
Terms as a flat vector
A term is one allocation, with nodes in child-before-parent order and every reference an index. No cycle is representable, and a walk over a nested triple term is index arithmetic.
STEP 03

Read a document

Implement Event_Sink. All three callbacks are abstract, so you must override each one. Usually only On_Quad carries logic; this collector overrides the other two as null procedures and stores each statement in a Dataset, the set covered in step 09.

a sink that collects
type Collector is limited new RDF.Turtle_Parsers.Event_Sink with record
   Data : RDF.Datasets.Dataset := RDF.Datasets.Empty;
end record;

overriding procedure On_Graph_Declaration
  (Target : in out Collector;
   Graph  : RDF.Quads.Graph_Name;
   Span   : RDF.Turtle_Parsers.Source_Span) is null;

overriding procedure On_Diagnostic
  (Target : in out Collector;
   Value  : RDF.Turtle_Parsers.Parse_Diagnostic) is null;

overriding procedure On_Quad
  (Target : in out Collector;
   Value  : RDF.Quads.Quad;
   Span   : RDF.Turtle_Parsers.Source_Span)
is
begin
   RDF.Datasets.Insert (Target.Data, Value);
end On_Quad;

Create a parser, feed it, and finish it. Finish reports whether the document was well formed; you must call it.

reading a whole document
Sink   : Collector;
Parser : RDF.Turtle_Parsers.Parser :=
  RDF.Turtle_Parsers.Create
    (Source_Name => "example",
     Base_IRI    => "http://example.org/",
     Syntax      => RDF.Turtle_Parsers.Turtle_Syntax);
begin
   RDF.Turtle_Parsers.Feed (Parser, Document, Sink);
   if RDF.Turtle_Parsers.Finish (Parser, Sink)
      = RDF.Turtle_Parsers.Parse_Succeeded
   then
      ...
   end if;
STEP 04

The four supported syntaxes

One parser reads all four. You choose the syntax at Create; it cannot change mid-document, because the syntax decides what a token is.

Turtle_Syntax
Prefixes, collections, property lists, and the RDF 1.2 term syntax. One graph: a graph block is a syntax error.
TriG_Syntax
Turtle plus named graph blocks. This is the default, because it is the syntax that can express everything a dataset holds.
NTriples_Syntax
One statement per line, with no prefixes and no abbreviation. Nothing is retained between lines, so a stream of any length adds nothing to what the parser holds.
NQuads_Syntax
N-Triples with a fourth position for the graph name.

Is_Line_Based distinguishes the line-based syntaxes from the abbreviating ones. The distinction matters when you decide how much input to buffer.

testing for a line-based syntax
if RDF.Turtle_Parsers.Is_Line_Based (Syntax) then
   ...
end if;
STEP 05

Read named graphs

Every event carries its graph name through Quads.Graph, so a sink never has to remember which block it is inside.

reading a named graph
procedure Note (Statement : RDF.Quads.Quad) is
   Where : constant RDF.Quads.Graph_Name :=
     RDF.Quads.Graph (Statement);
begin
   case RDF.Quads.Kind (Where) is
      when RDF.Quads.Default_Graph_Kind    => ...
      when RDF.Quads.IRI_Graph_Kind        => ...
      when RDF.Quads.Blank_Node_Graph_Kind => ...
   end case;
end Note;

On_Graph_Declaration runs when a block names a graph. Use it when you mirror the document's structure rather than its content.

STEP 06

Feed input that arrives in pieces

Call Feed as often as you like. A split may fall anywhere: inside an escape, inside a literal, or between the bytes of a UTF-8 sequence. The result does not depend on where it fell.

feeding a byte at a time
for Index in Document'Range loop
   RDF.Turtle_Parsers.Feed (Parser, Document (Index .. Index), Sink);
end loop;
Status := RDF.Turtle_Parsers.Finish (Parser, Sink);

The Notation3 and SPARQL conformance harnesses parse every document they accept twice, whole and one byte at a time, and require the two to agree: 1,160 Notation3 documents and 317 SPARQL queries in the current corpora. The RDF parser's own test suite parses each of its documents both ways. Those parsers take chunks through Flyology_N3.Parsers.Create and Flyology_SPARQL.Parsers.Create, though they emit nothing as they read: a formula is a term and a query is a tree, and neither is usable until it closes.

STEP 07

Read a diagnostic

A Parse_Diagnostic is a value, not a message to match on. It carries what went wrong, the production the parser was reading, and where the failure occurred.

reporting a failure
overriding procedure On_Diagnostic
  (Target : in out Collector;
   Value  : RDF.Turtle_Parsers.Parse_Diagnostic)
is
   Where : constant RDF.Turtle_Parsers.Source_Span :=
     RDF.Turtle_Parsers.Span (Value);
begin
   Put_Line
     (RDF.Turtle_Parsers.Source_Name (Value)
      & ":" & Where.Start_Line'Image
      & ":" & Where.Start_Column'Image
      & ": " & RDF.Turtle_Parsers.Code (Value)'Image
      & " in " & RDF.Turtle_Parsers.Production (Value)'Image);
end On_Diagnostic;

The Source_Span reports both ends, in bytes as well as lines and columns, so a caller can underline the offending region without rescanning the input.

STEP 08

Set limits and yield

Every limit is declared, and none of them is a guess. Reaching one produces a diagnostic like any other, not an exception and not a crash.

limiting a parse
Limits : constant RDF.Turtle_Parsers.Parse_Limits :=
  (Maximum_Bytes       => 1_024 * 1_024,
   Maximum_Quads       => 1_000,   --  zero means no limit
   Maximum_Nesting     => 64,
   Maximum_Token_Bytes => 64 * 1_024,
   Maximum_Prefixes    => 256);

The parser can hand control back during a long input. Pass a Work_Checkpoint and the parser calls it every Work_Checkpoint_Interval units of work. A lightweight task yields inside that callback, so the library needs no knowledge of the scheduler.

what a parse cost
Cost : constant RDF.Turtle_Parsers.Work_Statistics :=
  RDF.Turtle_Parsers.Work (Parser);
--  All Work_Count: Bytes_Fed, Bytes_Scanned,
--  Tokens_Scanned, Statements_Parsed,
--  Checkpoint_Calls, Maximum_Pending_Bytes,
--  Maximum_Pending_Tokens

Terms are immutable and are copied on every path that carries one, so their nodes are shared and reference counted rather than held by value: copying a term is an atomic increment, not a walk over its nodes. If your program uses no abort statements and no asynchronous transfer of control, adding the restrictions below lets the compiler drop the abort deferral it otherwise emits around every controlled finalization, which on a parse-heavy workload measured about 4 per cent here. They are configuration pragmas, so they are yours to set rather than the library's.

optional, in your own configuration pragmas file
pragma Restrictions (No_Abort_Statements);
pragma Restrictions (Max_Asynchronous_Select_Nesting => 0);
pragma Restrictions (No_Asynchronous_Control);
STEP 09

Build terms and datasets

Constructors enforce every rule, so you cannot build an invalid term. A predicate is typed as an IRI rather than a term, which makes a literal in predicate position a compile error.

building terms
Named    : constant RDF.Terms.Term :=
  RDF.Terms.IRI_Term
    (RDF.IRIs.From_UTF_8 ("http://example.org/ada"));
Unnamed  : constant RDF.Terms.Term := RDF.Terms.Blank_Node ("b1");
Plain    : constant RDF.Terms.Term :=
  RDF.Terms.Literal ("flight", RDF.Terms.String_Datatype);
Directed : constant RDF.Terms.Term :=
  RDF.Terms.Directional_Literal
    ("flight", "en", RDF.Terms.Left_To_Right);

RDF 1.2 triple terms nest, and the nesting is bounded. Terms.Depth and Terms.Node_Count report what a term costs, and the constructors refuse anything past the model's limits.

A dataset has set semantics: inserting the same statement twice leaves one. Iteration is deterministic, so two runs over the same data visit it in the same order. Iterate takes a procedure that accepts one Quad.

working with a dataset
procedure Tally (Statement : RDF.Quads.Quad) is
begin
   Count := Count + 1;
end Tally;

RDF.Datasets.Insert (Data, Statement);
RDF.Datasets.Delete (From => Data, Value => Statement);
RDF.Datasets.Iterate (Data, Tally'Access);

Total  : constant Natural := RDF.Datasets.Length (Data);
Graphs : constant Natural := RDF.Datasets.Graph_Count (Data);

Iterate_Graph visits one named graph, and Delete removes a single statement.

STEP 10

Write a document

A writer turns a dataset back into text. Bind the prefixes the output should use with Bind, then serialize. Data here is the dataset the collector filled in step 03.

writing Turtle
Prefixes : RDF.Turtle_Writers.Prefix_Map :=
  RDF.Turtle_Writers.No_Prefixes;
begin
   RDF.Turtle_Writers.Bind (Prefixes, "", "http://example.org/");
   Put (RDF.Turtle_Writers.To_Turtle (Data, Prefixes));
the output
@prefix : <http://example.org/> .

:ada :born "1815-12-10" ;
    :wrote :notes .

:notes :about "flight"@en .

To_TriG writes named graphs as blocks. If the dataset holds a named graph, To_Turtle raises Constraint_Error, because Turtle cannot express one. NQuads_Writers.Write_Quad writes one statement per line; canonical output and dataset comparison are built on it.

STEP 11

Compare two graphs

Two documents that differ only in their blank node labels denote the same thing, and comparing their text will not tell you so. Canonicalization decides it: every blank node is labeled from its position in the graph rather than from what it was called.

comparing two graphs
if RDF.Canonicalization.Is_Isomorphic (Left, Right) then
   ...
end if;

Put (RDF.Canonicalization.To_Canonical_NQuads (Left));
the canonical form
_:c14n0 <http://example.org/q> "x" .
_:c14n1 <http://example.org/p> _:c14n0 .

Is_Isomorphic and To_Canonical_NQuads are the function forms, and they raise Work_Limit_Error when the bound is reached. The procedure form reports instead, and it also yields the identifier map when you need to say which node in the output was which node in the input. The digest is part of the request because RDFC-1.0 admits two digests, and the issued labels differ between them.

canonicalizing with the map
RDF.Canonicalization.Canonicalize
  (Value     => Data,
   Output    => Text,
   Labels    => Issued,
   Status    => Status,
   Algorithm => RDF.Canonicalization.SHA_256);
STEP 12

Encode a term in binary

The binary term format is a compact, self-delimiting encoding for a term or a statement. Use it when a value goes into a cache or onto a wire rather than into a document.

round-tripping a term
Encoded  : constant String := RDF.Codecs.Encode (Original);
Restored : constant RDF.Terms.Term :=
  RDF.Codecs.Decode_Term (Encoded);
STEP 13

Read Notation3

Notation3 expresses statements RDF cannot. A formula is a Model.Term, so a rule is an ordinary statement whose subject and object are formulas. Flyology_N3.Parsers.Parse reads a whole document.

a rule in Notation3
Rule : constant Flyology_N3.Model.Term :=
  Flyology_N3.Parsers.Parse
    ("@prefix : <http://example.org/> ." & ASCII.LF
     & "{ ?who :wrote ?what } => { ?what :by ?who } .",
     "http://example.org/");

The result is the document as a formula, so walking it is the same at every level. Model.Statement_Count and Model.Statement_At walk one level, and Model.Term_Kind says what each term is.

walking a formula
for Index in 1 .. Flyology_N3.Model.Statement_Count (Rule) loop
   declare
      Fact : constant Flyology_N3.Model.Statement :=
        Flyology_N3.Model.Statement_At (Rule, Index);
   begin
      case Flyology_N3.Model.Kind
             (Flyology_N3.Model.Subject (Fact)) is
         when Flyology_N3.Model.Formula_Kind  => ...
         when Flyology_N3.Model.Variable_Kind => ...
         when Flyology_N3.Model.List_Kind     => ...
         when Flyology_N3.Model.RDF_Kind      => ...
      end case;
   end;
end loop;
STEP 14

Read SPARQL

The crate reads and writes queries back as documents through Flyology_SPARQL.Parsers.Parse and Flyology_SPARQL.Writers.To_SPARQL. Nothing evaluates them: a query here is something to check, format or inspect.

reading a query back out
Parsed : constant Flyology_SPARQL.Syntax.Query :=
  Flyology_SPARQL.Parsers.Parse (Text);
begin
   Put (Flyology_SPARQL.Writers.To_SPARQL (Parsed));

A Syntax.Query is the same flat vector the RDF terms use, so a walk over its Syntax.Node_Reference values is index arithmetic over one allocation. Syntax.Where_Clause is the root of the pattern.

inspecting the tree
Where : constant Flyology_SPARQL.Syntax.Node_Reference :=
  Flyology_SPARQL.Syntax.Where_Clause (Parsed);

for Index in 1 ..
  Flyology_SPARQL.Syntax.Child_Count (Parsed, Where)
loop
   declare
      Node : constant Flyology_SPARQL.Syntax.Node_Reference :=
        Flyology_SPARQL.Syntax.Child (Parsed, Where, Index);
   begin
      if Flyology_SPARQL.Syntax.Kind (Parsed, Node)
         = Flyology_SPARQL.Syntax.Triple_Node
      then
         ...
      end if;
   end;
end loop;

Coverage is the whole query language: the four query forms, the pattern and expression grammars, property paths as far as sequence, alternative, inverse and the three cardinality operators, subqueries, VALUES, EXISTS, dataset clauses, aggregates and the RDF 1.2 term syntax. The parser also rejects what the grammar alone would admit: a projection that repeats a name, an AS that takes a name already in scope, a variable projected past a GROUP BY that dropped it. Deliberately absent: SPARQL Update, federated-query specifics beyond recognising SERVICE, and evaluation of any kind.