Add the crates to an Alire application
Configure the Flyology organization index, then let Alire resolve the crates as ordinary dependencies. flyology_rdf depends only on flyology_iri; flyology_n3 and flyology_sparql depend on it for terms and for the shared scanner.
alr index --reset-community
alr index --add=git+https://github.com/flyology-ada/alire-index.git \
--name=flyology --before=community
alr with flyology_rdfAdd flyology_n3 or flyology_sparql the same way when you need them.
[[depends-on]]
flyology_rdf = "*"
flyology_n3 = "*" -- optional
flyology_sparql = "*" -- optionalflyology_rdf has no runtime dependency and no tasking. Its only scheduler-related feature is an optional cooperative-yield callback, so it runs the same in a plain sequential program and inside a lightweight-task runtime.
How the parts fit together
The pipeline has four parts, and only the middle two concern you.
- One scanner
- The scanner decides token boundaries in one place for all three grammars, with a dialect switch. Deciding them twice produces a family of conformance bugs: a scanner that recognizes
PREFIXby looking ahead misreads a prefixed name whose prefix isprefix. - A parser you feed
- The parser holds at most one partial token and the statement it is assembling. It does not hold the document, and it does not hold what it has already given you.
- A sink you write
- Events arrive as they complete: a graph declaration, a statement, a diagnostic. You decide what happens to them: collect, stream onward, or count and discard.
- Terms as a flat vector
- A term is one allocation, with nodes in child-before-parent order and every reference an index. No cycle is representable, and a walk over a nested triple term is index arithmetic.
Read a document
Implement Event_Sink. All three callbacks are abstract, so you must override each one. Usually only On_Quad carries logic; this collector overrides the other two as null procedures and stores each statement in a Dataset, the set covered in step 09.
type Collector is limited new RDF.Turtle_Parsers.Event_Sink with record
Data : RDF.Datasets.Dataset := RDF.Datasets.Empty;
end record;
overriding procedure On_Graph_Declaration
(Target : in out Collector;
Graph : RDF.Quads.Graph_Name;
Span : RDF.Turtle_Parsers.Source_Span) is null;
overriding procedure On_Diagnostic
(Target : in out Collector;
Value : RDF.Turtle_Parsers.Parse_Diagnostic) is null;
overriding procedure On_Quad
(Target : in out Collector;
Value : RDF.Quads.Quad;
Span : RDF.Turtle_Parsers.Source_Span)
is
begin
RDF.Datasets.Insert (Target.Data, Value);
end On_Quad;Create a parser, feed it, and finish it. Finish reports whether the document was well formed; you must call it.
Sink : Collector;
Parser : RDF.Turtle_Parsers.Parser :=
RDF.Turtle_Parsers.Create
(Source_Name => "example",
Base_IRI => "http://example.org/",
Syntax => RDF.Turtle_Parsers.Turtle_Syntax);
begin
RDF.Turtle_Parsers.Feed (Parser, Document, Sink);
if RDF.Turtle_Parsers.Finish (Parser, Sink)
= RDF.Turtle_Parsers.Parse_Succeeded
then
...
end if;The four supported syntaxes
One parser reads all four. You choose the syntax at Create; it cannot change mid-document, because the syntax decides what a token is.
Turtle_Syntax- Prefixes, collections, property lists, and the RDF 1.2 term syntax. One graph: a graph block is a syntax error.
TriG_Syntax- Turtle plus named graph blocks. This is the default, because it is the syntax that can express everything a dataset holds.
NTriples_Syntax- One statement per line, with no prefixes and no abbreviation. Nothing is retained between lines, so a stream of any length adds nothing to what the parser holds.
NQuads_Syntax- N-Triples with a fourth position for the graph name.
Is_Line_Based distinguishes the line-based syntaxes from the abbreviating ones. The distinction matters when you decide how much input to buffer.
if RDF.Turtle_Parsers.Is_Line_Based (Syntax) then
...
end if;Read named graphs
Every event carries its graph name through Quads.Graph, so a sink never has to remember which block it is inside.
procedure Note (Statement : RDF.Quads.Quad) is
Where : constant RDF.Quads.Graph_Name :=
RDF.Quads.Graph (Statement);
begin
case RDF.Quads.Kind (Where) is
when RDF.Quads.Default_Graph_Kind => ...
when RDF.Quads.IRI_Graph_Kind => ...
when RDF.Quads.Blank_Node_Graph_Kind => ...
end case;
end Note;On_Graph_Declaration runs when a block names a graph. Use it when you mirror the document's structure rather than its content.
Feed input that arrives in pieces
Call Feed as often as you like. A split may fall anywhere: inside an escape, inside a literal, or between the bytes of a UTF-8 sequence. The result does not depend on where it fell.
for Index in Document'Range loop
RDF.Turtle_Parsers.Feed (Parser, Document (Index .. Index), Sink);
end loop;
Status := RDF.Turtle_Parsers.Finish (Parser, Sink);The Notation3 and SPARQL conformance harnesses parse every document they accept twice, whole and one byte at a time, and require the two to agree: 1,160 Notation3 documents and 317 SPARQL queries in the current corpora. The RDF parser's own test suite parses each of its documents both ways. Those parsers take chunks through Flyology_N3.Parsers.Create and Flyology_SPARQL.Parsers.Create, though they emit nothing as they read: a formula is a term and a query is a tree, and neither is usable until it closes.
Read a diagnostic
A Parse_Diagnostic is a value, not a message to match on. It carries what went wrong, the production the parser was reading, and where the failure occurred.
overriding procedure On_Diagnostic
(Target : in out Collector;
Value : RDF.Turtle_Parsers.Parse_Diagnostic)
is
Where : constant RDF.Turtle_Parsers.Source_Span :=
RDF.Turtle_Parsers.Span (Value);
begin
Put_Line
(RDF.Turtle_Parsers.Source_Name (Value)
& ":" & Where.Start_Line'Image
& ":" & Where.Start_Column'Image
& ": " & RDF.Turtle_Parsers.Code (Value)'Image
& " in " & RDF.Turtle_Parsers.Production (Value)'Image);
end On_Diagnostic;The Source_Span reports both ends, in bytes as well as lines and columns, so a caller can underline the offending region without rescanning the input.
Set limits and yield
Every limit is declared, and none of them is a guess. Reaching one produces a diagnostic like any other, not an exception and not a crash.
Limits : constant RDF.Turtle_Parsers.Parse_Limits :=
(Maximum_Bytes => 1_024 * 1_024,
Maximum_Quads => 1_000, -- zero means no limit
Maximum_Nesting => 64,
Maximum_Token_Bytes => 64 * 1_024,
Maximum_Prefixes => 256);The parser can hand control back during a long input. Pass a Work_Checkpoint and the parser calls it every Work_Checkpoint_Interval units of work. A lightweight task yields inside that callback, so the library needs no knowledge of the scheduler.
Cost : constant RDF.Turtle_Parsers.Work_Statistics :=
RDF.Turtle_Parsers.Work (Parser);
-- All Work_Count: Bytes_Fed, Bytes_Scanned,
-- Tokens_Scanned, Statements_Parsed,
-- Checkpoint_Calls, Maximum_Pending_Bytes,
-- Maximum_Pending_TokensTerms are immutable and are copied on every path that carries one, so their nodes are shared and reference counted rather than held by value: copying a term is an atomic increment, not a walk over its nodes. If your program uses no abort statements and no asynchronous transfer of control, adding the restrictions below lets the compiler drop the abort deferral it otherwise emits around every controlled finalization, which on a parse-heavy workload measured about 4 per cent here. They are configuration pragmas, so they are yours to set rather than the library's.
pragma Restrictions (No_Abort_Statements);
pragma Restrictions (Max_Asynchronous_Select_Nesting => 0);
pragma Restrictions (No_Asynchronous_Control);Build terms and datasets
Constructors enforce every rule, so you cannot build an invalid term. A predicate is typed as an IRI rather than a term, which makes a literal in predicate position a compile error.
Named : constant RDF.Terms.Term :=
RDF.Terms.IRI_Term
(RDF.IRIs.From_UTF_8 ("http://example.org/ada"));
Unnamed : constant RDF.Terms.Term := RDF.Terms.Blank_Node ("b1");
Plain : constant RDF.Terms.Term :=
RDF.Terms.Literal ("flight", RDF.Terms.String_Datatype);
Directed : constant RDF.Terms.Term :=
RDF.Terms.Directional_Literal
("flight", "en", RDF.Terms.Left_To_Right);RDF 1.2 triple terms nest, and the nesting is bounded. Terms.Depth and Terms.Node_Count report what a term costs, and the constructors refuse anything past the model's limits.
A dataset has set semantics: inserting the same statement twice leaves one. Iteration is deterministic, so two runs over the same data visit it in the same order. Iterate takes a procedure that accepts one Quad.
procedure Tally (Statement : RDF.Quads.Quad) is
begin
Count := Count + 1;
end Tally;
RDF.Datasets.Insert (Data, Statement);
RDF.Datasets.Delete (From => Data, Value => Statement);
RDF.Datasets.Iterate (Data, Tally'Access);
Total : constant Natural := RDF.Datasets.Length (Data);
Graphs : constant Natural := RDF.Datasets.Graph_Count (Data);Iterate_Graph visits one named graph, and Delete removes a single statement.
Write a document
A writer turns a dataset back into text. Bind the prefixes the output should use with Bind, then serialize. Data here is the dataset the collector filled in step 03.
Prefixes : RDF.Turtle_Writers.Prefix_Map :=
RDF.Turtle_Writers.No_Prefixes;
begin
RDF.Turtle_Writers.Bind (Prefixes, "", "http://example.org/");
Put (RDF.Turtle_Writers.To_Turtle (Data, Prefixes));@prefix : <http://example.org/> .
:ada :born "1815-12-10" ;
:wrote :notes .
:notes :about "flight"@en .To_TriG writes named graphs as blocks. If the dataset holds a named graph, To_Turtle raises Constraint_Error, because Turtle cannot express one. NQuads_Writers.Write_Quad writes one statement per line; canonical output and dataset comparison are built on it.
Compare two graphs
Two documents that differ only in their blank node labels denote the same thing, and comparing their text will not tell you so. Canonicalization decides it: every blank node is labeled from its position in the graph rather than from what it was called.
if RDF.Canonicalization.Is_Isomorphic (Left, Right) then
...
end if;
Put (RDF.Canonicalization.To_Canonical_NQuads (Left));_:c14n0 <http://example.org/q> "x" .
_:c14n1 <http://example.org/p> _:c14n0 .Is_Isomorphic and To_Canonical_NQuads are the function forms, and they raise Work_Limit_Error when the bound is reached. The procedure form reports instead, and it also yields the identifier map when you need to say which node in the output was which node in the input. The digest is part of the request because RDFC-1.0 admits two digests, and the issued labels differ between them.
RDF.Canonicalization.Canonicalize
(Value => Data,
Output => Text,
Labels => Issued,
Status => Status,
Algorithm => RDF.Canonicalization.SHA_256);Encode a term in binary
The binary term format is a compact, self-delimiting encoding for a term or a statement. Use it when a value goes into a cache or onto a wire rather than into a document.
Encoded : constant String := RDF.Codecs.Encode (Original);
Restored : constant RDF.Terms.Term :=
RDF.Codecs.Decode_Term (Encoded);Read Notation3
Notation3 expresses statements RDF cannot. A formula is a Model.Term, so a rule is an ordinary statement whose subject and object are formulas. Flyology_N3.Parsers.Parse reads a whole document.
Rule : constant Flyology_N3.Model.Term :=
Flyology_N3.Parsers.Parse
("@prefix : <http://example.org/> ." & ASCII.LF
& "{ ?who :wrote ?what } => { ?what :by ?who } .",
"http://example.org/");The result is the document as a formula, so walking it is the same at every level. Model.Statement_Count and Model.Statement_At walk one level, and Model.Term_Kind says what each term is.
for Index in 1 .. Flyology_N3.Model.Statement_Count (Rule) loop
declare
Fact : constant Flyology_N3.Model.Statement :=
Flyology_N3.Model.Statement_At (Rule, Index);
begin
case Flyology_N3.Model.Kind
(Flyology_N3.Model.Subject (Fact)) is
when Flyology_N3.Model.Formula_Kind => ...
when Flyology_N3.Model.Variable_Kind => ...
when Flyology_N3.Model.List_Kind => ...
when Flyology_N3.Model.RDF_Kind => ...
end case;
end;
end loop;Read SPARQL
The crate reads and writes queries back as documents through Flyology_SPARQL.Parsers.Parse and Flyology_SPARQL.Writers.To_SPARQL. Nothing evaluates them: a query here is something to check, format or inspect.
Parsed : constant Flyology_SPARQL.Syntax.Query :=
Flyology_SPARQL.Parsers.Parse (Text);
begin
Put (Flyology_SPARQL.Writers.To_SPARQL (Parsed));A Syntax.Query is the same flat vector the RDF terms use, so a walk over its Syntax.Node_Reference values is index arithmetic over one allocation. Syntax.Where_Clause is the root of the pattern.
Where : constant Flyology_SPARQL.Syntax.Node_Reference :=
Flyology_SPARQL.Syntax.Where_Clause (Parsed);
for Index in 1 ..
Flyology_SPARQL.Syntax.Child_Count (Parsed, Where)
loop
declare
Node : constant Flyology_SPARQL.Syntax.Node_Reference :=
Flyology_SPARQL.Syntax.Child (Parsed, Where, Index);
begin
if Flyology_SPARQL.Syntax.Kind (Parsed, Node)
= Flyology_SPARQL.Syntax.Triple_Node
then
...
end if;
end;
end loop;Coverage is the whole query language: the four query forms, the pattern and expression grammars, property paths as far as sequence, alternative, inverse and the three cardinality operators, subqueries, VALUES, EXISTS, dataset clauses, aggregates and the RDF 1.2 term syntax. The parser also rejects what the grammar alone would admit: a projection that repeats a name, an AS that takes a name already in scope, a variable projected past a GROUP BY that dropped it. Deliberately absent: SPARQL Update, federated-query specifics beyond recognising SERVICE, and evaluation of any kind.