The box plot in this repository is not drawn from a spreadsheet: it is drawn from a graph, and every number in it can be asked for directly. The queries below are editable and run in your browser.
There is no SPARQL endpoint and no server. The graph is fetched as a
static file and parsed in your browser with rdflib under
Pyodide, so nothing you type leaves your machine and an archived copy of
this repository stays queryable without a service being kept alive for it.
The first visit downloads roughly 10 MB of Python runtime.
Every row of the box plot, in the order the figure draws them. The interval bounds come out of OWL-Time rather than from a column: the time-span has a beginning and an end, each of which is an instant holding a numeric position on a named scale. That indirection is what lets the calendar reading and the arithmetic position disagree about negative years without either being wrong.
No lado: and no time: term appears below, only CIDOC CRM — and it still returns all 41 findspots with their dates. That is the point of the materialised crosswalk: most triplestores perform no RDFS entailment, so the inferred rdf:type triples are written out explicitly rather than left to a reasoner that may never run. Load the vocabulary alone and this query returns nothing.
The findspots whose interval begins before the year 1. Note that numericPosition and inXSDgYear differ by one: xsd:gYear counts astronomically, so 1 BC is the year 0, whereas the source database counts historically and the arithmetic must stay on an unbroken number line. Rather than silently picking one, the graph carries both and declares its own time reference system for the number line.
Findspots that are tightly dated (qInterval above 0.55) while showing almost no die repetition (qRepetition below 0.2) — each stamp occurs once or nearly once. These are exactly the sites that a single combined quality score would destroy: multiply the two and Rougham or the Keltengraben at Vindonissa drop to zero, although their dating is among the sharpest in the set. The two axes measure different things and are deliberately never merged.
Bregenz contributes four separately dated findspots; Colchester, London, Vindonissa and La Graufesenque two each. This is why the dating hangs on the findspot rather than on the site, and why the URI fragment is a hash of the findspot name taken per site: a site identifier alone could not tell the four Bregenz contexts apart, and they are dated differently.
Every site, with its ancient name and its Pleiades identifier where the source has one. The gaps are the interesting part: roughly a third of the sites carry no Pleiades link and most carry no Latin name. Absence here is recorded as absence — no placeholder is invented — so this query doubles as a worklist for the concordance.
Everything the method rests on, in one picture. Forty-one dated assemblages as bars, and the five independently dated events as vertical lines across them: two Augustan camps, the naval base at Velsen, Vesuvius, Inchtuthil. The bars that cross a line are drawn darker — those are the assemblages the model puts in circulation at a moment somebody else dated for us. The five reference assemblages must each cross their own line, because that is the criterion that fixed tau, and Pompeii's crosses by a fifth of a year. A future revision that broke the calibration would show here as a bar sliding off its line, before any number changed in the documentation.
Five events dated by something other than samian — coins, dendrochronology, a volcano, the historical record — and the assemblage each of them bounds. Every one of these intervals must contain its own terminus, because that is the criterion that fixed tau = 6 in the first place, so this query is also a self-test. margin is how much room the model had to spare at the nearer edge: Pompeii is the tightest, which is why it is the reference that binds. Velsen is marked withheld — it appears in the figures but was kept out of the criterion, because activity continued on the site and the terminus may bound the base rather than the assemblage.
The findspots whose modelled interval contains AD 79. Not a claim that anything here was destroyed by the eruption — it is the other way round: an independently dated instant is used as a probe, and these are the assemblages the model puts in circulation at that moment. The map is the point. A horizon scattered across the provinces is what a well-dated ceramic phase looks like; one that clustered would suggest the dating follows excavation history rather than the material.
The whole 41 by 5 grid, counted. Each dating ends before an independent terminus, contains it, or begins after it. Because a terminus is an instant and not an interval, these are OWL-Time's own time:before, time:inside and time:after rather than Allen relations, which hold only between proper intervals. The shape of the table is a sanity check on the corpus: a terminus near the start should have almost nothing before it, one near the end almost nothing after it.
The one number the two queries below do not show between them. Every pair of the 41 datings stands in exactly one relation on the modelled interval, so the first row is all 820 pairs. The second counts only those whose relation is unchanged when the intervals are widened to the full span of their potters' date ranges. Just under half. That ratio is the uncertainty of the relative chronology, expressed as a count rather than as a caveat — and it is the reason time:interval* carries the smaller set.
Every pair of datings stands in exactly one of Allen's thirteen relations, computed on the modelled interval [eff_start, eff_end] and published under lado:possibly*. Complete means countable: the numbers below add up to all 820 pairs of 41 findspots, so a missing relation is a missing pair rather than a silence that could mean anything. Note how lopsided the result is — two thirds of all pairs are simple precedence, which is what a corpus spread over three centuries looks like.
No lado: predicate appears below — only OWL-Time's own interval relations. It works, and it deliberately returns FEWER pairs than the query above. OWL-Time relations are sharp: there is no way to say "before, probably". So only the relations that survive a more generous reading of the bounds are asserted with them, and the rest live under lado:possibly* where the hedge is visible. This is not OWL-Time falling short. It is a vocabulary without a notion of uncertainty saying exactly as much as it can support — and the gap between the two counts is the uncertainty, made countable.
The pairs that are ordered on the modelled intervals but not once the evidence is read generously. The columns are the bounds the relation is actually computed from, and the gap between them — sorted with the narrowest first, because those are the claims that fail soonest. Read the last two columns beside it: the modelled intervals are a year apart while the potters' date ranges overlap by decades. That is the whole distinction in one line, and it is why these pairs carry lado:possiblyBefore and not time:intervalBefore.
Coordinates, interval and colour for each findspot. The point sits on the discovery site rather than the findspot, because that is where the source records it — the four Bregenz contexts share one coordinate, and a map that wants them apart has to spread them itself rather than pretend the data distinguishes them. The colour is presentation and is published as such, on the plot row alongside the other "visual only" quantities, never on the time-span.
A colour in a dataset is usually an assertion nobody can check. These can be: each axis publishes its ramp stops, its interpolation and the domain the values were stretched onto, so any hex value in the graph can be recomputed from the graph alone. Watch lado:domainBasis — an observed domain is read off this corpus, which means adding a findspot can change the colour of one that did not change at all. The chronological axis avoids red-green on purpose: late material is not worse material.
The parameters are in the graph, not only in the paper. Each dating is generated by an activity, and that activity records the plan it followed — so a consumer can ask what k, τ, t₀ and the era convention were, and which datemax values the source query excluded, without reading any prose. τ and lado:referenceLength are unrelated quantities that happened to carry the same value until τ was calibrated: τ governs how fast the interval tightens as stamps accumulate, t₀ is the fixed length against which qStart and qEnd are read. lado:calibrationBasis says how τ was obtained, and lado:calibratedAgainst names the findspots it was fitted to — try following those IRIs. The excluded datemax values are an inherited filter whose justification is still open; publishing them is how that stays visible.
Vocabulary: rdf/lado_dating_extension.ttl, extending CIDOC CRM, OWL-Time and PROV-O. The bundle carries the vocabulary and a materialised CRM crosswalk, so it can be loaded into a triplestore on its own.
A quarto-live version of
this page, for reuse as teaching material, is in qmd/.