Publishing outward: harvesting and structured data

Four historical front pages fanned out on a dark ground, with a cyan feed line running from them to the right past panels reading “Set — from a browse facet”, “Record — one per item” and “Public page — read correctly”, under chips reading “Harvesting over OAI-PMH” and “Structured data on public pages”.

The catalogue leaves the building on the standards the sector already speaks. Records are exposed over OAI-PMH, in Dublin Core and in the Europeana Data Model, one record per item and dates in the standard form. Harvesters walk the feed page by page; the same feed opens in a browser as a readable page, so anyone can check what is actually being published without writing a client to do it.

The sets come from the way you already browse. An administrator chooses which fields become sets, and each distinct value of a chosen field is one. The sets a harvester sees are the facets your readers see, rather than a second structure maintained by hand and drifting from the first.

One mapping, two audiences. Each content type is mapped once — for search engines, for Dublin Core, for the Europeana model — and that mapping produces both the harvest response and the structured data embedded in the public page, so a partner catalogue and a search engine read the record the same way. The editors take a target type, an optional excerpt with a maximum length, fixed values such as a rights-statement address, and per-property visibility, so a field that should not leave the archive does not leave it.

And nothing is exposed until you choose it. A fresh installation publishes nothing at all: no collection is ticked, and no collection is harvested until someone ticks it.

What else goes out, and in what form, is on the platform page.

See it on your own material.

A guided demonstration on our demo server, with a few items from your collection if you wish.