The platform

The platform underneath the intelligence

Discovery is the visible half. Underneath it, MediaINFO Digital Library is what an institution runs every day: how material is organised, described, viewed, worked with, released and published outward. Six things an archive has to do — and the platform does all six on your own infrastructure.

One record laid flat, with six layers rising off it — one for each of the six things the platform does. The record is a front page of 10 November 1919 reporting the eclipse results that confirmed Einstein’s prediction, with a 1921 portrait of Einstein printed in its third column. Organize threads a collection tree in from the left and the page docks into it, filling three field chips: collections and sub-collections, content types, and the fields your cataloguers already use. Enrich outlines the page’s four article zones, boxes the face in the portrait and matches it to a person record, and ticks a transcript under the middle column — every chip it draws flagged as derived. Display dives into the headline until the type is sharp. Interact draws a selection around the report, names it, and slides a second viewer in beside it that pans in step. Reuse peels a download preset off the record — a web copy, with resolution limits and watermarks applied per preset. Share runs a feed line out of the frame toward a partner’s mark and leaves a structured-data chip behind. The six layers then settle into one stack: Organize, Enrich, Display, Interact, Reuse, Share.

Organize

Structure that mirrors the institution, not the software.

  • Collections and sub-collections that follow your own arrangement, browsable independently of the search index.
  • Content types — named field schemas per class of material, carrying text, number, date, taxonomy and related-item fields you define.
  • Bulk import at scale: batch jobs driven by a spreadsheet, with saved blueprints and filename or folder-path variables that title a whole import at once.
  • Access scoped per collection — who may view, download and manage each one — with a pre-flight report that previews the impact before you enable it.
  • Single sign-on through your own identity provider (SAML and OIDC), and directory integration where it is deployed.
  • Auto-login for reading-room terminals: visitors from listed addresses are signed in as a chosen account — pair it with a read-only one.

Item record

The front page of a 1932 newspaper, shown as the thumbnail of a catalogue record.
Content type
Newspaper page
Collection
Newspapers → 1932 → March
Date
19 March 1932 Gregorian · Hijri
Language
English
Access
Public · download by preset
Original
Retained as received

The fields are yours: each content type carries the schema your cataloguers already use.

Enrich

Understanding added on top of the object — and marked as added.

  • Text recognition of a class of its own: printed or handwritten, in left-to-right and right-to-left scripts, from newspapers and books to catalogue cards, forms and manuscripts — reading the kinds of material archives actually hold, and handing back every word with its position, so a search is highlighted where it occurs and a paragraph can be selected off the image.
  • Latin scripts, including historical blackletter, and Arabic — printed or handwritten, in both reading directions; where a collection's script or hand is not yet covered, the recognition can be taught it and the coverage added to an existing installation.
  • Layout analysis that reads a page into its articles — headings, text blocks, bylines, captions and pictures grouped into stories — each carrying the order a reader reads it in, in either direction, so a newspaper arrives as its articles and Arabic extracts as written; the newspaper's name, issue date and issue number are read off the masthead and dateline too, in Hijri or Gregorian with the other calendar's date worked out beside it, so a bulk of scanned issues files itself; an enhanced article pass is available when enabled.
  • Speech transcribed with timings from audio and video, producing searchable captions and subtitles.
  • Faces detected and matched to person records; public-people packs, available in certain countries and regions, provide automatic name labelling, with aliases, of thousands of public figures.
  • Descriptive tags, captions and scene analysis written back to the asset record — and shown as machine-derived.
  • Automated translation of extracted text, so one catalogue can be searched and read in more than one language.
  • Video indexed by what actually changes on screen, so a recording can be found by a frame in the middle of it; a licensed video analysis adds chapters that cover the whole recording, each titled, summarised and seekable.
  • Pages edited in place — rotate, insert, reorder, delete — then the workflow re-run on only the pages that changed.
  • Processing queued in the background, with schedules that hold heavy work until off-peak hours and usage visible to administrators.

Workflow · newspaper page Queued

  1. Recognise the text on every pageFull text, searchable and selectable on the page itself
  2. Analyse the layoutEach article becomes a segment of its own, with its reading order; the newspaper's name and issue date read off the page
  3. Enrich each articleA title, a summary, the author from its byline, named entities and tags, written back and shown as machine-derived — when licensed
  4. Detect and recognise facesNamed automatically where public-people packs are available in your country and licensed; otherwise named by your own team, and those names stay private to your collection
  5. Describe and tag the picturesA caption and tags, written back and shown as machine-derived
  6. Flag duplicates and copied textEvery picture and every text checked against the archive for duplicates and copied passages — when enabled
  7. Translate the extracted textOne or more target languages, kept beside the original

Administrators assemble the steps a content type needs, and re-run them on a single page when something changes. Steps illustrative.

Display

One viewer for every kind of material in the collection.

  • Deep zoom on high-resolution images and scanned pages, served from a built-in image API, with no pre-tiled derivatives to generate first.
  • Images, video, audio, documents and 3D models open in the same viewer, with the metadata panel beside them.
  • Faceted navigation by format, date, content type, collection and any custom field your administrators mark as facetable.
  • A date timeline for ranged filtering, and keyword highlighting that shows exactly where a query matched.
  • Interface and catalogue in multiple languages, right-to-left where the material calls for it.
  • The interface language and its translations are set per installation; language packs are managed in the admin panel.
  • Dates entered, displayed, searched and filtered in both Hijri and Gregorian calendars, per user.

The whole mission — 1,406 photographs, July 1969.

  • Magazine 40 · 127 frames
  • AS11-40-5903
  • Buzz Aldrin · person record
  • Neil Armstrong, reflected

The canvas is worked with the keyboard once it is held: the arrow keys move across it, plus and minus zoom, Home returns to the whole mission, and Escape — or the Release control — lets the page go on.

The 1,406 photographs of the Apollo 11 mission, July 1969, laid out in film-magazine order on one square canvas 38 frames across. The chapter descends from the whole mission into magazine 40 — the 127 frames of the surface walk — then into the frame AS11-40-5903, Buzz Aldrin standing beside the lunar module, and finally into his visor, where Neil Armstrong and the module are reflected. Four further frames are drawn at close range: AS11-40-5875, AS11-40-5878, AS11-40-5886 and AS11-44-6552. Every photograph is a NASA original in the public domain.

Interact

Material you can work with, not only look at.

  • Two or more items side by side, each panel an independent viewer, with a sync toggle that links panning and zooming.
  • Rectangular selections drawn on the canvas — named, listed, focused, shared or downloaded — alongside the editor's annotation tools.
  • Extracted text overlaid as selectable regions: hover to read a zone, click to copy it.
  • A translation overlay that swaps a page into another language while keeping its layout intact — and never replaces the original.
  • Click-to-zoom that flies to the article, shot or page you clicked and switches the sidebar to what is known about it.
  • More like this from any result, any photograph, region or article — with a score for each match and the filters still working inside the neighbours.
  • Find duplicates inside whatever you have searched and compare them side by side before you decide: the same file imported twice needs no AI; rescans, crops and copies of a picture at the strictness you choose; and copied passages between documents, marked on both — plagiarism review of a submitted thesis, or which title ran a story first — available when enabled for an installation.
  • Keyboard navigation in the viewer, for people who are in the archive all day.
Compare
The detail under flat light: a walled town built around a great oval arena crowded with tiny figures, every engraved line reading black on white.

The same detail of an engraved map, scanned twice — under raking light, which shows the relief and condition of the sheet, and under flat light, which brings out the printed line. Drag the divider; with Sync on, the two panes pan together.

Reuse

The material leaves in the form the work needs.

  • Download presets defined by your administrators: the original file, a web-optimised preview, or a custom preset per content type.
  • Format, resolution, watermark and licensing conditions set per preset — watermarks applied to derivatives on download.
  • Saved searches that store a query and its filters together, personal to each account.
  • A page builder for editorial pages that embed viewers, collections and freeform content alongside media.
  • Composites and word clouds assembled from an item's own pages and text.
  • Citations attached to assets and managed centrally.

Download

  • OriginalArchival master, as received · staff only
  • Web copyLong edge limited · watermark applied at render time
  • Study PDFSelected pages · licence terms carried with the file

Administrators decide what may leave and under what terms; the reader only sees the presets published for that content type.

Share

Publishing outward on the standards the sector already uses.

  • Harvesting over OAI-PMH in Dublin Core and the Europeana Data Model, with one record per item.
  • Mapping schemas per content type that translate your fields into each standard, edited in the admin panel.
  • Structured data on public pages, built from the same mappings, so catalogues and search engines read your records correctly.
  • Crawler pages and a sitemap regenerated on a schedule, with a robots file that allows those and nothing else.
  • A standard image API, so partner sites and aggregators can render your images in their own viewers.
  • Nothing is exposed until you choose it — a fresh installation publishes nothing, and you tick the collections that may be harvested.

One record, three renditions

Item The Sydney Bridge opened — Wonder of the world · 19 March 1932

  • Dublin Coredc:title · dc:date · dc:creator · dc:subject · dc:rightsHarvested over OAI-PMH
  • Europeana Data Modeledm:ProvidedCHO · ore:Aggregation · edm:rightsHarvested over OAI-PMH
  • Schema.orgname · dateCreated · creator · licenseEmbedded in the public page

One mapping per content type per standard, edited in the admin panel — so a harvester and a search engine each receive the record you intended.

Integrate and migrate

Everything the platform does is reachable over its HTTP APIs — library, search, assets, processing, access and system — and an interactive reference comes with the licence.

Bulk import is designed for migrating large existing collections or automating regular ingest from external systems, so a catalogue export or a nightly delivery arrives the same way a single upload does.

Collections that arrive with their text and structure already produced — METS/ALTO packages, with or without a BagIt wrapper — come in with both kept rather than recomputed, and the repository's storage can be moved later with a dry run and a verified copy.

Outward, the same record is harvested over OAI-PMH in Dublin Core and EDM and its images served through the IIIF Image API, so aggregators and partner sites read your records and render your images in their own systems.

Formats

50+ file formats, including RAW camera formats and GLB 3D models.

Handled natively across six media categories, with technical metadata — dimensions, duration, codec, colour profile — read on ingest. For preservation-grade holdings, lossless JPEG 2000 and retention of the original source file alongside every derivative.

  • ImagesJPEG, JPEG 2000, PNG, TIFF, BMP, GIF, WebP, HEIC, AVIF, JPEG XL, SVG — and RAW camera formats
  • VideoMP4, MOV, AVI, MKV, WebM, WMV, MPEG, MXF, transport streams and more
  • AudioWAV, FLAC, MP3, AAC, OPUS, AIFF, OGG and more
  • DocumentsPDF, DOCX, PPTX, XLSX and the OpenDocument and iWork families
  • 3D modelsGLB
  • PackagesBagIt archives, and OCR interchange in ALTO and METS/ALTO

Deployment

Runs where your archive lives.

On your own infrastructure, on your own terms — including with no way out to the network at all. Storage on a local filesystem, on object storage, or on a network share, with more than one backend active at once.

  • Runs on your infrastructure

    The platform is installed on your own servers, in your own data centre or private cloud. Your material, your metadata and your users stay inside your perimeter.

    • A guided installation onto hardware you control
    • Roles, permissions and audit under your administrators
    • Standards out: harvesting over OAI-PMH in Dublin Core and Europeana metadata, standard image delivery, structured data on public pages
  • Air-gapped when it must be

    An installation can be run with no outbound connection at all. The resident agent then never contacts the vendor; releases arrive as signed offline bundles.

    • No outbound check-in, by configuration
    • Upgrades delivered as offline bundles you carry in
    • AI processing can be confined to a segment with no route out, which your own network monitoring can verify
  • Upgrades in place, rolls back on its own

    Updates are offered, never applied unattended: a release is applied when an administrator asks for it, health-checked, and reverted automatically if it does not prove itself.

    • Every stage journalled, from snapshot to commit
    • A rollback restores the containers and the configuration
    • It does not restore your data — the database, the repository storage and the search index are never reverted by an upgrade, so take your own backup first

Security, deployment and data →

See the platform on your own material.

A guided demonstration on our demo server, with a few items from your collection if you wish.