Many Roads
Use case
Editorial photograph for Archive

Archive

Auction houses, museums, libraries, and historical collections sit on decades of records: past sale catalogs, accession ledgers, photographs, correspondence. The collection is a real asset; what's been hard is making it usable. Scanning is the easy part. Turning scans into something queryable, structured, and worth publishing has historically taken more cataloguer-years than the institution had.

Engagement
Digitization that makes the archive a public asset

We build the digitization pipeline that takes physical archive materials (past auction catalogs, accession records, photographs, paper provenance) and turns them into a structured, searchable, public-facing catalog. AI does the work that used to gate the project: cleaning up OCR on imperfect scans, extracting metadata from semi-structured text, identifying items across photographs, suggesting categorization. Specialists review where the AI's confidence is low; the rest ships through. The archive that used to live in boxes becomes something a researcher, a buyer, or a member of the public can actually use.

What we can build here

Capabilities

OCR cleanup on imperfect scans

Old paper, faded ink, tight typography, water damage: the documents archives actually hold are nothing like the clean PDFs OCR was built for. AI cleans up the OCR output, fills in the gaps from context, and flags the spans where the cleanup is uncertain. What comes out is text a researcher can search.

Metadata extraction from semi-structured records

Auction catalogs, accession ledgers, exhibition records: semi-structured documents where the structure is implicit, never machine-readable. AI reads each entry, extracts the structured fields (item, date, lot number, buyer, price, provenance), and writes them into a queryable record. The historian’s question (“every Italian Renaissance bronze sold at this house between 1900 and 1940”) becomes a query, not a year-long research project.

Identifying items across photographs

An archive of unlabeled photographs is mostly a mystery. AI can match items across photos, identify recurring subjects, link a depicted object to a separate written record. The archive becomes a network instead of a pile.

Discovery surfaces a public can actually use

A digital archive that nobody can navigate isn't an asset. We build the search-and-browse layer on top (semantic search, faceted filtering, item-detail pages, related-item suggestions) so researchers, buyers, and the public find what's there. The institution's archive becomes a real digital presence, not just a directory of scanned PDFs.

Questions

FAQ

How accurate is the AI extraction?
Accurate enough to publish where confidence is high; flagged for human review where it isn't. Each extracted field carries a confidence score; cataloguers review the low-confidence ones. Accuracy on published items lands well above the bar institutions hold themselves to.
What about copyright and access restrictions on the underlying materials?
We design the access model with the institution: what's public, what's behind a researcher login, what's available only on request. The AI processing happens within the same access perimeter; nothing leaks to a wider tier than the source allows.
Do we have to scan everything before AI can help?
No. We can start with what's already digitized and work backwards into the physical archive over time. Many engagements begin with the most-requested materials (past sale catalogs, accession records) and expand. The pipeline treats new scans the same as old.
Also relevant in