
Archive
Auction houses, museums, libraries, and historical collections sit on decades of records: past sale catalogs, accession ledgers, photographs, correspondence. The collection is a real asset; what's been hard is making it usable. Scanning is the easy part. Turning scans into something queryable, structured, and worth publishing has historically taken more cataloguer-years than the institution had.
We build the digitization pipeline that takes physical archive materials (past auction catalogs, accession records, photographs, paper provenance) and turns them into a structured, searchable, public-facing catalog. AI does the work that used to gate the project: cleaning up OCR on imperfect scans, extracting metadata from semi-structured text, identifying items across photographs, suggesting categorization. Specialists review where the AI's confidence is low; the rest ships through. The archive that used to live in boxes becomes something a researcher, a buyer, or a member of the public can actually use.
Capabilities
Old paper, faded ink, tight typography, water damage: the documents archives actually hold are nothing like the clean PDFs OCR was built for. AI cleans up the OCR output, fills in the gaps from context, and flags the spans where the cleanup is uncertain. What comes out is text a researcher can search.
Auction catalogs, accession ledgers, exhibition records: semi-structured documents where the structure is implicit, never machine-readable. AI reads each entry, extracts the structured fields (item, date, lot number, buyer, price, provenance), and writes them into a queryable record. The historian’s question (“every Italian Renaissance bronze sold at this house between 1900 and 1940”) becomes a query, not a year-long research project.
An archive of unlabeled photographs is mostly a mystery. AI can match items across photos, identify recurring subjects, link a depicted object to a separate written record. The archive becomes a network instead of a pile.
A digital archive that nobody can navigate isn't an asset. We build the search-and-browse layer on top (semantic search, faceted filtering, item-detail pages, related-item suggestions) so researchers, buyers, and the public find what's there. The institution's archive becomes a real digital presence, not just a directory of scanned PDFs.