
The Infinite Archive: AI in Libraries and Information Science
A stewardship-first guide to AI for metadata, OCR, discovery, preservation access, and reference service that protects provenance and reader privacy.
Read MoreZharfAI Team

Libraries operate an information supply chain: acquire and license resources, identify them, describe them, connect names and subjects, record holdings and rights, expose them through discovery, and help a user reach an authoritative item. AI can accelerate parts of that chain, but fluent text cannot repair a missing edition statement, a broken holdings feed, or an invented citation.
This February article takes a deliberately operational angle: catalog records, authority control, linked data, discovery indexes, digitization queues, and retrieval evaluation. The newer companion, AI in libraries and information science, examines the wider stewardship questions around provenance, privacy, copyright, preservation, procurement, and community governance. The two pieces should be read together, not treated as interchangeable versions.
A “book” can mean a work, expression, edition, physical copy, licensed ebook, digitized surrogate, or record imported from another institution. A model that does not distinguish these levels will merge what users and librarians need separated.
Define the operational event before automation: create a preliminary record for a newly acquired item; reconcile an incoming name against an authority file; suggest a subject term; identify probable duplicate records; extract text from a digitized page; expand a query; or rank resources already available to the patron.
Name the responsible unit, source systems, accepted output, review threshold, and reversible action. “Improve discovery” is not measurable. “Increase successful full-text access for known-item searches without reducing subject-search coverage” is.
A bibliographic record must support identification, selection, access, inventory, sharing, and later correction. Descriptive prose is only one part. Structured elements such as title proper, variant title, creator role, edition, publication statement, extent, language, identifiers, carrier, series, notes, subjects, and holdings interact with catalog rules and local policy.
The Library of Congress standards portal maintains specifications and resources including MARC, BIBFRAME, authority-description, and digital-library standards. These standards make exchange possible; they do not guarantee that a record is correct or culturally neutral.
AI-generated fields should carry source, model version, time, confidence or review status, and the evidence span when available. Keep transcribed information separate from normalized access points and inferred subjects. A model should not silently “correct” an unusual title page, historical spelling, or creator statement.
Names vary across scripts, transliteration systems, pseudonyms, married names, initials, and languages. Organizations change names; conferences recur; geographic boundaries change. String similarity can propose a match, but identity resolution needs dates, affiliations, titles, language, identifiers, relationships, and negative evidence.
A safe matcher returns candidates and explains which fields support or contradict each one. It should abstain when two people remain plausible. Merging distinct creators contaminates every downstream search, citation, and knowledge graph; splitting one creator reduces recall but is usually easier to repair.
Evaluate authority work on difficult cases, not only popular authors with rich records. Include non-Latin scripts, sparse local creators, historical organizations, homonyms, and name changes. Librarians should be able to reject a match and preserve the rationale so the same error is not proposed repeatedly.
In December 2025, OCLC announced AI-generated suggestions for Dewey, Library of Congress Classification, and Library of Congress Subject Headings in its cataloging tools. The described workflow keeps catalogers in control: suggestions can be accepted or ignored.
That operational pattern is more defensible than autonomous assignment. Classification and subject analysis depend on the whole resource, the edition, collection policy, vocabulary version, and community context. Historical headings can contain bias; a model trained on shared records may reproduce both good practice and legacy harm.
Measure top-k suggestion recall, exact and near acceptance, editing time, inappropriate-term rate, and performance by language, format, subject, and collection. A saved minute is not a success if the resulting term makes a community less discoverable or mischaracterizes a work.
MARC remains central to the representation and exchange of bibliographic data, while BIBFRAME is a Library of Congress initiative for bibliographic description in a linked-data environment. Transition is not a one-click format conversion.
Mappings can lose local fields, punctuation logic, notes, relationships, identifiers, script pairing, and holdings context. Round-trip tests should compare what survives from MARC to BIBFRAME and back, and which semantics become more explicit or less recoverable.
AI can propose mappings or enrich links, but deterministic crosswalks and validation rules should handle stable structure. Keep source records and conversion versions. Publish a reconciliation queue for unmapped values, ambiguous roles, and broken URIs rather than forcing every record into a confident graph.
A search result is useful only if the patron can obtain the item. Knowledge bases, link resolvers, authentication, license dates, coverage ranges, local holdings, and platform changes determine whether “available online” is true.
The NISO Open Discovery Initiative recommended practice addresses transparency in index-based discovery, including content coverage, metadata, source identity, fair linking, usage, and responsibilities among providers and libraries. It is a consensus recommended practice, not a legal mandate or formal product audit.
AI can detect suspicious coverage gaps, inconsistent dates, or unusual link-failure patterns. It should not fabricate entitlement or override a license record. Evaluate click-to-access success, false availability, false absence, resolver failures, and time to repair by provider and collection.
Embeddings and language models can connect a user’s concept with records that use different words. Hybrid retrieval combining keyword, fielded search, controlled vocabulary, identifiers, and semantic similarity is often safer than replacing established search.
The interface should show why an item matched: title or subject term, full-text passage, related name, cited work, or semantic expansion. Users need filters for date, language, format, collection, availability, and source. Known-item queries, identifiers, quotations, and exact names should not be degraded in pursuit of conversational search.
A generated answer must cite records that exist and distinguish metadata from full text. If only an abstract or catalog note is available, the system should not imply it read the item. For enterprise parallels, see AI for knowledge management and search.
One relevance score does not describe library use. Build test sets for known-item, topical, exploratory, citation, local-history, multilingual, course-reserve, accessible-format, and “no result should be returned” queries.
Judge precision, recall, normalized rank, source diversity, availability accuracy, citation validity, and task completion. Include users with screen readers, right-to-left scripts, transliteration needs, low bandwidth, and limited database experience. Measure whether a user reaches and understands a source, not whether they remain in a chat interface.
Watch for popularity feedback loops. Click data can amplify heavily promoted or English-language items while burying rare, local, or newly digitized collections. Use editorial and collection-based sampling to ensure the long tail remains discoverable.
Optical character recognition enables full-text search across scanned pages, but accuracy varies with script, typeface, layout, bleed-through, marginalia, paper damage, tables, and image quality. A clean paragraph generated from a damaged page may be a hallucinated repair.
Keep the preservation master, access image, raw OCR, corrected text, layout coordinates, language and model version as distinct objects. Link text to page regions so a user can inspect the image. Mark machine-generated or corrected text and retain previous versions.
Evaluate character and word error rate on a stratified sample, plus named entities, dates, page numbers, and search success. Averages can hide failure in newspapers, Persian or Arabic script, Fraktur, handwriting, or mixed-language material. AI for archival documents covers deeper archival-document workflows.
Models can rank material for scanning by demand, condition signals, uniqueness, format risk, rights status, or catalog completeness. That rank is a planning aid, not an automatic statement of cultural value.
Digitization queues should include preservation risk, community significance, teaching and research need, donor restrictions, privacy, copyright, accessibility, equipment, metadata capacity, and long-term storage. Collections that lack usage data should not be treated as low value; historical exclusion often produces low visibility.
For condition assessment, computer vision can flag tears, stains, mold-like patterns, warped bindings, or color change. A flag is not a conservation diagnosis. Staff should inspect the item, apply handling protocols, and involve a conservator where risk is material.
Translation, transliteration, query expansion, and multilingual embeddings can help patrons cross language boundaries. They can also merge distinct names, erase register, mistranslate a subject, or privilege a dominant-language description over an original one.
Preserve original-script metadata and show the transliteration scheme. Store translated fields as derived values with language, model, and review status. Search should let the patron see and correct the expansion, and results should not depend on creating a personal profile.
Evaluate separately by language and script, including zero-result reduction, inappropriate expansion, and retrieval of local-language resources. Use librarians and community reviewers for consequential vocabularies rather than relying only on back-translation metrics.
A discovery assistant can explain catalog syntax, suggest search terms, locate a database, or assemble citations from verified records. It should route complex research interviews, rare-material questions, sensitive topics, inaccessible resources, and unresolved uncertainty to a librarian.
Ground answers in the library’s approved systems and expose citations. Block invented call numbers, holdings, URLs, quotations, or access conditions. A generated citation must be checked against the source record and, for scholarly use, against the item when possible.
The IFLA Statement on Libraries and Artificial Intelligence provides international professional principles concerning AI and libraries. It is guidance, not national law, and implementation still depends on institutional mission and local rights.
Queries, borrowing, reading history, annotations, accessibility requests, and reference transcripts can reveal sensitive interests. Collection and license agreements may also restrict full-text indexing, retention, model training, or transmission to vendors.
Use the minimum patron data, separate identity from operational logs, set short retention where feasible, control staff and vendor access, and provide non-personalized search. Do not use licensed or restricted content to train a general model without authority.
For every indexed collection, record legal or contractual basis, permitted processing, display limits, snippets, user population, retention, deletion, and incident process. Rights are collection- and jurisdiction-specific; the model should enforce a current policy table, not infer permission from web accessibility.
A robust sequence is:
Bulk writes need sampling and thresholds by field. A model might safely normalize a language code under a rule yet require review for a creator merge or subject heading. Never let a vendor update overwrite local notes or community-supplied description without a recoverable comparison.
For cataloging, track accepted suggestions, edits, reversals, time, and error by field and collection. For identity resolution, track false merges and splits. For discovery, measure task success, coverage, availability, citation validity, language performance, accessibility, and long-tail exposure.
Monitor index freshness, holdings latency, broken links, OCR coverage, out-of-domain queries, librarian escalations, privacy incidents, complaints, and correction time. Re-test after model, vocabulary, crosswalk, index, vendor, or interface changes.
The release record should include input snapshot, standards and vocabulary versions, model and prompt, validation set, field-level permissions, rights constraints, approvers, index version, rollback, and review date. Useful library AI makes the path from query to evidence clearer; it does not replace that path with a persuasive answer.
Sources reviewed and status checked on 2026-07-30:

A stewardship-first guide to AI for metadata, OCR, discovery, preservation access, and reference service that protects provenance and reader privacy.
Read More
How enterprises can build permission-aware AI search with governed sources, provenance, measurable retrieval quality, controlled answers, and useful feedback.
Read More
Meeting AI is moving from transcripts to multimodal memory that understands slides, decisions, action items, sentiment, and follow-through.
Read MoreSee the daily briefing and the operational guides. This page is an archive note, not an invitation to start a project.