Book Metadata for LLM Training

ISBNLab is LLM-ready. Our catalog of more than 83 million book edition records is normalized, deduplicated, and linked, so AI teams can train and evaluate models on clean bibliographic data without spending months aggregating and reconciling it themselves.

Ask About Access

Factual bibliographic data, from publicly accessible sources

The catalog is primarily made up of factual bibliographic information about books, such as ISBNs, titles, authors, publishers, publication dates, formats, and page counts. It's aggregated from publicly accessible sources across the web, and we keep a record of where each edition's data came from.

We provide metadata only, never the text of books. We do not claim copyright in the underlying facts themselves. What we offer is access to them in one place, together with our own work on top of them: the classifications we apply and the way records are selected, reconciled, organized, and linked.

What LLM-ready means here

Normalized

One consistent schema for every record: cleaned titles, split subtitles, and standardized dates, formats, and languages.

Reconciled

When sources disagree, we rank and reconcile them into a single best record per edition instead of shipping duplicates.

Linked

Every record is keyed by ISBN, with ISBN-10 and ISBN-13 cross-converted, editions grouped under their work, and series in reading order.

Classified

BISAC genres and subject headings give you ready-made labels for classification and topic models.

What's in the catalog

83.7M

Edition Records

31.1M

Authors

3.8M

Publishers

2.6M

Subjects

Included fields
  • ISBN-10, ISBN-13, and EAN identifiers
  • Title and subtitle
  • Authors
  • Publisher and imprint
  • Publication date
  • Format, page count, and dimensions
  • Language
  • BISAC genres and subjects
  • Series name and position
  • Links to other editions of the same work
Not included
  • Book descriptions and flap copy
  • Excerpts and review quotes
  • Cover images
  • The text of books

These are creative works owned by publishers and authors. We don't provide them at all, on any plan.

What teams use it for

  • Entity resolution and catalog matching

    Teach models to recognize that two listings are the same book, or two editions of the same work, using real ISBN, title, and author variation.

  • Grounding and retrieval

    Give assistants a reliable bibliographic reference so answers about books, authors, and series are grounded in records instead of guesses.

  • Evaluation sets

    Build benchmarks that check whether a model gets publication facts, series order, and author attribution right.

  • Classification

    Train genre and subject classifiers against BISAC labels applied across tens of millions of editions.

What you're paying for

You're paying for access: the ability to query factual bibliographic data, aggregated from publicly accessible sources, up to your plan's limit. You aren't buying the data, and we do not claim copyright in the underlying facts themselves. Your plan covers the service we built around them: finding the records, reconciling conflicts between sources, normalizing everything into one schema, and linking editions together.

We don't provide the text of books, and we don't provide copyrighted material such as descriptions or cover images.

  • Query results may be used for model training and internal use.
  • Query results may not be republished or resold.
  • Query volume and available fields are agreed for each project.
  • Rights holders can ask us to remove records through our removal process.

How access works

Your plan sets how many queries you can make against the ISBNLab API. Training workloads usually need far more than our standard plans allow, so we size query volume to each project.

Want to try it first? The Free plan includes 1,000 API calls a day.

A server answering API queries

Training data FAQ

Yes. Training data access runs on the same API and catalog, with query volumes sized for training workloads and terms that allow model training on the results.

No. Descriptions, excerpts, review quotes, and cover images are creative works owned by publishers and authors. We don't provide them at all. Access covers factual metadata only.

No. Query results can be used for model training and internal use, but publishing or reselling them isn't permitted.

Yes. The Free plan includes 1,000 API calls a day. For training volumes, contact us with your use case.

See our DMCA and Removal Requests page, or email dmca@isbnlab.com. When we approve a removal or exclusion request, the change applies to the website and the API, including data made available for training purposes.

Tell us what you're training

Access is sized to each project. Share your use case, the fields you need, and rough query volume.

Contact Us

Related Articles from our Blog

What is an EAN and How Does it Relate to an ISBN?

If you’ve ever purchased a product online or in a store, you’ve likely come across a barcode. Behind these barcodes lies a standardized system that makes modern commerce possible. Two key codes often encountered are the EAN (European Article Number) and ISBN (International Standard Book Number). But what are they, how are they connected, and why are they so crucial? Let’s explore.

Read More