Book Metadata for LLM Training
ISBNLab is LLM-ready. Our catalog of more than 83 million book edition records is normalized, deduplicated, and linked, so AI teams can train and evaluate models on clean bibliographic data without spending months aggregating and reconciling it themselves.
Ask About AccessFactual bibliographic data, from publicly accessible sources
The catalog is primarily made up of factual bibliographic information about books, such as ISBNs, titles, authors, publishers, publication dates, formats, and page counts. It's aggregated from publicly accessible sources across the web, and we keep a record of where each edition's data came from.
We provide metadata only, never the text of books. We do not claim copyright in the underlying facts themselves. What we offer is access to them in one place, together with our own work on top of them: the classifications we apply and the way records are selected, reconciled, organized, and linked.
What LLM-ready means here
Normalized
One consistent schema for every record: cleaned titles, split subtitles, and standardized dates, formats, and languages.
Reconciled
When sources disagree, we rank and reconcile them into a single best record per edition instead of shipping duplicates.
Linked
Every record is keyed by ISBN, with ISBN-10 and ISBN-13 cross-converted, editions grouped under their work, and series in reading order.
Classified
BISAC genres and subject headings give you ready-made labels for classification and topic models.
What's in the catalog
Edition Records
Authors
Publishers
Subjects
Not included
- Book descriptions and flap copy
- Excerpts and review quotes
- Cover images
- The text of books
These are creative works owned by publishers and authors. We don't provide them at all, on any plan.
What teams use it for
-
Entity resolution and catalog matching
Teach models to recognize that two listings are the same book, or two editions of the same work, using real ISBN, title, and author variation.
-
Grounding and retrieval
Give assistants a reliable bibliographic reference so answers about books, authors, and series are grounded in records instead of guesses.
-
Evaluation sets
Build benchmarks that check whether a model gets publication facts, series order, and author attribution right.
-
Classification
Train genre and subject classifiers against BISAC labels applied across tens of millions of editions.
How access works
Your plan sets how many queries you can make against the ISBNLab API. Training workloads usually need far more than our standard plans allow, so we size query volume to each project.
Want to try it first? The Free plan includes 1,000 API calls a day.
Training data FAQ
Tell us what you're training
Access is sized to each project. Share your use case, the fields you need, and rough query volume.
Contact Us