01 / The model
A language engine.
A reading companion.
A large language model, or LLM, is a neural network trained to learn patterns in language. Given a prompt, it generates text a small unit—a token—at a time.
What it can do
Explain a difficult passage, compare ideas across novels, suggest reading paths, or help invent an original society. Training adjusts numerical parameters called weights; using the trained model to answer is called inference.
For science fiction, a service might explore how different books imagine first contact, machine consciousness, generation ships, or the politics of a distant colony.
What it cannot guarantee
Fluent language is not proof. A model can invent a quotation, misattribute a story, confuse editions, or blend fictional science with established science. It may also reproduce text it encountered during training.
A useful literary service separates evidence from interpretation, cites the edition it used, labels speculation, and says when its sources do not support an answer.
Design for discovery. Start with questions such as "How do these three authorized stories portray isolation?" Give readers source-backed analysis and paths to the books, with spoiler controls where useful.
02 / The books
"Someone gave me thousands
of old science-fiction ebooks."
Can those simply be added to an LLM?
Technically, they can be processed.
That does not establish the right to use them.
Do not bulk-upload an unverified collection. Possessing a copy does not itself transfer copyright. Establish the source, the rights status, and the permitted uses of each work and edition before ingestion. A donor's assurance is evidence to investigate, not a substitute for rights clearance. [1: Ownership of copies]
A / VERIFIED PUBLIC DOMAINCheck the work and territory.
Copyright may have expired, but "old," "out of print," and "free online" are not proof. U.S. terms depend on factors including creation, publication, and authorship; older works can require historical research. [2]
A new translation, introduction, or illustration may remain protected even if the underlying story is public domain. Check the precise edition. [3]
B / LICENSED OR PERMISSIONEDMatch permission to the use.
These works usually remain copyrighted. Obtain permission from the current owner or authorized representative; an author, estate, or publisher may control the needed rights. The name in an old copyright notice may no longer be the owner. [4]
As a practical contract checklist, specify storage, indexing, model-provider processing, training, excerpts, commercial use, territories, duration, and deletion obligations.
C / COPYRIGHTED OR UNCERTAINHold pending review.
Copyrighted works without suitable permission—and works with unresolved status—should stay outside the production corpus until you establish an applicable license or legal exception.
U.S. fair use depends on purpose, the nature of the work, the amount used, and market effect. Creative novels and uses that substitute for reading them require particular care. No single factor resolves the question. [5]
Public domain is territorial.
A work's status may differ across countries. Assess the laws relevant to your users and processing locations; one archive's label is not worldwide clearance. Public availability on the internet is also different from public-domain status. [6: WIPO guidance]
Exceptions differ, too.
EU Directive 2019/790 has distinct text-and-data-mining provisions. Article 3 concerns qualifying research bodies and cultural heritage institutions; Article 4 covers lawfully accessible material subject to conditions, including appropriate rights reservations. Check national implementation and the particular use. These provisions are not blanket permission to display books. [7: Articles 3–4]
Recommended intake decision: approved public-domain text → catalog and ingest; suitable license → ingest within its scope; unknown or unsuitable rights → quarantine and investigate. Treat this as a risk-management workflow, not a legal determination.
02A / Source discovery
Can Archive.org supply
your science-fiction library?
Yes—as a place to discover candidate texts and bibliographic evidence. An item being available there does not automatically authorize its use in an LLM service.
Discover first. Clear each edition.
The Internet Archive is a useful starting point for finding digitized books and magazines. Its search guide explains metadata and full-text searches. Look for science-fiction titles, authors, publication dates, and specific editions.
The Archive expressly says it does not guarantee an item's copyright status or the accuracy of rights information on item and collection pages. Treat uploader-supplied labels as leads to verify. Internet Archive: Rights.
Access is different from reuse.
Public-domain or suitably licensed texts may be candidates for your approved corpus. Downloadability, a borrowing option, or an "old magazine" label alone does not establish permission for indexing, provider uploads, training, or public excerpts. Apply the same rights review described above.
Do not treat a whole collection as cleared because some items are usable. For magazines, review individual stories, illustrations, and editorial contributions as well as the issue. Copyright in a contribution can be distinct from the collective work. 17 U.S.C. § 201(c).
Archive.org intake checklist: save the item URL and identifier, exact filename, edition details, uploader/source, access date, file hash, and a copy of any rights statement. Verify the statement for your territories and intended uses; record the evidence and approval. Only then ingest permitted files. Check OCR quality against the scan and preserve page references for citations. These are recommended workflow steps, not an Archive-issued license.
03 / The method
Choose how the books
inform the answers.
"Add books to an LLM" can mean several different things. For a cited literary research service, start with retrieval; change the model only when evaluation shows a need.
| Approach | What happens | Best fit | Tradeoff |
|---|
| Prompt / context | Supply a permitted passage with a question. The base model's weights do not change. | A small, one-off reading task. | Context length and input cost limit scale. Provider retention settings still matter. |
|---|
RAG Retrieval-augmented generation | Search a separate collection; send selected passages to the model when answering. | Book-grounded answers, citations, and a library that changes. | Retrieval can miss evidence. The index, stored text, and excerpts need rights controls. |
|---|
| Fine-tuning | Further train an existing model, changing weights with curated examples or domain text. | Consistent classification, response structure, or specialized task behavior. | It is not a dependable citation system. Removal of learned material is harder than deleting an index entry. |
|---|
| Training from scratch | Train new weights on a very large corpus, then evaluate and adapt the model. | A well-funded research or model-development program. | High data, computing, and expertise needs. Thousands of novels alone are not a practical general-purpose foundation-model plan. |
|---|
Technical background: Microsoft's RAG and fine-tuning guide and AWS's comparison. The recommended starting approach here is an engineering judgment.
RAG is not a copyright shortcut.
It usually avoids changing the answering model's weights, but a pipeline still copies, stores, retrieves, and sends text. A numerical embedding does not erase the rights questions surrounding those steps. Assess ingestion and outputs separately. The U.S. Copyright Office discusses RAG and training in its AI report. [8]
Training is not automatically fair use.
The Copyright Office's Part 3 report offers a fact-dependent analysis of AI training, not a universal exemption. Its website currently identifies Part 3 as a May 2025 pre-publication report; it is policy analysis, not a binding court ruling. Obtain current legal review before relying on an exception. [8]
03A / Available services
Tools for imagining
your next world.
These available fiction-oriented AI services can support science-fiction projects. They are not exclusively science-fiction models, and their listed features do not establish a licensed, comprehensive SF research corpus. Checked September 16, 2026.
WRITING & REVISIONA fiction-writing assistant for developing and drafting stories. Its genre settings explicitly include space opera and science fiction, helping steer the tone and conventions of generated outlines and prose.
Possible SF use: develop an original first-contact outline, then revise scenes and character interactions.
Official genre documentation ↗
STORY & WORLD CONTEXTAn AI storytelling service with text generation and a Lorebook. Lorebook entries can supply context about characters, places, and concepts when activation keys match.
Possible SF use: keep notes on invented factions, planets, and technologies available while drafting. This helps provide context; it does not guarantee consistency or factual accuracy.
Official Lorebook documentation ↗
INTERACTIVE FICTIONAn AI text-adventure service. Custom scenarios provide starting prompts, Plot Essentials, and Story Cards to shape an interactive adventure.
Possible SF use: create an original starship expedition and explore how the story responds to player choices. It is a storytelling experience, not a scholarly source.
Official scenario documentation ↗
Choose by purpose: drafting and revision → Sudowrite; story generation with world notes → NovelAI; interactive adventures → AI Dungeon. These are suggested applications of documented features, not comparative test results. For cited questions about a cleared book collection, use the RAG architecture in this guide. Before subscribing or uploading manuscripts, review current plans, privacy, retention, and content-use terms.
04 / The build
From cleared collection
to working service.
A practical starting architecture: an existing LLM, a curated library, a rights-aware search layer, and answers that can be checked against their sources.
01 / INGESTCleared ebooksRights record → clean text → versioned passages
02 / RETRIEVESearch & permissionsUser question → allowed collection → ranked evidence
03 / GENERATELLM + evidenceQuestion + passages + answer policy → draft response
04 / VERIFYChecked answerCitation and excerpt checks → reader + audit record
Define a narrow first service.
Choose thematic comparison, reading recommendations, or cited Q&A. Start with a small reviewed collection—for example, 50 cleared titles—and a test set of real reader questions. Define spoiler handling and how the system separates fictional claims from scientific facts.
Clear and inventory the books.
Record provenance and rights evidence before uploading files to an external provider. Verify who supplied them and how they were obtained. Review each edition and license, including commercial restrictions, attribution, territorial limits, and whether the intended processing is allowed.
Prepare versioned text.
Parse EPUB or other formats, correct OCR errors, remove duplicate editions, and separate stories from covers and editorial matter. Keep chapter and passage locations. Split approved text into coherent chunks; a few hundred tokens is a starting experiment, not a fixed rule.
Build storage and search.
Keep approved originals in private object storage and rights records in a database. Create keyword and embedding indexes linked to stable passage IDs. Use a vector-capable database or managed search service. Carry territory, expiry, and allowed-use metadata into every retrieval record.
Assemble the answer pipeline.
A backend authenticates requests where needed, applies rights and user-access filters before retrieval, combines keyword and semantic search, and reranks candidates. Send a limited evidence set to an API-hosted or self-hosted LLM. Treat retrieved text as untrusted data, not executable instructions.
Ground and check the response.
Require source IDs for substantive book claims; resolve them to real titles, authors, editions, and passage locations. Check that cited text supports the answer. If evidence is missing, ask for clarification or abstain. Apply quotation, similarity, and extraction controls before showing the result.
Evaluate before release.
Test retrieval accuracy, citation support, invented facts, spoiler behavior, latency, and cost per answer. Include attempts to extract chapters, bypass territory filters, or reconstruct books over multiple requests. Use held-out questions and human review by readers familiar with the collection.
Operate a reversible library.
Keep API secrets on the server. Review provider data-use, retention, training, and subcontractor terms. Set budgets and rate limits. Offer a rights-complaint channel, suspend disputed content promptly, and propagate removals through text, indexes, caches, and backups under a documented retention policy.
The minimum catalog record / one per edition
- Identity
- Work ID, edition ID, title, author, translator, publisher, language, publication date and place, chapter locators.
- Provenance
- Source URL or supplier, acquisition date, acquisition basis, original file hash, and preserved evidence.
- Rights decision
- Status by jurisdiction, rights holder, public-domain reasoning or license reference, reviewer, review date, and uncertainty.
- Use controls
- Allowed indexing / provider processing / training / display; attribution text; excerpt policy; commercial scope; territories; expiry.
- Lineage
- Text version, parser and embedding model versions, chunk IDs, index version, training-run IDs if applicable, and removal history.
Keep the audit trail connected. Link each answer to its retrieved passage IDs, corpus version, model version, policy version, and checks. Record approvals and changes. Minimize personal data in logs and restrict access; an audit trail should explain a decision without retaining everything forever.
05 / The guardrails
Make respect for the text
part of the system.
The following are practical risk-reduction measures. They do not guarantee legality or replace clearance of the underlying uses.
Quotations: no magic word count
There is no universally safe number of words or percentage under U.S. fair use. Even a short passage can matter qualitatively. Attribution is useful but does not replace permission or an applicable exception. Evaluate the purpose and context; follow license limits where relevant. [9: Fair-use FAQ]
As a product policy, quote only what supports the analysis, distinguish quotation from paraphrase, and link to the lawful source. An internal excerpt cap is a control—not a legal safe harbor. Closely paraphrasing a book is not a dependable workaround.
Prevent the service from reconstructing books
Decline requests for full chapters, large verbatim passages, and "continue from this line" when the necessary rights are absent. Detect repeated requests that cumulatively extract the same work. Combine overlap checks against the corpus with rate limits, per-work excerpt accounting, and human review of flagged answers.
Check outputs from the base model as well as retrieved passages. Prompt instructions alone can fail; similarity filters can miss paraphrases or block legitimate content. Test and refine both.
Make licenses operational
Store the exact license text and version. Translate restrictions into retrieval and display rules. A reading or download license should not be assumed to authorize AI training, and permission for RAG should not be assumed to cover fine-tuning. Request written clarification for ambiguous uses.
Keep rights-review records with contracts and correspondence. Verify the permission giver's authority and obtain the scope you need. [4: Obtaining permission] These contract and recordkeeping steps are recommended practice.
Plan for expiration, disputes, and model changes
Disable affected retrieval records when rights expire or a credible issue arises. Track copies and derived indexes so deletion can be verified. For fine-tuned weights, deleting the original training file does not undo training; plan for retraining or retiring affected model versions when necessary.
Before using fine-tuning, demonstrate a measurable gap that prompting and RAG do not solve. Prefer purpose-built, permissioned task examples. Keep training datasets and evaluations versioned so later reviews are possible.
06 / Sources
Read the original signals.
Primary legal sources and official technical documentation. Links were checked on September 16, 2026; laws, litigation, and provider terms can change.
- 17 U.S.C. § 202 — Ownership of a copyU.S. Copyright Office · Copy ownership and copyright are distinct.
- How long does copyright protection last?U.S. Copyright Office · Duration and historical complexity.
- Circular 14 — Derivative works and compilationsU.S. Copyright Office · New material and underlying works.
- How to obtain permissionU.S. Copyright Office · Finding and contacting rights holders.
- Fair Use Index — About fair useU.S. Copyright Office · Four factors and case-specific analysis.
- Copyright: frequently asked questionsWIPO · Territorial protection and public-domain distinctions.
- Directive (EU) 2019/790, Articles 3–4EUR-Lex · Text and data mining exceptions and conditions.
- Copyright and Artificial Intelligence, Part 3U.S. Copyright Office · Official report hub; Part 3 is listed as pre-publication.
- Fair-use FAQU.S. Copyright Office · No automatic quotation allowance.
- Augment LLMs with RAG or fine-tuningMicrosoft Learn · Technical overview.
- Comparing RAG and fine-tuningAWS Prescriptive Guidance · Engineering tradeoffs.
- Build advanced RAG systemsMicrosoft Learn · Ingestion, retrieval, and evaluation patterns.
General information, not individualized legal advice. This guide does not clear any particular book collection or establish that a proposed use is lawful. For a real service, have qualified counsel assess the relevant jurisdictions, provenance, licenses, intended processing, and outputs—especially when relying on an exception rather than permission.