AI-assisted fantasy writing workspace with a chapter workflow from planning through Vellum

About this experiment: I’m exploring how AI can assist with writing a fantasy novel while also maintaining the growing body of world lore needed to keep a long story—and potentially future books—consistent. The manuscript and worldbuilding are stored as structured Markdown in GitHub, with AI used for specific planning, drafting, review and verification tasks. I’m documenting the workflow here from a technical and practical perspective, while keeping the fiction itself separate.

I have been looking at how far I can use AI to help with two connected problems: writing a fantasy novel (and potentially a series of novels), and maintaining the much larger body of world lore that grows around it. Characters, places, cultures, history, organisations and timelines need to remain consistent not just across one manuscript, but across books I may not write for years.

That experiment has gradually moved away from “give the model a prompt and ask it to write a chapter”. My current setup looks much more like a small software delivery pipeline.

The manuscript and the world canon live in GitHub. Everything is stored primarily as Markdown: chapters and production documents on the manuscript side, and structured lore entries on the worldbuilding side. Lore files carry metadata for things such as type, region, chronology and tags, making them easier to search, cross-reference and supply selectively as context. The lore expands as I write and becomes reusable canon for future novels rather than disposable notes for the current book.

The repository also contains the book plan, chapter briefs, continuity state, style guidance and production status. The AI does not decide what to do next from conversational memory: repository state does. GitHub is therefore more than version control here; it is the durable source of truth for both the manuscript and the world behind it.

This is still experimental. I am not claiming it is the right way to write a novel, or even that every stage will survive another book. But if you are an author wondering how to use ChatGPT or other AI tools for a long novel without losing continuity, voice or control of the worldbuilding, this is the system I am currently testing. It has become a useful way of thinking about where AI is strong, where it is unreliable, and where adding another model call simply adds cost and complexity.

My AI novel writing workflow at a glance

AI-assisted novel writing workflow showing planning, drafting, review, revision, verification and final handoff

The current chapter flow is roughly:

Locked plan
→ story-ready chapter brief
→ first draft
→ consolidated review
→ editorial triage
→ one bounded revision
→ fresh verification
→ periodic/triggered independent QA
→ usage checkpoint
→ paragraph-architecture pass
→ DOCX/Vellum handoff
→ production complete
→ later human prose review

That may look elaborate, but an important constraint is hidden inside it: there is normally only one substantive automated rewrite. The pipeline is designed to stop models endlessly polishing other models.

Git is the memory

GitHub repository architecture showing a novel manuscript and structured world canon used as selected AI context

One of the most useful decisions was to stop treating the chat session as the source of truth.

For an ongoing manuscript, conversational memory is too soft. Instructions evolve. Decisions made several chapters ago matter again. A character may be carrying an object introduced 30,000 words earlier. A reveal may be known to the author but deliberately unavailable to the viewpoint character.

So the repository holds the durable state. A progress file identifies the next incomplete gate. A pipeline-state file records the production position. Chapter briefs and approved previous text provide local context. Separate references hold style, continuity and information-control rules.

The practical consequence is that I can start a new session and say, in effect, “run the next gate”. The system reconstructs the task from the repository rather than requiring me to remember the correct prompt sequence.

This feels much closer to working with a build system than maintaining one heroic prompt.

Planning is a gate, not part of drafting

Before prose generation, the chapter gets a compact production brief. It defines the entry state, scene purpose, causal progression, viewpoint, continuity constraints, information that must not yet be revealed, and the state the chapter must reach by the end.

I keep planning separate because I do not want the drafting model solving structural problems opportunistically while writing prose. If the plan contains a real contradiction, the correct result is a stop for a human decision, not a plausible invention that silently becomes part of the book.

I also deliberately scope context. The writer gets what the chapter needs rather than unrestricted access to every piece of worldbuilding. More context is not automatically better context, particularly when the story depends on controlling what a viewpoint character or reader can know.

Where it seems useful, I also manually sense-check a chapter brief or completed chapter with Claude. This is not a mandatory pipeline gate and it does not automatically change the manuscript. I use it as an independent second opinion when I want to challenge an assumption, check alignment or ask a specific question from outside the main workflow.

Draft once, then review independently

The first draft is currently produced with GPT-5.6 Sol at medium reasoning effort. It writes against the chapter brief and a fiction style guide.

After that I freeze the draft and run a separate consolidated review. Earlier versions of the workflow experimented with more reviewer fan-out, but the current system uses one review covering several lenses: story and causality, character and viewpoint, canon and continuity, momentum, prose and dialogue, and natural rhythm.

The reviewer diagnoses. It does not rewrite.

This separation has become one of the core principles of the setup. A model that can immediately “fix” everything it notices has a strong tendency to turn subjective preferences into edits. Diagnosis first gives me an intermediate evidence layer.

Not every criticism deserves a fix

The editorial stage classifies review findings into five buckets:

  • BLOCKER — something that prevents the chapter progressing.
  • FIX — a verified defect worth changing.
  • WATCH — a legitimate concern better judged later by a human or in rendered pages.
  • SUBJECTIVE — a defensible preference, but not a defect.
  • INVALID — a suggestion that conflicts with higher-authority decisions or would make the text worse.

Only accepted BLOCKER and FIX findings feed the normal revision.

This sounds like administrative machinery, but it solves a very practical AI problem: models are excellent at producing confident criticism. They are less reliable at knowing whether every criticism should cause a change.

The revision is therefore bounded. Preserve unaffected material, fix the verified problems, and stop. The default is one draft plus one substantive automated revision rather than an open-ended review/rewrite loop.

A simple example of triage

Imagine a reviewer flags this: “Character A refuses to enter the city, but the chapter never explains why. Add the reason here.” The reviewer has deliberately been given only the context needed for its review rather than every planning document. The criticism therefore sounds plausible, but the chapter brief explicitly says the reason must remain hidden until a later reveal.

In that case I would classify the finding as INVALID, not FIX. The reviewer has correctly noticed missing information but has drawn the wrong editorial conclusion because a higher-authority story constraint requires the omission. The action is therefore simple: no manuscript change.

That is the point of the triage layer. A review finding is evidence to evaluate, not an instruction to obey.

The editor does not verify its own work

AI writing model roles showing GPT-5.6 Sol, GPT-6 Astra, GPT-5.6 Terra, Claude and author oversight

After revision, I run a fresh verification pass. Its job is not to find exciting new improvements. It checks whether the accepted findings were actually resolved, whether the revision introduced continuity or story regressions, and whether the chapter still fulfils its original brief.

This matters beyond fiction. “I made the requested change” and “the resulting artefact is still correct” are different questions. Software development has understood this for a long time. AI workflows benefit from separating them too.

I tried a more expensive independent model

One experiment added a genuinely independent late prose audit using GPT-6 Astra. It received a frozen clean chapter and the craft standard, but none of the previous reviews or editorial notes. The intention was to avoid anchoring it on what the main pipeline already believed was wrong.

The independent reader found some material issues the normal pipeline had missed, which justified the experiment. The problem was cost.

On one measured end-to-end chapter run, the overall pipeline exhausted a full five-hour allowance, consumed 16 percentage points of the weekly allowance and used 169 purchased credits. In other words, that experimental run did not stay inside the subscription allowance: it spilled into paid usage. I could not attribute all of that precisely to individual models, so it would be wrong to claim the independent audit alone caused the cost. But the whole workflow had become too expensive to justify running its heaviest QA on every chapter.

That experiment changed the architecture. Cross-model prose auditing is now periodic or triggered: at selected checkpoints, on explicit request, or when a serious prose concern survives normal verification.

This is probably the most useful lesson from the pipeline so far. More AI is not automatically a better AI workflow. An extra model has to earn its place.

Cold reads are deliberately context-light

I also use periodic cold reads. These are different from continuity review.

For these periodic cold reads I currently use GPT-5.6 Terra. It receives the chapter with minimal context and judges it more like published fiction: would I continue reading, do the characters remain distinct, do transitions work, is the emotional weight landing, and does the prose show cumulative signs of artificiality?

Giving a reviewer less information can be useful. A heavily contextualised agent can become very good at explaining why a chapter satisfies the specification while missing the simpler question of whether the chapter is enjoyable to read.

Some checks should be deterministic

Not everything needs another language-model judgement.

The final manuscript handoff includes a paragraph-architecture pass where wording and punctuation are frozen. Only paragraph boundaries may change. Afterwards the pipeline verifies that the non-whitespace text is identical and that the word count has not changed.

The Markdown is then normalised and converted to DOCX with Pandoc and a small Lua filter. The output is checked for text equality, paragraph counts, empty paragraphs and rendering problems. I then copy the chapter from the DOCX into Vellum. The DOCX is not another manuscript authority; its practical purpose is to preserve paragraph formatting during that handoff.

I like this boundary because it separates subjective work from mechanical work. A model may be useful for deciding whether paragraph composition reads naturally. Whether the conversion accidentally dropped text is not an aesthetic question and should not depend on a model saying “looks fine”.

Human review moved later, not away

An earlier version of the process effectively encouraged detailed human approval chapter by chapter. That became disruptive to forward progress.

The current system allows a chapter to reach a production-complete state after its automated gates, while explicitly treating that as working-manuscript authority rather than publication approval. Detailed human prose review can then be batched across several chapters or an act.

This is now a working manuscript rather than a paper design: at the time of writing, 19 chapters have reached the pipeline’s production-complete state. I do not yet have a useful percentage for how much later human prose review changes those chapters, and I would rather measure that than guess. I have started treating that as an outcome metric worth collecting as the manuscript progresses.

Automation can establish that a chapter is sufficiently coherent to support writing the next one. It cannot establish that I am happy to publish it.

Human judgement still outranks the automated reviewers. If I reject a model’s stylistic preference, the pipeline should not quietly reintroduce it three stages later.

Cost is part of the architecture

The pipeline also has an explicit usage checkpoint before finalisation. I check the live usage meters and decide whether to continue. The system is not allowed to wander deliberately into purchased-credit usage just because another gate exists.

I also use ordinary ChatGPT conversations for as much of the work as practical and reserve Work sessions and more expensive model passes for stages where they add enough value. I could automate much more of the pipeline, but at the moment part of the experiment is seeing how far I can take it within a normal ChatGPT Plus subscription rather than designing a system that assumes effectively unlimited model usage.

That constraint has an unexpected benefit: it creates a cadence. Usage limits discourage endless polishing and force me to decide which checks actually matter before spending another expensive pass.

This is less glamorous than model selection, but it has become part of the design. A production workflow has to optimise for quality per unit of time and cost, not theoretical maximum quality if every possible reviewer is run on every artefact.

What I think is working

The pipeline is still changing, but several principles have survived repeated revisions:

  • Keep durable state outside the conversation.
  • Separate planning, generation, diagnosis, editing and verification.
  • Scope context instead of assuming more context is always better.
  • Do not let every review observation mutate the manuscript.
  • Limit automatic rewrite loops.
  • Use independent models where independence has measurable value, not by default.
  • Use deterministic checks for deterministic properties.
  • Treat cost and human attention as pipeline constraints.
  • Keep final publication judgement human.

The broader experiment for me is no longer “can AI write fiction?” That question is too coarse to be useful.

The more interesting question is: what combination of state, constraints, specialised model roles, deterministic checks and human decisions can make AI useful across a manuscript without gradually losing control of the work?

My current pipeline is one answer to that question. I expect it to keep changing as the manuscript exposes the next failure mode.

How I’m Using AI to Write a Fantasy Novel: My Current Workflow
Tagged on:             

Leave a Reply

Your email address will not be published. Required fields are marked *