Personal R&D · Python · Local LLMs · since 2023
Teaching a local model the Unreal docs
I've been running language models on my own machine since 2023, fully offline. The project I'm proudest of takes the entire Unreal Engine 5.6 API documentation export, more than half a million HTML pages, and turns it into something a local model can actually answer questions from.
Replaying real entries from the evaluator log and the cleaner's output.
Why bother
Ask a general model about a specific Unreal API and it will happily make something up. Point it at the real docs and it can answer from the source, and I can check the answer. Running it locally means nothing leaves my PC, and it works with no internet at all.
The catch is that the docs aren't written for a model. Most of the export is navigation, menus, theme toggles, and stub pages with a title and nothing else.
How I got here
- 2023Running private document Q&A locally with open-source tools, first on the CPU, then a GPU build.
- 2024My own Python scripts around local Llama models, pulling text out of PDF, Word, and HTML files for ingestion.
- 2025A full pipeline for the Unreal Engine 5.6 docs, below.
The pipeline
Five Python scripts with files in between, so I can check the output of any stage before the next one runs.
- Evaluateskip the junk pages
- CleanHTML to Markdown
- Plangroup and size
- Mergebundle by topic
- Embed & testtune retrieval
One page, before and after
The real "Has Tag" page from the export, and what the cleaner made of it. Struck-out lines are the noise it removes: 25 of 42 lines.
# Has Tag
_Path: BlueprintAPI/GameplayTags/Has Tag_
Check if the tag container has the specified tag
Target is Blueprint Gameplay Tag Library
## Inputs
- **Tag Container** (Gameplay Tag Container Structure (by ref)): Container to check for the tag
- **Tag** (Gameplay Tag Structure): Tag to check for in the container
- **Exact Match** (Boolean): If true, the tag has to be exactly present, if false then TagContainer will include it's parent tags while matching
## Outputs
- **Return Value** (Boolean): True if the container has the specified tag, false if it does not
The other three steps
Evaluate
Before converting anything, a scanner reads every page and keeps or skips it: too few real words, or pure boilerplate, and it's out. Every skip is logged with a reason, which turned a guess into a plan.
Plan
Hundreds of thousands of tiny files are bad for retrieval, so pages get grouped by topic (everything under BlueprintAPI_GameplayTags is one group) and sized. The plan is a CSV, not files, so it's cheap to check and rerun.
Merge
Following the plan, in parallel. Each topic becomes one or more documents, capped at 500 KB, with every original page kept as its own section with its source path.
Embed & test
Everything runs locally: Llama 3.1 8B on the GPU for answers, bge-m3 embeddings, and a LanceDB vector store. I first shipped it in AnythingLLM in 2025, then rebuilt the index to fix what testing showed. Results below.
What happened when I tested it
The same questions, asked three ways. For each one: did retrieval hand the model the right doc page, and was the answer right? Click any cell to read the actual answer.
Run on my machine in October 2026 with the original workspace settings: Llama 3.1 8B, bge-m3 embeddings, top 8 chunks, temperature 0.3. Same six questions for every column. Answers are trimmed.
No docs: 0 of 6. Fluent, confident, and wrong every time, including made-up node names.
Merged bundles: right page 4 of 6, right answer 3 of 6. This is what I first shipped in 2025. My chat logs from then show the same failure: broad questions landed, but a small node like Has Tag got buried inside a big bundle and never made the shortlist.
One chunk per node: right page 6 of 6, right answer 4 of 6. I rebuilt the index with every node page as its own chunk, plus a keyword match on node names. The keyword match mattered: by vector search alone, Has Tag ranked #40, and with the name match it ranked #2.
The two misses are worth keeping. A broad "list every timer node" question did better with bundles, because a bundle keeps a whole category together, so granularity cuts both ways. And on Get Actor Location the right pages came back, but the model still misread them. Retrieval fixes "can't find it"; it doesn't fix "misread it". The next step is to retrieve at both levels, so specific questions get the node page and broad questions also get the category around it.
What I took from it
- Most of the work is the data. Running a model locally is the easy part. Getting half a million pages into clean, well-named, sensibly sized pieces was the actual project.
- Test with real questions, then rebuild. Six questions and a scorecard showed exactly what was broken, which beats guessing at settings.
- Measure before you convert. Evaluating every page first, and logging every decision, showed where the real content was.
- Granularity matters, both ways. One node per chunk found every specific page; bundles buried them. But bundles handled the broad "list everything" question better.
- Small, checkable steps. The same reason RetroEngine has a debug view for every stage of its CRT.