A first benchmark
Across four runs of 216 turns each on a pinned build, with the same seeds for both models, the LoRA cut head-hopping (writing the player's words or actions) in what the player sees from 24% of replies to 6%, and kept tense consistent in 99% of replies against 46% for the base model. In the model's own output, before the app touches it, the drop is smaller: 12% to 7%.
It also found problems: replies run short (47 words on average against 125), the longest setting met the paragraph minimum only 21% of the time, it repeats phrases more, and it broke one "address the player as Captain" rule 43% of the time against 9%. The testing also caught a bug in my own app's text cleanup, since fixed. Those problems are on the list to fix.
How it was run
Scripted player, deterministic scoring, a rented GPU, full proof folder with hashes, and a methodology page that lists the failed attempts and the known limits. It is a small four-run pilot; read it as a feasibility result, not a final claim.
The full methodology and results will be published with the release.
The story of how it was built
From a Canva sketch and a chat window to a tested frontend, in the order it happened, including what went wrong. The full write-up will be published with the release.