Letting a model write the words, never the numbers

milestone narrative, llm, media

Procedural text has a ceiling and you hit it fast. Swap a few adjectives, plug in a scoreline, and by the fourth opposition briefing the manager reading it can see the template through the prose. It stops being atmosphere and becomes Mad Libs.

So mid-July went on words. There's a narrative queue now, with a composer behind it, generating opposition briefings, match reports, news articles, richer staff reports, post-match press wording, and the weekly and end-of-season copy that turns up in your inbox. Admin has Narrative and LLM settings plus metrics, so I can actually see whether the model path is live or whether everything's coming out of the fallback.

The rule I set and will not move on: the model writes the words, never the numbers. Match outcomes, finances, player development — all of it is decided before any prose exists, and the generated text describes what already happened. Nobody is going to get a better result because they got a better sentence. Wit is not something you buy in this game.

The other rule is that procedural generation stays a first-class path, not a panic button. If a key is missing, if a call times out, if the provider is having a day, you get the plain paragraph. An empty press conference is far worse than a boring one.

Then, a few days later, I threw out how I was running it.

The first version ran a model in my own cluster. Which, on the hardware I've got, was a lovely idea for about a week. Memory trimming, model pulls, all of it elbowing the game for room on a machine that also has a season to simulate. Elegant on a diagram, embarrassing in practice.

So the narrative path is an OpenAI-compatible call to a hosted provider now — a DeepSeek model by default, with base URL, key, and model overridable from admin if I want to try something else. That let me raise queue concurrency and drop the off-peak gating I'd added purely to stop inference from fighting the tick. The local-model plumbing came out of Compose and the cluster config entirely. No point keeping a second way to do it that I've already decided against.

It's not free and it's not mine. That's the trade: I've swapped a resource problem I couldn't win for a dependency and a bill. Given that the fallback is always there, I'll take it.

Mildly annoying leftover: old local model names are still sitting in admin state from the earlier setup, so the code has to know to ignore names it no longer believes in. Every migration leaves something like that behind. 😄

← All entries