Moving the heartbeat to Go

milestone architecture, go, reliability

In a game where the server moves the date, the clock is the product. That's the whole difference between the Football Manager style — you press continue, the world waits for you — and the Trophy Manager style, where Tuesday happens whether you logged in or not. I chose the second one. Which means every manager in the world is trusting a background process they will never see.

And that process runs on a small machine that gets restarted. Deploys, power, me poking at it.

If a tick fires twice, the world starts inventing things. Two training sessions from one week. A match result published twice, so your inbox tells you about a game you already read about. Nobody reports that as a bug; they just quietly stop believing the game.

So mid-July was about moving the heartbeat into dedicated Go binaries — a worker and a scheduler — and then making restarts uneventful.

The migration itself was deliberately unexciting. The Go side speaks the exact same Redis job shapes Nest was already using, so nothing about the contract changed. Match simulation (running the engine in-process by default now, rather than shelling out), game-day phases, finance, transfers, training, scouting, narrative, inbox fan-out, cup draws, season lifecycle — all ported, with structured logging and the Go lint and modernize gates on top. Nest stayed the answer key throughout: if Go disagreed with Nest about what a job should do, Go was wrong.

The load-bearing decision was queue ownership. Both runtimes are physically capable of reading the same queue, and two consumers on one queue isn't a migration, it's a haunting. So ownership is explicit config, queue by queue. Slower than one flag. Also the only version of this where rolling back is a thought rather than an incident.

Then the reliability half, which landed in the same window and is really the point of all of it:

Startup policies, so a booting worker knows whether it's allowed to do anything yet. Atomic claims and leases on simulation and broadcast, so exactly one process owns a match while it's being played. Idempotent notifications built on shared dedupe keys, so the same event can't become two inbox items. A delivery ledger for the awkward case where something external was probably sent but we can't prove it. And enqueue that tells you the difference between new work accepted and that's already in flight, instead of cheerfully queueing a duplicate.

Put together: running it again stopped meaning doing it twice.

The cost is two languages for a while, which is genuinely tedious — every side effect has to be ported and checked rather than reasoned about. But the shape was now obvious. Nest was becoming the door. The engine room had moved.

← All entries