After DarkKranti
BarrowspireA 2D multiplayer game and the Go microservices that keep it honest.

Thirty times a second, the whole world

How Barrowspire's game-service turns a keypress into a velocity and a tick into one full snapshot per player, and the three places its goroutines share memory without agreeing who owns it.

Waxing gibbous, 97% lit · waxing gibbousBy Kranti · 8 min readgo · concurrency · websocket · ecs · gamedev

Every world in Barrowspire, the hub and each escape run, is a handful of goroutines around a time.Ticker. Thirty times a second the ticker fires. The world moves everything, then serializes all of itself and sends each player their own copy. No deltas, no acks, no interest management. The root CLAUDE.md says the loop runs at 60 ticks a second. GameFrameRate says 30, and the constant is the one that runs.

This week I traced one keypress in and one frame out, and ran the race detector on a copy of the service while a hub ticked and players came and went. It found two problems in about a second. Reading the serializer turned up a third. None has bitten a player that I know of. I think that's luck.

The path, goroutine by goroutine

Three goroutines on the way in, three on the way out. Every hop, with its buffer:

Hop Goroutines Hands off through Buffer
ServeConnectedPlayer, read pump 1 per connection serverChan 10
messageHub.Run 1 per process session.MessageCh 100
manageClientMessages 1 per world writes components directly n/a
manageGameLoop 1 per world one goroutine per player per tick n/a
PushMessageToChannelQueue caller's msgChan[conn] 10
writer in setupClientWriter 1 per connection conn.WriteJSON n/a

Each world also runs manageEliminations and manageEndSession, so four long-lived goroutines per world. ADR-0015 caps the process at 50 concurrent players, and matchmaking pairs them two at a time (NewQueueService(2)). Worst case is 25 runs plus the hub, 26 ECS worlds at 30 Hz. The hub holds 40 at most (HubOccupancyCap).

The message hub is one goroutine for the whole process. It routes each message by the server's own record of where the player is, never by the payload1, into that world's MessageCh. It's also a single point of stall. If one world's 100 slots ever fill, the hub blocks on the send and every connection in the process queues behind it.

Input is latched, not queued

game-service/internal/game/session.gogo
func (s *Session) handleMove(playerID uuid.UUID, vx, vy float64) error {
	s.mu.RLock()
	// get specific player entity
	playerEntityID, ok := s.playerIDToEntitiesID[playerID]
	s.mu.RUnlock()
 
	if !ok {
		return fmt.Errorf("PlayerEntityID doesn't exist for playerID: %s", playerID)
	}
 
	playerEntity, ok := s.EntityManager.GetEntity(playerEntityID)
	if !ok {
		return fmt.Errorf("Player entity doesn't exist for id %s", playerID)
	}
 
	playerVelocityComponent, ok := playerEntity.GetComponent(ecs.ComponentTypeVelocity)
	if !ok {
		return fmt.Errorf("Players Velocity Component doesn't exist for entity ID: %s", playerEntity.ID)
	}
 
	component := playerVelocityComponent.(*components.VelocityComponent)
 
	component.VX = vx
	component.VY = vy
 
	return nil
}

That's the whole of a move once it reaches a world. The client sends a direction, the server writes it onto the velocity component, and nothing is queued. Whatever VX and VY hold when the ticker fires is what the tick uses, so five moves inside one 33 ms window collapse into the last. Position stays the server's job (VX * Speed * deltaTime, speed 200), so a client can't claim to be somewhere it isn't.

The catch is that handleMove runs on manageClientMessages while MovementSystem and the serializer read the same fields on manageGameLoop. The ECS has locks, just not here. EntityManager has an RWMutex over its entity map, and every Entity has one over its component map. Neither covers the fields inside a component. GetComponent hands back a pointer under a read lock, releases the lock, and after that the pointer is anyone's.

text
WARNING: DATA RACE
Read at 0x00c000132780 by goroutine 10:
  .../internal/serializer.(*StateSerializer).SerializeBackendState()
  .../internal/game.(*Session).broadcastFullState()
      internal/game/session.go:899
  .../internal/game.(*Session).manageGameLoop()
      internal/game/session.go:507
Previous write at 0x00c000132780 by goroutine 9:
  .../internal/game.(*Session).handleMove()
      internal/game/session.go:957
  .../internal/game.(*Session).manageClientMessages()
      internal/game/session.go:263

On arm64 and amd64 an aligned float64 store doesn't tear, so the likely symptom is small. A tick reads VX from one message and VY from the next, and you drift diagonally for 33 ms. It's still a race.

The second one is worse. broadcastFullState ranges over s.playerEntityIDToPlayerID without taking s.mu, while Admit writes that map under s.mu from the message hub's goroutine. In Go, a map written during iteration isn't a panic you can recover from. It's fatal error: concurrent map iteration and map write, and it takes the whole process down with all 26 worlds in it. The hub is exactly where people join mid-tick. Two weeks ago we moved the hub's guards into the world itself, because "a guard that every caller has to remember is one a caller will eventually forget." The lock is in the right place for writers now. The tick reads around it.

The tick

With the debug block and the metrics trimmed, the loop looks like this.

game-service/internal/game/session.gogo
func (s *Session) manageGameLoop() {
	ticker := time.NewTicker((1 * time.Second) / time.Duration(constants.GameFrameRate))
	defer ticker.Stop()
	for {
		select {
		case <-ticker.C:
			entities := s.EntityManager.GetAllEntities()
			deltaTime := 1.0 / float64(constants.GameFrameRate)
 
			// residents amble; a run has none, so this is the hub's alone
			if s.worldType == types.WorldTypeHub {
				wanderSys := systems.NewWanderSystem()
				wanderSys.Update(deltaTime, entities)
			}
			movementSys := systems.MovementSystem{MapWidth: s.mapWidth, MapHeight: s.mapHeight}
			movementSys.Update(deltaTime, entities)
			projectileSys := systems.NewProjectileSystem(s.EntityManager)
			projectileSys.Update(deltaTime, entities)
			interactionSys := systems.InteractionSystem{}
			interactionSys.Update(entities)
			eliminationSys := systems.EliminationSystem{}
			eliminationSys.Update(deltaTime, entities, s.ID, s.eliminationCh)
			rulesSys := systems.RulesSystem{}
			rulesSys.Update(deltaTime, entities, s.endSessionCh)
 
			err := s.broadcastFullState(entities)
			if err != nil {
				slog.Error("Error broadcasting state", "error", err)
				continue
			}
		case <-s.stopChan:
			return
		}
	}
}

deltaTime is a constant, never measured, so the step is fixed.2 I like that under load. A time.Ticker drops ticks for a slow receiver rather than queueing them. If a tick takes longer than 33.3 ms, the next one just arrives late, and the world runs in slow motion. It doesn't try to catch up and spiral. TickDuration goes to OpenTelemetry on every tick, so you can see it happen.

Fixed-step isn't the same as deterministic, though. GetAllEntities copies out of a map, so entity order changes every tick. That matters in MovementSystem's spatial hash, which stores one entity per 40 px cell (entitiesMap[key] = entity). Two movers in one cell means the last one written wins, and the other doesn't move that tick. The hub spawns everyone at the same point, (1000, 800). I put an idle delver there and walked a second one 40 px out of the shared cell six times. That should take 6 ticks. It took anywhere from 6 to 8, depending on which way Go's map iteration fell.

A whole world per player, per tick

game-service/internal/game/session.gogo
func (s *Session) broadcastFullState(entities []*ecs.Entity) error {
	ctx := context.Background()
	backendState, err := s.stateSerializer.SerializeBackendState(ctx, s.ID, s.worldType, entities)
	if err != nil {
		slog.Error("Failed to serialize state", "error", err)
		return err
	}
 
	clientStates := make(map[uuid.UUID]*types.ClientGameState)
	for _, playerID := range s.playerEntityIDToPlayerID {
		clientStates[playerID] = s.stateSerializer.FormatStateToClientState(backendState, playerID)
	}
 
	s.stateSerializer.PutBackendState(backendState)
 
	for playerID, clientState := range clientStates {
		go func(pID uuid.UUID, cState *types.ClientGameState) {
			s.sender.SendStateToPlayer(pID, cState)
		}(playerID, clientState)
	}
 
	return nil
}

Line 10 is the unlocked range from earlier. The world serializes once into a BackendGameState borrowed from a sync.Pool, then builds one ClientGameState per player (you as current_player, everyone else as other_players). It returns the pooled state and fires a goroutine per player to deliver. A full hub is 40 goroutines per tick, 1,200 a second, each living just long enough to push a pointer into a channel.

Full snapshots make everything downstream forgiving. A lost frame is replaced 33 ms later by a complete one, so nothing needs retransmitting. The client eases toward server positions rather than snapping, which is how 0d3292e fixed the hub judder. Delvers had been teleporting every 30 Hz tick on a 60 Hz screen.

I checked the aliasing with the real serializer. Build a frame, put the state back, open a door, serialize again, and look at the old frame.

text
tick N  frame, before tick N+1: door open=false
same pooled object reused: true
tick N  frame, after  tick N+1: door open=true

On a healthy connection the writer encodes the frame long before the next tick. Only a connection that has fallen behind sees it, which brings us to the last hop.

One writer per socket, and what it drops

Gorilla's websocket allows one concurrent writer per connection, so each connection gets a channel and one goroutine ranging over it into conn.WriteJSON. Anything that talks to a client goes through here.

game-service/internal/gameserver/server.gogo
func (s *Server) PushMessageToChannelQueue(playerID uuid.UUID, msg interface{}) error {
	conn, exists := s.GetConnFromPlayer(playerID)
	if !exists {
		return fmt.Errorf("player %s not found", playerID)
	}
 
	s.mu.RLock()
	ch, ok := s.msgChan[conn]
	s.mu.RUnlock()
 
	if !ok {
		return fmt.Errorf("message channel not found for player %s", playerID)
	}
 
	// non-blocking send to prevent slow clients from blocking
	select {
	case ch <- msg:
		return nil
	default:
		return fmt.Errorf("message channel full for player %s", playerID)
	}
}

The channel holds 10 messages, which is 333 ms of frames at 30 Hz. Past that, the send fails and the frame is dropped. The goroutine in broadcastFullState throws the error away, and for state frames that's the right call. The next snapshot supersedes it.

The trouble is that the lane isn't only for state. world_entered, the one message whose whole job is to tell the client it has changed worlds, goes through the same 10 slots. A client that's 10 frames behind when its run resolves drops that message the same way it drops a frame. Since 962a06b every state frame carries world_type as well, which gives the client a second chance to notice. The transition message itself still has no delivery guarantee.

What changes first

None of this needs a rewrite.

  • Drain input inside the tick. Have manageGameLoop empty MessageCh with a non-blocking loop at the top of each tick and drop the separate manageClientMessages goroutine. Then one goroutine writes components, and input really is applied per tick.
  • Read the roster under the lock. GetPlayerIDs already copies the IDs under RLock. broadcastFullState should call it.
  • Stop lending out the pool's slices. Copy them in FormatStateToClientState, or drop the pool. It saves a few allocations per tick and is the only reason frames can tear.
  • Give control messages their own lane, or at least a blocking send with a timeout, so a scene change can't lose a coin toss with a frame.

make test runs go test -v ./... -count=1, without -race. That's how all three of these got this far.

Notes

  1. FS-29KSH §Requirements 16. A client can't address a world it isn't in, and routing errors don't echo a session id back. ↩

  2. There's a test that drives the real ticker, and it fails. TestSession_GameLoopAppliesMovement_Integration sets a velocity, sleeps 1.2 s, and expects a single 0.81 px step. It gets 189 px, about 35 ticks' worth. A tick you can call by hand would make that testable without a clock. ↩