Skip to content
All posts
Self-Healing Operations

A tractor breaks down mid-storm. Now what?

A machine goes down with properties still on its list. On a static route those downstream driveways are stranded. On a self-healing fleet the work flows back to the swarm within minutes.

By The Snowmass Team

It is 3 a.m., the snow is still falling, and a hydraulic line lets go. One of your machines is done for the night. The operator radios it in. Now you have a problem that is bigger than one broken vehicle: you have a list of properties that were supposed to get serviced by that machine, an SLA clock running on each of them, and a dispatcher who has to figure out — right now, half-awake — who picks up the slack.

How your operation answers that question is the clearest test of whether your dispatch is static or dynamic.

On a fixed route, a breakdown strands work

When each operator runs a pre-built list, a breakdown orphans everything downstream on that list. Those properties do not belong to anyone else. To cover them, the dispatcher has to manually decide which other operators can absorb which orphaned properties, radio each of them, and hope the redistribution does not blow someone else's window in the process. It is triage, by hand, during the worst possible moment to be doing math.

And it depends entirely on the dispatcher noticing quickly. If the operator does not call it in — if the machine simply goes quiet — the stranded properties can sit unassigned until someone realizes coverage has a hole in it.

On a self-healing fleet, the work flows back

A dynamic dispatcher treats a breakdown as a routine reallocation, not a crisis. There are two mechanisms, and they work together.

1. The explicit recall

When the operator reports the breakdown, a dispatcher recalls the machine. That releases every property it had claimed back into the open pool with a reason stamped on it. On the very next scoring pass, the best-placed available vehicle claims each freed property as its highest-value next stop. Nobody redraws a route. The fleet simply re-optimizes around one fewer machine.

2. The silent-vehicle safety net

The recall assumes someone reports the problem. Real storms are messier — a machine gets stuck, a phone dies, an operator forgets to call it in. So a background job runs continuously and watches for silence. If a vehicle stops reporting for roughly five minutes past its claim window, the system assumes it can no longer do the work and expires its claims automatically. The stranded properties become claimable again without any human noticing first.

The moments a fixed route fails hardest — a breakdown, a no-show, a GPS dropout — are exactly the moments a self-healing fleet is designed to absorb.

Won't two machines grab the same freed-up property?

This is the fear that makes people distrust automatic reassignment, and it is worth addressing directly. When a property frees up and two vehicles both want it, the claim is a single atomic database operation. Exactly one machine wins the row; the other instantly sees the property is taken and moves to its next-best option. There is no window where both are heading to the same driveway. Double-booking is not unlikely — it is structurally impossible.

What the dispatcher does instead

The point of all this is not to remove the dispatcher. It is to change what they spend the storm doing. Instead of manually re-carving routes over the radio every time something breaks, they watch a live board that surfaces the real problems — a silent machine still holding work, a property with no live vehicle able to reach it — and they intervene with judgment where judgment actually helps. The mechanical redistribution happens on its own.

Key takeaways

  • On a fixed route, a breakdown strands every downstream property and forces manual, mid-storm triage that depends on the dispatcher noticing fast.
  • A recall releases a broken machine's properties back to the swarm; the next best-placed vehicle claims each one automatically.
  • A background reaper expires the claims of a silent vehicle after about five minutes, so coverage self-heals even when nobody reports the problem.
  • Atomic claims make double-booking a freed property impossible, not just unlikely.

Frequently asked questions

What happens to a broken-down machine’s remaining properties?
A dispatcher recalls the machine, its claims are released, and the open properties re-enter the pool immediately. The next best-placed available vehicle claims each one on its next scoring pass — no manual re-routing over the radio.
What if nobody notices the breakdown right away?
A background job watches for silence. If a vehicle stops reporting for five minutes past its claim window, its claims expire automatically and the properties become claimable again, so coverage self-heals even before a human intervenes.
Can two machines end up fighting over the same freed-up property?
No. Every claim is a database-atomic operation — exactly one vehicle wins the row and the other simply moves to its next-best property. Double-booking a driveway is structurally impossible.

Ready before the next storm

Run a simulation scaled to your fleet, or talk to the team about the season.