Reliability & edge state

What "Offline" Actually Means

A café's wifi drops for ninety seconds during the lunch rush. "Just queue the sale and retry later" sounds like a one-line fix. It isn't — and treating it like one is how a POS ends up losing sales.

A café's wifi drops for ninety seconds during the lunch rush. What should the till do?

"Just queue the sale and retry later" sounds like a one-line fix. It isn't, and treating it like one is how you end up with a POS that either loses sales or gets stuck the first time something weirder than a clean disconnect happens. Here's what building a real offline path for FranchiseTech's till actually involved.

The actual queue in the till UI: two unsynced sales, each with a plain-language explanation and a manual retry — not just a spinner.

navigator.onLine lies

The obvious first move is to check the browser's online/offline events and queue sales when offline. The problem: navigator.onLine mostly means "the device has a network interface up," not "the device can reach your server." A phone on a hotel wifi captive portal, or mobile data with a carrier-side routing hiccup, reports online: true and then every request to your API times out anyway.

So the queue doesn't trust that flag alone. There's a real probe behind it — a HEAD request to a static asset on the app itself, with a short timeout, cached for fifteen seconds so a flaky connection doesn't trigger a probe storm:

export async function probeServerOnline(): Promise<boolean> {
  if (!navigator.onLine) return false;           // fast path, skip the request entirely
  if (cachedRecently) return cached.online;      // don't hammer on rapid calls
  try {
    await fetch("/favicon.ico", { method: "HEAD", signal: abortAfter(5000) });
    return true;
  } catch {
    return false;
  }
}

That distinction — "my network interface is up" vs. "I can actually talk to home" — is the difference between a queue that engages exactly when it should and one that either queues sales unnecessarily on a fine connection or, worse, tries to submit a sale into the void on a connection that only looks fine.

Not every failure is a network failure

The more interesting problem showed up once the queue existed: what do you do when a checkout request throws? The naive answer — "any error means offline, queue it" — is wrong often enough to matter, because at least three genuinely different things can make that request fail:

  1. The connection is actually down or unreachable. Correct move: queue it, retry automatically when the probe says we're back.
  2. The deployed app was updated mid-shift, and the cashier's browser tab is still running a bundle that references a server action which no longer exists on the new deploy. Retrying — or queuing — doesn't fix this; the tab needs a reload to pick up the new bundle. Queuing it would just mean the reloaded page inherits a stale, doomed entry.
  3. The server genuinely rejected the sale — a discontinued product, a VAT configuration problem, something that will fail identically on the hundredth retry as it did on the first. Queuing this quietly and retrying forever hides a real problem behind a spinner that never resolves.

The classifier checks these in a specific, deliberate order — stale-bundle detection first, since reloading also incidentally fixes a transient network blip, then network-error pattern matching, and only then falls through to "this is a real failure, surface it":

export function classifySaleFailure(err: unknown): SaleFailureAction {
  if (isStaleServerActionError(err)) return "reload";
  if (isRetryableNetworkError(err)) return "queue";
  return "fail";
}

Getting this ordering right mattered more than any individual check — the wrong priority turns a two-second reload into a mysteriously stuck queue entry, or turns a real business-logic rejection into an infinite silent retry loop.

The queue has to fail loudly before it fails silently

The queue itself lives in localStorage, capped at 200 entries. When it's full, enqueueOfflineSale doesn't quietly drop the oldest unsynced sale to make room — it throws, and blocks the checkout instead:

Max 200 entries; enqueueing past that throws OFFLINE_QUEUE_FULL instead of dropping an older, unsynced sale — losing a recorded sale silently is worse than blocking checkout.

That's a real trade-off, not a default: blocking a sale is a bad moment for a cashier, but it's a visible bad moment they can escalate. Silently discarding a recorded sale to make room for a new one is invisible until someone reconciles the till at close and a transaction has simply vanished — by which point there's no way to recover it. Given the choice between an annoying failure and an undetectable one, the queue is built to fail the annoying way, on purpose.

Once a queued sale does sync, it isn't just "done" — it moves into a pending_fiscal state, because this product operates under Romania's fiscal-receipt rules and the legally required receipt can't be auto-printed for a sale that happened while offline; a human has to print it once the till's back online. And if the server comes back and permanently rejects a queued sale rather than timing out, it moves to a terminal needs_attention state and stops being retried automatically — so a bad entry doesn't sit there silently re-failing on every reconnect for the rest of the shift.

The common thread

Every one of these decisions — the real connectivity probe, the failure-order classifier, the hard cap that fails loud, the explicit needs_attention state — comes from the same place: refusing to let "it's probably fine, just retry" stand in for actually knowing what happened. A POS system that's wrong about whether a sale went through is worse than one that's slow, and I'd rather write the extra state-handling code than find that out from a business owner's till not balancing at the end of the night.

Built as CTO of FranchiseTech, a HORECA/POS platform for hospitality businesses in Romania.

← More writing