The Simulator

It does not model FrogNet. It runs FrogNet.

Simulate the system and you have two pieces of code written by the same person on the same afternoon holding the same misconception. They will agree. So FrogSim simulates the environment — the machine, the wire, the store, the origin — and runs the shipping code inside it.

§01 · The seams

Where the substitution happens, and why there.

One rule placed every boundary: cut where FrogNet already speaks to something outside itself. Not one layer up, where it is convenient. Not one layer down, where it disappears.

simulated atruns for realwhy the seam is there
the machineadvertise(blob=…)capability probe output, per nodeThe parameter already existed for exactly this. Hardware facts go in; the election scores them.
the tuple storeHTTP, at api.phpfrognet_tuples, frognet_role_elect, every role handler, hosts.pyfrognet_tuples does not speak SQL, it speaks HTTP. So the seam is an HTTP server honouring that contract, and everything behind it runs untouched over real sockets.
the wirenetwork namespace per nodediscovery merge, promotion, route installation; the proxy; the daemonEach node gets its own routing table and its own /etc/hosts via a private mount namespace. Without the mount namespace every node resolves databasehost to the same address and the whole point is lost.
the originbehind the semantic proxythe proxy and daemon, unmodifiedThe proxy has never cared what serves its upstream.
node identityone function, rebound per node—The only reimplementation in the harness. Identity derives from the machine’s interfaces, which is right on a node and wrong when fifty share a container. Left alone, fifty nodes write one capability row, the store folds them into one, and the election looks like it is working.
§02 · A run

Forty-eight nodes. Three access points, fifteen networks each, a full tunnel mesh.

convergence
314 merge events, quiesced, 6.5 per node
idealized sweep
4 cycles — printed alongside, because a sweep always terminates and always looks tidy. A wide gap between the two is the finding. Printing only the tidy one is how you never see it.
reachability
0 of 2,256 ordered pairs unreachable
election
one database host, agreed by all 48 independently
telemetry
144 readings, verified from a separate process
coverage
17 components, from 2,282 captured log lines

Shape costs, and the cost is legible:

star 198 merge events
snowflake 226
snake 283
ring 283
random 314 (seed 1234)

A star is cheapest — every leaf one hop from the authority. A snake is dearest — news walks. Ring lands exactly on snake, and the reason is the interesting part: a ring’s closing edge is peer adjacency, not an uplink. An access point is the root of its LAN and cannot be a guest of its own station. At the association layer a ring and a snake are the same tree.

That was learned the hard way. The generator’s first version treated the closing edge as an uplink and an access point quietly became a client of one of its own attached networks. The topology diagram caught it. Every number measured before that fix was measuring a network that could not exist.

And the average is not the shape. Access point: 52 merges, 99 ms mean. Leaf: 150 merges, 14 ms. Seven to one — an access point carries fifteen attached networks and the tunnel mesh and walks all of it; a leaf has an uplink and a horizon. Both numbers are useful and the mean of them is neither.

§03 · Oracles

Every check is paired with a control that must make it fail.

An oracle that cannot fail proves nothing. Two degradations are reproduced on purpose:

empty store
The tuple space is replaced with a stub that answers nothing. The election falls to its degraded floor and elects by address instead of capability — which is what the old behaviour looked like.
shared identity
A real store, with node identity left process-wide. Five nodes collapse into one row.

The check that matters most is the one telling those apart: the same code, against the same /etc/hosts, must produce different answers from an empty store and a populated one — the address on one, the hardware on the other. A harness returning the same answer either way is not exercising the election at all, and would pass every day while proving nothing.

Coverage is proven, not claimed. Each component registers the log marker it emits and the file that emits it, and counts as exercised only if that marker appears in the run’s own captured output. No marker, no credit — listed as not exercised, with the file named. The registry checks itself against the source every run, so a log line that moves is reported as a stale registry rather than becoming false coverage. It has already caught one. And absence is not always a gap: capability advertisement logs only on failure, so silence there means every advertise succeeded. Each such entry carries the reason its absence is expected. A coverage report that cannot tell “did not run” from “ran silently” is not a coverage report.

§04 · What it found

Five defects, and two of them share a shape.

A fleet is long past its first day. A simulator starts there on every run. Two of these are the same fault wearing different clothes: a file the system depends on that no release ever ships — every running node had it, acquired by hand at some point nobody remembers, and every new node would not have.

Two tables with no creator anywhere in the shipped tree

The code read them. A maintenance script truncated them. The only statement that created them lived in a file the release build excludes. Every running node had them, made by hand once and carried forward. Every new node would have failed on its first semantic request and stayed failed — unable to learn a template because it could not write one.

A fleet cannot report this. A fleet is made of nodes that already got past it.

A systemd target that no release ever shipped

Two services are ordered behind a target reached once discovery settles and names resolve. The gate service exists. Both consumers name the target. The packaging step that collects unit files matched services, timers, paths and mounts — not targets. systemd treats a want on a missing unit as a soft failure, logs a line, and continues.

So the ordering has never once been enforced, on any node, and every unit file reads as though it were.

A convergence validator that stopped walking at twelve hops

On a fifteen-deep chain it reported correct routes as unreachable. 357 false failures.

Zero when the limit was raised.

The route plane logging to nobody

The merge loop called into routing with a logger that discarded everything, and a diagnostic flag was read once at import so setting it later did nothing.

Run log 57 lines → 2,282. Coverage 3 components → 17. No code changed.

Media host election that elected nobody

The capability blob was missing the two fields the handler gates on, so every candidate scored below the floor and the name was removed from /etc/hosts.

Not that no host qualified — that the question was never properly asked.

§05 · Where it stops

Demonstrated, and not.

demonstrated
Discovery convergence under the real notification mechanism. Route promotion and installation. Election of the database host from published capability, agreed independently across the fleet. FrogNet Memory reads and writes with real freshness semantics. Telemetry round-tripped through a separate process. Station re-association after losing an access point.
demonstrated
Packet delivery over a real kernel — one network namespace per node, its own routing table, twenty topologies converging and carrying all-pairs traffic, 220 ordered pairs on the largest. Where the WireGuard module is unavailable the tunnels fall back to the userspace implementation, same protocol, same keys.
demonstrated
Semantic compression between two nodes: 1,129,267 bytes of payload left as 216,252 on the wire, measured from the daemon’s own log rather than from anything the test reported about itself. A request aimed at your own node is correctly served locally and never crosses the daemon — it takes two.
demonstrated
Real applications over time. The monitor’s own fetch code, unmodified, reads a value another node wrote. The A/V phone places a call and the far relay registers the peer. A value written before a partition is still correct after the heal, and a write attempted from the cut-off side is simply absent — not merged, not in conflict.
not demonstrated
Concurrency. Merges run one at a time in a single process, so every race is unreachable by construction.
partly demonstrated
A lossy wire. The transport tier carries the full bearer vocabulary — latency, jitter, bandwidth, outage probability and duration, MTU, and asymmetric forward and reverse paths — and the real SotF ladder has been driven down and back over it under contested-RF conditions. What is missing is the join: a topology edge cannot yet carry those parameters, so a degraded bearer can be modelled on one link but not attached to the New York leg of a five-site pond.
demonstrated
Agreement with hardware. All twenty topologies carry baselines recorded from real-kernel runs and stamped source: hardware; the offline engine reproduces every one, and committed route tables matched byte for byte. Measured per-edge round-trip times from those runs replace the model’s guessed values.
not demonstrated
Prediction. Agreement is not prediction. The offline engine reproduces what hardware did; no run has yet predicted something a real box then confirmed, which is the difference between a faithful harness and a qualified simulator.

None of that is hedging. A simulator that overstates its coverage is worse than no simulator, because the green run becomes the reason nobody looks.

The chapter → Demonstrations → Bring an oracle →