APEvidence DeskRecruiter route ↗

OBSERVATORY / EXACT RECORDS / 2026-08-09

Metrics with
an escape hatch.

Every result carries the commit, reproduction command, environment, and a reason not to overgeneralize it. Different units stay separate; there is no meaningless leaderboard across systems.

Cross-commit regression tape

First and current results stay visible together. A regression is recorded even when the implementation’s preferred story would rather omit it.

FAULTLINE / higher IS BETTER

3-node replication

3.4% lower on the release-candidate rerun; retained as visible variance, not rewritten away.

2026-07-315bdba04746,301.1 ops/sBASELINE
2026-08-09a5dd098720,980.5 ops/sCURRENT
FAULTLINE / higher IS BETTER

5-node replication

4.3% lower in the second local run; protocol source is unchanged and machine noise remains plausible.

2026-07-315bdba04413,382.6 ops/sBASELINE
2026-08-09a5dd098395,782.9 ops/sCURRENT
KERNELARENA / lower IS BETTER

TensorForge JavaScript gzip

1.3% growth buys the live cross-shape KERNELARENA matrix; tracked as an explicit bundle cost.

2026-08-09ba2beaf72.85 KBBASELINE
2026-08-0970a5c8573.81 KBCURRENT

10 auditable records

Values were reproduced locally or taken from committed repository output. Headline numbers are paired with the regime where they stop being useful.

01

3-node replication

FAULTLINE
746301.1 ops/s5bdba04

Apple M2 Pro, 16 GB, in-process simulator

$ make clean test benchmark

No sockets, serialization, fsync, scheduling, persistence, snapshots, or membership changes.

02

5-node replication

FAULTLINE
413382.6 ops/s5bdba04

Apple M2 Pro, 16 GB, in-process simulator

$ make clean test benchmark

A protocol microbenchmark, not production database throughput.

03

5-node failover

FAULTLINE
11 logical ticks5bdba04

Deterministic virtual transport

$ make clean test benchmark

Logical ticks are not wall-clock milliseconds.

04

IMM localization error

SIGNALROOM
14.4 OSPA loc9a1092b

5 paired seeds, configured ±6 degree/s model bank

$ python bench/benchmark.py --quick

The bank loses on straight motion and at 12 degree/s.

05

40 s VIO ATE

SIGNALROOM
0.041 m RMSEf38714f

3 seeds, 73.3 m repeating figure-eight

$ python bench/benchmark.py --quick

The repeating path implicitly re-observes landmarks; hover is a documented failure.

06

SAR range IRW error

SIGNALROOM
0.02 %8304157

200 MHz bandwidth, 200 m aperture

$ python bench/benchmark.py --quick

Synthetic point targets and time-domain backprojection.

07

AAD delta relative error

MARKETWIRE
5.56e-16 relative errorc26ef7a

Black-Scholes analytic comparison

$ make bench

Pathwise AAD returns zero for a sharp digital payoff; smoothing is required.

08

50-input AAD speedup

MARKETWIRE
57.6 x vs bump-and-revaluec26ef7a

Apple M2 Pro local run; 26.5 ms AAD vs 1524.5 ms bump-and-revalue

$ make bench

One workload and machine; not a general performance claim. AAD tape cost was 2.00 times pricing.

09

Inventory-skew dispersion

MARKETWIRE
12 inventory SD82c75a1

20,000 steps across 12 seeds; inventory-skew strategy

$ make bench

The strategy reduced inventory dispersion from 88.1 but also reduced P&L from 20,853 to 15,448.

10

Compiler/runtime tests

KERNELARENA
9 passing testsba2beaf

Vitest compiler and runtime suite

$ npm test -- --run && npm run build

Browser timing depends on the visitor's GPU and is never presented as a universal speedup.

Future runs append records only when commit, machine, command, and output are retained. Changing an old result in place would destroy the provenance trail.