# Experiments

This file holds **metric definitions** and **results from real runs only**. No illustrative numbers.

## Metric definitions

Timeline: the synthetic corpus has a native 30 FPS timeline; `t` is the scene time of a frame.
Under real-time replay, frame `t` is released to the pipeline at wall-clock `T0 + t / time_scale`.

| Metric | Definition | Kind |
|---|---|---|
| `critical_event_recall` | detected GT events / total GT events. A GT event is a maximal interval where a person's foot point is inside the restricted polygon. "Detected" = at least one alert on that camera emitted in `[start − lead, end + grace]`, lead = 0.5 s (detector box bottoms cross the line slightly before the analytic foot point), grace = 2 s. | MEASURED |
| `alert_precision` | alerts attributable to a GT event / all alerts. An alert is attributable if it falls inside some GT event's `[start − lead, end + grace]` window on the same camera. | MEASURED |
| `alert_latency_ms` (mean/median/p95) | wall-clock(alert emitted) − wall-clock(scheduled release of the first in-zone frame of that GT event). Only the first alert per GT event counts; an early alert (within the lead) counts as 0 ms. | MEASURED only if `time_scale == 1.0`; otherwise UNAVAILABLE |
| `cloud_bytes_total` | bytes sent + received through the cloud path (JPEG crop + response payload). | MEASURED |
| `frames_processed_per_s`, `frames_dropped` | throughput of the detector worker; frames discarded by the sampler backlog policy. | MEASURED |
| `cpu_percent_mean`, `rss_mb_peak` | psutil samples at 2 Hz over the run (process tree). | MEASURED |
| `gpu_util_percent`, `gpu_mem_mb` | NVML when available. | UNAVAILABLE locally · MEASURED on Nebius |
| `cloud_rtt_ms` (p50/p95), `cloud_timeouts` | observed round-trips through the injector. | MEASURED |
| `compute_pressure_observed` | detector per-frame time under pressure ÷ baseline. | DERIVED |

Invariants (hard) for `missions/restricted-zone.yaml`: `critical_event_recall ≥ 0.95`,
`alert_latency_p95_ms ≤ 2000`. Objectives (soft): maximise `alert_precision`, minimise
`cloud_bytes_total`, minimise `cpu_percent_mean`.

Verification means: *the configuration passed these checks on the stated corpus*. It does not mean
universal safety.

## Corpus statistics

`quick` tier, seed 1, scene hash `63fc2da84236`, 180 s per camera, 4 cameras
(`failsafe corpus stats --tier quick --seed 1`):

| stat | value |
|---|---|
| ground-truth zone events | 57 (A: 33, B: 8, C: 8, D: 8) |
| distractor tracks | 92 (44 near-miss, 48 wanderer) |
| event duration (s) min / p25 / median / p75 / max | 0.167 / 0.7 / 1.2 / 2.3 / 3.933 |
| events shorter than 0.5 s / 0.2 s | 13 / 3 |
| person height (px) min / p25 / median / p75 / max | 45.3 / 63.4 / 88.3 / 128.5 / 199.0 |
| occluded events (overhead, 15–25 %) | 26 |

Offline detectability of these events with YOLOv8n (conf 0.4, 3 sampled frames per event; recall
ceiling, no latency): 640 px 0.93 · 480 px 0.84 · 320 px 0.54. See DECISIONS D-016.

## Results

The full matrix (all numbers measured, generated by `failsafe report` from `artifacts/experiments/`)
is in `artifacts/phase1_report.md` (24 experiments); the summary below cites it. Corpus: synthetic `quick` seed 1,
scene `63fc2da84236`, 57 ground-truth events. Real-time replay, local backend, YOLOv8n on CPU,
YOLOv8m stand-in confirmer in a child process. Scoring rule: match-window [start−0.5 s, end+2 s].

### Findings (Phase 1)

1. **The landscape is non-trivial.** Of the named probe configurations, some pass the mission
   under a given failure condition and others do not, and the passing set changes with the
   condition. Under `healthy`: `island`, `island_5fps`, `island_bg5` pass (recall 0.965);
   `normal` fails (0.947 — 3 events lost to confirmer rejections at 2 fps background sampling plus
   525 dropped frames); `island_480` (0.737), `island_320` (0.596), `survival` (0.298),
   `normal_overload` (0.912, 1,672 drops) fail.
2. **Resolution is the dominant recall knob.** Single-knob pairs, healthy: 320→480→640 px gives
   recall 0.596→0.737→0.965 (offline detectability probe: 0.54→0.84→0.93). Person height is the
   mechanism: ≥130 px people are found at every resolution, 40–80 px people only at 640.
3. **Critical-camera FPS 15→5 costs nothing** on this corpus (0.965→0.965, drops 90→11). The
   short-event tail (5 events < 0.2 s) is caught by the 0.5 s lead / 2 s grace windows and by the
   fact that the detector fires on any in-zone frame. FPS is a cheap degradation here.
4. **Demand above edge capacity costs recall through dropped frames.** Measured sustainable
   throughput ≈ 21–23 fps at 640 px (detector 47 ms/frame in-pipeline). `normal_overload`
   (30 fps demand) drops 1,672 frames → recall 0.912, p95 1.68 s; `island_bg5` (30 fps) drops
   1,173 but still reaches 0.965 with p95 0.49 s because `drop_oldest` keeps the newest frames.
5. **Slow ≠ dead — measured.** Same `normal` configuration: `wan_offline` → recall 1.000,
   p95 0.84 s, PASS; `wan_zombie` (accepts requests, answers after 5–10 s) → recall 1.000,
   p95 **5.05 s**, FAIL; `wan_severely_slow` (~1.5 s RTT) → recall 0.930, p95 2.95 s, FAIL.
   A cloud that fails fast is handled by the fail-over path immediately; a cloud that is merely
   slow makes the pipeline wait out its timeout on every candidate.
6. **The timeout is a discoverable threshold.** `normal_short_timeout` (1 s instead of 3 s) under
   `wan_zombie`: recall 0.982, p95 1.58 s, PASS — the cloud path is kept and the mission holds.
   Under `wan_offline` it is the best configuration measured (recall 1.000, p95 0.43 s), and
   under `healthy` it passes too (recall 0.982, p95 0.90 s) — one configuration that survives
   three conditions, and the only cloud-confirmed configuration that passes when healthy.
7. **Fail-closed alerting is catastrophic without a policy.** `normal_drop_on_fail` (a failed
   confirmation drops the candidate) scores recall 0.000 under both offline and zombie clouds.
   This is the "system without Failsafe" baseline for the demo.
8. **Compute pressure breaks island mode at both 640 and 480.** With 8 competing busy processes
   (observed detector slowdown 2.6× at 640, 3.7× at 480), `island` → 0.737 (2,088 drops),
   `island_480` → 0.632 (1,799 drops), `island_5fps` → 0.807 (516 drops, p95 2.08 s),
   `survival` → 0.298. **No probe configuration passes under severe compute pressure**: this is a
   genuine no-verified-mode condition, exactly the case the runtime must fail closed on (and the
   case a search strategy has to explore further — e.g. 10 fps × 640 with background cameras off).
9. **Bandwidth alone (2 Mbps) does not break `normal`**: recall 0.965, p95 1.46 s, PASS. Each
   confirmation moves ≈28 KB; at 2 Mbps that is ≈0.1 s of transfer.
10. **Precision: no measurable cost of island mode on this corpus** (all experiments 1.000 after
    the lead-tolerance fix; probe 0/163 near-miss frames, 0/87 empty frames). Null result, D-022.
11. **Cloud confirmation can cost recall.** `normal × healthy` lost 3 events (one 158 px / 1.0 s)
    because the confirmer rejected a boundary candidate and, at 2 fps background sampling, no
    further frame of the event remained. `degraded × healthy` (cloud on, background at 1 fps)
    loses 4 events the same way (0.930). `normal × wan_offline`, where every candidate fails
    over to local alerting, scored 1.000.

### Caveats

- One seed, one corpus tier, one laptop. Run-to-run variation is not yet quantified (planned:
  3 seeds × repeated runs on Nebius).
- The confirmer is a heavier local detector, not a VLM; its rejections are geometric, not
  semantic. Precision findings will change once the confirmer judges authorised vs unauthorised
  persons (D-022 plan).
- No real-video holdout has been run yet (tooling exists: `FileVideoSource`, `corpus label`).
- Compute pressure is CPU contention; on a GPU edge device the mechanism differs.

## Phase 2 — repeatability (run-to-run noise)

`failsafe repeatability` over 3 runs per point (2026-08-27). Runs flagged † were measured while
other CPU work ran on the same laptop (D-028) and are excluded from the cache's first choice.

| scenario | config | runs | recall | p95 latency (ms) | frames dropped | pass rate |
|---|---|---|---|---|---|---|
| healthy | island | 3 | 0.965 ± 0.000 | 1050 ± 161 | 144 ± 61 | 3/3 |
| healthy | naive_default | 3 | 0.959 ± 0.017 | 975 ± 103 | 339 ± 140 | 1/3 |
| healthy | normal | 3 | 0.959 ± 0.022 † | 1077 ± 230 | 469 ± 290 † | 2/3 |
| wan_zombie | island | 3 | 0.947 ± 0.025 † | 1037 ± 132 | 276 ± 234 † | 2/3 |
| wan_zombie | naive_default | 3 | 1.000 ± 0.000 | 5077 ± 18 | 135 ± 46 | 0/3 |
| wan_zombie | normal | 3 | 0.994 ± 0.008 | 1765 ± 192 | 158 ± 47 | 2/3 |
| compute_severe | island | 3 | 0.708 ± 0.022 | 2039 ± 216 | 2244 ± 113 | 0/3 |

Reading: uncontended repeats agree to ±1 event (1.75 %) in recall and ≈ ±150–250 ms in p95
latency; the zombie-cloud latency violation (5.08 ± 0.02 s) is extremely stable. Verdicts that
sit within one event of the 0.95 threshold flip between runs — hence the robust/marginal margin
tiers (D-029). `normal × wan_zombie` p95 1.77 ± 0.19 s passes on recall every time but one of its
three clean runs (2.03 s) exceeds the 2000 ms invariant. Under the D-037 admission rule (every
clean run must pass) that configuration is therefore **not** admitted under `wan_zombie`; the
compiled zombie mode is the cloud-off `island` configuration (4 clean runs, all passing,
marginal). An earlier compiler admitted `normal` here on the strength of its first run alone and
labelled it robust; that was wrong and is recorded in D-037.

### The verification bar (D-037)

A configuration is *verified* for a condition only when:

1. it has at least **two clean runs** under that condition — runs not flagged by the contention
   check (D-028), pooled across corpus seeds and repeats (a pass on each of three seeds is three
   clean runs; one failing seed spoils the pool exactly like one failing repeat);
2. **every** clean run passes the mission invariants;
3. its tier is that of the **worst** clean run: *robust* only if every run clears the noise floor.

The recall, p95 and precision recorded for a mode are the worst observed across its clean runs,
never the best, and the policy file records the rule it was compiled under. A single passing run
is listed as `insufficient_evidence` with the experiment to repeat; a scenario with no stored
result is `untested`; one where every candidate failed a clean run is `refuted`, and `exhaustive`
when the whole 72-configuration space was tried.

The 16-run repeat campaign of 2026-09-06 (19:08–19:58 UTC) ran while Spotlight was indexing the
artifact directory; every replay was starved (detector 2–3× slower than baseline, thousands of
dropped frames), so the whole campaign was quarantined under
`artifacts/experiments/stale/contaminated-2026-09-06/` and re-run behind the idle gate (D-039,
D-041): `failsafe calibrate` records the host's isolated detector speed, and a replay does not
start while the host reads more than 1.25× slower than that. A second campaign on the idle host
then showed runs that followed a compute-pressure replay reading 1.30–1.49× and failing, even for
the cloud-off `island` configuration under `wan_timeout`, where the scenario cannot be the cause;
heat on a passively cooled laptop is contention too (D-043). Since then the gate reads the detector
under sustained load (the last third of ~300 frames), calibration refuses a machine that is
uniformly slower than its own record, and pressure scenarios run last with a 90 s cooldown after
each. None of the quarantined runs entered the evidence.

### The matching window decides nothing (D-045)

The 2.0 s post-event grace had no recorded justification (D-021 covers only the 0.5 s lead).
`scripts/grace_sensitivity.py` rescored every stored synthetic result, read-only, at three windows:

| scenario | runs | grace 0.5 s: pass / recall / precision | grace 1 s: pass / recall / precision | grace 2 s: pass / recall / precision |
|---|---|---|---|---|
| bandwidth_2mbps | 10 | 1 / 0.914 / 1.000 | 1 / 0.914 / 1.000 | 1 / 0.914 / 1.000 |
| compute_severe | 79 | 0 / 0.533 / 0.997 | 0 / 0.533 / 0.998 | 0 / 0.533 / 0.998 |
| healthy | 28 | 10 / 0.848 / 0.994 | 10 / 0.849 / 0.995 | 10 / 0.849 / 0.995 |
| offline_compute_severe | 8 | 0 / 0.452 / 0.971 | 0 / 0.452 / 0.971 | 0 / 0.452 / 0.971 |
| wan_offline | 8 | 6 / 0.860 / 1.000 | 6 / 0.860 / 1.000 | 6 / 0.860 / 1.000 |
| wan_severely_slow | 6 | 2 / 0.929 / 1.000 | 2 / 0.929 / 1.000 | 2 / 0.929 / 1.000 |
| wan_slow | 10 | 2 / 0.867 / 1.000 | 2 / 0.867 / 1.000 | 2 / 0.867 / 1.000 |
| wan_timeout | 11 | 3 / 0.922 / 1.000 | 3 / 0.922 / 1.000 | 3 / 0.922 / 1.000 |
| wan_zombie | 85 | 18 / 0.755 / 0.998 | 18 / 0.755 / 0.998 | 18 / 0.755 / 0.998 |

**Verdict flips between windows: 0 of 245 results.** Alerts either land while the person is still
in the zone or not at all; the grace only absorbs sub-second jitter. The window stays at 2.0 s so
stored results remain comparable, and any value from 0.5 s up would have compiled the same policy.
(Counts include the quarantined runs, which were still in the cache at the time of this study,
D-046; the flip count is what matters and it is unaffected.)

### The VLM confirmer vs the geometric rule (D-049)

`VlmConfirmer` puts a real NVIDIA multimodal reasoning model (`nemotron-3-nano-omni`) in the
cloud-confirmer seat: it is shown the frame with the zone polygon and the person's box drawn on, and
asked whether the feet are inside. On 30 detected people from the **real warehouse holdout**
(`scripts/vlm_confirm_eval.py`, verdicts cached in `artifacts/vlm_cache`), it agreed with the
geometric foot-in-polygon rule on **21 (70%)** — 13 both-inside, 8 both-outside. Median live call
**6.8 s**.

The 9 disagreements are systematic, not noise: **7** are the VLM calling a person "in the zone" when
the foot-point is just *outside* the polygon — its stated reasoning judges whole-body overlap with
the area rather than the strict foot point, so the two diverge near the boundary; **1** was a
transient API error, handled fail-safe as "unconfirmed" (the local decision stands); the rest are
boundary cases. This is the argument for the trust boundary, made concretely: a 70% agreement with a
boundary-sensitive rule is fine for a confirmer sanity-check but unacceptable for the arbiter, which
is exactly why the **deterministic geometric evaluator, not the VLM, decides pass/fail**
(non-negotiables 2–3). The 6.8 s call is also why the VLM is nowhere near the real-time decision
path. Reproduce: `uv run python scripts/vlm_confirm_eval.py`.

## Phase 3 — search strategies on the `wan_zombie` grid (72 configurations, measured 2026-08-27)

Full table: `docs/compare_wan_zombie.md` (generated by `failsafe compare`). Space: critical fps
{15, 5} × background fps {2, 1, 0} × resolution {640, 480} × cloud {off, on/3 s, on/1 s} × indexing
{on, off}. One real-time run per configuration (≈3.5 h wall-clock, laptop otherwise idle).

- Feasible: **12 / 72**; Pareto-optimal (capability ↑, CPU ↓, cloud bytes ↓): 10.
- Best verified mode (lexicographic): critical 15 fps, background 2 fps, 640 px, **cloud on with a
  1 s timeout**, indexing on — capability 1.00, recall 0.982, p95 1.58 s. Every feasible
  configuration is 640 px; no 480 px configuration passes recall; no cloud-on/3 s configuration
  passes latency.
- Strategy comparison by replay against the same grid (D-026):

| strategy | evaluations to first feasible | capability of what it found | best in space |
|---|---|---|---|
| grid | 72 (exhaustive) | 1.00 | 1.00 |
| random | 5.5 ± 4.7 (p90 12; analytic (N+1)/(k+1) = 5.6; 1000 orders) | 0.73 (mean) | 1.00 |
| greedy (bottleneck heuristic) | 2 | 1.00 | 1.00 |
| llm:mock (greedy moves through the LLM interface) | 2 | 1.00 | 1.00 |

Reading: under a zombie cloud the bottleneck is one knob (the cloud timeout), so a heuristic that
diagnoses "latency violated while waiting on the cloud" reaches the frontier in two evaluations,
and random search needs ~3× as many evaluations *and* settles for 27 % less capability on average.
This is the easy case for reasoning; the harder test is `compute_severe`, where Phase 1 found no
passing probe and the grid is being measured now. Nemotron has not been run yet (no API key by
choice); when it is, it replays against this same grid.

## Phase 3 — the `compute_severe` grid: a measured NO VERIFIED MODE (2026-08-27, analysed 2026-09-05)

Same 72-configuration space, one real-time run each, 8 competing busy-loop processes (observed
detector slowdown 2.6–4.6×):

- Feasible: **0 / 72**. Best recall achieved anywhere in the space: **0.807** (required ≥ 0.95);
  best p95 latency 452 ms (latency is satisfiable — recall is not: the detector simply cannot
  process enough frames).
- Greedy walks the intended ladder — shorten timeout → drop cloud → drop indexing → reduce
  background FPS to 0 → reduce critical FPS — in 7 evaluations and terminates with
  "no degradation move left". It does not promote a least-bad configuration; the outcome is an
  explicit `NO VERIFIED MODE`, now grounded in exhaustive measurement rather than 4 probes.
- The compiled policy (`artifacts/resilience-policy.yaml`) lists the condition as unverified with
  these ceilings; the runtime activates the conservative fallback, labels it UNVERIFIED and
  escalates (D-032).

Reading: this scenario is the honest boundary of the mission on this hardware. A useful follow-up
for the mission owner is a *relaxed* invariant set for extreme compute pressure (e.g. "critical
camera only, recall ≥ 0.8, latency ≤ 2 s"), which Failsafe could verify as a separate mission —
but that is the operator's decision, not something the system invents.

## Phase 5–6 — compiled policy + deterministic runtime, live cut-WAN demo (2026-09-05)

`failsafe demo --timeline "healthy:0,wan_offline:50,healthy:120"` — same corpus (quick seed 1,
57 events), same timeline, real-time replay, real YOLOv8n edge detector, stand-in cloud confirmer.
`naive` keeps the designer default statically (`naive_drop_on_fail`); `failsafe` runs the compiled
policy (`artifacts/resilience-policy.yaml`) through the deterministic runtime (health-ping probes
every 2 s, degrade after 3 consistent probes, recover after 5, fail closed).

```
== naive (no policy)
     0.0–  50.0s healthy            recall 0.857 (18/21)    p95 789 ms    false 
alerts 0  modes -
    50.0– 120.0s wan_offline        recall 0.000 (0/26)     p95 n/a       false 
alerts 0  modes -
   120.0– 180.0s healthy            recall 1.000 (10/10)    p95 694 ms    false 
alerts 0  modes -
  overall recall 0.491, p95 1184 ms → FAIL

== failsafe (with policy)
     0.0–  50.0s healthy            recall 0.952 (20/21)    p95 795 ms    false 
alerts 0  modes normal
    50.0– 120.0s wan_offline        recall 1.000 (26/26)    p95 389 ms    false 
alerts 0  modes offline-cloud
   120.0– 180.0s healthy            recall 1.000 (10/10)    p95 1126 ms   false 
alerts 0  modes normal
  → t=   2.0s - → normal   verified: verified mode for healthy × normal
  → t=  33.0s normal → normal   UNVERIFIED: NO VERIFIED MODE AVAILABLE for 
healthy × moderate — fallback `normal` (verified only under healthy); escalate
  → t=  38.0s normal → normal   verified: verified mode for healthy × normal
  → t=  53.0s normal → offline-cloud   verified: verified mode for offline × 
normal
  → t= 124.0s offline-cloud → normal   verified: verified mode for healthy × 
normal
  overall recall 0.982, p95 878 ms → PASS
```

Reading:
- WAN cut at t=50: the runtime classified OFFLINE and activated the verified `offline-cloud` mode
  **3 s** after the cut; every one of the 26 intrusions during the outage was alerted, p95 389 ms.
  The naive system missed all 26.
- WAN restored at t=120: verified recovery to `normal` after the 5-probe hysteresis (4 s).
- The brief `healthy × moderate` classification at t=33 is the honest edge of the compute
  classifier (detector slowdown estimate crossing 1.5×); the fallback config equals `normal`, so
  nothing changed operationally, and the transition is labelled UNVERIFIED rather than hidden.
- Overall (mixed-condition, informational): naive recall 0.491 FAIL; failsafe recall **0.982,
  p95 878 ms — PASS**. Per-condition verification remains the authoritative statement.
- Run-to-run: two earlier runs of the same demo scored failsafe 0.947 (marginal band, 3 healthy-
  phase misses); the offline phase was 26/26 in all three runs. Demo artifacts: artifacts/demo/.

## Phase 4 — Nemotron-guided search (NVIDIA-hosted `nemotron-3-super-120b-a12b`, 2026-09-05)

The Nemotron planner (via `LLMSearch`) proposes candidates from the bounded space; each is snapped
to the measured grid and looked up — so it competes against grid/random/greedy on identical
evidence (D-026), and its cost is its own reasoning calls, not new experiments. Endpoint: NVIDIA
Build during development (D-035); the model id is the same one Nebius Token Factory serves, so
submission runs re-point with one env swap.

### `wan_zombie`

| strategy | evaluations to first feasible | capability of what it found |
|---|---|---|
| grid | 72 (exhaustive) | 1.00 |
| random | 5.5 ± 4.7 (analytic 5.6; 1000 orders) | 0.73 (mean) |
| greedy | 2 | 1.00 |
| **Nemotron** (1 of 3 attempts completed) | **1** | **1.00** |

Three live `wan_zombie` attempts are recorded. Two produced **zero** evaluations: the first was
stopped after two consecutive replies that contained no JSON object, the second reached the round
limit with every candidate rejected as outside the search space (the integration fixes in D-035
followed). The table row is the third, completed attempt; the two rejections are the trust
boundary doing its job, not a rounding error, and they are counted here rather than dropped.

In that completed run Nemotron proposed the optimal mode (cloud on, 1 s timeout, full sensing) in its first candidate,
and its stated rationale reconstructed the discovered threshold from the evidence in the prompt:
*"Cloud timeout of 1 s ensures that even if the zombie cloud delays 5–10 s, the call times out and
triggers local alert."* It beat the greedy heuristic here (1 vs 2) by going straight to the 1 s
timeout instead of starting at 3 s and stepping down.

Caveats (honest): one scenario, one run; the prompt supplies the measured edge-capacity facts
(≈21–23 fps @640) and the fault descriptions, so the model is reasoning over given evidence, not
rediscovering physics. The measured grid is the ground truth it is scored against — the model
cannot make a failing config pass. A single reasoning call is also not free wall-clock: Nemotron 3
Super emits a long chain (thousands of completion tokens) per call.

### `compute_severe` — the honesty test

Nemotron ran 6 rounds proposing 15 distinct candidates; **none passed**, and it
terminated with `max rounds reached` → `NO VERIFIED MODE` (best recall achieved 0.807,
required 0.95 — matching the exhaustive grid's ground truth). It reasoned correctly about capacity
(dropping to 480 px "doubles detector throughput", cutting FPS keeps "demand well under capacity")
but every measured result still fell short, because under this CPU contention the detector cannot
sustain enough throughput at any resolution the mission's recall needs. The model did **not**
promote a least-bad configuration or claim an untested one would work — the guard rails and the
deterministic evaluator held. This is the LLM-in-the-loop version of the fail-closed property:
Failsafe (and Nemotron within it) can say "this cannot be done on this hardware."

| strategy | evaluations | outcome |
|---|---|---|
| grid | 72 | 0 feasible → NO VERIFIED MODE |
| greedy | 7 | 0 feasible → NO VERIFIED MODE |
| Nemotron | 15 | 0 feasible → NO VERIFIED MODE |
