Specter PoE unit test — S1 + S2

Updated 2026-08-14 after storage catch-all. Cisco 192.168.86.21 · VMS 192.168.86.54:8788. Cameras left UP. MQTT storage_health=OK is not overlay truth. No secrets.

Why S2 storage was silent — then the catch-all

You did not miss a toast. Nothing in the existing watchers treats camera overlay as a first-class alert.

WatcherCadenceWhat it actually watchesWhy 97% overlay was invisible
Platform Watcher (launchd, running)12s tick · edge probes 30s · disk 5mHost / on this Mac, MQTT keys present, go2rtc from 87.xdisk_host is the Mac disk (failCount 0). edge_s1/s2 circuits are open because this Mac cannot L2-probe streams — 11k fails, not storage.
Soft Watcher (canvas)12sStream repair via proxy MCPNo overlay probe. Escalates blue-LED / firmware-replay only.
MQTT devices/*/health~health periodrecorder.storage_healthStill published OK on S1 and S2 at 14:42 CDT after S2 had been 97%. Field is untrusted.
Edge recorder pruneretention / --prune-firstoldest chunk_*.h265Does not delete pre_event_thermal_* (the 42 GB bomb) or 13 GB logcat.log.
Reconcile LaunchAgent90s (if loaded on 86.54)ghost_core missing → runc startNo-op when pipeline is up and disk is dying.

So MQTT was green, the host Watcher was yellow for the wrong reason, and no playbook pruned thermal pre-event files. That is the gap.

New catch-all scripts/fleet-catch-all.py (LaunchAgent com.ghostprotocol.specter-fleet-catchall):

Not a 30s Grok/Hermes loop — Watcher doctrine is deterministic probe → playbook → verify. An LLM can explain after the fact; it must not be the 15s decision path.

Live health — 2026-08-14 14:12 CDT

Read-only jump + VMS. Both last-known oakapps still running. Grades are for this snapshot, not the boot clocks.

Nominal ops

S1S2
Uptime / load8h09 · 2.26 2.30 2.218h09 · 2.28 2.14 2.20
CID / runcaaf26107… running · 162465acd4e6… running · 1719
PIDsagent 1771 · core 1812 · go2rtc 17751794 · 1903 · 1802
Layerlayer1_rfdetr_thermal_safe · ENCODE_ONLY=0 · H.265 1280×720 · H265_FPS=5 · RF_DETR_STRIDE_FPS=5 · 1500 kbps · keyframe 2s
Canarytoken matchtoken match
MQTT 86.54shadow1 ~19sshadow2 ~8s
VMSfleet_grade 90 · readiness B/84 · Flask warn (host, not camera)

Ops: A− — last-known up. Not A: SoC warm/critical + H.265 FIFO stalls.

RF-DETR

S1S2
DLCghost_protocol_transit.dlc 135 MB + rfdetr_smallsame
.latest_dets.jsonage 0.2s · n=0 · persons=0 · 69.6°C normalage 0.2s · n=0 · persons=0 · 78.8°C warm
.rfdetr_statuskey=value (not JSON) · last_detection now · 0/minsame
ghost_core SoC78–90.7°C CRITICAL spikes79–82°C warm · VIDEO_FIFO drops ~355k
sysfs / MQTT thermal63°C / mqtt 73.1°C — different sensors56°C sysfs
Persons seenMQTT health person_count=3 same minute (empty scene at sidecar snap is ok)0 this snap

RF-DETR: B+ — file + live writer + MQTT persons. Not A: S1 SoC critical, warm stride×2, status file not JSON.

Storage

S1S2
/overlay at snap62% · 48/77 GB97% · 74.7/77 GB · 2.3 GB free
/overlay after prune26% · 20.1/77 GB · 56.9 GB free
Today 2026/08/142075 .h265 · 5 .tmp2534 .h265 · 1 .tmp
pre_event_thermal_*1321 files · 14.6 GB2536 files · 42.3 GB (~20 MB / 10s)
justified chunk_*756 · 11.4 GBpresent (~60s); drowned by pre_event
root /88% · 608 MB free61%

Storage: S1 C+ · S2 snap F → after prune B. Deleted 2607 pre_event_thermal_* (43.5 GB, kept 30) while oakapp stayed up. df stayed 100% until logcat.log (13.3 GB) was truncated in place and caches dropped — then 26% / 56.9 GB free. Ghost still 1794/1903. Do not PoE-cut; open .tmp remains.

Recordings / FPS

Target from env is 5 fps. ffprobe reports 25/1 with no duration — ignore. Trust chunk sidecar + NAL / frame_seq.

MeasureS1S2
Closed/open proofchunk_1786734726_980 sidecaropen chunk_1786734751_199 dets.jsonl
Wall duration60.38 s (target 60)~55.5 s dets window
Pictures / ffprobe frames412 = nb_read_frames 412
Sidecar fps tag15.0 (stale vs env 5)
Measured wall fps6.82~4.86 (seq 142522→142792)
Session encoded_frames / 8.15 h~6.9 fps~4.85 fps
RF-DETR samples38 dets lines / ~14 s on next chunk152 lines on open tmp
Justified / VPS+IDRyeswriting

Encode: B — continuous justified H.265. Actual ~5–7 fps, not 15 and not 25. Thermal stalls keep it off the sidecar 15 tag.

Run4 — 6:00am CDT one-shot simultaneous power-on

Scheduled one-shot. Login OK · adminEnable=1 both · T0 1786705338. 0 W right after enable (SOC boot). No reconcile. Good CIDs only: S1 aaf26107…4453 · S2 65acd4e6…9628c. New PIDs S1 ghost 1771/1812 (runc 1624) · S2 1794/1903 (runc 1719). Layer layer1_rfdetr_thermal_safe · ENCODE_ONLY=0. Scheduler 019ffea791e1 deleted after enable so it cannot fire again.

GateS1S2
T_ping / T_ssh00:52 / 00:3100:52 / 00:52
T_layer / T_config00:31 · token match00:52 · token match
T_ghost / T_rfdetr00:5200:52
T_rec_start (first H.265)01:5901:59
T_rec_stable02:4202:41
T_nominal02:42 · 3 min02:41 · 3 min
Overlay now41%40%
Open .h265.tmp51

CBS watts after detect=3 (post-nominal samples; milliwatts ÷ 1000). S1 sawtooth continues. S2 ~8–10 W.

T+gi1 S1gi2 S2Aligns with
00:000 W · en=1 det=20 W · en=1 det=2enable; SOC boot
03:384.6 W7.9 Wpost T_nominal
04:035.7 W9.1 W
04:277.7 W10.5 WS2 peak this set
04:516.7 W10.3 Wwatch sample end

T_mqtt at 00:31 is stale — VMS still listed shadow1+shadow2 grade 90 across the overnight halt. Do not treat the device-name list as a fresh clock. Real pipeline proof is ghost_core + new H.265 (chunk_1786705383 / chunk_1786705389).

Cameras left UP. No halt. No PoE-cut. JSONL: docs/grading/poe-unit-test-s1-s2-run4.jsonl.

Run3 — simultaneous power-up (prior practice)

Both ports were 0 W / adminEnable=2. Enabled together (adminEnable=1, status 0). No reconcile. Good CIDs only. New PIDs S1 1763/1846 · S2 1699/1862.

GateS1S2
T_ping / T_ssh00:3700:38
T_layer / T_config00:37 · token match00:38 · token match
T_ghost / T_rfdetr00:6001:00
T_rec_start (first H.265)02:0402:03
T_rec_stable02:4702:46
T_nominal02:47 · 3 min02:46 · 3 min
Overlay now41%72%
Open .h265.tmp45

CBS watts after detect=3 (full 5 min watch). S1 is a sawtooth, not the old flat ~4.6 W note. S2 sits ~8–13 W.

T+gi1 S1gi2 S2Aligns with
01:067.0 W8.9 Wghost_core just up
01:266.8 W11.7 WRF-DETR warm
01:457.8 W11.9 Wpre first chunk
02:046.1 W7.9 WT_rec_start
02:247.1 W12.3 W
02:437.5 W8.6 WT_nominal
03:023.6 W8.6 WS1 dip
03:225.9 W12.8 WS2 peak
03:412.9 W8.3 WS1 low
04:007.2 W10.6 W
04:206.2 W11.2 W
04:393.4 W10.4 W
04:583.1 W9.3 Wwatch end

T_mqtt at 00:37 is stale — VMS still listed shadow1+shadow2 grade 90 from before the halt. Do not treat the device-name list as a fresh clock. Real pipeline proof is ghost_core + new H.265.

When it is safe to begin shutdown

Now is a good time to start the graceful script — both units are past T_rec_stable. It is not safe to PoE-cut this second: 4/5 temp chunks are still open.

WindowDo thisDo not
T=00:00 → T_ghost (~60 s)Wait. OS / oak-agent coming up.Halt or PoE-cut — messy files, possible blue-LED next boot.
T_ghost → T_rec_start (~60–124 s)Wait unless emergency. Pipeline opening first file.PoE-cut while encoder starts.
After T_rec_stable (now)Run scripts/edge-safe-shutdown.sh S1 then S2 (or both). Script: runc kill → wait ghost gone (~4–5 s) → shutdown -h → wait SSH dead (~6–9 s) → then Cisco adminEnable=2.Skip the script and pull PoE. Halt leaves ~4–6 W until PoE off.
Emergency / no SSHCisco adminEnable=2 only. Expect orphan .tmp.Call that a clean stop.

Rule: inform the camera first, cut power last. The camera must close H.265. Cisco 0 W is step 4, not step 1. VMS 86.54 stays up.

Run3 shutdown executed 2026-08-14 after T_nominal. Parallel halt, then PoE.

StepS1S2
Pre PIDs1763 / 18461699 / 1862
Pre overlay41%72% (1968 chunks)
ghost gone+4 s+8 s
Prune1938 deleted, keep 30; df still 72% (space elsewhere or delayed reclaim)
SSH dead+9 s+14 s
Pingdeaddead
PoE after halt, before cut2.3 W3.8 W
PoE adminEnable=20 W0 W

How to tell the camera to shut down — options and M8

Anything that ends in edge-safe-shutdown.sh (or the same runc → sync → shutdown -h on-device) is a good trigger. PoE off is never the trigger.

TriggerWhere it livesFit
SSH / ops scriptThis Mac → jump → edge-safe-shutdown.shWhat we use in lab. Proven <10 s ghost-gone.
VMS / MQTT / proxy MCP86.54 tells the edge “prepare halt”Best fleet remote. Needs a small on-device listener (not built yet).
Watcher / reconcileDesired-state “stopped”Good for scheduled maintenance, not ignition-off.
M8 Button 1–3On the camera. Luxonis set_btn_callback / GPIO IRQBest physical lab button. Hold 2 s → same halt. LED = shutting down. Buzzer optional.
M8 GPIO (ignition / ACC)16× 3.3 V GPIO, ≤50 mA total; 2× 3.3 V 450 mA sensor railsVehicle key-off. Use an optocoupler — do not put 12 V on the pin. Falling edge = start halt.
M8 CAN / M8 CAN AdapterSocketCAN 2.0A/B to 1 Mbps. Example ID 0x123 on buttonBus “ignition off” / J1939-style frame. Adapter is CAN-only ($99); exclusive with USB-C host and FSync splitter.
M8 RS2321× serial to PLC / ECUFine if the vehicle already speaks RS232, not first choice.
M8 relays4× SPDT latching 16 A / 400 VAC — outputsNot an input. After halt they can drop accessory loads. Do not switch camera PoE with them.

M8 Controller Box (PR1, $249, powered from OAK4 M8 5 V): CAN, RS232, 16 GPIO, 4 relays, 3 buttons, 3 LEDs, 2× USB-A, buzzer. Examples run on the OAK4 via u2if — not host-driven. After shutdown -h the box dies with the camera, so it cannot cut Cisco PoE. Sequence: M8 trigger → software halt → VMS sees MQTT/SSH drop → Cisco adminEnable=2 (or ignition-switched midspan).

Recommended vehicle wiring: ignition ACC → optocoupler → GPIO; optional CAN frame as a second vote; Button 1 as a manual override. Same handler as the lab script. Strobe is only partially supported on PR1; ADC/I2C/SPI passthrough is not yet. On-device listener: scripts/m8-shutdown-listen.py (button-1 hold 2s / GPIO20 falling / CAN 0x100|0x123 data0=0x01). Copy onto the OAK after next boot; --dry-run first. Missing box exits M8_ABSENT.

Efficiency notes from run3

  1. Simultaneous gi1+gi2 on is as clean as serial and ~half the wall time. Use serial only when isolating a unit.
  2. After a clean halt, ping/SSH returns in ~37 s — faster than a mid-run hard cut story.
  3. Long pole this run was first H.265 (~02:04), not RF-DETR (~01:00). T_nominal is rec_stable, not ghost.
  4. Do not trust VMS mqtt_devices across a halt — names linger. Clock new files / new PIDs.
  5. CBS wattage is a free boot oscilloscope. Detect=3 + 6–8 W (S1) / 8–13 W (S2) ≈ pipeline up.
  6. S2 overlay 72% — prune keep-30 before the next long soak, not during boot.
  7. Open .tmp count is the real “don’t cut yet” flag. Script first, watts last.
  8. Enable-only (already 0 W) is cheaper than another 60/30 s drain when practicing boot.

Premium grades — last-known → nominal (run2)

Bar: proven 0W cut, same CID/layer/canary, RF-DETR with ghost_core, MQTT on 86.54, no reconcile. Premium time is ≤1 min; 3 min is lab-good.

CriterionPremium meansEvidenceGrade
Last-known restoreSame CID, layer, not encode-only, no hollow, no rewritelayer1_rfdetr_thermal_safe, ENCODE_ONLY=0, good CIDs, no reconcileA
Config persistDurable /data/config canarytoken poe-run2-20260814T014243Z on marker + profileA
RF-DETR resumeT_rfdetr after T=00:00S1 01:01 · S2 00:59 (n_dets=0 allowed)A−
Time to nominalPrefer ≤1 minS1 01:23 · S2 01:20 · 3 min bucketB+
Cut hygieneMid-hold 0W + new PIDsBoth units after write-path fixA
MQTT / VMSFresh on 86.54 onlyshadow1+shadow2 grade 90A−
Test craftSerial, one portRun2 after false-pass run0A−
Overlay headroomAuto prune before 95%S1 31% / S2 52%; prune not automatic on cycleB+
Preview JPEGOut of pass barS2 :5001 can timeout while H.265 is fineN/A

S1 recovery: A− · S2 recovery: A− · Session / ops craft: A− (B+ until write-path and password were fixed).

Latest outcome (run2) — last-known config + RF-DETR

UnitPoE cut?Config persisted?T_rfdetrT_nominalBucket
S1 Yes · 0W mid-hold · PIDs 1709/1888 → 1632/1835 Yes · token match 01:01 01:23 3 min
S2 Yes · 0W mid-hold · PIDs 1682/1820 → 1806/1913 Yes · token match 00:59 01:20 3 min

Both auto-returned last-known without reconcile. RF-DETR resumed with ghost_core (same poll). First H.265 appeared before ghost_core. MQTT on 86.54: shadow1+shadow2, grade 90.

Config settings under test (run2 canary)

SettingS1S2After boot
Good oakapp CIDaaf26107-e3d3-4a9d-9aff-a923d888445365acd4e6-58ee-4a38-ab37-e03e8ab9628csame CIDs running
Layerlayer1_rfdetr_thermal_safe · ENCODE_ONLY=0unchanged
Camera profile (unchanged encode)v6_5fps_rf1_720ptransit_1080p_15fps_rf5left as-is
Canary tokenpoe-run2-20260814T014243Zsurvived both boots
Marker file/data/config/poe-unit-marker.jsontoken identical
camera_profile.poe_unit_tokensame token written before draintoken identical
RF-DETRon (not encode-only)T_rfdetr ≈ T_ghost
MQTT / VMS86.54:1883 / :8788 (not 86.46, not 87.x)both shadows listed
Overlay at plant31%52%no prune (under 80%)
S1 hollow CIDs1a8e8356… / 3659a901… never startnot started

Metrics — run2 vs run1

MetricS1 run1S1 run2S2 run1S2 run2
PoE mid-hold0W / admin=20W / admin=20W / admin=20W / admin=2
T_ping / SSH00:3500:4000:52 (down at 00:32)00:59 ping / 00:37 ssh
T_ghost (both procs)00:5501:0100:5200:59
T_config (canary)00:4000:37
T_rfdetrnot gated01:01not gated00:59
T_rec_start00:3500:4000:5200:37
T_rec_stable01:1501:2301:3401:20
T_mqtt (86.54)~00:3500:40~00:5200:37
T_nominal01:1501:2301:3401:20
Bucket3 min3 min3 min3 min
Reconcile needed?nononono

T_rfdetr − T_ghost ≈ 0 on both. Model load is not the delay; ghost_core trails the agent by ~20s.

Run2 S1 timeline (unix T0=1786671857)

T+EventWhat happened
−60sdraingi1 adminEnable=2. mid-hold 0W / 0 mA.
00:00power_onadminEnable=1, still 0W — SOC coming up.
00:40T_ping T_ssh T_layer T_config T_rec_start T_mqttToken poe-run2-20260814T014243Z present. ghost_agent only. MQTT lists shadow1.
01:01T_ghost T_rfdetrghost_core PID 1835. ENCODE_ONLY=0 → RF-DETR path live.
01:23T_rec_stable T_nominalH.265 advancing. Last-known config + RF-DETR. 01:23 · 3 min.

Run2 S2 timeline (unix T0=1786671993)

T+EventWhat happened
−30sdraingi2 9.9W → 0W adminEnable=2. S1 left up (1835).
00:00power_onadmin=1, 0W — booting.
00:37T_ssh T_layer T_config T_rec_start T_mqttToken survived. ping still false. ghost_core not yet.
00:59T_ping T_ghost T_rfdetrFull pipeline + RF-DETR. New PIDs 1806/1913.
01:20T_rec_stable T_nominal01:20 · 3 min.

Run1 baseline (same day, no canary)

UnitT_nominalNotes
S101:15T0=1786669300 · PIDs 2536/2586 → 1709/1888 · overlay 29%
S201:34T0=1786669586 · PIDs 1780/1812 → 1682/1820 · overlay 48%

Run1 proved the CBS write path after false-passes (adminEnable=0 and expired-password read-only). Run2 added config persistence + RF-DETR clocks on that same path.

Issues found and mitigations

#IssueMitigation
1Cisco password expired → read-only session. Bastion login looked like “Bad User or Password”.Operator reset. Writable login from VMS 86.54 (statusCode 0).
2adminEnable=0 does not cut CBS PoE. Early cycles left PIDs alive (false-pass).adminEnable 2=off, 1=on. Mid-hold 0W required before T=00:00.
3POST XML without action="set" → “Wrong format of XML detected”.POST {PoEPSEInterfaceList} with action="set". In cisco_switch_api.py.
4ghost_agent up ~20s before ghost_core.Nominal waits for both. Run2 T_rfdetr tied to ghost_core.
5Nested SSH truncates /health JSON.Bastion python3 urllib: True shadow1,shadow2 90.
6This Mac (87.37) has no L2 to 86.x.Switch HTTP from 86.54; edges via jump; MQTT only on 86.54.
7S1 overlay 100% (58GB H.265) blocked oakapp on the unplanned POE.Run2 overlays 31%/52%. Reconcile prunes at ≥95%. Keep-60 pre-cycle if ≥80%.
8Config might have been overlay-only (lost on reboot).Canary on /data/config survived both run2 boots.

Top 10 — cut power-restored → nominal (after run2)

  1. CBS write path action="set" + adminEnable 2 — proven run1 and run2.
  2. Boot-pin good CID — both units auto-started; no reconcile.
  3. Rotate H.265 before overlay ≥95% (not hit this run; still the 2026-08-13 killer).
  4. LaunchAgent on 86.54: ping-up + no ghost_core by T+90s → runc. Scaffold: scripts/install-edge-reconcile-launchagent.sh.
  5. NTP in parallel with oakapp — best-effort added to reconcile-edge-desired-state.sh.
  6. Default start = non-TTY runc (oakctl hub password blocks).
  7. Judge MQTT only on 192.168.86.54 — monitor does this now.
  8. ghost_core trails agent ~20s (S1 00:40→01:01, S2 00:37→00:59) — don’t wait on preview JPEG.
  9. go2rtc preview watchdog (S2 :5001 can timeout while H.265 is fine).
  10. Pre-warm RF-DETR/DLC — not the bottleneck: T_rfdetr − T_ghost ≈ 0.

How to repeat

# Plant canary, then from VMS 86.54 after writable Cisco login:
# /data/config/poe-unit-marker.json + camera_profile.poe_unit_token
python3 /tmp/cisco-poe-cycle-lan.py gi1 60   # mid_hold must be 0W
python3 scripts/poe-unit-monitor.py --device S1 --t0 <unix> \
  --expect-token poe-run2-… --jsonl docs/grading/poe-unit-test-s1-s2-run2.jsonl
python3 /tmp/cisco-poe-cycle-lan.py gi2 30
python3 scripts/poe-unit-monitor.py --device S2 --t0 <unix> \
  --expect-token poe-run2-… --jsonl docs/grading/poe-unit-test-s1-s2-run2.jsonl

Raw: docs/grading/poe-unit-test-s1-s2.jsonl (run1) · docs/grading/poe-unit-test-s1-s2-run2.jsonl (run2) · docs/grading/poe-unit-test-s1-s2-run3.jsonl (run3 simultaneous) · docs/grading/poe-unit-test-s1-s2-run4.jsonl (run4 6:00am CDT one-shot). Detail twin: poe-unit-test-s1-s2-run2.html.

Safe shutdown (2026-08-14) — not a plug pull

Order: stop oakapp (close H.265) → sync → shutdown -h now → wait SSH/ping dead → then Cisco adminEnable=2. Script: scripts/edge-safe-shutdown.sh. VMS 86.54 left up.

StepS1S2
Pre: ghost PIDs1632 / 18351806 / 1913
Pre: overlay41%72%
runc kill + deleteSTOP_ISSUEDSTOP_ISSUED
ghost_core gone+4 s+5 s
OS halt / SSH dead+6 s+9 s
Bastion pingdeaddead
PoE after halt (before cut)still ~6.1 Wstill ~4.0 W
PoE adminEnable=20 W0 W

Halt ≠ zero watts. Script allows 15–90 s for chunk close; this run closed in <10 s each. Serial both + PoE off ≈ 1 min. Do not PoE-cut while ghost_core is writing.

Catch-all 2026-08-14T20:34:38Z: S2 warn pre_event=202; mqtt storage_health=OK ignored (untrusted) · rc=0 PRUNED 172 KEPT 30 WAS 202 /dev/sda11 80761644 27331264 53413996 34% /overlay Warning: Permanently added 'ssh-remote.ghostprotocol.app' (ED25519) to the list of