Ghost Protocol · internal · rewritten in plain language · 13 August 2026 · fleet addendum 14 August 2026

What we built, what it proved, and what “NIST” actually means

Same FR facts as the first report. Section 10 is the compiled 14 Aug fleet addendum (PoE boot clocks, RF-DETR health, S2 overlay emergency, catch-all watcher). If you only read one FR page, read through section 3. If you only read one ops page, read section 10.

1. Picture the product (forget NIST for a minute)

Think of a coat-check, not a fingerprint lab.

1. Camera sees a person 2. We cut a still of the head area 3. We compare it to saved stills 4. We show a percent + “recommend confirm” 5. A human says yes, no, or new person

The compare in step 3 is a cheap visual fingerprint (a 16×16 brightness pattern). Same idea as “these two photos look kinda similar,” not “this is Dean with a measured error rate.”

Tonight on the desk:

18 automated tests around that loop also passed. Dean has not yet physically walked the loop for this write-up — an agent clicked Confirm on live crops.

2. There are two different “NISTs.” This is the whole confusion.

Road A — NIST FRTE

What it is. A US government algorithm contest. Companies mail NIST a face-matcher program. NIST runs it on secret photos in their lab. They publish a public scorecard: how often it misses, how often it mixes people up.

What it is not. Not a certificate for Specter. Not a certificate for a COTA bus. Not something you get by writing careful UI copy.

Analogy. Crash-testing a car model at a national lab. The lab scores the engine. It does not certify your taxi company because you installed that engine — and it certainly does not certify you if you are still using a bicycle.

We are here: no engine We have a bicycle (visual hash). We cannot enter this contest, and we cannot put someone else’s score on our jersey.

Road B — NIST AI RMF evaluation

What it is. A voluntary homework packet for any AI product. Four verbs: Govern, Map, Measure, Manage. It means: write the rules, name the risks, measure what you can, and have a human in charge when it matters.

What it is not. NIST does not grade this packet. There is no “pass.” Finishing it does not make you “NIST Compliant.”

Analogy. A kitchen safety checklist you fill out yourself. Doing it well is responsible. Framing the filled-out checklist as a health-department certificate is a lie.

We are here: checklist started We used the four verbs to design crop → confirm. We have not written the policy, the risk list, or a real measurement notebook yet.

Road A · FRTERoad B · AI RMF evaluation
Who runs the test? NIST, on their computers, with their photos Us (or a consultant we hire). NIST is not in the room.
What gets tested? A face-matching algorithm How we operate the whole product
What you get A public report card with error rates Documents, logs, and better habits
Do we have it? No. We do not even have an algorithm they would accept. Started. Not finished. Not a stamp.
Needed for COTA MVP? No. Crop → confirm is useful without it. The habits, yes. The phrase “NIST evaluated,” no.

3. Recommended path (researched 13 Aug 2026)

Road A — license, do not submit

Submitting ourselves means we become a biometrics vendor: Linux C/C++ API, no network, their sequestered photos, months of wait, an ML team. That is not our product. Our house docs already said we consume edge FR and do not claim a rank.

When Road A opens: put a vendor SDK on the Mac behind the same /fleet/faces/compare door. Glass stays labels + Confirm. Claim becomes: “Matcher is <vendor>, see their NIST FRTE report card dated ….” Never “Specter is NIST FRTE.”

Who to call first (not a NIST endorsement — NIST does not pick winners):

  • ROC (roc.ai) — first conversation. US company, on-prem / air-gap SDK, 30-day trial, built for integrators. Fits a host Mac and “rider faces do not go to the cloud.”
  • Paravision — second conversation. Current public 1:1 cards (e.g. paravision-020, Dec 2025). Modular SDK; liveness is a later add-on, not day one.
  • Not first: NEC (often national-scale bundles), open-weight AdaFace/InsightFace (cannot inherit someone else’s FRTE card; license/export mess).

S2’s nist_fr hook is quality / optional PAD — not a FRTE matcher. Do not confuse it with Road A.

Road B — succeed here first

NIST’s own Core says: after some Govern, most teams start with Map, then Measure / Manage. Practitioners warn: policy with no inventory governs imaginary systems; most teams write Govern+Map and then skip Measure. We will not skip Measure.

Success on Road B means we can show:

  1. A one-page AI inventory (what models we run: RF-DETR person, aHash compare, Hermes text — not “FR”).
  2. A frozen use case: named person of interest only; riders stay unnamed.
  3. A claims owner and the “do not say” list.
  4. Host audit of every Confirm / Not them / New person.
  5. A still-folder score (operational, never called FNMR).
  6. Dean standing in S1 then S2, filmed after it works.
  7. Wrong-Confirm playbook + Face 1 cleaned up.
  8. Glass that stays up on login.

When those eight exist, an OEM swap is a brain transplant, not a rewrite. That is how we “go back to Road A” without lying.

Order, in weeks, not slogans

  1. This week — Map: inventory + use-case paragraph + risk list (wrong Confirm, spoof, claim-creep, person-box crop).
  2. This week — Measure/Manage lite: you stand in frame; name or delete Face 1; keep Alert off.
  3. Next — Measure: write Confirm clicks to the Mac; run a Dean / not-Dean still folder; count blur rejects.
  4. Next — Govern: one-page watchlist policy; who may enroll / Confirm / Alert; how to delete.
  5. Then — Manage: wrong-Confirm playbook; pause-auto-find switch; durable :8081.
  6. Only if COTA asks for identification, not suggestions: ROC then Paravision eval on the same compare API. Still HITL. Still no live nameplate.

Sources: NIST FRTE/FATE (four-month submit rule, algorithm-not-product), AI RMF Core (start Map after Govern), AI RMF Playbook (voluntary, pick a few), vendor public SDK/FRTE pages for ROC and Paravision. Recheck cards before any SOW — ranks move.

4. So what are we allowed to say?

Say this

  • The camera sees people, not names.
  • Specter crops a person, compares the still to enrolled photos, and asks an operator to confirm.
  • Walking by does not alert.
  • We used NIST’s AI risk checklist (Govern / Map / Measure / Manage) to design that human step. This is not a NIST evaluation and not FRTE.

Do not say this

  • We recognize Dean in real time.
  • NIST-aligned face recognition.
  • Any FRTE rank, accuracy %, or “as good as NEC/ROC.”
  • Liveness, anti-spoof, or “fair across groups.” We did not measure those.

5. How much is built (the coat-check, today)

PieceIn English
Saved faces on the MacYesDean pack enrolls. Drafts exist. Alert defaults off.
Compare stillsYes, cheapSame photo = 100%. Similar brightness pattern ≈ 55%+ suggestion. Labeled “not FRTE.”
Skip junk cropsYesToo small or too blurry never reaches Review.
Review cardsYesConfirm / Not them / New person. Optional name. Quiet auto-find while live is open.
Confirm teaches the galleryYesTonight Dean went 4 → 5 stills. Next find can use that crop.
No “Dean is here” bannerYesThat was the point of this pass.
Live video unnamedYesNo boxes, no nameplates on Sight.
Real face detectorNoWe cut the top of a person box. Backs, hats, far-away people will fail or look Unknown.
Real face matcherNoNo embeddings. Cannot enter FRTE. Cannot quote error rates.
Written decision log on the MacNoClicks live in the browser. Host audit list is empty.
Glass that stays runningNoVite on port 8081 dies. Port 8080 is a different app.

What the tests proved — and did not

18 / 18 unit tests: identical photos score 1.00, weak looks do not suggest, blur is dropped, Confirm closes the card, drafts can be compared but do not count as a named identify hit.

Not proved: Dean standing in the aisle himself. The next crop after Confirm scoring higher. Weapons. That a 76% suggestion is “probably Dean” in any statistical sense — 76% is pattern closeness, not probability.

6. Road A leftover — NIST FRTE, step by step

  1. Stop and choose. Either (i) buy a vendor who already has a FRTE scorecard and keep our Confirm UI, or (ii) become the vendor and submit our own matcher. Most teams pick (i). Submitting as Ghost Protocol is a company decision, not a bugfix.
    Not started
  2. If we buy: pick a vendor with a current 1:1 and 1:N card. Put their SDK on the Mac. Our glass still only shows labels and Confirm. We may say “matcher is Vendor X, see their NIST FRTE card.” We may not say “Specter is NIST FRTE.”
  3. If we build: throw away aHash as the matcher. We need a real face embedding model. Same /fleet/faces/compare door; different brain behind it.
    Blocked — we do not have this
  4. Detect a face, not a torso. FRTE assumes a face. Our upper-person crop is not that.
    Blocked
  5. Keep templates on the host. Browser stays labels-only. Confirm still sits in front of any name. We already believe this; we just do not have real templates yet.
  6. Only if we submit ourselves: build NIST’s offline Linux program (no internet, their API). Sign their agreement. Send it on their form. They allow about one try per track every four months.
    Not started — do not do this to “finish NIST”
  7. Wait. NIST runs it. A public card appears with miss / mix-up rates. That card is about the algorithm, not the bus install.
  8. Still do not put names on the live aisle picture or auto-page when someone walks by. A scorecard does not change that rule.

Optional later, different contest: FATE (quality, morphs) and PAD (is this a real face or a printout). Our blur check is not that. S2 may not even have depth. Do not say “liveness.”

7. Road B leftover — our own NIST-style evaluation, step by step

Govern = write the rules

Who is allowed to enroll a face? Who may click Confirm? Who may turn Alert on? How long do we keep crops? How do we delete someone?

  1. One-page watchlist policy. Not written.
  2. One human owns slide language. Same “do not say” list every time. Informal only.
  3. Deletion / retention for the face folder on the Mac. Not written.

Already doing: no names on live video, drafts never alert, Alert off unless someone flips it.

Map = say what this is for, and what can go wrong

Write two sentences: “This is for a named person of interest in an aisle or depot. Ordinary riders stay unnamed unless an operator names them.”

  1. Freeze that use case. In our heads, not on a signed page.
  2. Risk list: wrong Confirm, phone/print spoof, too many Review cards, leftover Face 1, sales saying “NIST.” Not written down.
  3. Say out loud: we do not detect faces, we detect people and crop high. That is a known hole. Said in this report.

Measure = keep score on our desk (not NIST’s)

This is a notebook, not a leaderboard. Never call these numbers FNMR, FMR, or FRTE.

  1. Every Confirm / Not them / New person written on the Mac (who, when, score, decision). Empty today.
  2. A folder of stills: Dean, not-Dean, same JPEG. How often does 55% light up the right person? Not run.
  3. How often do live people fail the blur/size gate? Not counted.
  4. Two operators, same 20 cards — do they agree? Not run.
  5. After Confirm, does the next crop of the same person score higher? Not run.
  6. Dean actually stands in front of S1, then S2. Not run for this write-up.

Already have: the gate, the percent, the unit tests, one live evening (76% / 61%).

Manage = what we do when it goes wrong

  1. Wrong Confirm: click Not them, drop that still, leave Alert off. Not written as a playbook.
  2. Face 1 leftover: name it or delete it before any film. Still sitting in the gallery.
  3. A switch to pause auto-find and Alert without rebooting cameras. Not a single switch.

Already doing: human must Confirm; we do not auto-enroll riders; we do not page on walk-by.

The order that will not waste a week

  1. You stand in front of the cameras. Name or delete Face 1.
  2. Write Confirm clicks onto the Mac (tiny audit).
  3. One-page policy + risk list.
  4. Run the still-folder and blur-reject counts. Call them operational, never FRTE.
  5. Write the wrong-Confirm playbook.
  6. Only if COTA asks for “identification, not suggestions,” talk about buying a FRTE-ranked matcher.

Do not start embeddings, FRTE paperwork, or liveness to “complete NIST.” That is Road A, and we have not chosen to walk it.

8. Repair, update, weaknesses — in English

Fix so a demo does not die

  • Glass on :8081 keeps dying. :8080 is not the glass. Need a real start-on-login.
  • If someone edits the face server, restart it, or Compare can show the COTA banner instead of faces.
  • Watcher was yellow during the test. Unrelated, but it will be in the shot.
  • Standing still can create several “Looks like Dean” cards. Annoying, not an alert.

Weak on purpose, or weak and we should say so

  • Not a face matcher. Lighting, pose, hat, back of head — miss or collide.
  • Not a face detector. Upper person box only.
  • 76% is not “probably Dean.” It is “these two thumbnails share a brightness pattern.”
  • A printout can look similar. No spoof check.
  • No fairness numbers. Cannot claim we treat groups equally.
  • Two lists of people (host gallery + local profiles) still both show. Easy to think there are two Deans.

9. Tomorrow morning (FR desk — still valid)

  1. Face server on port 8788.
  2. Glass on 8081, not 8080.
  3. Faces → Load + enroll Dean pack if empty. Leave Alert off.
  4. Open live Sight. Stand in view.
  5. Faces → Review. Confirm / Not them / New person.
  6. Expect 55–80% for live you vs pack photos. 100% only if it is the exact same file.
  7. Fleet: catch-all LaunchAgent should already be running. Do not PoE-cut while .h265.tmp is open.

10. Compiled fleet addendum — 14 August 2026

PoE practice, 6:00am one-shot boot, live health, S2 disk emergency, and the catch-all that now watches overlay. VMS is 192.168.86.54:8788. No secrets. Twin clocks: poe-unit-test-s1-s2.html.

02:42S1 T_nominal after 6am PoE-on
02:41S2 T_nominal (same T=00:00)
26%S2 overlay after prune (was 97% / 100%)

10.1 Last-known restore (run4 · 6:00am CDT one-shot)

Scheduled subagent + at 06:00 backup. Both ports were 0 W / adminEnable=2. Enabled together. T=00:00 unix 1786705338. Canary poe-run2-20260814T014243Z survived. No hollow CIDs. No reconcile.

GateS1S2
Ping / SSH00:52 / 00:3100:52
ghost_core + RF-DETR00:5200:52
First H.26501:5901:59
T_nominal02:42 · 3 min02:41 · 3 min
Good CIDaaf26107… 1771/181265acd4e6… 1794/1903
Layerlayer1_rfdetr_thermal_safe · ENCODE_ONLY=0 · H.265 1280×720 · env H265_FPS=5

Safe halt (later the same night, then this morning’s 6am power-on): runc kill → ghost gone 4–8s → shutdown -h → SSH dead 9–14s → then Cisco adminEnable=2. Halt still draws watts. Do not PoE-cut while .h265.tmp is open.

10.2 RF-DETR + encode FPS (measured ~14:12 CDT, 8h09 up)

DLC present both (ghost_protocol_transit.dlc 135 MB). .latest_dets.json age 0.2s. Empty scene (n=0) is allowed; MQTT had person_count=3 on S1 the same minute. .rfdetr_status is live key=value, not JSON.

S1S2
Sidecar temp69.6°C normal78.8°C warm
ghost_core SoC78–90.7°C CRITICAL spikes79–82°C · VIDEO_FIFO drops ~355k
ffprobe fps25/1 — ignore25/1 — ignore
Sidecar fps tag15.0 (stale vs env 5)
Measured wall fps6.82 (412 NALs / 60.38s)~4.86 (seq over 55.5s)

RF-DETR B+ · encode B. Continuous justified H.265. Thermal stalls keep it off the 15 tag.

10.3 Why you saw no storage warning

WatcherWhat it watchesWhy 97% was silent
Platform Watcher (launchd, 11k ticks)Mac /, MQTT keys exist, go2rtc from 87.xdisk_host is this Mac. edge_s1/s2 circuits open because 87.x cannot L2-probe streams.
Soft WatcherStream MCP restartNo overlay probe.
MQTT storage_healthrecorder stringStill OK on both after S2 was 97%.
Edge retentionoldest chunk_*.h265Does not delete pre_event_thermal_* (~20 MB / 10s) or 13 GB logcat.log.
Reconcile agentghost_core missingNo-op while pipeline is up and disk is dying.

S2 today folder was 42.3 GB of pre_event_thermal_* vs justified chunks. That is the same class of overlay-full that previously left a blue LED after PoE.

10.4 What we did about it

  1. Manual S2 prune: 2607 thermal files (43.5 GB). df stayed 100% until logcat.log (13.3 GB) was truncated + caches dropped → 26% / 56.9 GB free. Ghost 1794/1903 stayed up.
  2. Catch-all first tick also cleaned S1: 1452 thermal + 4.4 GB logcat → overlay 64% → 38%.
  3. Shipped scripts/fleet-catch-all.py + LaunchAgent com.ghostprotocol.specter-fleet-catchall (running). MQTT 15s · SSH df 30s. Decide with no LLM. Mitigate prune-keep-30 + truncate logcat. Notify macOS. Status file ~/.specter-vision/platform/fleet-catchall-status.json.
  4. Playbook repair-edge-storage. Platform Watcher now reads that status file. It does not restart go2rtc for disk.
  5. M8 shutdown listener landed: scripts/m8-shutdown-listen.py (button hold / GPIO20 / CAN). Missing box exits M8_ABSENT.

Catch-ahead is 80% overlay / 200 pre_event / 512 MB logcat, not 97%. MQTT OK cannot hide a 90% df.

10.5 Still later (not a bake tonight)

Doctrine stays: deterministic probe → playbook → verify. No 15s Grok loop.