Skip to content

Validation / Methodology

Methodology, evidence, and the road to proof

The full record behind the read, written by the side trying to falsify the product. It scores every dimension of validity honestly. It reports the failures next to the passes. It lists the pre-registered program and the owner gates that would earn the word impact. Pre-registered means the pass marks were published before any study ran, and those studies have not run.

What you can rely on today

Five things are true right now, before a single further study runs. Each is stated at the strength the evidence supports, and every number below traces to a named source.

A reliable, reproducible engine

One scoring core, ported and checked to the integer: 78 of 78 cells match. It holds steady when the input is nudged. It is deterministic, so the same footage returns the same read, every time.

We measure what we can see, and say so

The engine reads the room at group level: where attention points, motion, expression geometry, and how in sync the room is. The read is aggregate only. That means room totals. No one person is scored, no per-person profile is built, and no one is re-identified. No emotion is inferred.

Private by construction

Footage is processed, then deleted as soon as the analysis finishes. No faces are kept. Only the aggregate numbers are kept. There is no record of any one person, because none is ever created.

In step with the science and the regulators

The signals rest on published work in facial action and behaviour measurement. The engine infers no emotion. Regulators hold that emotion recognition has no scientific basis, and we agree. That position is set out in EU AI Act Article 5(1)(f). A runtime gate fails closed on emotion-framed output.

Proven in the open, not asserted

Every claim here is stated at the strength the evidence supports, with named sources. The harder claim is that a read predicts a real outcome. We hold that one behind a pre-registered program rather than assert it. Pre-registered means the design and the pass marks are published before the study runs. We hold the word until it is earned.

What we are not claiming yet

That a Proof of Impact score predicts a real-world outcome you would call impact. Today that is a measured appearance correlate: it travels with a thing, and it is not that thing. So it is a hypothesis under test, not a proven result. The full validity scorecard is below, unsoftened. So is the pre-registered program that will test it. Pre-registered means the pass marks were published first, and those studies have not run.

What the word Proof can and cannot mean here

The product is named Proof of Impact. Here is the honest split between what is proven, what is only measured as an appearance correlate, and what is still a hypothesis.

Proven

That the software is reliable and reproducible. There is one scoring core, ported faithfully, matching on 78 of 78 cells. It is deterministic, and it is stable when the input is nudged. This is about code. It is not about what the scores mean.

Measured

Behavioural appearance correlates: where faces point, motion, expression geometry, and room synchrony. A correlate travels with a thing and is not that thing. These reads are aggregate and at group level, so no one person is scored. They measure appearance. They do not measure attention, emotion, or impact.

Hypothesized

That these correlates track, or predict, a real outcome anyone would call impact. No study links a score to an outcome yet, so this is a hypothesis, not a result. Until it is tested, impact is not proven.

Validity scorecard

Every dimension of validity, scored honestly. Not established means no supporting evidence exists yet. A Not established is never softened into a Partial. This is the spine of the dossier, and of the go or no-go on the landing-page link.

DimensionStatusWhy
Construct validityNot establishedThe engine measures appearance correlates, not the latent construct impact or engagement. The one external-criterion test of its expression substrate (FACS AU4/AU6/AU12 versus OpenFace) fails to converge (CCC below 0.6 on all 9 clips), so any expression-derived read is a behavioural correlate, not a calibrated construct measurement.
Content validityPartialThe aggregate attention, energy, and synchrony signals are grounded in cited behavioural-science literature and the product scopes its claims to correlates, which supports partial content validity. It is partial and not established because no independent content-expert panel has judged coverage, and the expression component inherits the failed FACS convergence.
Criterion and predictive validityNot establishedNo study links any score to a real outcome. The Altitude venue sessions carry no outcome measure, so this cannot move on that footage. A commissioned predictive protocol is published; until an outcome study runs and passes its pre-registered gate, this stays Not established. This is the dimension that would license the word impact, and it is not earned.
Ecological and external validityNot establishedAlmost all evidence is benchmark, studio, or semi-synthetic. On the single real-venue test the skin-tone proxy demonstrably collapsed. Evidence that does not transfer to the deployment domain is not validation of the deployment domain. The venue program is pre-registered but blocked on consent, compute, and a working GPU pipeline, so no venue study result exists yet.
Reliability (measurement stability)PartialInternal reliability is shown: internal consistency (sub-score correlations with expected signs) and stability to 5 percent landmark noise, plus deterministic, reproducible output. But split-half, camera-concordance, and test-retest reliability on venue footage are pre-registered and unrun, so measurement reliability in the deployment domain is not established. Reliability is not validity.
Software reliability and reproducibilityEstablishedThe shipped TypeScript core matches the canonical Python engine on 78 of 78 cells exact (the 5 core plus single-person indices), output is deterministic (0 variance over repeated runs), and the surfaces conform to 0 delta on the FACE-ONLY core at a shared frame rate. Two scope limits: parity does NOT cover the multi-person aggregate path (deriveMultiPersonScores) that produces the shipped room read, and body-dependent scores still diverge up to 10 points across surfaces because body runs only on web (D8, open). This is Established as software reliability and reproducibility ONLY. It is explicitly not construct, criterion, or ecological validity: two programs agreeing on a number is not the number being right.
Measurement invariance and fairnessNot establishedMST-E studio portraits show a real darker-skin detection deficit around the 20 px floor (studio, and the reported intervals overstate precision until a person-level recompute runs). On real venue footage the tone proxy collapsed entirely, so no venue fairness gap is measurable today. The mitigation is studio-validated only. Venue fairness is Not established and is a pre-registered study.
Adversarial resilienceNot establishedThe engine carries abstain and confidence logic by design, but it has not been adversarially tested against the pre-registered negative controls (empty room, temporally shuffled frames, a single looped still). Until those controls run and the engine is shown to abstain rather than emit confident reads on them, adversarial resilience is Not established.
IndependenceNot establishedEvery result here is vendor-run. There is no third-party or academic replication of any figure. Independence is Not established and the page claims no external validation until it exists.

Claim by claim

Each card states the claim at the strength the evidence supports. It gives the method. It gives the effect size, which is how big the measured difference is. It gives the honest uncertainty. It lists the threats a hostile expert would raise. It names the sources every number traces to.

The shipped single-person scoring math is a faithful, reproducible port of the canonical engine.

Established

Dimension: Software reliability and reproducibility

Evidence
The boost-app TypeScript scoring core matches the canonical Python engine on 78 of 78 cells, exact to the integer, across every canonical scenario and context (the 5 core plus 8 single-person indices). Output is deterministic: 0 variance over 10 repeated runs of the full pipeline. The three surfaces conform to 0 delta on the FACE-ONLY core at a shared frame rate.
Method
A parity shim runs the shipped core and the canonical Python engine on shared single-person fixtures and compares all 13 score fields cell by cell; determinism and conformance are measured by repeated and cross-surface runs on fixed input.
Effect size and uncertainty
78/78 exact match on the 5 core plus single-person indices; 0.0 variance; 0.0 FACE-ONLY cross-surface delta at a shared frame rate.
Threats to validity
This is software reliability and reproducibility, not validity. Two scope limits carry forward: (a) parity covers the 5 core and single-person indices only. The multi-person AGGREGATE scoring path (deriveMultiPersonScores), which is the shipped Proof of Impact room read, is NOT covered by the 78/78, and its parity is a separate, unrun exercise. (b) Conformance is 0 delta on the FACE-ONLY core; body-dependent scores diverge up to 10 points (Stress most) across surfaces because body extraction runs only on web (D8, open). Two programs agreeing on a number is not the number being right.

Sources: docs/validation/engine-parity.md, docs/engine/02-conformance.md, docs/validation/03-local-validation-results.md

The engine is internally consistent and stable to small perturbation, but its venue reliability is unmeasured.

Partial

Dimension: Reliability (measurement stability)

Evidence
Sub-score correlations carry the expected signs and strength (Composure to Stress r = -0.92, Momentum to Engagement r = +0.83). Under 5 percent landmark noise, 16 of 16 scores stay within 15 points, most moving 0 to 2. Split-half, camera-concordance, and test-retest reliability on real venue footage are pre-registered but not yet run.
Method
Internal consistency and noise sensitivity are computed on synthetic scenarios in the local validation suite; venue reliability studies are specified in the pre-registration and are pending.
Effect size and uncertainty
r = -0.92 and +0.83 on the two headline pairs; 16/16 scores stable within 15 points under 5 percent noise. Venue split-half ICC: not yet measured.
Threats to validity
Internal consistency and noise stability are reliability, not validity, and are measured on synthetic inputs. Deployment-domain reliability (split-half on real footage, camera concordance) is the open piece.

Sources: docs/validation/03-local-validation-results.md, docs/validation/altitude-preregistration.md

On a crowd detection benchmark, recall falls off below about 20 px and tiling roughly doubles crowd recall. This bounds detection on a benchmark, not on the venue.

Not established

Dimension: Ecological and external validity

Evidence
On WIDER FACE, tiled recall is 0.64 at 15 to 20 px, 0.78 at 20 to 30 px, and 0.88 at 60 px and up; tiling lifts dense-crowd recall from 0.19 to 0.43 at 101 or more faces per image. The room count is therefore a detected lower bound that undercounts dense rooms by roughly half.
Method
Recall at IoU 0.5 on a stratified 400-image sample (21,201 faces) of the WIDER FACE validation set, single-pass versus tiled, by face size and image density.
Effect size and uncertainty
Tiled recall 0.64 (15 to 20 px) to 0.88 (60 px and up); dense-crowd recall 0.19 to 0.43 with tiling. Faces cluster within images, so intervals should be image-clustered.
Threats to validity
WIDER FACE is a public benchmark, not the deployment venue, and has no demographic labels, so it supports no ecological-deployment or fairness claim. The size and density envelope is a benchmark measurement only.

Sources: docs/engine/crowd-scale-result.md

On real venue footage the skin-tone proxy collapsed, so a venue fairness gap cannot currently be read. This is a published failure.

Not established

Dimension: Measurement invariance and fairness

Evidence
On the Altitude footage every harvested face was classified into a single tone band (proxy_valid: false); median ITA around -71 to -73, far past any real skin value, because indoor GoPro exposure crushed luminance. Even a normalization pass did not restore a trustworthy tone axis.
Method
Pseudo ground truth built from the footage: detect large confident faces, tag tone by ITA proxy on the crop, then re-detect at rendered crowd sizes. The tone axis was checked for validity, not assumed.
Effect size and uncertainty
Populated tone bands: 1 (degenerate). Median ITA around -72, outside the plausible skin range. No venue per-tone gap is computable as-is.
Threats to validity
This is the honest negative result and the reason the venue fairness program exists. Studio MST-E fairness does not transfer here. Closing this is pre-registered Study 8, gated on consent and compute.

Sources: docs/engine/altitude-footage-fairness-result.md, docs/validation/altitude-preregistration.md

On studio portraits with ground-truth Monk labels, the detector loses darker faces faster at the crowd-critical size. This is studio, not venue.

Not established

Dimension: Measurement invariance and fairness

Evidence
On MST-E, at the 20 px floor darker-skin recall is 0.52 against 0.74 for lighter faces (gap about 0.22); parity arrives by 30 px, and all tones collapse together by 15 px. A CLAHE plus upscaling mitigation lifts darker 20 px recall to about 0.70 and halves the gap.
Method
Detect each MST-E portrait, tag by ground-truth Monk band, render at crowd pixel sizes, re-detect; recall per band per size over the full 1543-image run.
Effect size and uncertainty
Darker recall 0.52 vs lighter 0.74 at 20 px (image-level; person-level interval pending). Effective sample about 19 people, not 1543 images.
Threats to validity
Pseudoreplication: the reported intervals treat 1543 images as independent when there are about 19 people; a person-level recompute is required and gated on re-acquiring the licensed data. And this is studio, not venue: on venue footage the proxy collapsed. Not established for the deployment domain.

Sources: docs/engine/mste-fairness-result.md

Cross-camera identity does not hold when cameras see different people; the multi-camera claim is not shipped.

Not established

Dimension: Reliability (measurement stability)

Evidence
On 9 people with a synthetic flipped second view, the disjoint-population false-merge rate is 2.8 percent at the deployed threshold and up to 26 percent at the low end, and the hardest same-demographic pair is unseparable at any usable threshold. Detection recall (0.933) clears its floor; the identity claim does not.
Method
Full enumeration of all 252 identity-disjoint splits of the 9 people, EER threshold on a calibration split, reported on held-out; plus a live end-to-end run on Modal.
Effect size and uncertainty
Disjoint false-merge 2.8 to 26 percent (does not clear a 1 percent bound). Effective sample 9 people; the 252 splits are correlated, not 252 samples.
Threats to validity
Pseudoreplication (9 people, 252 correlated splits) and a synthetic second view (a horizontal flip, not a real angle). The honest recommendation is to not ship the multi-camera identity claim; it waits for labeled multi-view data.

Sources: docs/engine/sface-validation-result.md

The expression substrate does not converge with an independent FACS criterion, so expression-derived reads are correlates, not calibrated measurements.

Not established

Dimension: Construct validity

Evidence
Against OpenFace, AU4 and AU6 concordance (CCC) stays below 0.6 on all 9 clips; AU12 tracks direction well (Pearson up to 0.96) but fails calibrated CCC. Zero of nine clips pass the priority-AU gate.
Method
Concordance correlation (CCC, gate 0.6) between the engine AU substrate and OpenFace on the convergence corpus; the corpus itself rates only 3 of 9 clips as usable material.
Effect size and uncertainty
CCC below 0.6 on all 9 clips for AU4 and AU6. The one external-criterion test, and it fails.
Threats to validity
The corpus is weak, which confounds the result, but even on the usable clips AU4 and AU6 do not converge. The scores remain behavioural correlates, which the product already states.

Sources: docs/validation/03-local-validation-results.md

The multi-hour capability claim is not yet substantiated on venue footage; no continuous two-hour session even exists in the corpus.

Not established

Dimension: Reliability (measurement stability)

Evidence
The read-only inventory shows 10 distinct capture sessions across 2 days, with the longest continuous single recording at 50.2 minutes and 0 clip groups safely concatenable. A two-hour endurance run would be an assembled input, not a continuous session, and is blocked on compute and a working chunked GPU pipeline.
Method
ffprobe metadata inventory plus creation_time clustering (no frame decoded); the multi-hour pipeline study is pre-registered as an endurance and cost test.
Effect size and uncertainty
Longest continuous view 50.2 minutes; 0 of 13 clip groups concatenable; multi-hour run not executed.
Threats to validity
Endurance is not a validity claim about the read. The corpus cannot supply a genuine continuous two-hour single-audience recording, so any multi-hour input is a stitched endurance harness, labeled as such.

Sources: docs/validation/altitude-corpus-manifest.json, docs/validation/altitude-preregistration.md

The shipped read is aggregate attention, energy, and synchrony, framed as behaviour and geometry correlates, with no emotion inference.

Partial

Dimension: Content validity

Evidence
The aggregate room report emits only attention on the speaker, room synchrony, a session-phase energy arc, and behavioural pillars, with no per-person affect. An EU AI Act Article 5(1)(f) runtime gate fails closed on emotion-framed output in workplace and education contexts.
Method
Allowlist projection in the aggregate report plus a runtime emotion gate with unit tests; the framing is grounded in cited behavioural literature.
Effect size and uncertainty
No numeric effect size; this is a scoping and compliance posture, grounded but not validated by an external content-expert panel.
Threats to validity
Content grounding is not construct or criterion validity. The framing reduces legal and overclaiming risk; it does not prove the correlates measure the underlying construct.

Sources: docs/legal/access-compliance.md, app/what-it-measures/page.tsx

No score has ever been linked to a real outcome, so the product cannot honestly use the word impact yet.

Not established

Dimension: Criterion and predictive validity

Evidence
No criterion or predictive study exists. The Altitude sessions carry no outcome measure (session ratings, evaluations, net promoter, recall, behaviour). A predictive protocol is commissioned and published; it runs only if outcome data is found for these sessions.
Method
Gap statement. The commissioned predictive protocol is specified in the results document.
Effect size and uncertainty
No effect size exists because no outcome variable has been linked.
Threats to validity
This is the load-bearing gap behind the product name. Until a pre-registered predictive study passes, criterion and predictive validity is Not established and the page says so.

Sources: docs/validation/altitude-results.md, docs/validation/altitude-preregistration.md

Every result here is vendor-run; there is no independent replication.

Not established

Dimension: Independence

Evidence
No third-party or academic group has replicated any figure on this page. The program is designed to be replicable (pre-registered thresholds, named harnesses, committed artifacts), but replication has not happened.
Method
Gap statement.
Effect size and uncertainty
Zero independent replications.
Threats to validity
Vendor-run validation is the weakest form of assurance. The page claims no external validation until an independent review exists.

Sources: docs/validation/validity-scorecard.json

The engine has not been shown to abstain on negative controls; that test is pre-registered and unrun.

Not established

Dimension: Adversarial resilience

Evidence
The pre-registration requires the engine to abstain or return null on an empty room, a temporally shuffled session, and a looped still frame. These have not been run on venue footage. If the engine emits a confident room-energy or attention read on any of them, that is disconfirming and will be published.
Method
Pre-registered Study 10; not yet executed (gated on consent and compute).
Effect size and uncertainty
No result yet; required behaviour and pass or fail thresholds are fixed in the pre-registration.
Threats to validity
Abstain logic existing in code is not the same as passing an adversarial test. Until the controls run, adversarial resilience is Not established.

Sources: docs/validation/altitude-preregistration.md

How a skeptic would attack this, and our answer

The strongest version of the case against this product, stated by us, with our honest answer. If any answer reads as a dodge, it is a bug in this page.

You call it Proof of Impact, but you never measured impact on anyone.

Correct, and we say so at the top. The engine measures behavioural appearance correlates: where faces point, motion, expression geometry, and synchrony. A correlate travels with a thing. It is not the thing, and none of this is impact. Criterion and predictive validity is Not established, which means no supporting evidence exists yet. No score has been linked to a real outcome. The panel above reconciles the name.

Your evidence is benchmarks, studio portraits, and synthetic data, not the venue you deploy in.

Also correct, and it is why the venue program exists. On the single real-venue test the skin-tone proxy collapsed. A proxy is an indirect stand-in for the thing you want. Evidence that does not transfer to the deployment domain is not validation of that domain. So ecological validity is Not established. The venue studies are pre-registered, which means the design and the pass marks were published before any study ran. Those studies have not run.

Your ground truth is a human watching the same faces the engine watches, so agreement is circular.

We flag the circularity in the pre-registration. Coders judge appearance, and the engine measures appearance. So agreement is evidence about a shared appearance construct. It is not proof of attention, and it is not proof of impact. It is convergent evidence with a stated ceiling. It is never a criterion for impact.

Your samples are tiny and clustered: one event, one audience, a handful of people.

The effective sample is about 10 capture sessions inside ONE event with one recurring audience. The cross-camera work used about 9 people, and MST-E has about 19. Our reporting unit is the session or the person. We report the effective n in sessions and people, never in frames. The session-level cluster bootstrap and the multiplicity correction are the pre-registered method for the venue program. Multiplicity correction means an adjustment for testing many things at once. That program has not run. The studies that HAVE run report intervals at the image or split level. Each is flagged for a required, still-pending recompute that groups the data by person. The MST-E person-level recompute is one of those.

Two code paths matching to the integer is not science.

Agreed. The 78 of 78 parity result is about code. So are the determinism and conformance results. All three sit under software reliability and reproducibility, not validity. Two programs agreeing on a number is not the number being right. The build check fails if a real-world-validity claim rests only on parity.

Confounds: lighting, camera angle, seating depth, session content, and culture all move your signals.

They do. The sensitivity study varies the arbitrary knobs. The negative controls test whether the engine abstains on empty rooms and shuffled frames. To abstain is to return no number at all. The known-groups check ties signals to real events. None of these has run yet. Until they do, any read remains confounded by lighting, angle, seating, content, and culture.

No one independent has checked any of this.

True. Every figure is vendor-run: we ran it ourselves. Independence is Not established, and we claim no outside check. We built the program to be replicable. The thresholds are pre-registered, so we published them before any study ran. The harnesses are named and the files are committed. That is deliberate, so an outside group can check the work.

Regulators say emotion inference from faces has no scientific basis.

We agree with that position and do not infer emotion. The EU AI Act Article 5(1)(f) and Recital 44 cite the lack of scientific basis for emotion recognition. A runtime gate fails closed on emotion-framed output in EU and UK workplace and education contexts. The read is aggregate attention, energy, and synchrony. Aggregate means room totals, with no one person scored. Each is a correlate, not an emotion.

The venue-footage program

Real audiences in a real venue: the domain the current evidence lacks. The program is pre-registered before any result, so the thresholds cannot be moved to fit the data. It has not run: it is blocked on the owner gates below.

10

distinct capture sessions, one event, one recurring audience

0/10

pre-registered studies passed

50.2m

longest continuous view (no continuous two-hour session exists)

Effective sample size is about 10 distinct capture sessions inside ONE event with one recurring audience, not 39 files and not the millions of frames. Clip-number grouping over-merges (simultaneous multi-camera captures), so no continuous two-hour session exists in clean form.

Read the detail:docs/validation/altitude-preregistration.mddocs/validation/altitude-results.mddocs/validation/altitude-corpus-manifest.json

Status of the validation gate

Software reliability is established today. Three gates for the setting we ship into are still pending: ecological validity, criterion validity, and fairness. Each is pre-registered, so its pass marks were published before any study ran. Those studies have not run. So the stronger claim, that a read predicts impact, is not yet earned. This status is published in the open, never hidden.

  • Venue studies not passed (0/10 pass, rest pending or failed).
  • Consent basis for research analysis of the footage is not documented.
  • No deployment-domain validity dimension (ecological, criterion, or fairness) is Established.

This status comes from the committed results file, not a hand-set switch. That file is docs/validation/validity-scorecard.json. Artifact version v0.2.0, dated 2026-07-31.

Owner gates

These are not engineering choices. They are decisions and reviews only the owner and counsel can close, and they block the studies, the publication, and the word impact.

open

Consent basis for research analysis of the Altitude footage, plus retention and deletion schedule

Blocks: all venue studies

open

Counsel review and sign-off before publish

Blocks: publishing the dossier beyond an internal artifact

open

Independent third-party or academic review

Blocks: any claim of external validation

open

Whether outcome data exists for the Altitude sessions

Blocks: criterion and predictive validity

open

Compute budget for the multi-hour and full-corpus runs, and the go-live decision

Blocks: multi-hour and full-corpus studies

Expert appendix

Statistics posture. Our reporting unit is the session or the person. We report the effective sample size in sessions and people, never in frames or chunks. The venue program has a PRE-REGISTERED method: a session-level cluster bootstrap, plus a Benjamini-Hochberg correction for testing many things at once. That program has not run. Studies that have already run report intervals at the image or split level. MST-E is one of them. Each is flagged for a required, still-pending recompute that groups the data by person. We do not present these as person-level confidence.

Errata. We audited the earlier evidence, and we corrected it. Software parity, conformance, and determinism now sit under software reliability. They are not validity. The MST-E intervals were flagged for pseudoreplication. That is counting about 19 people as if they were 1543 images. A person-level recompute is the required next step. It is still pending. The 57-face MST-E pilot figures are superseded. We added effect-size framing to the tone and size comparisons. We also flagged that many things were tested at once. Each corrected document carries its own errata note.

No fabricated authority. This work carries no certification. It has had no peer review. It has had no independent audit. The engine infers no emotion. Regulators hold that emotion recognition has no scientific basis, and we agree.

Full record. The record behind every figure on this page sits in the source files. Those are the pre-registration, the corpus manifest, the full results, the validity scorecard artifact, and each corrected source document. Where this page and those files differ, the files rule.