Behavioural capture stack
What is read from the video?
- Inputs
- Any modern video: venue AV, GoPro, iPhone, webcam. No wearables.
- Computation
- Frames are sampled, not read one by one, so compute stays bounded on a long session. A 2 minute clip runs at 12 frames a second, a 10 minute clip at 2, an hour at one frame per 4 seconds. On each sampled frame MediaPipe FaceMesh reads 468 landmarks per detected face, MediaPipe Pose reads body keypoints, and Whisper-derived features read pitch, pace and loudness. Face embeddings are not stored. Only the scoring numbers survive the pass.
- Output
- One record per sampled frame: the face slot, the smile, surprise and neutral probabilities, and an attending flag. Attending is true when the head points at the stage. That means within 25 degrees up or down, and 30 degrees left or right.