Building a Cinematic Machine Vision HUD with YOLO, TAPIR, and MediaPipe
An exploration into the intersection of heavy machine learning and cinematic UI design, creating a 5-phase visual state machine HUD.

What happens when you combine state-of-the-art neural networks with the visual aesthetics of a 90s cyberpunk film? Over the past week, Tekano R&D has been exploring the intersection of heavy machine learning and cinematic UI design.
We built a real-time, zero-latency heads-up display (HUD) that doesn't just passively track a user—it interrogates them. By stacking YOLOv8 (for human silhouette segmentation), DeepMind's TAPIR (for temporal particle tracking), and MediaPipe (for dense 3D facial geometry), we created a 5-phase visual state machine.
When a subject steps into the camera frame, the system physically reacts:
- Target Acquisition: YOLO isolates the human silhouette, triggering a massive visual burst of outward-streaming matrix particles and pulsing edge-detection lines.
- Geometry & Emotion: MediaPipe maps 478 facial landmarks, analyzing the subject's smile ratio and calculating live emotional states.
- rPPG Vitals Scan: We implemented Remote Photoplethysmography (rPPG) to isolate micro-color changes in the subject's forehead. A secondary HUD panel drops down to the chest, rendering a live ECG line and real-time BPM (Heart Rate).
- Thermal Override: Once biometric lock is achieved, the background is preserved while the YOLO silhouette is dynamically hijacked by an aggressive, glowing INFERNO thermal color map.
- Interactive Telemetry: Opening the mouth triggers an interactive state, projecting tracking lasers from the irises to a physical X/Y/Z coordinate target on the screen.
This project was an experiment in making machine vision tools combined can perhaps feel like a machine thinking.