It fits the person in front of it
No two faces, cameras, or seating positions produce the same picture, so we do not ship one frozen model and hope. A short calibration fine-tunes the engine to the individual user, on the machine they are actually using.
It keeps learning after calibration
Real use tells the engine when it was probably wrong. We feed that back in to keep improving during normal use, instead of sending people back to a calibration grid every time something drifts.
It runs on about one CPU core
Capture, inference, and smoothing together. No GPU budget, no dedicated silicon, and enough headroom left over for the product the gaze signal is feeding.