Eye tracking used to mean a lab, a grant, and a graduate student. For short video it now means choosing the right rung of a four-step ladder. Here is each rung, with honest accuracy numbers.

When I tell creators they can eye track their videos, the first reaction is usually a mental image of electrode caps and university basements. That image is thirty years out of date. Eye tracking today is a ladder with four rungs, from free-and-rough to lab-and-rigorous, and the right rung depends on one thing: the question you are asking your video.
This guide walks the ladder honestly, including the accuracy numbers each rung actually delivers, because "we track eyes" marketing has a habit of skipping them.
Play your video and honestly notice where your eyes go, then play it at half attention while doing something else and notice again. This costs nothing and catches only the loudest problems, because you know where you are supposed to look. Your gaze is contaminated by authorship; a stranger's is not. Useful as a first filter, never as evidence.
Several research tools track gaze through a laptop webcam at roughly 2 to 4 degrees of accuracy. The catch for short video is context: your vertical video will be watched on phones, in feeds, at arm's length, and desktop viewing changes posture, distance, and frame size. Fine for websites; a mismatch for Reels.
Real viewers watch on their own phones while the front camera tracks gaze, calibrated once per viewer (Jeena samples gaze at 15 frames per second). Accuracy is region-level, which is the level creator questions live at, and the viewing context is the real one. A panel of 5 to 10 viewers costs around ten euros and returns within about a day.
Research-grade rigs resolve 0.3 to 0.8 degrees, enough to see which word was read. If your question is genuinely word-level (legal disclosures, package fine print), rent one. For "did they look at my face or my caption", the lab adds cost and artificiality without adding answerable precision.
Every short-video question I have ever seen a creator actually ask is a region question: face or background? product or texture? caption read or ignored? on-screen or gone? Region questions are answered identically well at rung 3 and rung 4, and rung 3 answers them on the device and in the posture your video will really face.
The rungs are not a quality ranking. They are a specificity ranking, and short video's questions bottom out at regions. Spending lab money on a Reel is like renting an electron microscope to check if your soup is too salty.
One: export the exact cut you intend to post, captions burned in, because you must test the exact frame the feed will show (in our aggregate data captions draw only about 7% of gaze, but a badly placed one can hijack an entire opener, so their presence and position change the video). Two: upload to Jeena and set the goal that matches your intent (Views, Sales, Pitch, Followers), since the analysis reads attention against your goal. Three: let the panel watch; 5 to 10 real viewers on their own phones, each calibrated, each finishing with a short impressions survey. Four: open the report, typically within a day.
Read it in this order: visibility map first (what was never seen, the fastest shock), then the attention heatmap scene by scene (what won each auction), then the wow-moments chart against your intended beats, then the three recommendations, which arrive with timestamps. The order matters because the visibility map recalibrates your ego before the heatmap tempts you to rationalize.
Expect region-level findings with one-line fixes: a caption that needs to move off a face, a backdrop that needs desaturating, a static stretch that needs one interrupt, an ending that needs to mirror the opener. In everything we have published, the fix was almost never "reshoot the video"; it was a crop, a caption move, or a re-cut of two seconds.
And test the fix, not just the diagnosis: the ten-euro price exists so that testing twice (before and after the change) is still cheaper than one boosted post. If your first test is on a hook, here is the dedicated hook protocol; if you want the theory of what you will see, start with how people actually watch short videos.
Upload your video to Jeena. Real viewers watch it on their phones with the front camera on, and the report shows an attention heatmap, a visibility map, a wow-moments chart, a summary of how viewers perceived it, and three concrete recommendations, typically within a day.
No "schedule a call." No sales rep. Upload, get your report.
You can approximate it: watch your own video while honestly tracking your gaze, then have a few strangers do the same and tell you where they looked. It is better than nothing and catches loud problems, but self-report is unreliable (people do not know where their eyes went) and your own gaze is biased by knowing the video. Free methods are a filter; measured gaze from real viewers is the evidence.
Fewer than intuition says. Attention problems recur: a caption blocking a face blocks it for everyone, so the recurring leaks surface within the first handful of viewers. A panel of 5 to 10 catches the findings that matter, which is also what keeps a test around ten euros instead of a five-figure lab study.
Region level, roughly 2 to 4 degrees of visual angle, which smartphone front-camera tracking delivers. That resolves the questions short video poses: face versus caption versus product versus background, and on-screen versus gone. Word-level precision (0.3 to 0.8 degrees, infrared lab rigs) is only needed for questions like which word of fine print was read, which short-form creative decisions almost never turn on.
Jeena is a neuromarketing platform for short-form video. Real people watch your video on their phone with the front camera on. Jeena captures their gaze direction, blink rate, eyebrow raises, and their impressions of the video in a short survey afterward. You receive an AI-powered report with an attention heatmap, a visibility map, a wow-moments chart, a summary of how viewers perceived the video, and three specific recommendations for making the video work harder.
Jeena uses smartphone front-camera gaze tracking. Each engager calibrates once, then watches your video. The platform records where their gaze lands frame by frame, flags moments of surprise from facial expression, and combines that with a short impressions survey afterward. The result is a per-second timeline of what real viewers actually looked at and felt, plus a summary of how they perceived the video overall.
A typical test costs around ten euros. See the pricing page for current rates.