Back to blog
GuideGuideJuly 28, 2026

Eye tracking for short videos: how it works, what it shows, what it costs

You cannot see where viewers look from a view counter. Here is the complete picture of eye tracking for short-form video: the technology, the workflow, the report, the price.

A soft-3D pink phone playing a creator video with a gaze heatmap over her face, a beam connecting the front camera to a floating eye, and a pink report card with charts beside it. The image conveys the front camera watching the watcher and producing a report.

Why I wrote this

I keep posting some version of the same wish: I wish more creators used eye tracking platforms for short videos before publishing. And every time, the same questions come back. Is that a real thing? Does it need a lab? Does it work for a fifteen-second video? What does it cost? This page is the complete answer, written once, so I can stop answering in fragments.

The short version: yes, it is a real thing; no, it has not needed a lab since front cameras got good enough; yes, short videos are exactly what it is best at; and a test costs around ten euros. The long version is below, including the honest parts about accuracy that most tool pages skip.

What eye tracking for short videos means

Eye tracking means recording where a viewer's gaze lands, moment by moment, while they watch something. Applied to short video, it answers the questions analytics structurally cannot: not how long people watched (retention tells you that), but what they were looking at while they watched, what they never saw at all, and what was on screen at the exact moment they left. A view counter records the verdict. Eye tracking records the deliberation.

For short-form specifically, the deliberation is brutally compressed: sharp human vision covers roughly two degrees of the visual field, viewers spend a handful of fixations per scene, and the swipe option is always one thumb-twitch away. Whether those few fixations land on your face, your caption, or your kitchen cabinets is frequently the whole difference between videos that perform and videos that do not, and it is invisible from the outside.

The three ways to eye track a video

  • Infrared lab rigs: maximum precision, minimum realism
    Research-grade trackers hit 0.3 to 0.8 degrees of accuracy, fine enough to see which word someone read. They also require a facility, per-participant sessions, and a viewing setup nothing like a phone in bed. Agencies use them for TV-ad research at four-to-five-figure study costs.
  • Webcam tools: browser-based, desktop-bound
    Computer-vision tracking through a laptop webcam, typically 2 to 4 degrees of accuracy. Good for websites and desktop UX. The mismatch for short video: your audience does not watch Reels on a laptop, and vertical video on a desktop monitor is not the real viewing situation.
  • Smartphone front-camera panels: the real context
    Real people watch your video on their own phones with the front camera tracking gaze (Jeena samples at 15 frames per second after a one-time calibration). Region-level accuracy, which is the level short-video questions live at, in the exact device, posture, and feed context the video will actually face.

What this measures at scale (our own audit)

Numbers from our audit of 88 real short videos (~530 eye-tracked viewer sessions): of on-screen gaze, 47.7% goes to the subject, 25.5% leaks to backgrounds (more than the 17.7% faces get), and captions draw just 7.4%. And the videos viewers watch with the most visible concentration are frequently the ones they rate lowest, staring marks effort, not enjoyment.

The full findings, including that last inversion, are in the Attention Paradox study. The point here: these are the kinds of facts a report surfaces about YOUR video specifically.

The honest accuracy paragraph

Front-camera tracking will not tell you which letter of your caption someone read; published webcam-class accuracy runs around 2 to 4 degrees of visual angle versus a lab's sub-degree precision. What it resolves reliably is regions: face versus caption versus product versus background, and on-screen versus gone.

Every finding that has ever changed one of my test videos was a region-level finding. A lab coat out-pulling a face. A caption sitting on a face. A backdrop eating a tutorial. If your question is "which pixel", rent a lab. If your question is "what are they actually looking at", the front camera answers it in the viewer's natural habitat.

How a test works, end to end

On Jeena the workflow is: upload the video, set the goal (Views, Sales, Pitch, Followers), and the platform routes it to a panel of real viewers, 5 to 10 people on the Basic tier. Each viewer calibrates their front camera once, watches on their own phone, and answers a short impressions survey afterward. Motion detection invalidates watches where someone put the phone down. The report typically lands within a day.

Five to ten viewers sounds small until you look at how attention findings distribute: the recurring problems (the ones worth fixing) show up in viewer after viewer, and five catches most of what ten confirms. This is the panel math that makes a ten-euro price possible.

What the report contains

  • Attention heatmap
    Your video overlaid with where the panel's gaze concentrated, frame by frame. This is where you discover what the loudest thing in each scene actually was. How to read one.
  • Visibility map
    The inverse view: your video dimmed wherever nobody was looking. The fastest way to find the expensive thing you filmed that no one ever saw.
  • Wow-moments chart
    Eyebrow-raise events from the front camera plotted on your timeline: where anyone actually reacted, and where your intended big beat passed in silence.
  • Three concrete recommendations
    AI recommendations grounded in the gaze data, with timestamps: move this caption off the face, interrupt this static stretch, mirror the opener at the close. Plus a summary of how viewers perceived the video and their survey impressions.

What it costs, and when it pays

A typical test on Jeena costs around ten euros (the pricing page has current rates). The comparison that matters is not against lab studies, it is against the cost of finding out the hard way: a post that spends your best idea on an attention leak you could have fixed for the price of lunch. In our published tests, the recurring leaks (props stealing claims, endings nobody watched, captions killing their own faces) were all one-line fixes once they were visible.

When is it worth it? Not for every daily post. The high-leverage moments are the expensive ones: a video you paid production money for, an ad about to take spend, a format you are about to commit a month to, a pitch. Anywhere the cost of being wrong exceeds ten euros, which is most places that matter.

Eye track your next short video

Upload your video to Jeena. Real viewers watch it on their phones with the front camera on, and the report shows an attention heatmap, a visibility map, a wow-moments chart, a summary of how viewers perceived it, and three concrete recommendations, typically within a day.

No "schedule a call." No sales rep. Upload, get your report.

Frequently asked

Is there an eye tracking platform for short videos?+

Yes. Jeena is an eye tracking platform built specifically for short-form video: real viewers watch your Reel, TikTok, or Short on their own phones with the front camera tracking their gaze, and you get an attention heatmap, a visibility map, a wow-moments chart, and three concrete recommendations before you post. A typical test uses 5 to 10 real viewers and costs around ten euros.

Can you do eye tracking without special equipment?+

Yes, that is what changed in the last few years. Smartphone front cameras plus computer vision now track gaze at region-level accuracy (roughly 2 to 4 degrees, versus 0.3 to 0.8 for infrared lab rigs). Region level is sufficient for the questions short video actually poses: face versus caption versus product versus background, and whether viewers were looking at the screen at all. No hardware beyond the phones viewers already own.

Does eye tracking work on videos as short as 15 seconds?+

Short videos are the best case for it, not an edge case. A 15-second video gives each viewer only a few dozen fixations total, so where they land matters enormously, and the whole timeline fits in one readable report. The findings are also the most actionable: with so few scenes, a single leak (a busy opener, a caption on the face, a dead ending) is a large share of the video, and a one-line fix moves a large share of the outcome.

What is Jeena?+

Jeena is a neuromarketing platform for short-form video. Real people watch your video on their phone with the front camera on. Jeena captures their gaze direction, blink rate, eyebrow raises, and their impressions of the video in a short survey afterward. You receive an AI-powered report with an attention heatmap, a visibility map, a wow-moments chart, a summary of how viewers perceived the video, and three specific recommendations for making the video work harder.

How does Jeena measure viewer attention?+

Jeena uses smartphone front-camera gaze tracking. Each engager calibrates once, then watches your video. The platform records where their gaze lands frame by frame, flags moments of surprise from facial expression, and combines that with a short impressions survey afterward. The result is a per-second timeline of what real viewers actually looked at and felt, plus a summary of how they perceived the video overall.

How much does it cost to test a video on Jeena?+

A typical test costs around ten euros. See the pricing page for current rates.