Eye tracking is almost 150 years old, and every generation of it was built to answer the same question: what are people actually looking at? Short video is just the newest, hardest version of that question.

I spend my days watching heatmaps of people watching short videos, and at some point I wanted to know where all this started. The answer surprised me: eye tracking is one of the oldest instruments in psychology, older than the IQ test, older than most of what we call marketing. And it has been reinvented roughly once a generation, each time because someone wanted the same thing we want today: not what people say they saw, but where their eyes actually went.
Full disclosure before we start: I build one of the modern front-camera eye-tracking platforms (Jeena), so I have a stake in the last chapter of this story. The first five chapters belong to people far more patient than any of us.
The story usually starts in Paris, in the laboratory of ophthalmologist Louis Emile Javal. Watching people read (by the low-tech method of observing their eyes via a mirror placed beside the page), his lab found something nobody expected: the eyes do not glide smoothly along a line of text. They jump. Short, violent hops with brief stops in between. Javal coined the term for those hops, "saccades", and the word stuck for the next century and a half.
There is a lovely historical wrinkle here: careful historians of science have shown that Javal himself did not run the measurements; the observations came from his colleague M. Lamare, and the credit migrated to Javal thanks to loose wording in Edmund Huey's classic 1908 book on the psychology of reading. Huey, for his part, built one of the first actual eye-tracking devices: a contraption involving a plaster-of-Paris cup on the cornea connected to a pointer. Uncomfortable is an understatement. The finding, though, was foundational: seeing is not a stream. It is a sequence of fixations, and everything that came after is an attempt to map them.
The next leap came from an educational psychologist at the University of Chicago named Guy Buswell. In 1935 he published "How People Look at Pictures", the first large-scale eye-tracking study of images, built on a device that finally did not touch the eye at all: a camera photographing a beam of light reflected off the cornea, with head position tracked via a chromium bead on the observer's glasses. Around 200 people looked at paintings, photographs, and posters while the camera recorded every fixation.
Two of Buswell's findings could headline a creator blog today. First, gaze clusters: viewers do not scan a picture evenly, their fixations pile up on a few regions of interest and ignore the rest. Second, the first seconds are special: early in viewing, different people look at nearly the same places, and only later do their paths diverge. Ninety years before anyone said "hook", the data already showed that the opening moments of looking are the most predictable, and therefore the most designable.
Then comes the most famous chapter: Alfred Yarbus, a Soviet physicist, working with devices that sound like a dare. His "caps", tiny suction devices attached directly to the eyeball, held miniature mirrors that let him record gaze with remarkable stability. With them he produced the single most reproduced figure in attention research.
Yarbus showed people Ilya Repin's painting "The Unexpected Visitor" and asked different questions before three-minute viewings: estimate the family's wealth; give the ages of the people; guess what they were doing before the visitor arrived; remember the clothes. Same painting, same eyes, radically different gaze maps. Asked about wealth, viewers scoured the furniture and the walls. Asked about ages, they went face to face to face. The image did not decide where people looked. The task did.
That result, published in his 1967 book "Eye Movements and Vision", is arguably the founding insight of everything my industry does: attention is not a property of the picture, it is a negotiation between the picture and what the viewer is trying to do. A viewer opening a feed to be entertained is running a task too. Your video is being read against it, frame by frame.
Yarbus's point in one line: where people look is not decided by the image. It is decided by what they came to do.
The modern lab era replaced suction caps with kindness: infrared corneal-reflection trackers that sit under a monitor and triangulate gaze from the glint on the eye. Accuracy became spectacular, typical research-grade systems land within roughly 0.3 to 0.8 degrees of visual angle, which is finer than the width of your thumbnail at arm's length. Advertising research, usability labs, and reading science ran on these machines for decades, producing classics like the Nielsen Norman Group's F-shaped reading pattern for web pages.
The catch was never accuracy. It was the room. Lab eye tracking meant recruiting people into a facility, one at a time, under lighting conditions and posture instructions nothing like a sofa and a phone. For a category of content that is watched exclusively on sofas and phones, the most accurate instrument in the building was also the least natural place to use it.
The newest chapter runs on hardware everyone already owns. Webcam and smartphone front-camera eye tracking uses computer vision instead of infrared glints, and the honest numbers matter here: published comparisons put webcam-based tracking at roughly 2 to 4 degrees of error versus the lab's sub-degree precision. You cannot read someone's exact word fixations with it. What you can do reliably is region-level tracking: face versus caption versus background versus product, which regions held the eye and when it left the screen entirely.
And region-level happens to be exactly the granularity short video lives or dies by. The questions that decide a video's fate (did they look at the speaker or the lab coat, the product or the acne closeup, the tutorial or the mountains behind it?) are questions about regions, not letters. The trade of the front-camera era is a fair one: give up the last degree of precision, gain real viewers on real phones in real feeds, at panel sizes and prices a solo creator can afford. That is the generation I work in.
Every previous instrument watched people study one thing: a page, a painting, a website. Short video inverts every assumption. The stimulus changes completely every few seconds, the viewer can dismiss it with a thumb, and the decision window is measured in seconds. Vision science adds a brutal constraint: sharp, detailed seeing comes from the fovea, which covers roughly the central one to two degrees of the visual field, under one percent of everything in front of you. On a phone-sized video, that means the viewer truly sees one small patch of your frame at a time, and their gaze budget is a few fixations per scene before the swipe verdict.
The stakes have moved beyond marketing, too. A 2025 study in npj Science of Learning (Li and colleagues) had people watch 25 minutes of randomly-fed short videos before a continuous film, tracking their eyes throughout, and found the short-video group's gaze desynchronized at the film's event boundaries and their memory of it fragmented. The scientists are now using eye tracking to study what short video does to us. Creators can use the same instrument to study what we do to viewers. Same 150-year-old question, pointed both ways: where are the eyes, really?
No single person. The founding observation, that reading eyes move in jumps called saccades, came out of Louis Emile Javal's Paris laboratory around 1879 (the measurements were made by his colleague Lamare). Edmund Huey built one of the first tracking devices around 1908. Guy Buswell built the first non-contact photographic tracker and published the first large picture-viewing study in 1935, and Alfred Yarbus's suction-cap recordings in the 1950s and 60s produced the field's most famous findings.
That gaze is task-driven. Showing viewers the same Repin painting under seven different instructions (estimate the family's wealth, give the ages, remember the clothes, and so on), Yarbus recorded radically different gaze maps for each question. The image alone does not determine where people look; the viewer's goal does. For video, the implication is that your content is always being read against what the viewer came to the feed to do.
Research-grade infrared trackers achieve roughly 0.3 to 0.8 degrees of visual angle; published studies put webcam and front-camera tracking at roughly 2 to 4 degrees. That is not enough to tell which word someone read, but it is enough for region-level questions: whether viewers looked at the face or the caption, the product or the background, and when they stopped looking at the screen at all. Those are the questions that matter for short-form video, which is why front-camera panels test Reels, TikToks, and Shorts at region level.