Sharp human vision covers roughly the central two degrees of the visual field. Everything else in your frame is a blur the brain politely fills in. Here is what that means for anyone making short video.

Here is the fact that reorganized how I think about every video I watch: sharp human vision, the kind that reads text and recognizes detail, comes from the fovea, a pit in the retina covering roughly the central one to two degrees of your visual field. Under one percent of what is in front of you is ever actually in focus. The rest, more than 99 percent of the field, is peripheral vision: coarse, colorblind at the edges, excellent at detecting motion and almost nothing else.
You do not experience the world this way because your brain runs a continuous reconstruction, stitching a stable, detailed-feeling scene out of two to four foveal samples per second. But the seams show the moment you measure where those samples actually land. A viewer "watching" your video is really pointing a two-degree spotlight at three or four spots per scene, and taking your word for the rest.
I build an eye-tracking platform for short video (Jeena), which means I spend my days looking at where those spotlights actually go. This piece is the science underneath what I see in the data every week.
Ninety years of eye-tracking research gives a fairly stable ranking of what pulls the fovea. Faces come first in orienting speed: human vision turns toward a face faster than toward anything else. Motion is second, peripheral vision exists largely to notice movement and drag the fovea toward it. High-contrast text, especially large text, reliably captures early fixations. And behind all of it sits the oldest finding in the field: Guy Buswell's 1935 picture studies showed that fixations cluster on a few regions of interest and that different viewers' first seconds look remarkably alike, before their paths diverge.
But here is where our own data corrects the textbook. When we audited 88 real short videos (the Attention Paradox study), faces did NOT dominate the accumulated gaze: of all on-screen attention, 47.7% went to the subject or product, 25.5% leaked to the background, and only 17.7% landed on faces. Orienting first is not the same as holding longest, and in real, busy short-video frames, backgrounds out-collect faces. Captions, the element creators lean on hardest, drew 7.4%.
Note what this implies for a busy frame: the attractors compete, and the accumulated winner is often not the fast one. A face, a moving background, a bold caption, and a product shot in one composition are four candidates for two or three fixations. Something loses, and the viewer will never know what they missed, because peripheral vision papers over the gap. The only party who finds out is you, after the video underperforms.
A viewer is not watching your frame. They are pointing a two-degree spotlight at three spots per scene and trusting their brain about the rest.
The two-degree budget stops being abstract the first time you watch it spend itself on the wrong thing. In tests we have published: a dermatologist's lab coat drew more gaze than her face (36% versus 24%) while she explained a skincare routine. An acne-texture closeup pulled 64% of gaze while the product being sold got 30%. A tutorial's scenic mountain backdrop out-competed the tutorial itself. A creator's own hook caption, placed across her face, got the fixations her face needed.
None of these are exotic failures. They are the two-degree window landing exactly where the ranking above predicts: on the highest-contrast texture, the brightest region, the face-like thing, whether or not that is what the creator needed seen. The frame did not have an attention problem. It had an attention auction, and the wrong bidder won.
Most of the classic research watched people study one stimulus at length. Short-form viewing is a different regime: the stimulus replaces itself every few seconds, and the viewer holds a standing option to swipe. The clearest evidence that this regime is genuinely different came in 2025, from a study by Li and colleagues in npj Science of Learning. Participants watched 25 minutes of randomly-fed short videos, then a continuous 21-minute film, with their eyes tracked throughout.
The random short-video group came out measurably different: their gaze desynchronized from other viewers' at the film's event boundaries, the moments where scenes turn, and their later recall of the film was more fragmented. Static-image memory was unaffected; the cost was specific to following continuous events. In plain terms: heavy random-feed watching seems to train the visual system out of the very synchronization that storytelling depends on. For creators the practical translation is sobering: your viewer arrives pre-fragmented, and the burden of re-synchronizing their attention at every beat of your video is on you.
At any instant, decide which single region should win the fixation and mute the competition: quiet the background, desaturate the decor, keep captions off the face. When everything bids, the brightest texture wins, and it is rarely your message.
A small inset running beside a talking head asks one spotlight to be in two places. Show the proof full-frame for a beat, then return to the face. Two seconds of one thing beats ten seconds of two things.
In our audit, viewers' gaze synchrony ran uniformly high (~0.80) on strong and weak videos alike: when eyes are on screen, everyone lands on the same spot at the same moment. Directing attention is largely a solved problem; the differentiator is what the synchronized eye finds there, and what the ear hears while it looks. Give each scene one landing point worth landing on.
Most of your frame is never seen. That sentence sounds like pessimism until you flip it: the two or three regions that ARE seen carry the entire video, which means design effort concentrated there pays absurdly well, and effort spent anywhere else is invisible by physiology. The viewers are not skimming out of disrespect. They are built this way, and they always were: 1935 viewers looking at paintings spent their fixations exactly as unequally.
The only new thing is that we can now watch the spotlight move on real phones, in real feeds, before a video is published, instead of finding out from the retention graph after. Where the eyes go was never guessable from inside your own head, because you know where they are supposed to go. That is precisely the knowledge your viewer does not have.
A small fraction at any moment. Sharp (foveal) vision covers roughly the central one to two degrees of the visual field, under one percent of it, and viewers make about two to four fixations per second. Over a few seconds of a scene, that is a handful of small patches; peripheral vision fills in the rest coarsely, tuned mainly to motion. Which patches get the fixations decides what the video effectively contains for that viewer.
Evidence says yes, and the difference persists after the app closes. A 2025 study in npj Science of Learning found that after 25 minutes of randomly-fed short videos, viewers' gaze desynchronized at a continuous film's event boundaries and their recall of it fragmented, while static-image memory was untouched. Short-form viewing is a regime of constant stimulus replacement under a standing swipe option, and it appears to train attention accordingly.
In orienting speed: faces first, then motion, then high-contrast elements like bold text, with fixations clustering on a few regions of interest (a finding that goes back to Buswell's 1935 picture studies). But accumulated gaze tells a different story: in our audit of 88 real short videos, subjects and products collected 47.7% of on-screen attention, backgrounds 25.5%, and faces only 17.7%. The attractors compete, and a busy frame holds an auction: a lab coat can outbid a face, a backdrop can outbid a tutorial, and clutter reliably outbids everything you meant to be seen.