Why eye tracking matters for AI world models

  • Blog
  • by Tobii
  • 6 min

Blue graphic image about AI

AI is getting better at recognizing and generating content. The next challenge is understanding how the physical world works. World models aim to help AI systems build representations of environments, objects, actions, and their consequences. But learning from video alone leaves out an important part of how humans experience the world: attention. 

A camera can record everything within its field of view. Humans don’t process a scene in quite the same way. When we walk into a room, assemble a piece of furniture, prepare a meal, or navigate a busy street, our eyes continuously select information from the environment. We look at one object and then another. We check something before reaching for it. We return our gaze to something that requires closer attention. These patterns contain information that conventional video does not explicitly capture. 

For developers working on world models, robotics, and embodied AI, eye tracking can add a layer of human attention data to first-person video — providing richer examples of how people perceive and interact with real environments.

Ai models - a person giving a flower to a robotic arm

What are world models? 

A world model is an AI model designed to build an internal representation of an environment and how it changes. Rather than simply recognizing that a cup, table, and hand appear in a video frame, for example, a useful world model might learn relationships between them: objects persist when we look away, actions have consequences, and certain visual information becomes relevant depending on the task being performed. 

This ability is particularly important for embodied AI — AI systems such as robots or intelligent devices that need to operate in physical environments. Training these systems requires large and diverse datasets showing the world in action. Video is therefore becoming an increasingly important source of training data, but video and human perception aren’t the same thing. 

From merely seeing the world to understanding it

Video can capture environments, objects, movements, and interactions, providing models with rich information about what happens in the physical world. But there is a fundamental difference between what a camera records and what a human actually pays attention to: a camera captures everything within its field of view. Eye tracking reveals where a person is actually looking. 

When combined with first-person video, gaze data adds information about how people visually navigate their surroundings — what attracts their attention, what they look at while performing a task, and how their gaze shifts as they interact with the world. 

Imagine recording a person making a cup of coffee with a head-mounted camera. The video might capture the coffee machine, cups, buttons, countertop, other appliances, the person’s hands, and dozens of objects in the surrounding environment. From the video alone, all of those pixels are available to the model. 

A person making coffee at a coffee station

But when adding eye tracking to the dataset, it can also reveal that the person looked at the display before pressing a button, glanced toward the cup while positioning it, checked the coffee level as it poured, and shifted attention to the next object before reaching toward it. The scene hasn’t changed. The information available about the human interacting with it has. 

For world models, this provides another layer of context to the actions captured on video. Rather than only observing what happened, the dataset can contain information about the visual attention that accompanied an action — helping to create richer representations of how humans perceive and interact with the physical world. 

Recognizing nuances: Novice vs. Expert 

Experts and beginners don’t necessarily look at the same things. An experienced technician may immediately inspect a component that a novice overlooks. A skilled operator may monitor several indicators in a particular sequence. Someone learning a task may spend longer searching for the information they need. 

These differences can be valuable when collecting demonstrations of human tasks for AI systems, but many of the nuances can’t be captured from conventional video alone. Two people may ultimately perform the same action, while taking very different visual paths to get there. 

Gaze data provides another dimension to these demonstrations by showing what different individuals pay attention to. This can make it possible to capture not only the actions of experts, but also some of the visual behaviors that accompany their expertise. 

An expert might look at a component differently than a novice, wearable eye trackers captures the differences.

Capture multimodal field data via Tobii Glasses X

World models need rich data from the real world — across different people, environments, situations, and tasks. And the physical world is inherently multimodal. We don’t experience our surroundings through vision alone: we see, hear, move, orient ourselves, and direct our attention simultaneously

Tobii Glasses X provides a wearable platform for capturing several of these streams of real-world data from the human perspective. 

At its core, Tobii Glasses X combines eye tracking with first-person scene video, capturing both what is happening around a person and where that person is looking. But the device also contains additional sensors that can provide valuable context for AI training:

  • Stereo audio from two microphones

  • Accelerometers

  • A gyroscope capturing movement

  • A magnetometer that provides orientation information

The benefit for AI developers is not simply the availability of several types of data. It is the ability to capture these different signals together, during the same real-world interaction and from the same human perspective. 

Consider, again, a person making a cup of coffee. Scene video provides the visual environment and the sequence of events. Eye tracking shows where the person directs their attention. Audio could capture the sound of the machine, a spoken instruction, or other events in the surrounding environment. Movement data can add information about how the wearer moves their head while searching for an object or performing the task, while orientation data provides additional context about how they are positioned and move through the environment. 

Together, these signals provide a richer representation of the experience than any one of them could provide independently. 

The same principle becomes even more relevant in complex environments. Imagine an experienced technician diagnosing a malfunctioning machine. First-person video can record what the technician sees and does. Eye tracking can reveal which gauges, components, or warning indicators receive attention and in what sequence. Audio can capture an unusual sound from the machine or spoken communication with a colleague. Motion and orientation data can provide further context about how the technician moves while inspecting different parts of the equipment. 

For developers of world models and embodied AI, multimodal datasets like these can provide richer examples of the relationship between environment, attention, sound, movement, and action. 

Tobii Glasses X used in training tasks in manufacturing, operations, assembly and other complex tasks.
Tobii Glasses X used in training tasks in manufacturing, operations, assembly and other complex tasks.

Capturing the complexity of the real world 

Another important aspect of wearable data collection is the ability to move beyond controlled environments. 

Real-world environments are complex and unpredictable. People move differently, approach tasks in different ways, encounter distractions, respond to sounds, and adapt their behavior to changing circumstances. AI systems intended to operate in these environments ultimately need to account for this variation. 

Because Tobii Glasses X is wearable, data can be collected while people naturally move through and interact with their surroundings. This opens opportunities to capture human-centered training data across everything from everyday activities to manufacturing, maintenance, logistics, robotics, and other complex real-world tasks. 

It also creates opportunities to collect demonstrations from different people and levels of expertise. Instead of teaching an AI system from a single standardized example, developers can potentially capture multiple ways of approaching the same problem — together with the visual attention, movement, audio, and environmental context associated with each approach. 

Multimodal data does not automatically tell an AI system why a person behaved in a particular way, nor does every AI application need every available data stream. But combining these signals can give developers a richer dataset from which to identify relationships and patterns that may not be visible in video alone. 

Adding the human perspective to world models 

In conclusion: World models represent a shift from AI that primarily recognizes or generates information toward systems designed to build richer representations of the physical world and how it changes. 

Video can show AI what happens. Human demonstrations can show what we do. Eye tracking can add where we direct our attention. And additional sensor data — including audio, movement, and orientation — can provide further context about the experience and environment in which those actions take place. 

For world models and embodied AI, this can add an important dimension to training data: not simply recording the physical world but capturing more of the context around how humans perceive, navigate, and interact with it. 

 

Want to know more about eye tracking?

Contact our sales experts to learn how you can use eye tracking to enhance your research, study or training.

Continue learning about eye tracking and AI

Sign up for our newsletter

Sign up for our newsletter

Register for our newsletter to get Tobii’s latest blog posts and insights delivered to your inbox. Explore our articles focusing on real‑world behavior, attention, and the innovations shaping tomorrow’s technology.