Capture multimodal field data via Tobii Glasses X
World models need rich data from the real world — across different people, environments, situations, and tasks. And the physical world is inherently multimodal. We don’t experience our surroundings through vision alone: we see, hear, move, orient ourselves, and direct our attention simultaneously.
Tobii Glasses X provides a wearable platform for capturing several of these streams of real-world data from the human perspective.
At its core, Tobii Glasses X combines eye tracking with first-person scene video, capturing both what is happening around a person and where that person is looking. But the device also contains additional sensors that can provide valuable context for AI training:
Stereo audio from two microphones
Accelerometers
A gyroscope capturing movement
A magnetometer that provides orientation information
The benefit for AI developers is not simply the availability of several types of data. It is the ability to capture these different signals together, during the same real-world interaction and from the same human perspective.
Consider, again, a person making a cup of coffee. Scene video provides the visual environment and the sequence of events. Eye tracking shows where the person directs their attention. Audio could capture the sound of the machine, a spoken instruction, or other events in the surrounding environment. Movement data can add information about how the wearer moves their head while searching for an object or performing the task, while orientation data provides additional context about how they are positioned and move through the environment.
Together, these signals provide a richer representation of the experience than any one of them could provide independently.
The same principle becomes even more relevant in complex environments. Imagine an experienced technician diagnosing a malfunctioning machine. First-person video can record what the technician sees and does. Eye tracking can reveal which gauges, components, or warning indicators receive attention and in what sequence. Audio can capture an unusual sound from the machine or spoken communication with a colleague. Motion and orientation data can provide further context about how the technician moves while inspecting different parts of the equipment.
For developers of world models and embodied AI, multimodal datasets like these can provide richer examples of the relationship between environment, attention, sound, movement, and action.