TwelveLabs launches Pegasus 1.6 video understanding model, expands beyond media to physical AI

[DigitalToday reporter Chi-gyu Hwang] Multimodal video understanding AI company TwelveLabs said on Tuesday it has launched Pegasus 1.6, a vision-language model designed to understand egocentric data.

Egocentric data is first-person video footage captured from a person’s point of view. It is typically collected by filming workers performing specific actions while wearing action cameras or smart glasses.

Egocentric data is drawing attention for robot training. But TwelveLabs said securing first-person video alone does not complete robot training data. It said suitable scenes must be selected, actions in the video divided into segments, and details identified, including what tools or objects the hands interacted with and what the outcomes were.

While the process has relied heavily on human labor, Pegasus 1.6 supports the egocentric video understanding capabilities needed to automate it.

The model can understand a person’s actions in video together with surrounding objects and the preceding and following context, then divide tasks into segments and describe them. For example, in a scene where a worker picks up a specific part, uses a tool to work on it, and then moves the finished object, it can structure the actions, objects that appear, and the sequence of actions.

To help apply such video understanding to building robot training datasets, Pegasus 1.6 supports five core functions: Action Segmentation and Labeling, Dense Caption Labeling, Quality Scoring, Search and Curation, and Consent and Compliance Flagging.

TwelveLabs CEO Jae-sung Lee (이재성) said the direction the company has consistently pursued has been enabling machines to understand how the real world works through video. He said physical AI is a natural next step in that direction. He added that most of the experiences people acquire in the real world, such as correcting something when it slips, have not been recorded in a form machines can learn from. He said Pegasus 1.6 will turn video into structured knowledge so physical AI and robotics companies can use real human experiences for learning.

TwelveLabs has applied general-purpose video understanding technology that grasps actions and events in video, as well as preceding and following context, across markets including media, entertainment, sports and the public sector. More recently, it is expanding into industrial AX areas applied to physical AI, including robots, and manufacturing sites.

Keyword

#TwelveLabs #Pegasus 1.6 #egocentric data #vision-language model #Physical AI
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.