Pegasus 1.5 by Twelve Labs is a video intelligence model that converts raw video into searchable, AI-ready metadata at massive scale. It ingests multimodal data through a single pipeline at approximately 60x real-time speed, allowing users to index an hour of video in a minute and process over 10,000 hours per day. The platform enables natural language search across entire video libraries, locating specific actions, scenes, dialogue, and emotions without manual tagging. It supports use cases such as content discovery, compliance scanning, highlight creation, and insight generation. Designed for organizations working with video at scale, it serves sectors including media and entertainment, advertising, and government. The model achieves a +13.1% improvement over Gemini 3.1 Pro on multimodal prompting and offers APIs, SDKs, and integrations for developers.
Key Benefits
- Ingests multimodal data at ~60x real-time speed
- Indexes an hour of video in a minute
- Supports natural language search across video libraries
- Achieves SOTA composite accuracy