What Twelve Labs does
TwelveLabs develops video-native multimodal foundation models and a platform that turns raw video archives into searchable, structured, and action-ready representations via APIs. The company’s core product is a “video understanding platform” that combines (1) video embeddings for semantic retrieval and (2) a video reasoning layer to generate structured, timestamped understanding of what happens in a video, designed for deployment by developers and enterprises that hold large volumes of video content.
More
TwelveLabs positions its approach as “genuine multimodality” for video—rather than language/image models that look at sampled frames—so customers can search for specific actions, scenes, dialogue, and other events across long-running footage without manual tagging. Business model: TwelveLabs sells developer/enterprise access to its video foundation models and workflow tooling primarily through APIs (with SDK support) and also offers integration/distribution paths via cloud marketplaces such as Amazon Bedrock. Current strategic position (as of late 2026): TwelveLabs has expanded from “video understanding models” into a more full-stack, agentic video intelligence framing, and it has emphasized production deployment and scaling on AWS infrastructure. In July 2026, it announced a $100M Series B co-led by NEA and NAVER Ventures (with Amazon participation) to build toward “Video Superintelligence.”