TwelveLabs is making its Marengo video embedding model available inside Amazon Bedrock Managed Knowledge Bases, giving enterprises a managed way to search video, audio and images using natural-language queries. The integration could remove one of the biggest barriers to enterprise video AI: building and maintaining the ingestion, embedding and vector-search infrastructure needed to make unstructured media searchable.
TwelveLabs is bringing its Marengo Embed 3.0 model deeper into Amazon Web Services’ AI stack, allowing customers to build semantic search over video and other media through Amazon Bedrock Managed Knowledge Bases rather than assembling their own retrieval pipeline.
The announcement is significant because video remains one of the more difficult forms of enterprise data for AI systems to search. Unlike text documents, video combines images, motion, spoken language, music, sound effects, on-screen text and temporal context. Finding a specific moment traditionally requires multiple processing steps before a search system can retrieve the relevant clip.
AWS has already expanded Amazon Bedrock Knowledge Bases to support multimodal retrieval across text, images, audio and video. With Marengo Embed 3.0, TwelveLabs adds a specialized video embedding model designed to represent media for semantic search and retrieval.
The practical difference is considerable.
A conventional video search system might require developers to extract frames, transcribe speech, create embeddings, store vectors in a database and then build a retrieval layer that connects all those components. AWS’s managed approach handles ingestion, processing, embedding and retrieval within the knowledge-base workflow.
That means an organization with thousands of hours of footage can potentially ask a question such as “show me the penalty kick in the second half” and retrieve the relevant section based on semantic meaning rather than an exact keyword or manually assigned metadata.
Marengo Embed 3.0 is designed specifically for this type of multimodal retrieval. AWS documentation says the model can process video, audio, images and text and generate embeddings for similarity search. For video, the model can separately represent visual information, audio and transcribed speech, or combine multiple modalities into a fused embedding. It can process videos up to four hours in length and return clip-level or asset-level embeddings.
The result is a shift from video metadata search to video understanding.
That distinction matters across several industries. Broadcasters and sports organizations can locate highlights without manually tagging every scene. Film and television studios can search archival footage. Corporate training teams can retrieve particular demonstrations from instructional videos. Financial institutions and government agencies can search recorded material while keeping data within their AWS environment.
The security architecture is also part of the proposition. TwelveLabs says customers can keep their data within their own AWS account, an important consideration for regulated organizations handling sensitive media.
For Amazon, the integration strengthens Bedrock’s position as an enterprise AI development platform rather than simply a place to access foundation models. Bedrock Knowledge Bases already automates major portions of the retrieval-augmented generation pipeline, including ingestion, chunking, embedding and retrieval. Multimodal support extends that architecture beyond conventional text-centric enterprise search.
It also illustrates an increasingly important trend in enterprise AI: the model is becoming one component of a larger data-access system.
Gartner has warned that enterprises’ AI-readiness efforts increasingly depend on their ability to manage unstructured data, arguing that data-management platforms must make such information available for discovery, retrieval and AI use.
Video is an especially important example of that problem. Organizations have accumulated enormous archives of footage, but much of that information remains effectively trapped because conventional search systems cannot understand what happens inside the files.
TwelveLabs is not alone in addressing the problem. AWS itself offers Amazon Nova Multimodal Embeddings, while other AI infrastructure providers are developing multimodal retrieval and video-understanding systems. The competitive question is therefore shifting from whether enterprises can perform semantic video search to which platform can make it accurate, scalable, secure and economical enough for production use.
TwelveLabs’ advantage is specialization. Its Marengo model is built around video understanding rather than being a general-purpose text embedding model adapted to media. AWS, meanwhile, supplies the managed infrastructure and enterprise distribution.
The partnership could therefore be more important than the model announcement alone. By embedding Marengo into Bedrock’s managed knowledge-base architecture, TwelveLabs moves from asking customers to build video intelligence infrastructure toward making video intelligence a consumable component of an existing AI platform.
The early Iconik integration illustrates the commercial opportunity. Backlight’s Iconik media asset management platform is using Marengo through Amazon Bedrock Managed Knowledge Base to offer TwelveLabs-powered semantic search to its customers.
That could matter beyond media companies. As multimodal AI becomes a standard part of enterprise knowledge management, organizations will increasingly expect their AI assistants and agents to retrieve information from presentations, images, recordings and video—not just PDFs and databases.
For developers building enterprise AI applications, the implication is straightforward: the next generation of RAG systems will increasingly need to understand what happened in a video, not merely what its transcript says.
TwelveLabs and AWS are betting that making that capability managed and accessible will accelerate adoption.
The bigger technology trend is even broader. As enterprises move toward multimodal AI and AI agents, search is becoming a foundation layer for machine reasoning. Before an agent can act on a company’s video archive, it first needs a reliable way to find the right evidence.
Marengo’s arrival inside Amazon Bedrock is another step toward making that retrieval layer available without requiring every enterprise to build it from scratch.
Market Landscape
The enterprise AI market is moving from text-first retrieval toward multimodal knowledge systems capable of searching across documents, images, audio and video.
AWS introduced general availability for multimodal retrieval in Bedrock Knowledge Bases in January 2026, enabling organizations to build RAG applications across multiple content types.
The addition of Marengo gives AWS customers another specialized embedding option for video. The competitive landscape now includes hyperscaler-native multimodal models, specialized video-intelligence companies and vector-search infrastructure providers.
The differentiator will increasingly be retrieval accuracy, multimodal understanding, enterprise security, latency and total cost—not simply the underlying LLM.
Top Insights
- Video becomes searchable data: Marengo turns visual scenes, speech, audio and other media signals into embeddings that can support semantic retrieval.
- Managed infrastructure lowers barriers: Bedrock handles ingestion, indexing and retrieval, reducing the engineering required to build a production video-search pipeline.
- Multimodal RAG is expanding: Enterprise knowledge systems are moving beyond text toward unified retrieval across documents, images, audio and video.
- Specialized models still matter: TwelveLabs focuses specifically on video understanding, while AWS provides the cloud infrastructure and managed retrieval layer.
- AI agents need richer evidence: Better multimodal retrieval gives enterprise agents access to information previously buried inside large video archives.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI












