What is VLM?
VLM (Vision-Language Model) is a model trained to connect visual and linguistic representations so it can retrieve, classify, localize, understand, or generate content across images or video and text.
Quick Facts
| Full Name | Vision-Language Model |
|---|---|
| Created | 2021 (CLIP by OpenAI), 2023-2026 (production multimodal LLMs) |
| Specification | Official Specification |
How It Works
Architecture and Capability Boundary
A VLM may align separate image and text embeddings, inject visual features into language-model layers, or serialize visual patches as tokens for a generative decoder. These choices affect retrieval, captioning, visual question answering, localization and text generation differently. Video support additionally requires a documented frame or clip sampling policy; a model that accepts video bytes does not necessarily preserve event order, small text or every frame. Treat supported modalities, resolutions, coordinate systems and output schemas as versioned capabilities, not properties implied by the VLM label.
Input and Evidence Contract
A reproducible request records the original asset digest, media type, orientation, page or frame range, resize and crop transforms, model and processor revisions, prompt, decoding settings and output schema. Coordinates must declare whether they refer to original pixels, normalized values or transformed model input. For extraction or high-impact decisions, require each claim to point to source regions and validate those regions independently. Resizing or compression can remove small text and defects, so optimize media only after measuring task accuracy rather than assuming lower resolution is harmless.
Evaluation, Hallucination, and Security
[OmniDocBench](https://arxiv.org/abs/2412.07626) shows why document parsing needs evaluation across document sources, layout categories, attributes and both pipeline and end-to-end systems. A VLM release should likewise separate OCR, chart, counting, spatial, temporal, grounding, abstention and downstream task metrics, with realistic capture degradation and language slices. [OWASP LLM01](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) treats multimodal content as a prompt-injection surface: keep tool authorization outside the model, isolate retrieved media, test hidden visual instructions and require human review for medical, legal, financial or safety-critical use.
Key Characteristics
- Connects visual and language representations, but may be contrastive, generative, or hybrid
- Uses model-specific image tiling, video sampling, visual tokenization, and fusion contracts
- Can support captioning, retrieval, VQA, localization, extraction, or generation in different combinations
- May hallucinate objects, text, relationships, counts, coordinates, or temporal order
- Requires task- and slice-specific evaluation rather than one aggregate multimodal score
- Treats visual content as untrusted data when used inside agents or tool workflows
Common Use Cases
- Answering questions about images while returning region-level evidence
- Extracting text, tables, charts, and fields from documents with validation
- Searching image-text collections with aligned representations
- Understanding screenshots for bounded, approval-gated interface assistance
- Analyzing sampled video events with explicit temporal coverage and uncertainty
Example
Loading code...Frequently Asked Questions
How is a VLM different from a multimodal model or multimodal LLM?
VLM names the broad vision-language family, including dual encoders that retrieve or classify and generative systems that produce text. Multimodal model is broader because it can combine audio, text, video, actions or other modalities without a vision-language task. Multimodal LLM usually refers to a generative language-model-centered VLM, so the terms overlap but are not exact synonyms.
Is a VLM a replacement for OCR or computer vision models?
Not automatically. A VLM can read text and describe regions, but dedicated OCR, layout analysis, detection or segmentation may provide better coordinates, deterministic schemas and debuggable confidence signals. Compare an end-to-end VLM with a modular pipeline on the same business cases, including degraded images and abstentions.
How do VLMs process images and video?
Common systems encode image patches, transform those features into representations the language component can consume, then fuse visual and text context before prediction. Other systems use dual encoders or layer-wise cross-attention. Video models additionally sample or compress frames, so temporal coverage depends on the processor and cannot be inferred from the model name.
How should a team evaluate a VLM?
Define the exact task and score perception separately from downstream decisions. Measure OCR, tables, charts, counting, spatial or temporal reasoning, localization, hallucination, abstention, robustness, latency and cost on representative slices. Preserve source regions so reviewers can distinguish an incorrect observation from an incorrect business rule.
Can a VLM safely control tools from screenshots or documents?
Only behind external controls. Visual content can contain prompt injection, stale state or misleading overlays. The runtime should authenticate the actor, validate current application state, authorize exact arguments, limit side effects and require approval where needed; the VLM's visual interpretation is a proposal, not authorization.