2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2025-05-05 15:57:53 -07:00
2026-02-26 19:41:21 -08:00

FastVLM

Fast vision-language model architecture research. Part of the Zen LM ecosystem.

License

Overview

FastVLM explores efficient architectures for vision-language models, focusing on reducing computational overhead while maintaining strong multimodal understanding.

Features

  • Efficient vision-language model architecture
  • Reduced computational overhead vs standard VLMs
  • Strong multimodal understanding
  • Research reference implementation
  • zen-vl — Zen vision-language models
  • jin — Multimodal understanding framework
  • Zen LM — Full model family

License

See LICENSE file.

Part of the Zen LM ecosystem by Hanzo AI

S
Description
This repository contains the official implementation of "FastVLM: Efficient Vision Encoding for Vision Language Models" - CVPR 2025
Readme
11 MiB
Languages
Python 81.5%
Swift 17.2%
Shell 1.2%