Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Molmo is a family of vision-language AI models released by the Allen Institute for AI (Ai2), not a single chatbot. Its original September 2024 release stood out for visual grounding: a model could describe an image and indicate where relevant objects or details appeared. Ai2 made model weights, code, a demo and substantial training and evaluation resources available, though “open” does not mean every component or dataset has identical terms.

What Ai2 released in September 2024

Ai2 announced Molmo on September 25, 2024, describing it as a family of open vision-language models. The initial release included four variants: MolmoE-1B, Molmo-7B-O, Molmo-7B-D and Molmo-72B. The announcement included a public demo, inference code, model weights and a technical report. Ai2 later described releases of the PixMo dataset family and training and evaluation code. Read Ai2’s dated announcement for the original scope and timeline.

The models combine an image encoder with a language model. Ai2 emphasized its approach to training data: human annotators provided detailed image captions through speech-based descriptions, alongside examples involving 2D pointing. Ai2 presented this as an alternative to depending on outputs distilled from proprietary vision-language models. That description reflects Ai2’s account of its design; it does not establish that every component or source of training data was unrestricted.

What made the original Molmo notable

It could ground answers in image regions

Many vision-language models answer questions about an image in text. Molmo’s distinctive demonstration was that a model could also point to relevant areas in the image. That grounding can make an answer easier to inspect: a response about an object or detail is accompanied by a visual cue showing where the model found it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pointing is not a guarantee that an interpretation is correct. It is a way to expose the model’s visual reference, which readers can compare with the image and the accompanying answer.

It came in different scales and configurations

The 1B, 7B and 72B labels indicate different model scales, while the two 7B names represent distinct variants rather than interchangeable labels. The original release should therefore be understood as a set of model choices, not one fixed capability or hardware profile. Compare the exact variant and its current instructions before choosing one.

What Ai2’s performance claims mean

Ai2 reported strong results across academic benchmarks and human evaluations, and compared Molmo with both proprietary and open systems. These are claims from the release, tied to the evaluations available at that time—not an independent or permanent ranking. The announcement does not provide one universal score that captures performance across all tasks.

When comparing Molmo with another model, match the task, model version, benchmark and evaluation conditions. A ranking on one image benchmark does not establish which model is best for video, multi-image reasoning or a particular deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How Molmo has expanded since the original release

The original announcement focused on image understanding. Ai2’s current Molmo page describes a broader family that includes Molmo 2 variants at 4B, 8B and 7B O sizes, with support for video and multi-image understanding as well as image tasks. Ai2 highlights features such as pointing, tracking, counting and dense captioning. These are current-family capabilities, not features to retroactively attribute to every model in the 2024 release. See Ai2’s current Molmo page for its model and documentation links.

What “open” does—and does not—tell you

For Molmo, “open” points to meaningful public access to model artifacts and development resources, but it is not a substitute for checking the terms that apply to a specific use. The model, its components and its training data may have different conditions. This distinction matters especially for commercial deployment.

Ai2 says Molmo 2 is licensed under Apache 2.0 and intended for research and educational use under its Responsible Use Guidelines. It also notes that some third-party datasets used for Molmo 2 are restricted to academic and non-commercial research use. The model license alone therefore does not answer every question about whether a proposed use is permitted. Review Ai2’s Molmo 2 release information, the exact model card, and applicable dataset terms before relying on the model in a product or service.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Trying Molmo or running it locally

Ai2 links to a playground and model downloads from its current Molmo page. For local use, first select the specific variant and task, then follow that artifact’s current installation and inference instructions. Hardware needs depend on the variant, software setup, numerical precision and workload.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One community discussion records a user report of running Molmo-7B-D in bfloat16 on a 24GB RTX 4090. That is a single configuration report, not an official minimum requirement or a guarantee for other variants, inputs or software versions. See the community report in context.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.