one year on
Meta releases Llama 3.2 with vision and edge-device capabilities
The new model family includes 11B and 90B vision LLMs, plus lightweight 1B and 3B text-only models optimized for on-device use, alongside the first Llama Stack distributions.
Meta today released Llama 3.2, a major update to its open-source model family that adds vision capabilities and introduces lightweight models designed to run on edge devices and mobile phones. The release includes two vision-language models at 11 billion and 90 billion parameters, as well as text-only models at 1 billion and 3 billion parameters that support a 128,000-token context window.
The 1B and 3B models, created through pruning and distillation from the larger Llama 3.1 8B model, are small enough to fit on-device while retaining strong performance on tasks like summarization, instruction following, and tool use. Meta says these models are enabled on day one for Qualcomm and MediaTek hardware and optimized for Arm processors. The vision models are designed as drop-in replacements for Llama 3.1 text models, adding image understanding capabilities without sacrificing existing language performance.
Alongside the models, Meta released the first official Llama Stack distributions — standardized APIs for inference, tool use, and retrieval-augmented generation — with support from partners including AWS, Databricks, Dell, Google Cloud, and NVIDIA. On-device distribution runs via PyTorch ExecuTorch, while single-node distribution is available through Ollama.
The announcement was immediately met with enthusiasm, with developer Simon Willison saying he is ‘absolutely amazed at how capable the new 1B model is’ given the 1.3GB download. The thread also saw skepticism about the vision models’ competitiveness, with some users noting that alternatives like Qwen-2-VL and Molmo appear to offer better performance on visual benchmarks. Meta acknowledged that the vision models were compared against Claude 3 Haiku and GPT-4o-mini.
The record
Amazed at the 1B model's capabilities given its 1.3GB size, noting it handled a 128K-token codebase summarization surprisingly well.
Questioned the vision models' competitiveness, suggesting Qwen-2-72B and Molmo as better open alternatives, and noted that Meta compared Llama 3.2 to Claude 3 Haiku and GPT-4o-mini rather than stronger models.
Argued that benchmarks are unreliable due to training on validation sets, and claimed the 11B model is even with Molmo 72B in A/B testing, also praised the novel tokenization adapter for efficiency.
One year later — open only if you can handle spoilers
Llama 3.2’s lightweight models became a popular choice for on-device AI applications, though the vision models faced stiff competition from Qwen2-VL and Molmo in the open-source community. The Llama Stack API gained modest adoption, particularly among enterprise users seeking standardized deployment.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy