Meta's open-weight 8B-parameter language model, Transformer with GQA, 128K context window, multilingual.
Context window
128K
tokens
Parameters
8B
parameters
Release date
23 July 2024
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud📱 On-device
Overview
Access & deployment
DownloadAPIHosted
LocalCloudOn-device
Weights: Open weights
Key parameters
📏 Context: 128K
🧩 Parameters: 8B
✓ Tools · ✓ Fine-tuning
📥 Input: text
Technical specification
Context window
128K
tokens
Parameters
8B
parameters
Knowledge cutoff
31 Dec 2023
Knowledge boundary
License
Llama 3.1 Community License
Hardware requirements
~16 GB VRAM in BF16; with 4-bit quantization runs on consumer GPUs (8-12 GB) and high-end laptops.
Features:✓ Tool use✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
textcode
Capabilities and applications
Native model capabilities
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Synthetic data generation
Generating synthetic datasets that preserve the statistical properties of the original — used for model training, testing, and privacy protection.
Category: structured_generation
Benchmark results
6 benchmarks
MMLU
accuracy · 5-shot
66.7%
📄 Meta Llama 3.1 model card (pretrained)
Base pretrained model.
MMLU-Pro
accuracy · 5-shot, CoT
37.1%
📄 Meta Llama 3.1 model card (pretrained)
Base pretrained model.
ARC-Challenge
accuracy · 25-shot
79.7%
📄 Meta Llama 3.1 model card (pretrained)
Base pretrained model.
WinoGrande
accuracy · 5-shot
60.5%
📄 Meta Llama 3.1 model card (pretrained)
Base pretrained model.
BIG-Bench Hard (BBH)
average / exact match · 3-shot, CoT
64.2%
📄 Meta Llama 3.1 model card (pretrained)
Base pretrained model.
DROP
F1 · 3-shot
59.5points
📄 Meta Llama 3.1 model card (pretrained)
Base pretrained model.
Technical architecture
Core Architecture
Model Form
Training Techniques
