Meta Llama Runs locally

Llama 3.2 8B

Running on a laptop, classification, drafting

128K tokens
Context window
8K tokens
Max output
Local
Input per 1M
Local
Output per 1M
late 2023
Knowledge cutoff
text
Modalities

Last verified 20 May 2026 Report a correction

What it tends to refuse

Depends on the tuning

About 5 GB at 4-bit, so it fits on a machine with 16 GB of RAM. Keep prompts to one task and the output format simple.

Standing rules for this model go in Modelfile, not in every message. See the config files.

Written for Llama

All 6