Meta Llama Runs locally

Llama 3.3 70B

Private data, offline work, fine-tuning

128K tokens
Context window
8K tokens
Max output
Local
Input per 1M
Local
Output per 1M
late 2023
Knowledge cutoff
text
Modalities

Last verified 20 May 2026 Report a correction

What it tends to refuse

Base weights refuse very little; behaviour depends on the instruct tuning you run

At 4-bit quantisation this needs roughly 42 GB plus room for the context, so two 24 GB cards or one 48 GB card. Follows multi-part instructions less reliably than hosted models.

Standing rules for this model go in Modelfile, not in every message. See the config files.

Written for Llama

All 6