Quali modelli IA puoi eseguire sul tuo computer?
Scrivi com'è il tuo computer o inserisci i dati che conosci. La tabella completa resta disponibile anche se JavaScript non viene caricato.
Cosa puoi eseguire
Nei Mac con chip Apple non c'è VRAM separata: processore e grafica condividono la stessa memoria.Il motore locale confronta pesi, cache KV e memoria disponibile. Nessuna IA decide i numeri del verdetto.
Scrivi il tuo computer per vedere quali modelli puoi eseguire.
- Pesi
- Cache KV
- Runtime e buffer
- Totale stimato
Quantizzazione consigliata:
Velocità stimata:
Catalogo dei modelli locali
Memoria approssimativa per pesi, cache KV e overhead di esecuzione con 4K token. È un riferimento iniziale, non una garanzia di prestazioni.
| Modello | Famiglia | Parametri | Contesto max | Memoria approssimativa per quantizzazione |
|---|---|---|---|---|
| SmolLM2 135M Instruct | SmolLM | 130 M | 8K |
|
| SmolLM2 360M Instruct | SmolLM | 360 M | 8K |
|
| Qwen2.5 0.5B Instruct | Qwen | 490 M | 32K |
|
| Qwen3 0.6B | Qwen | 750 M | 40K |
|
| Gemma 3 1B Instruct | Gemma | 1 B | 32K |
|
| Llama 3.2 1B Instruct | Llama | 1.24 B | 128K |
|
| Qwen2.5 1.5B Instruct | Qwen | 1.54 B | 32K |
|
| SmolLM2 1.7B Instruct | SmolLM | 1.71 B | 8K |
|
| DeepSeek-R1 Distill Qwen 1.5B | DeepSeek | 1.78 B | 128K |
|
| Qwen3 1.7B | Qwen | 2.03 B | 40K |
|
| Granite 3.1 2B Instruct | Granite | 2.53 B | 128K |
|
| Gemma 2 2B Instruct | Gemma | 2.61 B | 8K |
|
| Qwen2.5 3B Instruct | Qwen | 3.09 B | 32K |
|
| Llama 3.2 3B Instruct | Llama | 3.21 B | 128K |
|
| Phi-3.5 mini Instruct | Phi | 3.82 B | 128K |
|
| Phi-4 mini Instruct | Phi | 3.84 B | 128K |
|
| Qwen3 4B | Qwen | 4.02 B | 40K |
|
| Gemma 3 4B Instruct | Gemma | 4.3 B | 128K |
|
| DeepSeek Coder 6.7B Instruct | DeepSeek | 6.74 B | 16K |
|
| Mistral 7B Instruct v0.3 | Mistral | 7.25 B | 32K |
|
| DeepSeek-R1 Distill Qwen 7B | DeepSeek | 7.62 B | 128K |
|
| Qwen2.5 7B Instruct | Qwen | 7.62 B | 32K |
|
| Qwen2.5-Coder 7B Instruct | Qwen | 7.62 B | 32K |
|
| DeepSeek-R1 Distill Llama 8B | DeepSeek | 8.03 B | 128K |
|
| Llama 3 8B Instruct | Llama | 8.03 B | 8K |
|
| Llama 3.1 8B Instruct | Llama | 8.03 B | 128K |
|
| Llama 3.1 Nemotron Nano 8B v1 | Nemotron | 8.03 B | 128K |
|
| Granite 3.1 8B Instruct | Granite | 8.17 B | 128K |
|
| Qwen3 8B | Qwen | 8.19 B | 40K |
|
| Gemma 2 9B Instruct | Gemma | 9.24 B | 8K |
|
| Gemma 3 12B Instruct | Gemma | 12.19 B | 128K |
|
| Mistral NeMo 12B Instruct | Mistral | 12.25 B | 128K |
|
| Phi-4 14B | Phi | 14.66 B | 16K |
|
| DeepSeek-R1 Distill Qwen 14B | DeepSeek | 14.77 B | 128K |
|
| Qwen2.5 14B Instruct | Qwen | 14.77 B | 32K |
|
| Qwen2.5-Coder 14B Instruct | Qwen | 14.77 B | 32K |
|
| Qwen3 14B | Qwen | 14.77 B | 40K |
|
| Codestral 22B v0.1 | Mistral | 22.25 B | 32K |
|
| Mistral Small 3 24B Instruct | Mistral | 23.57 B | 32K |
|
| Gemma 2 27B Instruct | Gemma | 27.23 B | 8K |
|
| Gemma 3 27B Instruct | Gemma | 27.43 B | 128K |
|
| Qwen3 30B-A3B (MoE) | Qwen | 30.53 B | 40K |
|
| DeepSeek-R1 Distill Qwen 32B | DeepSeek | 32.76 B | 128K |
|
| Qwen2.5 32B Instruct | Qwen | 32.76 B | 32K |
|
| Qwen2.5-Coder 32B Instruct | Qwen | 32.76 B | 32K |
|
| Qwen3 32B | Qwen | 32.76 B | 40K |
|
| Mixtral 8x7B Instruct v0.1 (MoE) | Mistral | 46.7 B | 32K |
|
| Llama 3.3 70B Instruct | Llama | 70.55 B | 128K |
|
| Qwen2.5 72B Instruct | Qwen | 72.71 B | 32K |
|
Il valore può variare in base al runtime, al sistema operativo e alla configurazione.
Come leggere questi valori
La memoria parte dalla dimensione pubblicata dei pesi di ogni file GGUF. Poi si aggiungono la cache KV, che cresce con la lunghezza del contesto, e un margine per runtime e attivazioni.
La velocità è sempre una stima: dipende da banda di memoria, processore, backend e configurazione. Il motore interattivo calcolerà poi il verdetto usando i tuoi dati.
Domande frequenti
La VRAM è uguale alla RAM?
No. La VRAM appartiene alla scheda grafica ed è spesso il limite per caricare un modello sulla GPU. La RAM di sistema può aiutare con CPU o offload, ma di solito riduce la velocità.
Perché conta la dimensione del contesto?
La cache KV cresce quando il modello deve conservare più token della conversazione o del documento. Per questo un modello che entra a 4K può richiedere molta più memoria a 32K.
Dove posso scaricare questi modelli?
Scarica ogni modello dal repository ufficiale del progetto o dal canale ufficiale dello strumento che usi. DescargasIA non ospita installer, modelli o mirror.