Run AI on hardware you control
Keep your data local, and free yourself from rate limits.
Choose your local AI setup
Starter
A complete system with a 16GB card runs 7–20B models for chat, autocomplete, and transcription. The cheapest way in.
- Runs: gpt-oss-20b · Qwen3 14B · Whisper
- ~40+ tok/s on 8B models
- 650–750W PSU
Enthusiast
A 24GB card (3090/4090-class) is the sweet spot: 30B-class models entirely on the card, fast image generation, and the broadest tool support.
- Runs: Qwen 32B · SDXL & FLUX · coding agents
- ~20–35 tok/s on 30B models
- 850W PSU
Power
A 24GB+ NVIDIA flagship with 64GB+ system RAM: bigger context windows, MoE models spilling into RAM, and headroom for always-on work.
- Runs: 30B + image gen at once · big MoE via RAM · 24/7 agents
- ~25–40 tok/s on 30B models
- 1000W PSU (always-on adds heat & power cost)
Hybrid Local + Cloud
Keep the private, always-on work local. Burst to the cloud for more advanced tasks.
For scale: you read at roughly 5–10 tokens per second, and 20+ feels instant. Speeds are ballparks for 4-bit models, not benchmarks.
Want 70B at full speed? That takes about 40GB of VRAM, usually two 24GB cards. llama.cpp and vLLM split a model across cards well, but plan for a 1200W+ PSU and the case and motherboard room. We can't filter listings by combined VRAM, so shop single cards and pair them yourself.
Complete systems
See allWhat fits in how much VRAM
Smarter models need more memory to run. On a graphics card that memory is VRAM, and it's the single biggest factor in which models you can run and how fast they feel.
| Model size | Needs (4-bit) | Fits on (GPU) | Good for | |
|---|---|---|---|---|
7–9B (≈ GPT-4o mini) Qwen3.5 9B · Llama 3.1 8B · DeepSeek-R1 8B | ~5–6 GB | Any 8GB+ card (RTX 3050 / 4060) | Chat, autocomplete, basic summaries | Shop PCs → |
12–14B (≈ GPT-4.1 mini) Qwen3 14B · Gemma 4 12B · Phi-4 14B | ~9 GB | 12GB+ cards (RTX 3060 12GB / 4070) | Better chat, basic coding, lightweight tools | Shop PCs → |
24–32B (≈ GPT-4.1 / GPT-5 mini) Qwen3.6 27B · Mistral Small 24B · Gemma 4 31B | ~14–20 GB | 24GB cards (RTX 3090 / 4090) + 32GB system RAM | Coding agents, stronger reasoning, image workflows | Shop PCs → |
70B (≈ GPT-4o) Llama 3.3 70B · Qwen2.5 72B | ~40 GB | Two 24GB cards (two RTX 3090s), or spill to system RAM, much slower | Larger local models and advanced workflows | Shop PCs → |
MoE (Qwen, gpt-oss) (up to ≈ GPT-5 mini) Qwen3.5 35B-A3B · gpt-oss-120b | Splits across VRAM + RAM | Partial GPU offload on a 12GB+ card (RTX 3060 12GB) + 32–64GB system RAM, slower than on-card | Experimental users comfortable tuning | Shop PCs → |
Sizes assume 4-bit quantization (the standard compression that shrinks typical 16-bit model weights to roughly a quarter of their size for a small quality cost). Leave a couple GB spare for context. The headline giants, like DeepSeek and Kimi, weigh hundreds of GB even compressed, so they stay on servers. And since local models are 4–40GB files each, a 1–2TB NVMe fills faster than you'd think.
Cards that run the big models
See allWorried a used card was mined on?
Fair question. Every Jawa order is buyer-protected, condition is disclosed, and you can check seller ratings and reviews where available. If it's not as described, report it within 2 days of delivery to get your money back.
Why run AI locally?
The honest tradeoff: the biggest cloud models are still smarter on the hardest problems, and API pricing is cheap right now. Local wins on privacy, control, and always-on use. And it all runs on a few hundred watts at your desk, about what a gaming session draws, with no data center or cooling water involved. Plenty of people run both.
Privacy
Run local-only software and your prompts, files, and notes stay on your machine. Nothing has to leave your desk.
Ownership
Buy the hardware once and it's yours. Cloud pricing, limits, and terms change on someone else's schedule. A model on your desk doesn't.
No rate limits
No 'you've hit your limit, try again later.' No surprise price changes. The model is always available because it's on your desk.
Total control
Tinker, swap models, automate things, leave it running overnight. It's your machine, so do what you like with it.
Memory upgrades
See allFast storage
See allBuild it with people who've done it before
Get advice and support from our community of expert PC builders
Join the Discord