10,000+ gaming PCs bought and sold on Jawa!
For local AI

Run AI on hardware you control

Keep your data local, and free yourself from rate limits.

Featured PC build
Fully-local + Hybrid builds

Choose your local AI setup

Local AI · 16GB VRAM

Starter

$1,500–$2,500

A complete system with a 16GB card runs 7–20B models for chat, autocomplete, and transcription. The cheapest way in.

  • Runs: gpt-oss-20b · Qwen3 14B · Whisper
  • ~40+ tok/s on 8B models
  • 650–750W PSU
Local AI · 24GB+ VRAM

Enthusiast

$3,000–$6,000

A 24GB card (3090/4090-class) is the sweet spot: 30B-class models entirely on the card, fast image generation, and the broadest tool support.

  • Runs: Qwen 32B · SDXL & FLUX · coding agents
  • ~20–35 tok/s on 30B models
  • 850W PSU
Local AI · 24GB+ VRAM · 64GB RAM

Power

$4,000+

A 24GB+ NVIDIA flagship with 64GB+ system RAM: bigger context windows, MoE models spilling into RAM, and headroom for always-on work.

  • Runs: 30B + image gen at once · big MoE via RAM · 24/7 agents
  • ~25–40 tok/s on 30B models
  • 1000W PSU (always-on adds heat & power cost)
Hybrid

Hybrid Local + Cloud

Keep the private, always-on work local. Burst to the cloud for more advanced tasks.

Starting at ~$400

For scale: you read at roughly 5–10 tokens per second, and 20+ feels instant. Speeds are ballparks for 4-bit models, not benchmarks.

Want 70B at full speed? That takes about 40GB of VRAM, usually two 24GB cards. llama.cpp and vLLM split a model across cards well, but plan for a 1200W+ PSU and the case and motherboard room. We can't filter listings by combined VRAM, so shop single cards and pair them yourself.

Complete systems

See all
Turnkey rigs from verified sellers. Skip the build, start running.
Model size + VRAM

What fits in how much VRAM

Smarter models need more memory to run. On a graphics card that memory is VRAM, and it's the single biggest factor in which models you can run and how fast they feel.

Model sizeNeeds (4-bit)Fits on (GPU)Good for
7–9B (≈ GPT-4o mini)
Qwen3.5 9B · Llama 3.1 8B · DeepSeek-R1 8B
~5–6 GBAny 8GB+ card (RTX 3050 / 4060)Chat, autocomplete, basic summariesShop PCs →
12–14B (≈ GPT-4.1 mini)
Qwen3 14B · Gemma 4 12B · Phi-4 14B
~9 GB12GB+ cards (RTX 3060 12GB / 4070)Better chat, basic coding, lightweight toolsShop PCs →
24–32B (≈ GPT-4.1 / GPT-5 mini)
Qwen3.6 27B · Mistral Small 24B · Gemma 4 31B
~14–20 GB24GB cards (RTX 3090 / 4090) + 32GB system RAMCoding agents, stronger reasoning, image workflowsShop PCs →
70B (≈ GPT-4o)
Llama 3.3 70B · Qwen2.5 72B
~40 GBTwo 24GB cards (two RTX 3090s), or spill to system RAM, much slowerLarger local models and advanced workflowsShop PCs →
MoE (Qwen, gpt-oss) (up to ≈ GPT-5 mini)
Qwen3.5 35B-A3B · gpt-oss-120b
Splits across VRAM + RAMPartial GPU offload on a 12GB+ card (RTX 3060 12GB) + 32–64GB system RAM, slower than on-cardExperimental users comfortable tuningShop PCs →

Sizes assume 4-bit quantization (the standard compression that shrinks typical 16-bit model weights to roughly a quarter of their size for a small quality cost). Leave a couple GB spare for context. The headline giants, like DeepSeek and Kimi, weigh hundreds of GB even compressed, so they stay on servers. And since local models are 4–40GB files each, a 1–2TB NVMe fills faster than you'd think.

Cards that run the big models

See all
24GB and up: enough VRAM to run 30B-class models entirely on the card.

Worried a used card was mined on?

Fair question. Every Jawa order is buyer-protected, condition is disclosed, and you can check seller ratings and reviews where available. If it's not as described, report it within 2 days of delivery to get your money back.

Why run AI locally?

The honest tradeoff: the biggest cloud models are still smarter on the hardest problems, and API pricing is cheap right now. Local wins on privacy, control, and always-on use. And it all runs on a few hundred watts at your desk, about what a gaming session draws, with no data center or cooling water involved. Plenty of people run both.

Privacy

Run local-only software and your prompts, files, and notes stay on your machine. Nothing has to leave your desk.

Ownership

Buy the hardware once and it's yours. Cloud pricing, limits, and terms change on someone else's schedule. A model on your desk doesn't.

No rate limits

No 'you've hit your limit, try again later.' No surprise price changes. The model is always available because it's on your desk.

Total control

Tinker, swap models, automate things, leave it running overnight. It's your machine, so do what you like with it.

Memory upgrades

See all
32GB and up (DDR4/DDR5) with headroom to spill big MoE models into system RAM.

Fast storage

See all
1TB+ SATA, NVMe, and PCIe SSDs. Model files stack up 4–40GB at a time.

Build it with people who've done it before

Get advice and support from our community of expert PC builders

Join the Discord