10,000+ gaming PCs bought and sold on Jawa!
Start here

Running AI on your own computer, explained simply.

You don't need a data center or a computer science degree. With the right hardware and open tools, you can run useful AI models on your own machine.

Shop Local AI Builds
The simple version

What does “running AI locally” mean?

Running AI locally means the model runs on your computer instead of only through a cloud service. You download a model, open it in a local app, and use your own hardware to generate answers, images, transcripts, or summaries.

1

The model lives on your machine

You download it once, then run it from your own computer when you need it.

2

Your hardware does the work

The GPU, memory, CPU, and storage determine what fits and how fast it runs.

3

Cloud still has a place

Cloud tools are still better for some frontier tasks. Local shines when privacy, repeat use, and control matter.

Choose your path

Which hardware? NVIDIA, AMD, Intel, and Mac

All four can run local models. They just make different trade-offs on speed, memory, and how much setup you'll do.

NVIDIA

Coding agents · image generation · experimentation

The default choice for many local AI tools. CUDA support gives NVIDIA the broadest compatibility, strongest tutorials, and fastest path for most buyers.

AMD

Value shoppers · image generation · tinkering

AMD cards can offer lots of memory for the money. Compatibility is improving, but setup may require more patience depending on the tool.

Intel

Budget builds · experimental setups

Intel Arc cards and new chips are becoming more capable for local AI, but the ecosystem is still catching up compared with NVIDIA.

Apple Mac

Quiet workstations · large-model fit · simple setup

Apple Silicon shares memory across the system, which helps larger models fit. Macs are quiet and efficient, but usually slower than high-end NVIDIA rigs.

Planning to train or fine-tune your own models?

That's more advanced and CUDA-heavy, so lean NVIDIA with as much VRAM as you can get. Most buyers only run models, where more hardware paths are workable.

The key consideration

GPU memory decides what fits

Smarter models need more memory to run. On a graphics card (GPU) that's called VRAM. Sizes below assume 4-bit quantization (the standard compression that shrinks typical 16-bit model weights to roughly a quarter of their size for a small quality cost). Newer “mixture-of-experts” models bend the rule: they run usefully, if more slowly, on a midrange card plus plenty of system RAM, so don't rule out a smaller GPU to start.

MemoryModel sizeWhat it's good forExample cards
8–12 GBSmall (7–9B)Chat, transcription, note summaries, basic autocompleteRTX 3050 / 4060, or 3060 12GB
16 GBMedium (12–20B)Coding help, document search (RAG), image generationRTX 4060 Ti 16GB, 4070 Ti Super
24 GBLarge (30B-class)Coding agents, autonomous assistants, stronger reasoningRTX 3090, 4090
~40 GB (two cards, or a big Mac)Huge (70B)Near-frontier chat and coding, entirely localTwo RTX 3090s, or a big Mac

Used cards are the value play

If a card arrives and is not as described, Jawa buyer protection helps you send it back.

Good first cards

See all
A used 16GB card is the friendliest place to start running your first models.

Cards that run the big models

See all
24GB and up: enough VRAM to run 30B-class models entirely on the card.
Getting Started: Software

Start with friendly software

A few free apps cover almost everything. Start at the top and work down as you get comfortable.

SoftwarePlatformsWhat it's for
LM StudioWin · Mac · LinuxThe friendliest start: download a model and chat in a polished app.
OllamaWin · Mac · LinuxRun a model with one command. Great for scripting and automation.
JanWin · Mac · LinuxOpen-source, offline, ChatGPT-style desktop app.
Open WebUIBrowser (Docker)Self-hosted web interface that pairs with Ollama for a ChatGPT-like UI.
ComfyUIWin · Mac · LinuxNode-based studio for local image and video generation.
llama.cppWin · Mac · LinuxThe engine many apps are built on, for tinkerers who want full control.
Getting Started: Models

Try a model that matches the job.

Open-weight models live on Hugging Face. For local apps, look for GGUF files (the format LM Studio, Ollama, and llama.cpp use). The landscape moves fast, so treat these as starting points. Qwen3.8 27B is the current default pick for most people, since it fits on a single 24GB card at 4-bit. One caveat: the giant open-weight models in the headlines, like DeepSeek, Kimi, and GLM, weigh hundreds of gigabytes even compressed and need server hardware. The families below are the ones that run well at home.

If you want to…Models to try
General chat & assistantsQwen3.8 27B, Gemma, Mistral, LFM2.5 2.6B (runs on almost anything)
Coding & agentsQwen3.8 27B, gpt-oss, NVIDIA Nemotron 3.5 Lightning
Image & videoFLUX, Qwen-Image, Z-Image Turbo, Stable Diffusion (SDXL), Wan (video)
TranscriptionWhisper, faster-whisper
Vision & documentsQwen3.8 27B (built-in vision), Qwen-VL, Gemma

Complete systems, ready to run

See all
Verified-seller rigs. Skip the build and start running.

Terms you'll keep hearing

The local AI world has a lot of jargon. Here are the terms worth knowing before you shop.

Model
The AI itself. You download it once as a file, then run it as often as you like.
LLM
Large language model: the text-based AI behind chatbots, writing help, and coding tools, and the kind most people run at home.
Parameters
A model's internal settings, also called its weights, counted in billions (a '7B' model has 7 billion). More parameters buy more capability but demand more memory.
Token
A chunk of text, roughly ¾ of a word. A model’s speed is measured in tokens per second (tok/s).
Context window
How much text a model holds in mind at once, counting both your prompt and its reply. A bigger window fits more files or a longer conversation.
VRAM
The memory on your graphics card. It is the single biggest factor in which models you can run.
Unified memory
Apple Silicon's shared pool of memory for the CPU and GPU. It lets a Mac load unusually large models.
Quantization
Shrinking a model to fit in less memory. You trade a little quality for a lot of room.
Inference
Running a model to get an answer. It happens every time you send a prompt, and how fast it runs is your tokens per second.
Training
Teaching a model from scratch on huge amounts of data. It takes enormous compute, so almost nobody does it at home.
Fine-tuning
Adapting a finished model to your own data or task. It is lighter than training, but still leans on a strong NVIDIA GPU.
Open-weight model
A model whose trained weights are published, so anyone can download and run it. Qwen, Gemma, DeepSeek, and Kimi are examples.

A word on model sizes

You download each model once and keep it on disk, and the files are big. A single model is anywhere from about 4 to 40GB, and most people collect a few, so a roomy, fast SSD fills up quicker than you'd expect. To run, a model has to load into memory, ideally your GPU's VRAM, or system RAM when it spills over, so disk and memory together decide how much you can keep on hand and how much you can run at once.

Memory upgrades

See all
32GB and up (DDR4/DDR5) for headroom when models spill into system RAM.

Fast storage

See all
1TB+ SATA, NVMe, and PCIe SSDs, because model files stack up fast.

Build it with people who've done it before

Get advice and support from our community of expert PC builders

Join the Discord