Cloud APIs bill per token, log your prompts, and go down at the worst moment. A local LLM does none of that. Three tools dominate the space: Ollama, LM Studio, and llama.cpp. They share one engine and serve three different kinds of users. This guide shows you which one fits your hardware, your skill level, and your workflow, with working commands for each.
Quick Takeaways
- llama.cpp is the C/C++ inference engine underneath most local LLM tools. It gives maximum control and the best raw performance tuning.
- Ollama is a CLI and background service that wraps llama.cpp. It offers one-line model pulls and an OpenAI-compatible API at
localhost:11434. - LM Studio is a desktop GUI with a built-in model browser and a local server. It is the fastest path for non-developers.
- Most 7B-8B models run well on 16 GB RAM at Q4_K_M quantization. No GPU is required, but one makes responses far faster.
At-a-Glance Comparison
| Feature | llama.cpp | Ollama | LM Studio |
|---|---|---|---|
| Interface | CLI / C++ library | CLI + REST API | Desktop GUI + server |
| Difficulty | Advanced | Beginner-friendly | Easiest |
| License | MIT (open source) | MIT (open source) | Free, closed-source app |
| Model format | GGUF | GGUF (own registry) | GGUF (+ MLX on Mac) |
| API compatibility | OpenAI-style via llama-server |
OpenAI-style + native API | OpenAI-style server |
| Fine-grained control | Full | Moderate (Modelfile) | GUI sliders |
| Best for | Engineers, embedded, tuning | Developers, scripting, apps | Beginners, quick testing |
| Cost | Free | Free | Free |
Why Run an LLM Locally?
Four reasons drive most adoption.
- Privacy. Prompts and documents never leave your machine. This matters for legal, medical, and proprietary code.
- Cost. After the hardware, inference is free. There is no per-token billing.
- Offline access. Work on a plane, in a secure facility, or on a flaky connection.
- Control. Pick your model, quantization, context window, and sampling settings. Nobody deprecates the model under you.
The trade-off is quality and speed. A local 8B model will not match a frontier cloud model on hard reasoning. For summarization, drafting, extraction, code completion, and RAG over private files, it is often enough.
Hardware and Quantization Basics
Before choosing a tool, understand what limits you: memory, not raw compute.
How Much Memory Do You Need?
A model’s weights must fit in RAM or VRAM, plus room for the KV cache, which grows with context length. A rough rule: at Q4 quantization, budget about 0.6 GB per billion parameters, plus 1-2 GB overhead.
| Model Size | Q4_K_M Size (approx.) | Minimum Memory | Typical Hardware |
|---|---|---|---|
| 3B | ~2 GB | 8 GB RAM | Laptops, mini PCs |
| 7B-8B | ~4.5-5 GB | 16 GB RAM / 8 GB VRAM | Gaming laptops, M-series Macs |
| 13B-14B | ~8-9 GB | 16-24 GB | RTX 4070-class, 24 GB Macs |
| 32B-34B | ~19-20 GB | 32 GB | RTX 4090, 32 GB+ Macs |
| 70B | ~40 GB | 48-64 GB | Dual GPUs, 64 GB+ unified memory |
Understanding Quantization
Quantization compresses weights from 16-bit floats to fewer bits. It cuts memory use and speeds up inference, with a small quality cost.
| Level | Bits/Weight (approx.) | Quality Loss | Use Case |
|---|---|---|---|
| Q8_0 | ~8.5 | Near zero | Plenty of memory |
| Q5_K_M | ~5.7 | Very low | Quality-focused |
| Q4_K_M | ~4.8 | Low | Best default balance |
| Q3_K_M | ~3.9 | Noticeable | Tight memory |
| Q2_K | ~3.0 | High | Last resort |
Start with Q4_K_M. Move up to Q5 or Q6 only if you have spare memory and notice quality problems.
CPU vs GPU vs Apple Silicon
- NVIDIA GPUs use CUDA and offer the highest throughput per dollar.
- Apple Silicon uses unified memory, so a 64 GB Mac can load models that need multiple consumer GPUs on a PC. Metal acceleration is built in.
- AMD GPUs work through ROCm or Vulkan, with varying support by card and OS.
- CPU-only works but is slow. Expect a few tokens per second on 7B models.
llama.cpp: The Engine Underneath
llama.cpp is an open-source inference library written in C/C++. Ollama and LM Studio both build on it. It introduced the GGUF file format, which packs weights, tokenizer, and metadata into one file.
Strengths
- Highest degree of control over threads, batch size, GPU layer offload, KV cache quantization, and speculative decoding
- Runs on nearly anything: x86, ARM, CUDA, Metal, Vulkan, even a Raspberry Pi
- Newest features land here first
- Ships
llama-server, an OpenAI-compatible HTTP server
Weaknesses
- You manage model downloads, builds, and flags yourself
- Flag names and defaults change between releases
- No GUI
Install and Run
On macOS or Linux with Homebrew:
brew install llama.cpp
Or build from source with CUDA support:
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j
Run a model pulled directly from Hugging Face:
# -hf : download a GGUF from Hugging Face
# -ngl : number of layers offloaded to GPU (99 = all that fit)
# -c : context window in tokens
llama-cli -hf bartowski/Llama-3.1-8B-Instruct-GGUF:Q4_K_M -ngl 99 -c 8192
Serve an OpenAI-compatible API:
# --temp : sampling temperature
# --port : HTTP port
llama-server -hf bartowski/Llama-3.1-8B-Instruct-GGUF:Q4_K_M \
-ngl 99 -c 8192 --port 8080
Key Flags to Tune
| Flag | Purpose | Typical Value |
|---|---|---|
-ngl |
GPU layers offloaded | 99 for full offload |
-c |
Context length (tokens) | 4096–32768 |
-t |
CPU threads | Physical core count |
-b |
Prompt batch size | 512–2048 |
--flash-attn |
Faster attention, lower memory | Enable on supported GPUs |
-ctk / -ctv |
KV cache quantization | q8_0 to save VRAM |
Flag names evolve, so run llama-cli --help to confirm what your build supports.
Ollama: The Developer’s Default
Ollama runs as a background service. You pull a model by name and talk to it over a CLI or REST API. It handles download, GPU detection, and memory management for you.
Strengths
- One-line setup:
ollama run llama3.1 - Built-in model registry with sensible default quantizations
- Modelfile system for custom system prompts and parameters
- Native REST API plus OpenAI-compatible endpoints
- Models stay loaded in memory and unload after idle time
Weaknesses
- Less low-level control than raw llama.cpp
- Defaults can be conservative. For example, the default context window has historically been small, so set it explicitly for long documents
- Registry models are repackaged, so verify the exact quantization with
ollama show
Install and Run
Install on Linux:
curl -fsSL https://ollama.com/install.sh | sh
On macOS and Windows, download the installer from ollama.com. Then:
ollama pull llama3.1:8b # download the model
ollama run llama3.1:8b # interactive chat
ollama list # show installed models
ollama ps # show loaded models and memory use
Customize with a Modelfile
FROM llama3.1:8b
PARAMETER temperature 0.2
PARAMETER num_ctx 8192
SYSTEM "You are a concise technical assistant. Answer in plain English."
ollama create tech-helper -f Modelfile
ollama run tech-helper
Call the API from Python
import requests
resp = requests.post(
"http://localhost:11434/api/generate",
json={
"model": "llama3.1:8b",
"prompt": "Explain KV cache in two sentences.",
"stream": False,
"options": {"temperature": 0.2, "num_ctx": 8192},
},
)
print(resp.json()["response"])
Use the OpenAI SDK against Ollama’s compatible endpoint:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama") # key is ignored
out = client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Give me three uses for local LLMs."}],
temperature=0.3,
)
print(out.choices[0].message.content)
LM Studio: The Friendly GUI
LM Studio is a desktop app for Windows, macOS, and Linux. You search for models inside the app, click download, and chat. It shows memory estimates before you load a model, which prevents the most common beginner mistake: picking a model that is too large.
Strengths
- Built-in Hugging Face model browser with hardware compatibility hints
- Visual controls for temperature, context length, and GPU offload
- Local server mode with OpenAI-compatible endpoints
- Supports MLX models on Apple Silicon, which often run faster than GGUF on Macs
- Chat with documents (local RAG) without writing code
Weaknesses
- The application is closed-source (free for personal use; check current terms for commercial use)
- Heavier than a CLI tool
- Less suited to headless servers and automation, though a CLI exists
Setup Steps
- Download the installer from lmstudio.ai.
- Open the Discover tab and search for a model such as
Llama 3.1 8B Instruct. - Choose a Q4_K_M build that shows a green fit indicator.
- Click Download, then open the Chat tab and load the model.
- For API access, open the Developer tab and start the local server (default port 1234).
Call the LM Studio Server
from openai import OpenAI
client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")
out = client.chat.completions.create(
model="local-model", # use the model identifier shown in LM Studio
messages=[{"role": "user", "content": "Summarize what GGUF is."}],
temperature=0.2,
)
print(out.choices[0].message.content)
Head-to-Head: Which Tool Wins?
| Criteria | llama.cpp | Ollama | LM Studio |
|---|---|---|---|
| Setup time | 10-30 min (build) or 2 min (brew) | ~2 min | ~5 min |
| Learning curve | Steep | Gentle | Minimal |
| Raw speed potential | Highest (tunable) | Near parity | Near parity |
| Model discovery | Manual (Hugging Face) | Curated registry | Built-in browser |
| Headless/server use | Excellent | Excellent | Limited |
| Docker/production | Excellent | Good | Poor |
| Multi-model switching | Manual | Automatic | GUI-driven |
| Cloud-style API | Yes | Yes | Yes |
Speed differences between the three are usually small, because they share the same core kernels. Gaps come from default settings, not fundamental design. A tuned llama.cpp setup can pull ahead by a modest margin. Benchmark on your own hardware with your own prompts before drawing conclusions.
Quick Decision Guide
- You have never run a local model: Start with LM Studio.
- You want to build an app or script: Use Ollama.
- You need maximum performance, custom builds, or edge deployment: Use llama.cpp.
- You run a team server: Use llama.cpp’s
llama-serveror Ollama behind a reverse proxy.
Best Models to Run Locally
Model availability moves quickly. These families have consistently performed well in the 3B-14B range. Verify the latest releases on Hugging Face and each tool’s catalog.
| Family | Typical Sizes | Strength | License Type |
|---|---|---|---|
| Llama 3.x | 1B, 3B, 8B, 70B | General chat, broad support | Community license |
| Qwen 2.5 / 3 | 0.5B-72B | Coding, multilingual | Mostly Apache 2.0 |
| Gemma | 2B-27B | Efficient, strong for its size | Gemma terms |
| Mistral / Mixtral | 7B, MoE variants | Speed, instruction following | Apache 2.0 (most) |
| Phi | 3B-14B | Reasoning at small scale | MIT |
| DeepSeek (distills) | 7B-70B | Chain-of-thought reasoning | MIT (distills) |
Check each model’s license before commercial use. Terms differ significantly.
Practical Workflows and Real-World Use Cases
Workflow 1: Private Document Q&A (Local RAG)
Goal: Ask questions about confidential PDFs without sending data to the cloud.
Stack: Ollama + an embedding model + a vector store.
ollama pull llama3.1:8b
ollama pull nomic-embed-text # 768-dimension embeddings
pip install chromadb ollama
import ollama, chromadb
client = chromadb.Client()
col = client.create_collection("docs")
chunks = ["Refunds are processed within 14 days.", "Support hours are 9-5 EST."]
for i, text in enumerate(chunks):
emb = ollama.embeddings(model="nomic-embed-text", prompt=text)["embedding"]
col.add(ids=[str(i)], embeddings=[emb], documents=[text])
question = "How long do refunds take?"
q_emb = ollama.embeddings(model="nomic-embed-text", prompt=question)["embedding"]
context = col.query(query_embeddings=[q_emb], n_results=1)["documents"][0][0]
answer = ollama.chat(
model="llama3.1:8b",
messages=[{"role": "user",
"content": f"Answer using only this context:\n{context}\n\nQ: {question}"}],
options={"temperature": 0.1},
)
print(answer["message"]["content"])
Set temperature low (0.0-0.2) for factual retrieval tasks to reduce hallucination. Chunk documents at roughly 300-500 tokens with a small overlap.
Workflow 2: Local Coding Assistant
Point your editor extension at an OpenAI-compatible endpoint. Pair a code-tuned model such as a Qwen coder variant with LM Studio’s server or Ollama. Keep temperature at 0.1-0.3 for deterministic completions and set the context window to at least 8192 tokens so the model sees enough of your file.
Workflow 3: Prompt Engineering Template for Small Models
Small models need tighter instructions than frontier models. Use this structure:
ROLE: You are a [specific role].
TASK: [One clear action.]
CONSTRAINTS: Max [N] words. Output as [format]. Do not [unwanted behavior].
EXAMPLE INPUT: [sample]
EXAMPLE OUTPUT: [sample]
INPUT: {user_text}
Few-shot examples improve consistency more on 7B-8B models than on large ones. Ask for JSON output and validate it in code.
Workflow 4: Air-Gapped Enterprise Deployment
Teams in regulated industries run llama-server or Ollama inside a private network. A common pattern:
- Mirror GGUF files to an internal artifact store.
- Run the server in a container with GPU passthrough.
- Place an authenticated reverse proxy in front of it.
- Log usage internally for audit.
Common Problems and Fixes
| Problem | Likely Cause | Fix |
|---|---|---|
| Very slow output | Model running on CPU | Check GPU offload (-ngl, ollama ps) |
| Out-of-memory crash | Model or context too large | Use a smaller quant or reduce context |
| Model forgets earlier text | Context window too small | Raise num_ctx / -c |
| Gibberish output | Wrong chat template | Use the instruct variant and the default template |
| Port already in use | Another server running | Change the port or stop the other process |
Final Recommendation
Install LM Studio first if you want results in ten minutes. Move to Ollama when you start scripting or building apps. Learn llama.cpp when you need control that neither wrapper exposes. Because all three read GGUF files, you can switch tools without re-downloading everything.
Frequently Asked Questions
Is Ollama better than LM Studio?
Neither is universally better. Ollama suits developers who want a CLI, scripting, and an API. LM Studio suits beginners who prefer a graphical interface and visual model browsing. Both run the same GGUF models with similar speed.
Can I run an LLM locally without a GPU?
Yes. All three tools run on CPU. A 3B-8B model at Q4_K_M works on a laptop with 8-16 GB of RAM, but expect slower output than on a GPU. Smaller models and lower quantization improve CPU speed.
How much RAM do I need to run Llama 3 locally?
For the 8B model at Q4_K_M, plan for 16 GB of RAM (about 5 GB for weights plus overhead and context). The 70B model needs roughly 48-64 GB at Q4 quantization.
Is llama.cpp faster than Ollama?
Slightly, when tuned. Ollama runs llama.cpp under the hood, so baseline speed is close. llama.cpp lets you adjust threads, batch size, KV cache type, and other flags that can improve throughput on specific hardware.




