Skip to content

Ollama settings guide

Ollama (Local LLM Runner) · Local LLM · Open Source · Privacy

Chat / LLMFree tier
Token burn
Low burn

Token-efficient; free tiers go a long way. Last reviewed May 2026; plans and model names change often, so check the tool's own pricing page before you buy.

1

Run open models such as Llama, Mistral, Gemma and DeepSeek offline on your own machine; `ollama run <model>` downloads and starts one in a single command, with no API key.

2

On a 16GB laptop, start with small models (around 3–8B parameters); speed depends on your hardware, so test before relying on it.

3

Ollama's API is OpenAI-compatible — set base_url to http://localhost:11434/v1 in your app to switch from OpenAI to local models with zero code changes.

4

Use `ollama serve` to expose a local API server — connect Cursor, Open WebUI, or any OpenAI-SDK app to your local model, making all AI calls free and private.

5

Pick a quantisation that fits your RAM: Q4_K_M is a common balance of quality and speed.

6

Set system context in your Modelfile to create persistent custom personas — build a specialised coding assistant or India-context chatbot that loads with your rules every time.

7

Combine Ollama with AnythingLLM or Open WebUI for a full local ChatGPT-like interface — completely offline, works without internet, fully private for sensitive documents.

API cost estimate

Monthly cost at list prices (September 2026), before and after a token cut you choose.

Now
$1.92
₹161
After a 25% cut
$1.44
₹121
Saved per month
$0.48
₹40

Assumes about 600 input and 200 output tokens per request (low burn). Excludes caching discounts, search or request fees and taxes. Check the vendor's pricing page before budgeting.

Reveaxa guide reference: RVX-T-ollama