Ollama settings guide
Ollama (Local LLM Runner) · Local LLM · Open Source · Privacy
- Token burn
- Low burn
Token-efficient; free tiers go a long way. Last reviewed May 2026; plans and model names change often, so check the tool's own pricing page before you buy.
Run open models such as Llama, Mistral, Gemma and DeepSeek offline on your own machine; `ollama run <model>` downloads and starts one in a single command, with no API key.
On a 16GB laptop, start with small models (around 3–8B parameters); speed depends on your hardware, so test before relying on it.
Ollama's API is OpenAI-compatible — set base_url to http://localhost:11434/v1 in your app to switch from OpenAI to local models with zero code changes.
Use `ollama serve` to expose a local API server — connect Cursor, Open WebUI, or any OpenAI-SDK app to your local model, making all AI calls free and private.
Pick a quantisation that fits your RAM: Q4_K_M is a common balance of quality and speed.
Set system context in your Modelfile to create persistent custom personas — build a specialised coding assistant or India-context chatbot that loads with your rules every time.
Combine Ollama with AnythingLLM or Open WebUI for a full local ChatGPT-like interface — completely offline, works without internet, fully private for sensitive documents.
API cost estimate
Monthly cost at list prices (September 2026), before and after a token cut you choose.
- Now
- $1.92
- ₹161
- After a 25% cut
- $1.44
- ₹121
- Saved per month
- $0.48
- ₹40
Assumes about 600 input and 200 output tokens per request (low burn). Excludes caching discounts, search or request fees and taxes. Check the vendor's pricing page before budgeting.
Reveaxa guide reference: RVX-T-ollama
More Chat / LLM guides
Claude
Chat / LLM · Long context · Reasoning · Coding
Put standing context in a Project: upload reference files and write project instructions once, so you stop pasting the same background into every chat.
ChatGPT
Chat / LLM · General purpose · Multimodal
Fill in Custom Instructions (Settings → Personalization) with who you are and how you want answers. It applies to every chat, so you stop repeating it.
Gemini
Chat / LLM · Multimodal · Google Workspace
The free plan includes Gemini 3.6 Flash, with Gemini 3.1 Pro when capacity allows. Google AI Plus is ₹399/month and AI Pro ₹1,950/month in India (gemini.google, September 2026).