AI on your computer, with Ollama
The workbench can summarise a paper, answer a question from your own papers, check a claim against the source it cites and search by meaning, all without anything leaving your computer. The models that do it run in Ollama, a free program that runs AI models on your own machine. There is no account, no key and nothing to pay, and once a model is downloaded it works offline.
Ollama is optional. Without it the workbench still keeps, reads and searches your papers by their words, and you can use a hosted model from Anthropic, OpenAI or Google with your own key instead. With it, the AI is as private as the rest of your vault.
Built The workbench on your computer finds Ollama by itself. The workbench in the cloud cannot reach a model on your computer, and uses hosted models only.
Installing Ollama
Once, on each computer. Download it from ollama.com/download.
- Windows
-
Run OllamaSetup.exe. It installs for your account and runs in the background, with its
icon in the notification area, and starts with Windows. Or, in a terminal:
winget install Ollama.Ollama. - macOS
-
Open the download, move Ollama to Applications and open it once. It runs from the menu
bar and starts when you log in. Or, with Homebrew:
brew install ollama. - Linux
-
In a terminal:
curl -fsSL https://ollama.com/install.sh | sh. It installs Ollama as a service that starts with the computer.
To check it is running, open http://127.0.0.1:11434 in a browser: it says
Ollama is running. That address is on your computer only.
The models to download
In a terminal (PowerShell, or Terminal on a Mac), one command for each. Start with the first two: about 5 GB together.
- nomic-embed-text (about 270 MB)
-
Searching by meaning, for papers in English: the workbench's default.
ollama pull nomic-embed-text - llama3.1 (about 4.9 GB)
-
The model that writes: summaries, answers from your papers, claim checks. The workbench's default.
ollama pull llama3.1 - bge-m3 (about 1.2 GB)
-
Searching by meaning when your papers are in several languages, instead of nomic-embed-text.
ollama pull bge-m3 - llava (about 4.7 GB)
-
A model that can look at pictures, for redrawing a figure from a paper as a diagram or a table.
ollama pull llava - qwen2.5-coder (about 4.7 GB)
-
Trained on code, for having the workbench write and test code.
ollama pull qwen2.5-coder - qwen3-embedding:0.6b (about 640 MB)
-
Searching code by meaning, in repositories you bring in and code the workbench wrote.
ollama pull qwen3-embedding:0.6b
You need not remember these. The workbench's Settings shows the command for any model it is missing, with a button to copy it.
Choosing them in the workbench
In Settings, on the workbench on your computer.
- Searching: under Searching by meaning, the model is marked installed once Ollama has it. Choose it and press Use this model. Your papers are then indexed while the workbench is idle, and searches use it once that is done.
- Writing: under The model that writes, choose
ollama:llama3.1, or any model you have pulled, written asollama:and its name. Under Which model does what, you can give a job a model of its own: a picture-reading model for figures, a coding model for code. - If you started Ollama after the workbench, press Check again on the Searching tab.
Every call is on the record. Whatever a model is asked is written down before it is asked, with the model and its version, and kept in the workbench's Evidence, so an answer can always be traced to what produced it.
What your computer needs
Speed depends far more on the computer than on the workbench.
A model like llama3.1 needs about 8 GB of memory free while it works, and 16 GB in the computer is comfortable. A graphics card Ollama can use - NVIDIA, AMD, or any Apple silicon Mac - makes it many times faster.
Without one it still works, on the processor, but slowly. Measured on a laptop with no such graphics card: an answer from eight passages of your papers took about four minutes, and checking one claim against its source between one and five. Searching by meaning is quick either way, and the indexing it needs runs in the background.
The workbench gives a model up to ten minutes to answer, and sizes what it asks to fit the model: a request too long for the model is refused with the reason, never quietly cut short.
When it does not work
The workbench says what it found.
- "Ollama is not answering"
- Start Ollama, or install it. The workbench looks for it at
127.0.0.1:11434. - "Searched by words only"
- The search model is not installed yet, or your papers are still being indexed. Settings, on the Searching tab, says which.
- Very slow, or the computer runs out of memory
-
Close what you can, or use a smaller model:
ollama pull llama3.2is about 2 GB, and is chosen asollama:llama3.2. - Ollama on another port or computer
-
Set the
OLLAMA_HOSTenvironment variable before starting the workbench, for exampleOLLAMA_HOST=http://127.0.0.1:11500.
More is in the documentation, and questions are welcome in the forum.