What you get
A Python script that sends a message to a chat model running on your own computer and prints the reply, with no API key and no bill.
Prerequisites
- A Mac, Linux or Windows computer.
- Python 3.8 or newer.
- About 7.5 GB of free disk space for the model. The Ollama quickstart says the download is about 7.2 GB and the model library page lists 7.5GB.
- Memory: Ollama recommends 8 GB of available VRAM, or unified memory on a Mac. With less VRAM it can use system RAM, but replies may be slower.
- An internet connection for the install and the model download. After that the model runs on your machine.
Steps
-
Install Ollama.
macOS or Linux:
curl -fsSL https://ollama.com/install.sh | sh
Windows (PowerShell):
irm https://ollama.com/install.ps1 | iex
You can also download the installer from https://ollama.com/download.
-
Download the model. gemma4:e2b is the smallest Gemma 4 model in Ollama's library and the one the official quickstart uses for local runs.
-
Install Ollama's official Python library.
pip install ollama==0.6.3
-
Create chat.py.
from ollama import chat
response = chat(
model='gemma4:e2b',
messages=[
{'role': 'user', 'content': 'Why is the sky blue? Answer in two sentences.'},
],
)
print(response.message.content)
There is no API key to set. The library talks to the Ollama server on your machine, which listens on 127.0.0.1 port 11434 by default.
Test it
Make sure Ollama is running, then run:
You should see a short answer about the sky printed in your terminal. The first reply can take longer while the model loads into memory.
If you get a connection error, Ollama is not running. Start the Ollama app and run the script again.
If you get a model not found error, run ollama pull gemma4:e2b again and check the name matches exactly.