ollama/ollama-python

★ 10,541⑂ 1,176

Ollama Python library

About ollama/ollama-python

ollama/ollama-python is an open-source project on GitHub, mainly written in Python. Ollama Python library It currently holds 10,541 stars and 1,176 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Local & On-Device AI board.

GitHub Repository Details

Repository ollama/ollama-python · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

Ollama Python Library

The Ollama Python library provides the easiest way to integrate Python 3.8+ projects with Ollama.

Prerequisites

Install

pip install ollama

Usage

from ollama import chat
from ollama import ChatResponse

response: ChatResponse = chat( model='gemma4', messages=[ { 'role': 'user', 'content': 'Why is the sky blue?', }, ], ) print(response['message']['content'])

or access fields directly from the response object

print(response.message.content)

See _types.py for more information on the response types.

Streaming responses

Response streaming can be enabled by setting stream=True.

from ollama import chat

stream = chat( model='gemma4', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}], stream=True, )

for chunk in stream: print(chunk['message']['content'], end='', flush=True)

Cloud Models

Run larger models by offloading to Ollama’s cloud while keeping your local workflow.

Run via local Ollama

1) Sign in (one-time):

ollama signin

2) Pull a cloud model:

ollama pull gpt-oss:120b-cloud

3) Make a request:

from ollama import Client

client = Client()

messages = [ { 'role': 'user', 'content': 'Why is the sky blue?', }, ]

for part in client.chat('gpt-oss:120b-cloud', messages=messages, stream=True): print(part.message.content, end='', flush=True)

Cloud API (ollama.com)

Access cloud models directly by pointing the client at https://ollama.com.

1) Create an API key from ollama.com , then set:

export OLLAMA_API_KEY=your_api_key

2) (Optional) List models available via the API:

curl https://ollama.com/api/tags

3) Generate a response via the cloud API:

import os
from ollama import Client

client = Client(host='https://ollama.com', headers={'Authorization': 'Bearer ' + os.environ.get('OLLAMA_API_KEY')})

messages = [ { 'role': 'user', 'content': 'Why is the sky blue?', }, ]

for part in client.chat('gpt-oss:120b', messages=messages, stream=True): print(part.message.content, end='', flush=True)

Custom client

A custom client can be created by instantiating Client or AsyncClient from ollama.

All extra keyword arguments are passed into the httpx.Client.

from ollama import Client

client = Client(host='http://localhost:11434', headers={'x-some-header': 'some-value'}) response = client.chat( model='gemma4', messages=[ { 'role': 'user', 'content': 'Why is the sky blue?', }, ], )

Async client

The AsyncClient class is used to make asynchronous requests. It can be configured with the same fields as the Client class.

import asyncio
from ollama import AsyncClient

async def chat(): message = {'role': 'user', 'content': 'Why is the sky blue?'} response = await AsyncClient().chat(model='gemma4', messages=[message])

asyncio.run(chat())

Setting stream=True modifies functions to return a Python asynchronous generator:

import asyncio
from ollama import AsyncClient

async def chat(): message = {'role': 'user', 'content': 'Why is the sky blue?'} async for part in await AsyncClient().chat(model='gemma4', messages=[message], stream=True): print(part['message']['content'], end='', flush=True)

asyncio.run(chat())

API

The Ollama Python library's API is designed around the Ollama REST API

Chat

ollama.chat(model='gemma4', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}])

Generate

ollama.generate(model='gemma4', prompt='Why is the sky blue?')

List

ollama.list()

Show

ollama.show('gemma4')

Create

ollama.create(model='example', from_='gemma4', system='You are Mario from Super Mario Bros.')

Copy

ollama.copy('gemma4', 'user/gemma4')

Delete

ollama.delete('gemma4')

Pull

ollama.pull('gemma4')

Push

ollama.push('user/gemma4')

Embed

ollama.embed(model='gemma4', input='The sky is blue because of rayleigh scattering')

Embed (batch)

ollama.embed(model='gemma4', input=['The sky is blue because of rayleigh scattering', 'Grass is green because of chlorophyll'])

Ps

ollama.ps()

Errors

Errors are raised if requests return an error status or if an error is detected while streaming.

model = 'does-not-yet-exist'

try: ollama.chat(model) except ollama.ResponseError as e: print('Error:', e.error) if e.status_code == 404: ollama.pull(model)

GitHub Stars & Activity

10,541Stars
1,176Forks
0Open issues
PythonLanguage

GitHub Popularity

GitHub stars10,541
Forks1,176
Open issues0
Primary languagePython
License-
Stars gained today0
Created-
Last pushed-

Trending History

Trending statusnot on today's boards

Related AI Projects

1

open-webui / open-webui

Python★ 152,601⑂ 22,332
2

HKUDS / nanobot

Python★ 48,395⑂ 8,551
3

chatchat-space / Langchain-Chatchat

Python★ 38,648⑂ 6,265
4

1Panel-dev / MaxKB

Python★ 22,844⑂ 3,157
5

lss233 / kirara-ai

Python★ 19,027⑂ 1,836
6

AsyncFuncAI / deepwiki-open

Python★ 18,018⑂ 2,003
7

langbot-app / LangBot

Python★ 17,926⑂ 1,602
8

MODSetter / SurfSense

Python★ 16,169⑂ 1,538

More AI Rankings