baby's first test bench

ramming my way to a useful ollama

For the follow-on to using my gaming rig I have to provide an update to my 2026 February post. Pi is in WSL2, and 2026 February was direct to Windows, so I’ll install into WSL2.

sudo apt-get install zstd
curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen3.5:9b
ollama pull ministral-3:8b
ollama pull gemma4:12b
ollama pull qwen3.8:27b

Inference smoke test first. This loads the model as well so it’ll take some time.

curl -s http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ministral-3:8b",
    "messages": [{"role": "user", "content": "Respond with the single word: OK"}],
    "temperature": 0
  }' | jq

That failed initially and curl -s http://localhost:11434/v1/models returned {"object":"list","data":null} so no models. I had pulled models into ~. So I edited system ollama to see these models: sudo systemctl edit ollama.service and add the following between the comments.

[Service]
Environment="OLLAMA_MODELS=/home/syed/.ollama/models"

Then restart system ollama and give ollama access to the models.

sudo chmod +rx ~
sudo chmod -R +r ~/.ollama/models
sudo systemctl daemon-reload
sudo systemctl restart ollama

Now inference is cooking.

$ curl -s http://localhost:11434/v1/models
{
  "object": "list",
  "data": [ ... ]
}
$ curl -s http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ministral-3:8b",
    "messages": [{"role": "user", "content": "Respond with the single word: OK"}],
    "temperature": 0
  }' | jq
{
  "id": "chatcmpl-355",
  ...
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "OK"
      },
      "finish_reason": "stop"
    }
  ],
  ...
}

This does work in PowerShell as well meaning Windows also sees the local LLM inference.

PS > Invoke-RestMethod -Uri "http://localhost:11434/v1/models" | Select-Object -ExpandProperty data

id             object    created owned_by
--             ------    ------- --------
qwen3.8:27b    model  1788755951 library
gemma4:12b     model  1788755300 library
ministral-3:8b model  1788755017 library
qwen3.5:9b     model  1788754752 library


PS > $body = @{
>>     model = "ministral-3:8b"
>>     messages = @(@{ role = "user"; content = "Respond with the single word: OK" })
>>     temperature = 0
>> } | ConvertTo-Json
PS > (Invoke-RestMethod -Uri "http://localhost:11434/v1/chat/completions" -Method Post -ContentType "application/json" -Body $body).choices[0].message.content

OK

Now add the local providers to ~/.pi/agent/models.json.

{
  "providers": {
    "ollama": {
      "name": "Ollama",
      "baseUrl": "http://localhost:11434/v1",
      "api": "openai-completions",
      "apiKey": "ollama",
      "compat": {
        "supportsDeveloperRole": false,
        "supportsReasoningEffort": false
      },
      "models": [
        {
          "id": "qwen3.5:9b",
          "name": "Qwen 3.5 9B",
          "reasoning": true,
          "input": ["text", "image"],
          "contextWindow": 32768,
          "maxTokens": 8192,
          "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
        },
        {
          "id": "ministral-3:8b",
          "name": "Ministral 3 8B",
          "reasoning": false,
          "input": ["text", "image"],
          "contextWindow": 24576,
          "maxTokens": 8192,
          "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
        },
        {
          "id": "gemma4:12b",
          "name": "Gemma 4 12B",
          "reasoning": true,
          "input": ["text", "image"],
          "contextWindow": 16384,
          "maxTokens": 8192,
          "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
        },
        {
          "id": "qwen3.8:27b",
          "name": "Qwen 3.8 27B",
          "reasoning": true,
          "input": ["text", "image"],
          "contextWindow": 8192,
          "maxTokens": 4096,
          "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
        }
      ]
    }
  }
}

Now smoke test Pi. It passed, right? Right?

$ pi --model ministral-3:8b -p "1+1"
The result of **1 + 1** is **2**. 😊
$ pi --model gemma4:12b -p "1+1"
2
$ pi --model qwen3.5:9b -p "1+1"
1 + 1 = 2
$ pi --model qwen3.8:27b -p "1+1"
2

No. In the experiment Qwen 3.5 initially failed suspiciously early. Opening up the trajectory of the failed run showed the run failed due to max output length … at the wrong length.

...
{
  "role": "assistant",
  "model": "qwen3.5:9b",
  "content": [ {
    "type": "thinking",
    "thinking": "I see - each call overwrites the file. This is probably not ideal for a journal where you want to preserve previous entries on the same date. Let me update",
    "thinkingSignature": "reasoning"
  } ],
  "usage": {
    "cacheRead": 3998,
    "totalTokens": 4096,  # Way below 8192 set in my models.json.
    ...
  },
  "stopReason": "length",  # Whoa whoa why stop so early!?
  "rawStopReason": "length",
  ...
}

To re-test I ran Pi on Qwen 3.5 with a multi-turn prompt: pi --model qwen3.5:9b and then “Scan the internet to find the ten most common words in the SLavic language.” (yes with the typo). Watching ollama with watch -n 1 'curl -s http://localhost:11434/api/ps | jq' shows a curious parameter choice.

{ "models": [ {
  "name": "qwen3.5:9b",
  "model": "qwen3.5:9b",
  ...,
  "context_length": 4096  # <-- Where is this from!?
} ] }

The debug fix was to add Environment="OLLAMA_CONTEXT_LENGTH=32768" on a line in sudo systemctl edit ollama right after the line specifying where I downloaded LLMs. That means Pi’s models.json cannot control Ollama’s context length parameter num_ctx since OpenAI’s API cannot set the context size. What to do?

Use Modelfiles. They wrap base models with parameters via a composability UX coming from Makefile and Dockerfile: FROM <base> PARAMETER num_ctx <> ....

~/.ollama/modelfiles/qwen3.5

FROM qwen3.5:9b
PARAMETER num_ctx 32768
PARAMETER num_predict 8192

~/.ollama/modelfiles/ministral-3

FROM ministral-3:8b
PARAMETER num_ctx 24576
PARAMETER num_predict 8192

~/.ollama/modelfiles/gemma4

FROM gemma4:12b
PARAMETER num_ctx 16384
PARAMETER num_predict 8192

~/.ollama/modelfiles/qwen3.8

FROM qwen3.8:27b
PARAMETER num_ctx 8192
PARAMETER num_predict 4096

Then install all four. You will learn that to do this with the background ollama you must let ollama take the wheel and relinquish ownership of your own folder.

sudo chown -R ollama:ollama /home/syed/.ollama
ollama create qwen3.5 -f ~/.ollama/modelfiles/qwen3.5
ollama create ministral-3 -f ~/.ollama/modelfiles/ministral-3
ollama create gemma4 -f ~/.ollama/modelfiles/gemma4
ollama create qwen3.8 -f ~/.ollama/modelfiles/qwen3.8

Pi’s models will now change but only subtly. It’s the IDs: they do not have the parameter tag anymore.

{ "providers": { "ollama": {
  "name": "Ollama",
  "baseUrl": "http://localhost:11434/v1",
  ...
  "models": [ {
    "id": "qwen3.5:latest",  # <-- the name above in `ollama create`
    "contextWindow": 32768,
    "maxTokens": 8192,
    ...
  }, {
    "id": "ministral-3:latest",
    "contextWindow": 24576,  # <-- keep parameters here too for Pi to know
    "maxTokens": 8192,
    ...
  }, {
    "id": "gemma4:latest",
    "contextWindow": 16384,
    "maxTokens": 8192,
    ...
  }, {
    "id": "qwen3.8:latest",
    "contextWindow": 8192,
    "maxTokens": 4096,
    ...
  } ]
} } }

Re-test the smoke pi --model qwen3.5 -p "1+1" while looking at curl -s http://localhost:11434/api/ps | jq.

Every 1.0s: curl -s http://localhost:11434/api/ps | jq    strawberry-moon: Mon Sep  7 15:39:48 2026

{ "models": [ {
  "name": "qwen3.5:latest",
  "model": "qwen3.5:latest",
  ...
  "context_length": 32768  # <-- finally correct
} ] }

Repeat for the other models and now the configuration is respected. So every custom context window and output token setting needs to be in both Pi’s configuration and Ollama’s Modelfiles. API contracts still matter.

But at least now I’m flexing my local LLMs to their max.

Published by using 1036 words.