Skip to content
πŸͺΆPivra Desk

Choosing a model

A model on this computer, your own key, or Plus + Credits; which Gemma size fits 8 GB and 16 GB; and how to change it in Settings β†’ Model.

Updated 4 Oct 2026

Pivra Desk is not tied to one model. You choose where the model runs during setup, and you can change it any time in Settings β†’ Model.

The two main choices

Free and private Β· Runs on this computerFaster Β· With a Claude or OpenAI API key
AccountNone. Nothing to pay.An API key from Anthropic or OpenAI, paid as you go. A Claude Pro or ChatGPT Plus subscription does not include one.
Where your text goesNowhere. What you write never leaves this computer.From your computer to that provider, under your own account and their terms. Not through Pivra.
Good atReading, summaries, drafting, short tasks with tools.Long, multi-step work.
SpeedDepends on your computer. On an M2 Pro with 16 GB, Gemma 4 E2B writes about 78 tokens a second (measured 4 October 2026).Fast; depends on the provider.

Getting a key: in the provider's console, add some credit under Billing (five dollars is plenty to start), create a key under API keys, copy it, and paste it into Pivra Desk. Where do I get a … key? under each provider in Settings β†’ Model has the exact steps.

Which model on this computer?

Offline mode runs Google's Gemma 4 or a small Qwen through llama.cpp on your own machine. Pivra Desk downloads the file once from Hugging Face and checks it against a pinned checksum before using it.

ModelDownloadPivra Desk's own note
Gemma 4 E2B3.3 GBSmaller Gemma for 8 GB computers that are already busy. Faster than E4B, less careful.
Gemma 4 E4B5.2 GBThe everyday choice for 8 to 16 GB computers: reading, summaries, drafting, and short tasks with tools.
Gemma 4 12B7 GBNeeds a lot of memory: more careful answers for computers with 24 GB or more. On 16 GB it pushes other apps out and can slow to a crawl.
Qwen2.5 0.5B0.5 GBTiny and fast, for computers with little memory. Short, simple tasks only; it often gets tool use wrong. Automatic memory stays off while it is in use, because it is too small to tell lasting facts from its own replies.

Before it loads a model, the app checks the memory free right now and picks a context size that fits. If memory is tight it uses a shorter context; if it is short, the model still loads but crawls, and the app tells you to close other apps or choose a smaller model in Settings β†’ Model. You can keep more than one model downloaded and switch between them; You can download it again later if you remove one.

On an 8 GB computer

Start with Gemma 4 E2B. Close the browser tabs and apps you are not using before a long task. If it is still slow, Qwen2.5 0.5B answers simple questions quickly but is not reliable with tools. We are measuring 8 GB Macs and Windows laptops and will publish the numbers.

Advanced: other model services

Under Advanced: other model services you can use OpenRouter, Cloudflare AI Gateway, Vercel AI Gateway, or any OpenAI-compatible endpoint, including a local server on your own network: Ollama (http://127.0.0.1:11434/v1) or LM Studio (http://127.0.0.1:1234/v1). Pick the Model from the list, or choose Other… and type its name.

Plus + Credits

The Plus + Credits plan (US$15 or A$23 a month) includes monthly cloud model credits so you can use a cloud model without a key of your own. The in-app credits model is still being built; until it ships, Plus + Credits gives you everything in Plus, and this page will say when the credits are usable. If you need a cloud model today, your own key is the way. See Plans.

Which should I pick?

  • You want nothing to leave your computer: Runs on this computer, Gemma 4 E4B on 16 GB, E2B on 8 GB.
  • You hand the agent long, multi-step jobs: your own Claude or OpenAI key.
  • You already run Ollama or LM Studio: point Pivra Desk at it under Advanced.

Whatever you choose, approvals work the same way: the model proposes, and anything that sends, pays or changes something waits for you.

Was this helpful?

If you have questions or suggestions, email us at support@pivra.ai .