Skip to main content
Version: 2026.8.1

Choosing a local model

On a 16 GB graphics card, run qwen2.5vl:7b for photos and documents, and qwen3:14b or gemma4 for the builder. For chat, the free Gemini Flash Lite is still better than any local model.

Tested on 2 October 2026 against that day's Cobblr nightly, with Ollama 0.35.0, on a laptop with an RTX 5080 (16 GB of video memory). Every model started cold.

What to pick​

  • Photos and documents: qwen2.5vl:7b. It identified 6 of 6 photos and read 40 of 40 document fields with every check passing. It takes about 4 seconds a photo and 11 seconds a page, and it fits on the card (7.3 GB).
  • The builder: qwen3:14b or gemma4. Both score 0.957. qwen3:14b fits on the card at 11.7 GB. gemma4 is far smaller (3.4 GB) and answers a build in about 10 seconds. qwen3:30b-a3b scores highest (0.996) but does not fit on 16 GB and spills into system memory.
  • Chat: the free Gemini Flash Lite (7 of 8). No local model beat it. The best local result was 6 of 8, from a large model that took about 18 seconds a round. qwen3-coder:30b gets 5 of 8 at about 1 second a round.
  • One model for everything: qwen3.5. Photos 6 of 6, documents 39 of 40, chat 5 of 8, about 65 tokens a second, and it fits on the card. Its builder score is weak (0.333).
  • Two that fit on 16 GB together: qwen2.5vl:7b with qwen3.5 (13.9 GB in all), or qwen2.5vl:7b with gemma4 (10.7 GB).
Do not use gemma4 for photos

gemma4 reports that it can read images, but it cannot. Shown a printed label, it describes "an abstract background design", whether the photo is a PNG or a JPEG. It scored 0 on photos and documents. Use it for the builder and chat only.

The big 27 to 28B models are accurate but too slow on 16 GB. Two of them identified every photo and read every document field, but they spill into system memory and took from about 45 seconds to over four minutes a call. Every builder call ran past Cobblr's two-minute limit.

Smaller cards, and no card​

Video memoryWhat you get
16 GBThe pair above, both loaded, no pauses.
12 GBEach model fits on its own. A chat turn straight after a scan waits about 5 to 10 seconds while one model is swapped for the other.
8 GBThe photo model only. No chat model good enough to run actions fits, so leave chat on a hosted provider.
No graphics cardNot usable. A large model that only partly spilled into system memory slowed to about 9 tokens a second, and running entirely on the processor is slower still.

What a local model can and cannot do here​

The scan inbox works locally. When you scan, the inbox fills in the name, the picture and the details in the background, and you never wait on it. A local model takes several seconds a photo there, which nobody notices.

The camera's instant answer does not. The moment you point the camera, Cobblr has about a second to say what it sees. No local photo model came close on a home graphics card (several seconds a photo even on the 16 GB laptop above). That moment is answered by the barcode catalog, which takes a few hundredths of a second, or by a hosted model. A barcode on the thing is still the fastest way in.

Local buys free and private, not faster. Same cases, hosted: Gemini Flash Lite answers a chat round in about a second.

Two models to skip

llava is a 2023 model that older guides still suggest. Few machines have it, and if it is not installed, every photo fails. llama3.2-vision does not load on current Ollama at all ("unknown model architecture: 'mllama'"), so every request fails.

Accuracy​

ModelFits on 16 GB?ChatBuilderPhotosDocuments
qwen2.5vl:7byes, 7.3 GBno toolsno tools6/640/40, checks 4/4
minicpm-v:8byes, 6.4 GBno toolsno tools4/636/40, checks 4/4
qwen3.5yes, 6.6 GB5/80.3336/639/40, checks 4/4
gemma4yes, 3.4 GB4/80.9570/6, cannot see0/40, cannot see
qwen3:14byes, 11.7 GB4/80.957no visionno vision
qwen3-coder:30bno, 20.4 GB (71% on the card)5/80.956no visionno vision
qwen3:30b-a3bno, 20.4 GB (71% on the card)5/80.996no visionno vision
A 27 to 28B modelno, 17.5 GB (75% on the card)5/80, too slow6/640/40, checks 4/4
Another 27 to 28B modelno, 19.0 GB (59% on the card)6/80.333, mostly too slow6/640/40, checks 4/4

Two smaller community models were also run for chat and the builder. Neither beat the models above.

Speed on the test laptop​

Seconds for a typical call once the model is loaded, and how long the first call took to load it. These compare models with each other on the same hardware. Your own card will give different numbers.

ModelTokens a secondChat roundBuilder callPhotoDocument pageFirst load
qwen2.5vl:7babout 794.1 s10.8 s6 to 8 s
minicpm-v:8babout 1052.7 s5.1 s7 to 9 s
qwen3.5about 652.9 s32.2 s42.6 sabout 6 s
gemma4about 922.4 s10.0 s(cannot see)(cannot see)8 to 10 s
qwen3:14babout 397.2 s25.7 sabout 10 s
qwen3-coder:30babout 651.0 s7.5 s16 s
qwen3:30b-a3babout 628.0 s30.0 s20 s
A 27 to 28B model6 to 1040.8 sover 2 min195 s252 s24 s
Another 27 to 28B model9 to 1317.7 sover 2 min47 s99 s22 to 25 s

A blank cell is a job the model cannot do, or one that was not timed.

Hosted models, for comparison​

These run on the provider's hardware over the internet, so their speed is not comparable with the table above.

ModelChatBuilderPhotosDocumentsTypical call
Gemini Flash Lite (free)7/80.9966/639/40, checks 4/4chat 4.8 s, photo 9.0 s
Gemini Flash (free)no result: the free daily allowance ran out

How we test​

Each model does the jobs Cobblr gives it, with Cobblr's own prompts and tools, and is scored on the result:

  • Photos: six photos of everyday things. A photo scores when the model identifies it correctly.
  • Documents: four pages, two sale forms each as a clean scan and as a faint carbon copy. 40 fields to read, plus four checks that the totals and dates hang together.
  • Chat: eight requests to the assistant that need it to call the right tool with the right details.
  • The builder: nine descriptions of a workspace to build. The score runs from 0 to 1, and 0.4 or more is a pass.

A model is only given the jobs it has the features for: photos and documents need vision, chat and the builder need tool calls. A call that runs past Cobblr's own two-minute limit counts as a failure, because in use it would be one. A call lost to the network is asked again and left out of the totals.

Older scoreboards move to Past results when a newer one replaces this page.