Choosing a local model
On a 16 GB graphics card, run qwen2.5vl:7b for photos and documents, and
qwen3:14b or gemma4 for the builder. For chat, the free Gemini Flash Lite
is still better than any local model.
Tested on 2 October 2026 against that day's Cobblr nightly, with Ollama 0.35.0, on a laptop with an RTX 5080 (16 GB of video memory). Every model started cold.
What to pick
- Photos and documents:
qwen2.5vl:7b. It identified 6 of 6 photos and read 40 of 40 document fields with every check passing. It takes about 4 seconds a photo and 11 seconds a page, and it fits on the card (7.3 GB). - The builder:
qwen3:14borgemma4. Both score 0.957.qwen3:14bfits on the card at 11.7 GB.gemma4is far smaller (3.4 GB) and answers a build in about 10 seconds.qwen3:30b-a3bscores highest (0.996) but does not fit on 16 GB and spills into system memory. - Chat: the free Gemini Flash Lite (7 of 8). No local model beat it. The
best local result was 6 of 8, from a large model that took about 18 seconds a
round.
qwen3-coder:30bgets 5 of 8 at about 1 second a round. - One model for everything:
qwen3.5. Photos 6 of 6, documents 39 of 40, chat 5 of 8, about 65 tokens a second, and it fits on the card. Its builder score is weak (0.333). - Two that fit on 16 GB together:
qwen2.5vl:7bwithqwen3.5(13.9 GB in all), orqwen2.5vl:7bwithgemma4(10.7 GB).
gemma4 reports that it can read images, but it cannot. Shown a printed
label, it describes "an abstract background design", whether the photo is a
PNG or a JPEG. It scored 0 on photos and documents. Use it for the builder and
chat only.
The big 27 to 28B models are accurate but too slow on 16 GB. Two of them identified every photo and read every document field, but they spill into system memory and took from about 45 seconds to over four minutes a call. Every builder call ran past Cobblr's two-minute limit.
Smaller cards, and no card
| Video memory | What you get |
|---|---|
| 16 GB | The pair above, both loaded, no pauses. |
| 12 GB | Each model fits on its own. A chat turn straight after a scan waits about 5 to 10 seconds while one model is swapped for the other. |
| 8 GB | The photo model only. No chat model good enough to run actions fits, so leave chat on a hosted provider. |
| No graphics card | Not usable. A large model that only partly spilled into system memory slowed to about 9 tokens a second, and running entirely on the processor is slower still. |
What a local model can and cannot do here
The scan inbox works locally. When you scan, the inbox fills in the name, the picture and the details in the background, and you never wait on it. A local model takes several seconds a photo there, which nobody notices.
The camera's instant answer does not. The moment you point the camera, Cobblr has about a second to say what it sees. No local photo model came close on a home graphics card (several seconds a photo even on the 16 GB laptop above). That moment is answered by the barcode catalog, which takes a few hundredths of a second, or by a hosted model. A barcode on the thing is still the fastest way in.
Local buys free and private, not faster. Same cases, hosted: Gemini Flash Lite answers a chat round in about a second.
llava is a 2023 model that older guides still suggest. Few machines have it,
and if it is not installed, every photo fails. llama3.2-vision does not load
on current Ollama at all ("unknown model architecture: 'mllama'"), so every
request fails.
Accuracy
| Model | Fits on 16 GB? | Chat | Builder | Photos | Documents |
|---|---|---|---|---|---|
qwen2.5vl:7b | yes, 7.3 GB | no tools | no tools | 6/6 | 40/40, checks 4/4 |
minicpm-v:8b | yes, 6.4 GB | no tools | no tools | 4/6 | 36/40, checks 4/4 |
qwen3.5 | yes, 6.6 GB | 5/8 | 0.333 | 6/6 | 39/40, checks 4/4 |
gemma4 | yes, 3.4 GB | 4/8 | 0.957 | 0/6, cannot see | 0/40, cannot see |
qwen3:14b | yes, 11.7 GB | 4/8 | 0.957 | no vision | no vision |
qwen3-coder:30b | no, 20.4 GB (71% on the card) | 5/8 | 0.956 | no vision | no vision |
qwen3:30b-a3b | no, 20.4 GB (71% on the card) | 5/8 | 0.996 | no vision | no vision |
| A 27 to 28B model | no, 17.5 GB (75% on the card) | 5/8 | 0, too slow | 6/6 | 40/40, checks 4/4 |
| Another 27 to 28B model | no, 19.0 GB (59% on the card) | 6/8 | 0.333, mostly too slow | 6/6 | 40/40, checks 4/4 |
Two smaller community models were also run for chat and the builder. Neither beat the models above.
Speed on the test laptop
Seconds for a typical call once the model is loaded, and how long the first call took to load it. These compare models with each other on the same hardware. Your own card will give different numbers.
| Model | Tokens a second | Chat round | Builder call | Photo | Document page | First load |
|---|---|---|---|---|---|---|
qwen2.5vl:7b | about 79 | 4.1 s | 10.8 s | 6 to 8 s | ||
minicpm-v:8b | about 105 | 2.7 s | 5.1 s | 7 to 9 s | ||
qwen3.5 | about 65 | 2.9 s | 32.2 s | 42.6 s | about 6 s | |
gemma4 | about 92 | 2.4 s | 10.0 s | (cannot see) | (cannot see) | 8 to 10 s |
qwen3:14b | about 39 | 7.2 s | 25.7 s | about 10 s | ||
qwen3-coder:30b | about 65 | 1.0 s | 7.5 s | 16 s | ||
qwen3:30b-a3b | about 62 | 8.0 s | 30.0 s | 20 s | ||
| A 27 to 28B model | 6 to 10 | 40.8 s | over 2 min | 195 s | 252 s | 24 s |
| Another 27 to 28B model | 9 to 13 | 17.7 s | over 2 min | 47 s | 99 s | 22 to 25 s |
A blank cell is a job the model cannot do, or one that was not timed.
Hosted models, for comparison
These run on the provider's hardware over the internet, so their speed is not comparable with the table above.
| Model | Chat | Builder | Photos | Documents | Typical call |
|---|---|---|---|---|---|
| Gemini Flash Lite (free) | 7/8 | 0.996 | 6/6 | 39/40, checks 4/4 | chat 4.8 s, photo 9.0 s |
| Gemini Flash (free) | no result: the free daily allowance ran out |
How we test
Each model does the jobs Cobblr gives it, with Cobblr's own prompts and tools, and is scored on the result:
- Photos: six photos of everyday things. A photo scores when the model identifies it correctly.
- Documents: four pages, two sale forms each as a clean scan and as a faint carbon copy. 40 fields to read, plus four checks that the totals and dates hang together.
- Chat: eight requests to the assistant that need it to call the right tool with the right details.
- The builder: nine descriptions of a workspace to build. The score runs from 0 to 1, and 0.4 or more is a pass.
A model is only given the jobs it has the features for: photos and documents need vision, chat and the builder need tool calls. A call that runs past Cobblr's own two-minute limit counts as a failure, because in use it would be one. A call lost to the network is asked again and left out of the totals.
Older scoreboards move to Past results when a newer one replaces this page.