Skip to main content
Version: Next (nightly)

Choosing a provider

AI runs on a provider you connect. Cobblr integrates with several, from hosted APIs to a model on your own machine, and runs none of them itself. It installs nothing on your behalf: you bring the account, the key stays yours, and the calls go straight from your instance to whoever you chose.

To connect one:

  1. Pick from the list below. If your box can't run a model of its own, Gemini on the free tier is the easiest start.
  2. Add it on the AI settings page, or under your account as a personal connection you grant to chosen workspaces.

The providers Cobblr works with​

  • Google AI Studio (Gemini). A free Google key. No card, no cloud account. The easiest starting point if you can't run a model on your own hardware. See Gemini on the free tier for the one setting to get right.
  • Anthropic (Claude). Your Anthropic API key. Strong on the chat and vision assists.
  • OpenAI. Your OpenAI API key.
  • OpenRouter (one key, any model). One key that reaches many models behind a single gateway, and you name the model. Requests transit OpenRouter's infrastructure, which is the trade for the convenience, and it is stated on the key field rather than hidden.
  • Ollama (local). A model you run yourself with Ollama, reached at its URL. An optional bearer token covers a remote or proxied endpoint. When the URL is on your own machine or LAN, a How Cobblr reaches it option can route the call through your edge bridge instead of directly. Which local model to run for each job, measured on a 16 GB card, is on Choosing a local model.
  • OpenAI-compatible (LM Studio, vLLM, and similar). Any server that speaks the OpenAI v1 API. You give the base URL, and the key and model are optional depending on the server. It carries the same How Cobblr reaches it option for a URL on your own network.
  • Local AI (via edge bridge). Reaches a model on your own device through a live edge channel, for when Cobblr cannot reach the URL directly (the model runs on your laptop or LAN, not on the server). It carries no credential and stays inert until an edge agent connects. Set it up under personal connections.

Gemini on the free tier​

Free, no card, no cloud account to set up.

  • Set the model to gemini-flash-lite-latest. Copy that exactly, -latest included.
  • The free plan allows 500 requests a day, which covers ordinary use comfortably.
  • Let a batch work through at its own pace instead of starting everything at the same moment. Filing a big pile all at once can trip a per-minute limit, and the scans that trip it fail.
The free plan learns from what you send

Google uses what you send on the free plan to improve its models. That is a fair trade for a photo of a monitor or a book spine, and a poor one for anything private. It is your key and your call. A paid Google key, or a model you run yourself, does not do this.

Why -latest, and why Lite rather than Flash

Google retires models that have a version number in the name, and gemini-2.5-flash has already stopped working, so anything pinned to a number will break one day without warning. The -latest name keeps up on its own.

The stronger Flash model is free too, but allows only 20 a day, so Lite is the better default even though it is the smaller model.

Free plans change. This is right as of writing. The -latest name survives a model being retired. It cannot survive a free plan being withdrawn.

Hosted or local​

  • Quality and effort. A hosted API is the least setup and generally the strongest models. A local model is free to run and keeps data on your hardware, but you provide the machine and pick a model good enough for the job.
  • Where the data goes. A hosted provider receives the content each call needs (a photo, a description). A local model, direct or over the edge bridge, keeps that on your own network.
  • One key or many. A single provider is simplest. OpenRouter is one key across many models if you want to switch models without managing several keys, at the cost of routing through their gateway.

Per-capability models and budgets​

  • Some jobs have a row of their own on the AI page, so you can put them on a different model from everyday chat and see what each costs: Build a workspace from a description (run a handful of times, so worth a stronger model, and on the free Google tier it already defaults to the fuller one), First-look question and Split a photo into items (run often, fine on a cheaper or local model). Leave a row alone and it uses whatever the workspace would otherwise use.
  • Pick a model per job. A provider covers several jobs (chat, image classification, text extraction, and so on), and you can pick which model handles each, so a cheap model does the routine work and a stronger one handles the harder calls.
  • Set a monthly budget on a provider: a ceiling on what it is allowed to spend.
Pinning needs a provider the workspace owns

When its AI comes only from a shared connection, there is nothing to pin to, and the page says so inline rather than opening an edit that cannot save. A job whose pinned provider is later removed falls back to automatic, with a clear-pin action when you open it.

Which providers let Cobb act​

Reading your records and proposing changes needs a provider that supports tool calling.

  • Anthropic, OpenAI, OpenRouter, and any tool-capable OpenAI-compatible or local server all do, and Cobb runs its full loop with them.
  • A provider without tool calling still chats. Cobb drops to a simpler one-move-at-a-time mode.
  • How this AI runs tools is a choice on the Ollama, OpenAI-compatible, and Local-AI connections, because some local backends are agents that run tools themselves and only hand back text. Leave it on the standard setting for an ordinary model. Set it to "runs tools itself" to give that backend read-only access to the workspace, so Cobb can read your data through it. That path is read-only. To let Cobb make changes, connect a tool-calling provider.

Personal connections​

Hold the key on your own account instead of putting it on the workspace, and grant it to the workspaces you choose. The workspace gets the capability and never sees the credential, and the grant comes back whenever you take it back. This is how a shared workspace runs AI on one member's key without handing that key around. See Personal connections.

Turning it off​

  • Use AI in this workspace, on Configuration → AI, turns AI off for a single workspace. Owners and admins can flip it. A member's own personal connection still works there.
  • On a self-hosted install, the operator can also disable all built-in AI with one switch, regardless of what keys are set.
  • That hard floor, and the privacy detail of what each call sends, are on the AI settings page.