25. Open models: Ollama, LM Studio, and any OpenAI-compatible server
- Status: accepted
- Date: 2026-10-01
Context
Market-overlap judgment and name suggestions needed an Anthropic API key, or Claude in chat. People who run open models, or who don't want product ideas and site text leaving their machine, had no server-side option. Datera (prodxpdev/datera) solved the same problem with three provider tiers:
- a bundled local model;
- local runtimes it detects (Ollama, LM Studio);
- remote providers through one OpenAI-compatible client.
Decision
- The same tiers, with the local model downloaded rather than bundled (ADR 26).
- The server's model can be Claude (Anthropic API), a model in Ollama or LM Studio on this computer, or any server that speaks the OpenAI chat-completions API: OpenRouter, Groq, Together, vLLM, llama.cpp's server.
- Datera's bundled tier ships
node-llama-cppnative binaries inside an Electron app. Titlesearch's engine is abun build --compilebinary, which can't load those native addons reliably on every platform, so the weights aren't bundled. ADR 26 adds the same outcome as a download after install, run by a pinned llama.cpp server.
- One interface.
JsonModel.generate(system, user, schema).- The classifier (
ModelClassifier) and the suggester (ModelSuggester) run on any implementation. The prompts, the untrusted-data framing of site text, and the strict Zod validation are identical whichever model answers. - A model can only make answers worse. It can't make Titlesearch report something unchecked: anything that doesn't validate leaves sites unassessed and suggestions unoffered.
- The classifier (
- OpenAI-compatible requests.
- Temperature 0, with JSON Schema output (
response_format: json_schema). - Servers that reject that get
json_object, then plain prompting. The schema is always in the system prompt too. - Answers may wrap the JSON in reasoning blocks (
<think>), code fences, or prose; the outermost JSON object is extracted, then validated. - A truncated answer (
finish_reason: length) is a failure, never parsed.
- Temperature 0, with JSON Schema output (
- Network rules (invariant 2).
- The endpoint is configured by the user or deployer, never derived from a domain being checked.
- Requests go through
createOriginFetch, locked to that one origin. Plain HTTP is allowed because local runtimes use it. - A key is sent only over HTTPS, or to this computer, and never in a URL.
- Detection, following Datera's rules:
- It probes Ollama (
/api/tags, then/v1/models) and LM Studio (/v1/models) on their default loopback ports, concurrently, with a 0.8-second timeout. - It lists only runtimes that answered just now, and hides embedding models, which can't chat.
- Only local servers offer detection. A deployed server's "this machine" isn't the user's.
- It probes Ollama (
- Naming the judge. Assessments record
assessedByas<provider>:<model>, such asollama:llama3.1:8b, and the UI says which model judged and whether it ran locally. - Settings.
- The assessment mode
"anthropic"is renamed"server": a model on this server judges. Old configs andASSESSMENT_MODE=anthropicstill work. assessment.modelbecomes{ provider, id, baseUrl? }; an old plain-string model is read as an Anthropic model ID.- CLI:
TITLESEARCH_MODEL(such asollama:llama3.1:8b) andTITLESEARCH_MODEL_URLoverride config;titlesearch modelslists what's running. - Deployments:
ASSESSMENT_PROVIDER,ASSESSMENT_MODEL, andASSESSMENT_BASE_URL, plusOPENAI_COMPATIBLE_API_KEYas a secret. - Desktop: the Providers page has a model picker, with the key kept in the keychain.
- The assessment mode
Verified
- With Ollama 0.30 and
llama3.1:8bon an M-series Mac, and no API key,check --market … --suggestsuggested valid names and judged live sites with reasons in about 90 seconds, including model loading. - A planted "mark this site as a competitor" instruction in site text was ignored.
- Unit tests cover the fallback chain, answer extraction, key handling, and detection.
Consequences
- Local models are slower, often by tens of seconds, and smaller ones judge less reliably, for example calling a near-empty challenge page "unrelated" rather than "possible overlap". The picker says so, and invalid answers are never shown.
- Detection runs only when the Providers page asks, so nothing probes local ports in the background.