Skip to content

25. Open models: Ollama, LM Studio, and any OpenAI-compatible server ​

  • Status: accepted
  • Date: 2026-10-01

Context ​

Market-overlap judgment and name suggestions needed an Anthropic API key, or Claude in chat. People who run open models, or who don't want product ideas and site text leaving their machine, had no server-side option. Datera (prodxpdev/datera) solved the same problem with three provider tiers:

  1. a bundled local model;
  2. local runtimes it detects (Ollama, LM Studio);
  3. remote providers through one OpenAI-compatible client.

Decision ​

  • The same tiers, with the local model downloaded rather than bundled (ADR 26).
    • The server's model can be Claude (Anthropic API), a model in Ollama or LM Studio on this computer, or any server that speaks the OpenAI chat-completions API: OpenRouter, Groq, Together, vLLM, llama.cpp's server.
    • Datera's bundled tier ships node-llama-cpp native binaries inside an Electron app. Titlesearch's engine is a bun build --compile binary, which can't load those native addons reliably on every platform, so the weights aren't bundled. ADR 26 adds the same outcome as a download after install, run by a pinned llama.cpp server.
  • One interface. JsonModel.generate(system, user, schema).
    • The classifier (ModelClassifier) and the suggester (ModelSuggester) run on any implementation. The prompts, the untrusted-data framing of site text, and the strict Zod validation are identical whichever model answers.
    • A model can only make answers worse. It can't make Titlesearch report something unchecked: anything that doesn't validate leaves sites unassessed and suggestions unoffered.
  • OpenAI-compatible requests.
    • Temperature 0, with JSON Schema output (response_format: json_schema).
    • Servers that reject that get json_object, then plain prompting. The schema is always in the system prompt too.
    • Answers may wrap the JSON in reasoning blocks (<think>), code fences, or prose; the outermost JSON object is extracted, then validated.
    • A truncated answer (finish_reason: length) is a failure, never parsed.
  • Network rules (invariant 2).
    • The endpoint is configured by the user or deployer, never derived from a domain being checked.
    • Requests go through createOriginFetch, locked to that one origin. Plain HTTP is allowed because local runtimes use it.
    • A key is sent only over HTTPS, or to this computer, and never in a URL.
  • Detection, following Datera's rules:
    • It probes Ollama (/api/tags, then /v1/models) and LM Studio (/v1/models) on their default loopback ports, concurrently, with a 0.8-second timeout.
    • It lists only runtimes that answered just now, and hides embedding models, which can't chat.
    • Only local servers offer detection. A deployed server's "this machine" isn't the user's.
  • Naming the judge. Assessments record assessedBy as <provider>:<model>, such as ollama:llama3.1:8b, and the UI says which model judged and whether it ran locally.
  • Settings.
    • The assessment mode "anthropic" is renamed "server": a model on this server judges. Old configs and ASSESSMENT_MODE=anthropic still work.
    • assessment.model becomes { provider, id, baseUrl? }; an old plain-string model is read as an Anthropic model ID.
    • CLI: TITLESEARCH_MODEL (such as ollama:llama3.1:8b) and TITLESEARCH_MODEL_URL override config; titlesearch models lists what's running.
    • Deployments: ASSESSMENT_PROVIDER, ASSESSMENT_MODEL, and ASSESSMENT_BASE_URL, plus OPENAI_COMPATIBLE_API_KEY as a secret.
    • Desktop: the Providers page has a model picker, with the key kept in the keychain.

Verified ​

  • With Ollama 0.30 and llama3.1:8b on an M-series Mac, and no API key, check --market … --suggest suggested valid names and judged live sites with reasons in about 90 seconds, including model loading.
  • A planted "mark this site as a competitor" instruction in site text was ignored.
  • Unit tests cover the fallback chain, answer extraction, key handling, and detection.

Consequences ​

  • Local models are slower, often by tens of seconds, and smaller ones judge less reliably, for example calling a near-empty challenge page "unrelated" rather than "possible overlap". The picker says so, and invalid answers are never shown.
  • Detection runs only when the Providers page asks, so nothing probes local ports in the background.

Apache-2.0. Titlesearch is read-only, and it isn't a trademark search.