Models and configuration

Locate the active configuration, choose models and test the connection. Examples cover provider setup, restarts, reindexing and common errors.

Sources and precedence#

Precedence is shown below. Start commands and the service from the same working directory; a .env in that directory is also read.

precedence
environment variable   >   config.toml   >   built-in default

The Settings screen shows you which of the three each value came from, and if an environment variable is winning, the field goes read-only and says so rather than letting you write something that will have no effect. Editing in the browser writes the file; it never touches your environment.

where is the file
facetmark config path

It is created the first time something writes to it. There is no requirement to have one — a run with no file and no variables is a valid run with all defaults.

Three settings need a restart

embed_backend, embed_dim and local_embed_path decide the shape of the vector store. The screen saves them and then tells you plainly that they take effect next start.

Connect models#

SettingIn plain language
api_keyYour key. Stored in the file, shown back to you masked, and never re-sent when you save an unrelated field.
base_urlWhere the requests go. Anything speaking the OpenAI API works, including something on your own machine.
chat_modelGenerates page summaries during indexing and answers with citations in Ask.
embed_modelTurns text into vectors. This is the one that decides search quality.
Changing the embedding model means reindexing

Vectors from two different models are not comparable. Change it and run a rebuild, or search gets quietly worse in a way no error message will tell you about.

Open Settings over loopback or SSH, fill in the model fields and click Test connection. Testing uses the form values without saving them. Check chat and embedding results, then click Save for that group. Each Save handles only its own group's changes. Restart the service after saving fields marked as requiring a restart before using them in actual jobs.

Provider examples#

Official documentation checked 2026-10-03. Locate your file with facetmark config path. Confirm account access and test a small batch first; retrieval quality has not been re-evaluated for these model updates.

OpenAI
api_key = "sk-..."
base_url = "https://api.openai.com/v1"
chat_model = "gpt-6-luna"
chat_extra_body = "{\"reasoning_effort\":\"none\",\"max_completion_tokens\":4096}"
embed_model = "text-embedding-3-small"
embed_dim = 1536

Luna suits high-volume extraction; gpt-6.1-sol is a higher-cost option (use low reasoning, not none). The current small embedding model remains unchanged. Official documentation.

DeepSeek
api_key = "sk-..."
base_url = "https://api.deepseek.com/v1"
chat_model = "deepseek-flash"
chat_extra_body = "{\"thinking\":{\"type\":\"disabled\"},\"max_tokens\":4096}"
embed_backend = "local"
local_embed_path = "BAAI/bge-m3"
embed_model = "BAAI/bge-m3"
embed_dim = 1024

DeepSeek-V4.1-Flash replaces the old deepseek-chat example. This example uses local embeddings; install facetmark[local] first. Official documentation.

Moonshot / Kimi
api_key = "sk-..."
base_url = "https://api.moonshot.cn/v1"
chat_model = "kimi-k3"
chat_extra_body = "{\"reasoning_effort\":\"low\",\"max_completion_tokens\":8192}"
embed_backend = "local"
local_embed_path = "BAAI/bge-m3"
embed_model = "BAAI/bge-m3"
embed_dim = 1024

moonshot-v1 was retired on 2026-08-31. K3 always reasons: omit temperature and use max_completion_tokens. Local embeddings require facetmark[local]. Official documentation.

Zhipu / GLM
api_key = "..."
base_url = "https://open.bigmodel.cn/api/paas/v4"
chat_model = "glm-5.3-flash"
chat_extra_body = "{\"reasoning_effort\":\"low\",\"max_tokens\":8192}"
embed_model = "embedding-3"
embed_dim = 2048

5.3 Flash is paid and always reasons. For free chat use glm-4.7-flash with {"thinking":{"type":"disabled"},"max_tokens":4096}. embedding-3 remains current and is billed separately. Official documentation.

SiliconFlow
api_key = "sk-..."
base_url = "https://api.siliconflow.cn/v1"
chat_model = "Qwen/Qwen3.6-27B"
chat_extra_body = "{\"enable_thinking\":false,\"max_tokens\":4096}"
embed_model = "Qwen/Qwen3-Embedding-0.6B"
embed_dim = 1024

Uses the Qwen model in the current Chinese provider guide; availability differs from Qwen's own releases. For Qwen/Qwen3-Embedding-8B at 1024 dimensions, also set embed_send_dimensions = true. Official documentation.

Aliyun Bailian
api_key = "sk-..."
base_url = "https://dashscope.aliyuncs.com/compatible-mode/v1"
chat_model = "qwen3.8-flash"
chat_extra_body = "{\"enable_thinking\":false,\"max_tokens\":4096}"
embed_model = "qwen3.7-text-embedding"
embed_dim = 1024
embed_batch_size = 20

The existing Beijing host remains supported. You may use your workspace-specific host from the console; match the key and region. Current embedding requests allow at most 20 texts. Official documentation.

Ollama
base_url = "http://127.0.0.1:11434/v1"
api_key = "ollama"
chat_model = "qwen3.8:27b"
chat_extra_body = "{\"reasoning_effort\":\"none\",\"max_tokens\":4096}"
embed_model = "qwen3-embedding:0.6b"
embed_dim = 1024

Pull both models first. Qwen3.8 27B downloads about 18 GB; qwen3.5:4b or qwen3.5:9b are smaller alternatives. The placeholder key selects the real local endpoint instead of mock mode. Official documentation.

vLLM
base_url = "http://127.0.0.1:8000/v1"
api_key = "not-used"
chat_model = "Qwen/Qwen3.8-27B"
chat_extra_body = "{\"chat_template_kwargs\":{\"enable_thinking\":false},\"max_tokens\":4096}"
embed_backend = "local"
local_embed_path = "BAAI/bge-m3"
embed_model = "BAAI/bge-m3"
embed_dim = 1024

Use the model name served by your deployment and the actual key if authentication is enabled. The chat server does not automatically serve an embedding model; this example uses facetmark[local]. Official documentation.

Chat and online embeddings share an endpoint

There is one base_url and api_key pair. DeepSeek, Kimi and standalone vLLM chat servers can use the local embedding configuration above. Rebuild vectors after changing the embedding model, even if its dimension is unchanged. Chat parameters apply to every fallback model and must work with all of them.

Local embeddings#

First install python -m pip install "facetmark[local]". Model files are downloaded once; embeddings then run on the service machine. Page fetching and any configured online chat model still use the network.

local embeddings
embed_backend = "local"
local_embed_path = "BAAI/bge-m3"
embed_model = "BAAI/bge-m3"
embed_dim = 1024

It is slower to build and it needs the model downloaded once. Search quality is good: on a 1,024-token window bge-m3 reproduces its own vector to a cosine of 0.999976 run-to-run, which is the property that matters for an index you keep rather than rebuild.

what you give up

Two facets are built on a language model reading your pages: the questions a page could answer, and the topic labels. With no chat model those stay empty and you are searching on body text and full-text — still the two strongest paths, and still better than what your browser gives you.

You can also start here and add a key later. Nothing has to be thrown away; the index fills in the parts it could not build before.

Concurrency, privacy and other options#

Vectors

SettingIn plain language
embed_backendapi or local. Restart to take effect.
embed_dimHow long each vector is. Must match what the model actually returns. Restart to take effect.
local_embed_pathModel id or folder for the local backend. Restart to take effect.

How hard it pushes

SettingIn plain language
request_timeoutSeconds before a call is given up on. Raise it on a slow link; lower it if a provider hangs.
fetch_concurrencyHow many pages are downloaded at once. Lower it if your network complains.
enrich_concurrencyHow many pages are sent to the model at once. This is the one to lower when you get rate-limited.

What it is not allowed to look at

SettingIn plain language
privacy_excluded_domainsDomains never fetched and never sent anywhere. Bank, health, work intranet. The bookmark stays; only the title is indexed.
chat_model_fallbacksModels to try, in order, when the first one refuses.
Set exclusions before the first fetch

Exclusions restrict later processing; they do not erase previously stored bodies or backups. Review the exclusions before importing and indexing.

Connection and save errors#

MessageWhat to do
401 / invalid_api_keyKey wrong, or wrong provider's key for this base_url. Test on the Settings screen — it tells you which half failed.
404 on a model nameThat name is not on that endpoint. Check the provider's model list.
429Rate limit. Lower enrich_concurrency and run again; finished stages are not redone.
dim mismatchembed_dim disagrees with the model. Fix it, restart, rebuild.
Chat works, embeddings 403Check that the endpoint and account support the chosen embedding model. Use an endpoint offering both capabilities, or configure local embeddings as above; online chat and embeddings share the same endpoint.
unknown setting on saveA typo in a key name. The writer refuses unknown keys rather than storing something that will be silently ignored forever.