hermes.llm
Models
complete returns the reply text, the model name, and the provider. The default provider is a local Ollama server.
Which provider
HERMES_LLM_PROVIDER=openai selects the HTTP client below. Any other value, including an empty one, selects Ollama. Temperature is 0 for both. The system message asks for one SPARQL query and nothing else.
Source: llm.py.
Ollama
| Variable | Default |
|---|---|
HERMES_OLLAMA_BASE_URL | http://127.0.0.1:11434 |
HERMES_OLLAMA_MODEL | chosen from installed models |
With no model set, the client reads /api/tags and picks the first installed name that matches this order: llama3.1, llama3, mistral, phi3, qwen2.5, qwen. A name matches exactly or with a tag, as in llama3.1:latest. If none match, the first installed model is used. If none are installed, the error tells you to run ollama pull llama3.1.
The chat call uses /api/chat, stream false, and keep_alive of 10 minutes. The timeout is 180 seconds. A connection failure says to start ollama serve. A timeout suggests a smaller HERMES_OLLAMA_MODEL. A 404 says to ollama pull that model.
OpenAI-compatible
| Variable | Default |
|---|---|
HERMES_LLM_API_KEY | required |
HERMES_LLM_BASE_URL | https://api.together.xyz/v1 |
HERMES_LLM_MODEL | meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo |
The client posts to /chat/completions with a bearer token. The timeout is 40 seconds. A missing API key raises LookupError and points back at Ollama.