Full size image
Checking...

Start a conversation

Send a message or attach a file to begin chatting with your local LLM.

~0 tokens
Processing...

Settings

Appearance

Theme
Layout

Processing

Text Chunk Size (bytes)
PDF Pages per Chunk
Conversation History (messages)
Model Context
Relative KV-cache use: 8K β‰ˆ 25%, 16K β‰ˆ 50%, 32K = 100%. Actual total RAM depends on the model and quantisation. Larger contexts handle deeper documents but use more memory and can process prompts more slowly. Utility calls use 4K automatically.

API Configuration

Ollama Endpoint
Perplexity API Key
OpenAI API Key
Tavily API Key
Gemini API Key

Agentic Functions

Auto-title Threads
LLM names thread after first exchange
Auto-search (Tavily)
LLM decides when to search (single pass)
Auto-summarise Context
Summarise older messages when context fills
Auto Fact-check
Verify factual claims via Tavily
Use Memories
Inject relevant saved facts into prompts
βš— Experimental Off by default β€” may affect performance
Suggested Follow-ups
Show clickable follow-up questions
Learn Memories
Extract durable facts during idle time

WebGPU Models ○ Chrome / Edge only

Enable Qwen3.5 WebGPU models
Runs entirely in-browser, no Ollama needed
Fast: ~850MB. Quality: ~1.8GB and higher memory use. Downloads are cached; only one model occupies GPU memory at a time.
Model status Not loaded

Danger Zone

Clear all messages in thread
View & manage memories
Clear all memories
Reset entire database

Create New Thread

Remembered Facts

Encrypted facts are selected by relevance and a fixed prompt budget. Review, edit, or delete anything inaccurate.

Keyboard Shortcuts

Send message
Enter
New line
Shift Enter
New thread
Alt N
Search messages
Alt S
Open settings
Alt ,
Show shortcuts
Alt /
Toggle fullscreen
F11
Close modal
Esc