Self-hosted mem0 MCP: setup and the cold-start timeout
mem0 gives an agent long-term memory. The hosted version wants an API key and your data; the self-hosted version needs an embedding model, an LLM for fact extraction, and a vector store. I already had a Linux box on the LAN with a GPU and a pile of Ollama models, so everything but the MCP server itself runs there.
On the LAN box
Ollama as a user-level systemd service
(~/.config/systemd/user/ollama.service, enabled and running), bound to
0.0.0.0:11434 so other machines can reach it, reusing the ~47GB of models already on
disk. Pulled bge-m3 for embeddings. For fact extraction I used an existing llama3:8b
rather than mem0’s default qwen3:14b — the card is a 6GB 3060 Laptop, and the smaller
model fits without an extra download.
Qdrant as a Docker container (qdrant/qdrant, --restart unless-stopped), ports
6333/6334, with a generated API key. Worth confirming the key actually took effect
rather than assuming it:
curl -s -o /dev/null -w '%{http_code}\n' http://<lan-ip>:6333/collections
# 401
On the client machine
uv/uvx went in via the official installer. Homebrew had no prebuilt bottle for this
macOS version and started compiling Rust from source, which is not a thing worth waiting
for.
Then register the MCP server, pointed at the LAN box’s Ollama and Qdrant:
claude mcp add --scope user mem0 ...
That crashed on first launch with ModuleNotFoundError. The cause is a real bug in
elvismdev/mem0-mcp-selfhosted: its dependencies don’t pin mcp<2, so a fresh install
resolves mcp==2.1.1, which renamed the FastMCP API the code imports. Forcing the
older compatible version works around it:
--with "mcp<2"
After that, claude mcp list showed mem0: ... ✔ Connected.
Fix: Request timed out on connect
A few weeks later the server stopped connecting. Claude Code reported
mem0 (CONNECT_TIMEOUT): "Request timed out" at session start, and reconnecting via
/mcp failed the same way. Both backends answered HTTP 200, so this wasn’t the LAN box.
Cause: the server was launched with
uvx --from git+https://github.com/elvismdev/mem0-mcp-selfhosted.git \
--with "mcp<2" mem0-mcp-selfhosted
The git ref is unpinned. When the uv cache goes stale — new upstream commit, or cache
eviction — uvx re-fetches the repo and rebuilds the environment, which includes
compiling cryptography from source. A cold start took 66s; the MCP startup timeout is
shorter than that. A warm start took 2s, which is why it had worked for weeks.
Fix: install the server once as a persistent uv tool and run the binary directly, so startup no longer depends on the uvx cache:
uv tool install --from git+https://github.com/elvismdev/mem0-mcp-selfhosted.git \
--with 'mcp<2' mem0-mcp-selfhosted
Then, in ~/.claude.json under mcpServers.mem0, swap the launch command for the
installed binary (the env block stays unchanged):
"command": "~/.local/bin/mem0-mcp-selfhosted",
"args": []
Back up the config first — it holds every MCP server’s settings, not just this one.
Upgrading later: the tool no longer tracks upstream automatically.
uv tool upgrade mem0-mcp-selfhosted
If an upgrade drops the mcp<2 constraint, reinstall with the uv tool install ... --with 'mcp<2' command above plus --force.
Two things to know before copying this
The Ollama user service survives logout/login but not a reboot without an active
session. On a headless box, sudo loginctl enable-linger <user> is required — same trap
described in auto-starting a Hugo dev server at boot.
Nothing restricts LAN access beyond Qdrant’s API key: Ollama is open to anyone on the network, and the Qdrant key sits in plain text in the client’s MCP config. That is fine on a trusted home LAN and not fine anywhere else.