Self-hosted mem0 MCP: setup and the cold-start timeout

· 3 min read

mem0 gives an agent long-term memory. The hosted version wants an API key and your data; the self-hosted version needs an embedding model, an LLM for fact extraction, and a vector store. I already had a Linux box on the LAN with a GPU and a pile of Ollama models, so everything but the MCP server itself runs there.

On the LAN box

Ollama as a user-level systemd service (~/.config/systemd/user/ollama.service, enabled and running), bound to 0.0.0.0:11434 so other machines can reach it, reusing the ~47GB of models already on disk. Pulled bge-m3 for embeddings. For fact extraction I used an existing llama3:8b rather than mem0’s default qwen3:14b — the card is a 6GB 3060 Laptop, and the smaller model fits without an extra download.

Qdrant as a Docker container (qdrant/qdrant, --restart unless-stopped), ports 6333/6334, with a generated API key. Worth confirming the key actually took effect rather than assuming it:

curl -s -o /dev/null -w '%{http_code}\n' http://<lan-ip>:6333/collections
# 401

On the client machine

uv/uvx went in via the official installer. Homebrew had no prebuilt bottle for this macOS version and started compiling Rust from source, which is not a thing worth waiting for.

Then register the MCP server, pointed at the LAN box’s Ollama and Qdrant:

claude mcp add --scope user mem0 ...

That crashed on first launch with ModuleNotFoundError. The cause is a real bug in elvismdev/mem0-mcp-selfhosted: its dependencies don’t pin mcp<2, so a fresh install resolves mcp==2.1.1, which renamed the FastMCP API the code imports. Forcing the older compatible version works around it:

--with "mcp<2"

After that, claude mcp list showed mem0: ... ✔ Connected.

Fix: Request timed out on connect

A few weeks later the server stopped connecting. Claude Code reported mem0 (CONNECT_TIMEOUT): "Request timed out" at session start, and reconnecting via /mcp failed the same way. Both backends answered HTTP 200, so this wasn’t the LAN box.

Cause: the server was launched with

uvx --from git+https://github.com/elvismdev/mem0-mcp-selfhosted.git \
  --with "mcp<2" mem0-mcp-selfhosted

The git ref is unpinned. When the uv cache goes stale — new upstream commit, or cache eviction — uvx re-fetches the repo and rebuilds the environment, which includes compiling cryptography from source. A cold start took 66s; the MCP startup timeout is shorter than that. A warm start took 2s, which is why it had worked for weeks.

Fix: install the server once as a persistent uv tool and run the binary directly, so startup no longer depends on the uvx cache:

uv tool install --from git+https://github.com/elvismdev/mem0-mcp-selfhosted.git \
  --with 'mcp<2' mem0-mcp-selfhosted

Then, in ~/.claude.json under mcpServers.mem0, swap the launch command for the installed binary (the env block stays unchanged):

"command": "~/.local/bin/mem0-mcp-selfhosted",
"args": []

Back up the config first — it holds every MCP server’s settings, not just this one.

Upgrading later: the tool no longer tracks upstream automatically.

uv tool upgrade mem0-mcp-selfhosted

If an upgrade drops the mcp<2 constraint, reinstall with the uv tool install ... --with 'mcp<2' command above plus --force.

Two things to know before copying this

The Ollama user service survives logout/login but not a reboot without an active session. On a headless box, sudo loginctl enable-linger <user> is required — same trap described in auto-starting a Hugo dev server at boot.

Nothing restricts LAN access beyond Qdrant’s API key: Ollama is open to anyone on the network, and the Qdrant key sits in plain text in the client’s MCP config. That is fine on a trusted home LAN and not fine anywhere else.