Skip to main content
Supermemory integrates with LiveKit Agents, so a voice agent can remember a caller across sessions and use that context on the next turn. Recall runs before the model replies. Completed turns are stored automatically. The model can also search, remember, and forget explicitly. A Supermemory timeout or outage is skipped. It does not end the call.

Installation

Create a key at console.supermemory.ai. For a self-hosted API, pass base_url to SupermemoryLiveKit.

Scope

Memory is isolated by container tag. Use a stable caller id, the same one you use in the rest of your app, so a LiveKit call and a chat session share one profile. Container tags may only contain letters, numbers, _, -, and :, and must be 100 characters or fewer. Two ways to set it:
  • Pass container_tag yourself, usually from job metadata.
  • Let the plugin read the participant attribute supermemory_container_tag, or fall back to the participant identity.
Identities that are not valid container tags are sanitized to a stable tag. Set the attribute when the identity is an email, phone number, or SIP address and you already have memories under a different id. Set the attribute in the access token your backend issues. Do not grant the caller canUpdateOwnMetadata, or they could change the attribute and read another caller’s memory. One call is stored as a single document. Pass the LiveKit room name as session_id. The session id is the document’s custom id, so a later update with the same id appends to that call instead of creating another. Use one SupermemoryLiveKit per call. If you do rebind it, turns already captured stay with the caller who said them, and memory from the earlier caller is never shown to the new one.

Quick start

Create one SupermemoryLiveKit inside the job. A worker handles many calls, and each call needs its own scope.
preload puts the caller’s profile into the first turn, so the greeting can use it. attach stores the conversation after each agent reply, and stores anything left when the session closes, including a caller who hangs up mid-turn. SupermemoryAgent recalls inside llm_node before each reply and adds the memory tools. That covers voice turns, text turns from generate_reply or session.run, and LiveKit’s preemptive generation. Each user message is recalled once, and tool follow-ups reuse it. A runnable version, with a greeting that uses the caller’s memory, is in the example agent. agent_name turns on explicit dispatch, which is how job metadata reaches the agent. Remove it to join every new room automatically, for example when testing in the Agents Playground.

Your own agent

If you already subclass Agent, keep that class. Pass the tools in, and recall from llm_node.
Call memory.attach(session) before session.start. Recall in llm_node rather than on_user_turn_completed. Changing the turn context in that hook makes LiveKit discard its preemptive generation, which adds latency to every turn. Realtime models skip llm_node. With a realtime model, call await self.memory.on_user_turn_completed(turn_ctx, new_message) from on_user_turn_completed instead. That hook only runs when turn detection runs in your agent, not inside the model. See LiveKit’s external data guide. SupermemoryAgent picks the right hook for you. Turn capture still listens for conversation_item_added.

What gets recalled

Recall waits at most recall_timeout seconds (default 2). If the profile call is slower than that, or fails, the turn falls back to the profile loaded by preload at the start of the call, or proceeds without memory if there was none. In our tests most recalls took 0.4 to 1.4 seconds, with occasional calls near 3 seconds. Raise recall_timeout if you would rather wait than miss memory on a slow turn. Retrieved text is inserted immediately before the user message and is not stored back as something the agent said.

Tools

SupermemoryAgent adds these tools. They are scoped to the bound container tag, so the model cannot read or write another caller. remember stores a standalone fact and processes it immediately (dreaming="instant"), so it is usually recallable within a minute. If your organization has no balance for instant processing, it is saved on the default schedule instead. It is not appended to the call transcript. Automatic capture is what records the conversation.

When a call becomes recallable

Captured turns are written after each agent reply, but they become memories on Supermemory’s processing schedule:
  • With the default capture_dreaming="instant", each write is processed on its own and is usually recallable within a minute, so a caller who rings back is recalled from the last call. Each write bills one extra operation, and the call is written after every agent reply. If your organization has no balance for instant processing, the call is saved on the dynamic schedule instead.
  • With capture_dreaming="dynamic", related documents are processed together at no extra cost. In our tests this took 10 to 20 minutes, so a caller who rings back right away is not recalled from the last call yet.
  • In both modes, search_memories finds the raw call text as soon as it is stored, so the model can still look up a recent call.

Self-hosting