Repository navigation
[horizontal scaling] How to actually build session persistence in streamable http MCP server? #880
Description
Activity
Hi @karthich,
The sessions are currently stored in memory by the SDK. This is a limitation of this implementation, not a limitation of the protocol itself.
For a quick fix, if your use case permits, you could run the server in stateless mode.
If you're interested in knowing how this is implemented internally:
- This dictionary stores the sessions
- Then it's used to check incoming session ids here.
One way to fix this: SDK should (IMHO) allow the user to provide a custom
MemoryManagerclass which implements someMemoryManagerInterface(methods could be something like this:storeSession,retrieveSession,deleteSession). Then users of the sdk could provide their own implementation of the interface, for instance storing it in some DB. I could work on this over the weekend if some core maintainers agree it's a good idea (cc: @ihrpr @jerome3o-anthropic ).Disclaimer: I am not a core maintainer, but found this issue to be interesting.
Reacted by Michal Augustýn, Greg Sadetsky, Saransh Barua, lanyangyang, Sameer Jain, Adam Caviness, Anthony Chu, Akshay Ram Vignesh, Furquan Uddin, Chetan Jarande and 3 moreYes. I was hoping there was something that would let me do that but the implementation is pretty strong about what it is and there's no good way to communicate to the SDK to use a different way to store the sessions. I know its not part of this project but I was able to implement a Redis backed implementation of the EventStore class but alas no such thing for managing the session themselves.
I'm using sampling which only works in stateful mode as i noticed.
Session persistence is a new need now.+1 for this
As a temporary WA and checks, I did successfully make multi-workers for fastapi uvicorn to persist session id to redis by overriding the _handle_stateful_request() method in StreamableHTTPSessionManager by creating a custom class and creating a new session with same session id and save to dictionary if request goes to new worker, though duplicating server_instances dict in different workers this seems to work for tools, prompts and resources at client, but sampling still didn't work.
Override class:
class PersistentSessionManager(StreamableHTTPSessionManager): """ StreamableHTTPSessionManager with session ID persistence in Redis. """ def __init__( self, *args, redis_url="redis://localhost:6379/0", **kwargs ): super().__init__(*args, **kwargs) self.redis = aioredis.from_url(redis_url, decode_responses=True) async def _create_and_start_session( self, session_id: str ) -> StreamableHTTPServerTransport: http_transport = StreamableHTTPServerTransport( mcp_session_id=session_id, is_json_response_enabled=self.json_response, event_store=self.event_store, ) assert http_transport.mcp_session_id is not None self._server_instances[http_transport.mcp_session_id] = ( http_transport ) async def run_server( *, task_status: TaskStatus[None] = anyio.TASK_STATUS_IGNORED, ): async with http_transport.connect() as streams: read_stream, write_stream = streams task_status.started() await self.app.run( read_stream, write_stream, self.app.create_initialization_options(), stateless=False, # Stateful mode ) assert self._task_group is not None await self._task_group.start(run_server) return http_transport async def _handle_stateful_request( self, scope, receive, send, ): request = Request(scope, receive) request_mcp_session_id = request.headers.get(MCP_SESSION_ID_HEADER) # Try in-memory first if ( request_mcp_session_id is not None and request_mcp_session_id in self._server_instances ): transport = self._server_instances[request_mcp_session_id] LOGGER.debug( "Session found in memory, handling request directly" ) await transport.handle_request(scope, receive, send) return # Try Redis for session persistence if request_mcp_session_id is not None: session_exists = await self.redis.exists( f"mcp_session:{request_mcp_session_id}" ) if session_exists: LOGGER.info( f"Session {request_mcp_session_id} found in Redis (not in memory), recreating transport." ) async with self._session_creation_lock: http_transport = await self._create_and_start_session( request_mcp_session_id ) await http_transport.handle_request( scope, receive, send ) return # If no session ID provided, create a new session and save to Redis if request_mcp_session_id is None: async with self._session_creation_lock: new_session_id = uuid4().hex http_transport = await self._create_and_start_session( new_session_id ) await self.save_session_id(new_session_id) LOGGER.info( f"Created new session and saved to Redis: {new_session_id}" ) await http_transport.handle_request(scope, receive, send) return # Invalid session ID (not in memory, not in Redis) response = Response( "Bad Request: No valid session ID provided", status_code=400, ) await response(scope, receive, send) async def save_session_id(self, session_id: str): """ Save the session ID to Redis. """ await self.redis.set(f"mcp_session:{session_id}", "1", ex=360)
But sampling now get stuck at with error:
Request stream _GET_stream not found for message. Still processing message as the client might reconnect and replay.
This seems may not be simple implementation of persisting just session id in redis and recreating session with same id in different workers may not work for sampling atleast OR may be some fix needed for sampling to work.
I've identified that StreamableHTTPSessionManager supports custom EventStore configurations, which opens up possibilities for implementing persistence solutions.

Proposed Solution
We can develop a custom EventStore implementation to handle session persistence. I've begun exploring a MongoDB-based approach as a reference implementation.
class MongoEventStore(EventStore): """ MongoEventStore provides a MongoDB-based implementation of the EventStore interface for MCP session resumability. By implementing the EventStore interface, MCP can support various types of event stores for session persistence. To use a different backend, simply implement the EventStore interface accordingly. This class handles storing and replaying events for a given stream using MongoDB as the backend. It ensures connection management, event serialization, and efficient querying via indexes. """ def __init__( self, connection_string: str = "mongodb://localhost:27017", database_name: str = "mcp_sessions", collection_name: str = "events" ): self.connection_string = connection_string self.database_name = database_name self.collection_name = collection_name self._client: Optional[AsyncIOMotorClient] = None self._collection: Optional[AsyncIOMotorCollection] = None self._event_counter = 0 async def _ensure_connection(self): """Ensure MongoDB connection is established.""" if self._client is None: self._client = AsyncIOMotorClient(self.connection_string) self._collection = self._client[self.database_name][self.collection_name] # Create indexes to optimize query performance await self._collection.create_index([ ("stream_id", 1), ("event_id", 1) ]) await self._collection.create_index([ ("stream_id", 1), ("timestamp", 1) ]) async def store_event( self, stream_id: StreamId, message: JSONRPCMessage ) -> EventId: """Store an event into MongoDB.""" await self._ensure_connection() # Generate a unique event ID event_id = str(ObjectId()) # Serialize JSONRPCMessage message_dict = message.model_dump(by_alias=True, exclude_none=True) # Create event document event_doc = { "_id": ObjectId(event_id), "stream_id": stream_id, "event_id": event_id, "message": message_dict, "timestamp": datetime.utcnow(), "message_type": message.root.__class__.__name__ } # Insert into MongoDB await self._collection.insert_one(event_doc) return event_id async def replay_events_after( self, last_event_id: EventId, send_callback: EventCallback, ) -> StreamId | None: """Replay events after the specified event ID.""" await self._ensure_connection() # Find the timestamp of the last event last_event = await self._collection.find_one( {"event_id": last_event_id} ) if last_event: # Replay events after this timestamp query = { "stream_id": last_event["stream_id"], "timestamp": {"$gt": last_event["timestamp"]} } else: # If event ID not found, cannot replay # Need stream_id to replay from the beginning; return None for now return None # Query events in chronological order cursor = self._collection.find(query).sort("timestamp", 1) stream_id = None async for event_doc in cursor: # Reconstruct JSONRPCMessage message_dict = event_doc["message"] message = JSONRPCMessage.model_validate(message_dict) # Send event event_message = EventMessage(message, event_doc["event_id"]) await send_callback(event_message) # Record stream_id if stream_id is None: stream_id = event_doc["stream_id"] return stream_id async def close(self): """Close MongoDB connection.""" if self._client: self._client.close()
Current Status
- With the
MongoEventStore, I've written each session to MongoDB
- For the session resuming, the implementation is currently blocked by an upstream issue in the Python SDK:
Issue: Resumption of streamable HTTP session has potential for deadlock #860
It causes deadlock when resumes session, so I can not test the session resuming from Mongo.
Next Steps
Priority: I would try to resolve the session resume functionality first
Follow-up: Complete the MongoEventStore implementation and testing once the blocking issue is resolvedReacted by waferslove- With the
@qwertyuiop17u I built an event store backed by Redis as well. This gives me the ability to resume events across multiple app instances. However I cannot resume those events due to the in-memory storage of the sessions themselves. I'm really curious to hear what the maintainers think we should do here. It seems like stateless MCP server is the only option for resilient deployments.
Reacted by qwertyuiop17u, Mihovil Mandic, Nikhil Yadav, Michal Augustýn, Saransh Barua, Thomas Steinacher, Anthony Chu, Brandon Shar, Chetan Jarande and RyanLee@yadavnikhil It is a good attempt creating custom implementation of
StreamableHTTPSessionManagerwith redis as backend.Creating a new plain object of
StreamableHTTPServerTransportmay not work as it seems_request_streamsdictionary inStreamableHTTPServerTransporthas significance on handling and process request.The new object of
StreamableHTTPServerTransportclearly has an empty dictionary for_request_streamsand that's why you see the errorRequest stream _GET_stream not found for message. Still processing message as the client might reconnect and replay.
Hence I believe, not only "1" but also the state of
StreamableHTTPServerTransportwhich is represented by_request_streamsshould be stored in redis and re-assigned to the new object ofStreamableHTTPServerTransport.And that leads to the fact that
_server_instanceswill not be needed any more. Each request can re-create the transport object with the state retrieved from redis, handle the request, store the updated state back to redis, and dispose the transport object.So in the end it is not just a simple implementation of persisting just session id in redis and recreating session with same id in different workers.
Reacted by Michal Augustýn and Paul B.Correct. Do the maintainers have a plan to introduce the ability to store the session object in a backend at all? I have not seen them chime in yet.
I encountered the same problem. Has there been any progress on this issue?
+1 for thisSame problem - would be happy to see a more thought through solution
+1
+1
My use case is currently limited to tool use for which I implemented : https://github.com/bh-rat/mcp-db ; I built a session storage based on redis that allows to switch the session between multiple instances of MCP servers. If the session doesn't exist on the instance where the request is sent, I check if there is an active session with the session id in redis and then force create the streamable http session in this instance as well.
A cleaner approach I am/was considering is re-implementing the streamable http transport (with the support needed) but that seems like an overkill and I will have to adapt all current and future servers to it.
Note : As pointed above in a comment, the streams are still maintained at the session level - for which it will not work.
It would be good to get some official support around this..
Reacted by Daniel Tran and bjwswang+1
I've identified that StreamableHTTPSessionManager supports custom EventStore configurations, which opens up possibilities for implementing persistence solutions.

Proposed Solution
We can develop a custom EventStore implementation to handle session persistence. I've begun exploring a MongoDB-based approach as a reference implementation.
class MongoEventStore(EventStore):
"""
MongoEventStore provides a MongoDB-based implementation of the EventStore interface for MCP session resumability.
By implementing the EventStore interface, MCP can support various types of event stores for session persistence.
To use a different backend, simply implement the EventStore interface accordingly.This class handles storing and replaying events for a given stream using MongoDB as the backend. It ensures connection management, event serialization, and efficient querying via indexes. """ def __init__( self, connection_string: str = "mongodb://localhost:27017", database_name: str = "mcp_sessions", collection_name: str = "events" ): self.connection_string = connection_string self.database_name = database_name self.collection_name = collection_name self._client: Optional[AsyncIOMotorClient] = None self._collection: Optional[AsyncIOMotorCollection] = None self._event_counter = 0 async def _ensure_connection(self): """Ensure MongoDB connection is established.""" if self._client is None: self._client = AsyncIOMotorClient(self.connection_string) self._collection = self._client[self.database_name][self.collection_name] # Create indexes to optimize query performance await self._collection.create_index([ ("stream_id", 1), ("event_id", 1) ]) await self._collection.create_index([ ("stream_id", 1), ("timestamp", 1) ]) async def store_event( self, stream_id: StreamId, message: JSONRPCMessage ) -> EventId: """Store an event into MongoDB.""" await self._ensure_connection() # Generate a unique event ID event_id = str(ObjectId()) # Serialize JSONRPCMessage message_dict = message.model_dump(by_alias=True, exclude_none=True) # Create event document event_doc = { "_id": ObjectId(event_id), "stream_id": stream_id, "event_id": event_id, "message": message_dict, "timestamp": datetime.utcnow(), "message_type": message.root.__class__.__name__ } # Insert into MongoDB await self._collection.insert_one(event_doc) return event_id async def replay_events_after( self, last_event_id: EventId, send_callback: EventCallback, ) -> StreamId | None: """Replay events after the specified event ID.""" await self._ensure_connection() # Find the timestamp of the last event last_event = await self._collection.find_one( {"event_id": last_event_id} ) if last_event: # Replay events after this timestamp query = { "stream_id": last_event["stream_id"], "timestamp": {"$gt": last_event["timestamp"]} } else: # If event ID not found, cannot replay # Need stream_id to replay from the beginning; return None for now return None # Query events in chronological order cursor = self._collection.find(query).sort("timestamp", 1) stream_id = None async for event_doc in cursor: # Reconstruct JSONRPCMessage message_dict = event_doc["message"] message = JSONRPCMessage.model_validate(message_dict) # Send event event_message = EventMessage(message, event_doc["event_id"]) await send_callback(event_message) # Record stream_id if stream_id is None: stream_id = event_doc["stream_id"] return stream_id async def close(self): """Close MongoDB connection.""" if self._client: self._client.close()Current Status
- With the , I've written each session to MongoDB
MongoEventStore
* For the session resuming, the implementation is currently blocked by an upstream issue in the Python SDK:
Issue: [Resumption of streamable HTTP session has potential for deadlock #860](https://github.com//issues/860)
It causes deadlock when resumes session, so I can not test the session resuming from Mongo.
Next Steps
Priority: I would try to resolve the session resume functionality first Follow-up: Complete the MongoEventStore implementation and testing once the blocking issue is resolved
Can the MCP Server using the SSE protocol also have the same implementation? I checked sse.py, and it seems that it cannot be supported
- With the , I've written each session to MongoDB
+1 Same problem
10 remaining items
- marked Issue: Sticky Session / Session Affinity Not Working with Azure Web App + FastAPI +FastMCP + Gunicorn #1350 as a duplicate of this issue
on Oct 7, 2025 - changed the title
[-]How to actually build session persistence in streamable http MCP server?[/-][+][horizontal scaling] How to actually build session persistence in streamable http MCP server?[/+]on Oct 7, 2025 - added a commit that references this issue
on Jan 2, 2026 +1 to support stateful sessions in multiple pods
- added a commit that references this issue
on Jan 31, 2026 Hit this on ECS too. Sticky sessions via ALB cookie work but it's a tax on every load-balancer config and breaks when a task cycles mid-session.
The real fix is a pluggable SessionStore protocol so Redis or DynamoDB can replace the in-memory dict in StreamableHTTPSessionManager.
stateless_http=True isn't a substitute when tools need auth context or resumability across calls. Anyone working on a PR for the store-protocol path?
Looking a RC Spec session-id will be removed soon - https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/
It mean that we're going to lose at all the persistent session, I guessIf this is still open for contribution, I'd like to pick up the docs task. We operate a multi-replica streamable-HTTP MCP server in production (GKE behind a load balancer), so I can write this from operational experience rather than theory.
Proposed scope for a documentation page/section:
- Why naive horizontal scaling breaks:
_request_streamsinsideStreamableHTTPServerTransportis in-process state (stream id → active request queues), so persisting only the session id in Redis/a DB is insufficient — a follow-up request landing on a different replica has the session id but not the live streams. - Working deployment patterns, with trade-offs:
- session affinity / sticky routing at the LB (simplest; what we run),
- stateless mode for request/response-style servers that don't need server-initiated streams,
- event store + resumability (
EventStore) for replay across replicas.
- Forward-compat caveat: the 2026-07-28 RC removes session ids, so the doc should frame these as patterns for current-spec deployments and note what changes (per @alessandro308's point).
Happy to adjust scope to whatever fits the docs structure — and can have a PR up shortly once assigned.
- Why naive horizontal scaling breaks:
The v2 rework on
maintargets the 2026-07-28 protocol revision, where Streamable HTTP is sessionless by construction — noinitializehandshake, noMcp-Session-Id, each request a self-contained POST — so horizontal scaling needs nothing from the SDK: any replica behind a plain round-robin load balancer can answer (Deploy & scale). Clients on 2025-11-25 or earlier still land on the session-based leg of the same app, and that session record is still an in-process dict with no pluggable store, so those connections keep the trade-off this thread already found: sticky routing, orstateless_http=Trueat the cost of the server-to-client channel (Serving legacy clients, Protocol versions). The two things statelessness doesn't hand you across workers are on the deploy page too: a sharedRequestStateSecurity(keys=[...])plus the same server name for multi-round-trip tools, and your ownSubscriptionBusif change notifications have to cross replicas. For the sampling case raised above, 2026-07-28 retires the server-to-client back-channel in favour of the multi-round-trip return path (ResolveonMCPServer), so there's no stream left to pin to a worker. So v2 doesn't add the Redis/DB-backed session store proposed here — the modern path has no session to persist — and theEventStoreand session-manager workarounds people posted remain the practical answer for pre-2026 clients today.Closing this as it's been fixed in v2, which is now released. If you still run into this on v2, feel free to reopen.
I have a question here regarding how to build session persistence into an MCP server.
I am running my mcp server (FastMCP) in an ecs service with multiple tasks. I have built a redis event store to allow for resumability. The service sits behind an ALB that directs HTTPS traffic to it. Inspite of all of it, Im still observing intermittent
Bad Request: No valid session ID provided400 errors when I connect the mcpinspector.My suspicion is StreamableHTTPSessionManager doesn't seem to be able to store session-id data in an external data source. Therefore traffic from a client has to always hit the same task since the session id is tied to a server.
In an AWS architecture, how do I configure my server and client to accomplish this. Is using sticky sessions the only route here? Are streamable http sessions not horizontally scalable?