Why MCP context window tool overload starts before any work
Open a client, connect an account, and the tool list lands in the model's prompt before the user types anything. MCP (Model Context Protocol) is the open protocol an AI client uses to discover and call tools on a server. The client calls tools/list on connection, and the response goes into context. Each entry carries a name, a description and a JSON schema for every argument it accepts.
MCP context window tool overload is the point where that list costs more than it returns. The list is not a menu the model consults on demand. It is text in the prompt, present on every turn. Argument schemas are the expensive part. A single tool with nested objects and enums can run longer than the message the user typed.
Two costs, different in kind. The first is budget: context spent on tool definitions is context not spent on the ticket, the thread or the document. The second is selection. Once dozens of near-identical names are in view, the model has to pick. Two Notion workspaces connected to one person advertise the same operations under different account prefixes. Nothing in the names says which one holds the customer contracts.
Why one connected app is normally enough to trip the 30-tool threshold
Because the threshold counts catalog operations and connected tools together, not connected tools alone.
Elaichi serves two namespaces on one organization-wide endpoint, POST /mcp. The endpoint is the single address every client dials. Control-plane operations are named elaichi__{resource}__{operation} and cover connections, toolboxes, restrictions, members and the audit trail. A member's connected third-party tools are named {account}__{tool}. The control-plane catalog by itself is dozens of operations.
The count therefore starts well above zero before anybody connects anything. Add one account from the 450+ connectors in the Elaichi catalog and the combined total passes 30. Collapse is the normal case on this endpoint, not a mode reserved for large rollouts. Plan for it as the default shape of the tool list.
What the model sees after the connected tools collapse
Two meta-tools, search_tools and execute_tool, alongside the control-plane operations, which stay listed individually.
Only the connected half collapses. Control-plane operations remain addressable by name, and search_tools never returns one. Administration is a small, fixed, self-describing surface. Hiding it behind a search step would make the assistant worse at the thing it is asked to do most often.
The working loop changes shape. Instead of scanning a flat list, the model issues a query against intent, reads back a short ranked set, then calls one tool. It pays the schema cost only for what the search returned. The two meta-tools themselves cost very little to advertise.
How search ranks a collapsed tool list
Purely lexically, over three fields: the tool name, the description and the connector label. An exact name match scores 5, a name prefix in either direction scores 3, the connector label scores 2, the description scores 1.
A relevance floor sits on top. A tool must account for at least half the query's own IDF-weighted mass to be returned at all. Without a floor a search always returns something, and something from an app the user did not ask about is worse than nothing, because the model calls it. The incident that produced the floor is worth reading before you tune any queries.
Two smaller refinements decide more than they look like they should. Description stop-words come first: set, connection, frozen and more are dropped from description tokens, because they are scaffolding present on every merged or frozen tool. Before that, a query containing the word connection scored every merged tool the same. The tie broke on name order, and the model was handed an arbitrary account. Ties now break on codepoint order, never locale collation. Advertised names use exactly the character set locale-aware comparison reorders, and the ranked list is sliced to a limit.
Account labels stay scorable on purpose. Name two Notion accounts notion-legal and notion-marketing and a query mentioning legal ranks the right one. Name them notion and notion-2 and you have moved the ambiguity out of the model and into your own naming.
Does execute_tool create a second path into your systems?
No. execute_tool is only a naming indirection. It unwraps to the same tool name and the same arguments, then falls through the identical gates. There is no separate execution path and no privilege in the wrapper.
The tool:execute permission gates the whole endpoint ahead of every scope. Without it, tools/list comes back empty and a call returns an in-band error naming the missing permission. Guest, Auditor and Billing Admin do not hold it. OAuth supplies the second ladder: mcp:read, mcp:write, mcp:destructive and mcp:tools. OAuth is the sign-in standard that issues a scoped grant instead of handing over a password. A tool classified forbidden is reachable under no scope. mcp:tools does not replace the ladder, so a connected tool whose method is a delete still needs mcp:destructive.
Restrictions are then checked at four points against the same resolver: browse, connect, advertise and execute, plus a final check on the fully substituted outbound URL.
One limit, stated plainly. The prompt-injection write gate lives in the Elaichi agent window and does not apply to a raw tool call. What holds on the endpoint is RBAC per operation, the forbidden classification, output redaction, OAuth scope limits and full audit logging.
Restrictions and frozen parameters shrink the list before the model reads it
Tools withheld by a restriction are excluded from the advertised list and from the tool count. They were handed to nobody, so they spend no context and never appear in a search result.
A restriction is a rule about which connectors and which individual tools a target may reach. Targets are a role or a single user. restriction targets are role or user only; the organization default is the absence of a rule, which means allow everything; the organization default is the absence of any rule, which means allow-all. A user override replaces role rules outright instead of layering on top of them. Within the winning layer, allows union, blocks union, and blocks beat allows.
One trap is worth stating on its own. the allowlist stage engages on the presence of an allow rule, not its contents, so an allow rule that names nothing denies everything. An allow rule naming nothing denies everything.
Blocks match the advertised tool name or the operation pinned against the catalog at write time. Allows match the pinned operation only, because a displayed name is a label the governed party can edit. Why the two sides match differently is worth reading before your first allowlist.
Frozen parameters trim the schema rather than the list. Frozen keys are stripped from the advertised schema, so the model never sees them. Frozen values are merged over caller arguments at execution, so passing the key cannot un-freeze it. Precedence runs entry defaults, then caller arguments, then frozen params.
One timing note. Restriction and role changes resolve through a cache and take effect within about two minutes, on the MCP endpoint and in the console alike. Only grant revocation, member removal and suspension are effective on the next call.
What the audit trail records for a collapsed call
The execution, not the wrapper. one entry per tool-call attempt, succeeded or failed, naming the account actually reached, taken from the execution rather than from the intent.
That is the first question after an unexpected change: which of my two Notion workspaces did the agent write to. The record carries the operation and tool, the connection, the classification, whether it was approved, the outcome, and an error code only. Argument names and counts are logged. Argument values never are. A reviewer does not have to unwrap execute_tool to read what happened, and the free read-only Auditor seat means that reviewer does not cost a license.
What you give up by collapsing the tool list
Discovery becomes a search problem, and search can miss. The ranking is lexical, not semantic. A query phrased in words that appear in none of the tool name, the description or the connector label will not match. The relevance floor makes the honest outcome an empty result rather than a confident wrong one. Empty is the better failure. It is still a failure the user notices.
There is also a round trip. Search, then execute, is two calls where a flat list was one. For a model working through a long task, that is latency on every new kind of action it tries.
When a flat tool list is still the right answer
Overload is a function of breadth, and breadth is what a company-wide endpoint has by definition. A single-purpose MCP server with six tools and one user never approaches a threshold. Collapse would never engage, and the extra round trip would buy nothing. If that is your situation, the reasonable move is to wait, which is the argument in the case against adding a gateway yet.
The threshold starts to matter once a second team, a second account or a second client shows up. At that point the question stops being how many tools the model can hold. It becomes which tools this person should have been offered at all. Search does not answer that. Roles and restrictions do.
To see the shape of a governed endpoint, read how the control plane fits together, browse the connector catalog, or look at what each team does with it. More on the protocol itself sits in How MCP works.