# Build vs buy MCP server: pricing the ongoing work

> The build vs buy MCP server decision turns on recurring work: connector maintenance, OAuth refresh, audit tenancy, and tool discovery past 30 tools.

**TL;DR** The build vs buy MCP server decision is not settled by how fast you can stand up a first server. It is settled by five recurring obligations: connector maintenance against APIs you do not control, OAuth refresh and its failure states, per-organization audit tenancy, tool discovery past 30 tools, and an access model with roles, sharing and restrictions. Elaichi ships all five behind one organization-wide MCP endpoint, so the honest comparison is that recurring work against Gold at $15 per user per month.

## Where the build vs buy MCP server question gets decided

| Axis | Build (internal MCP server) | Buy (managed control plane, e.g. Elaichi) |
| --- | --- | --- |
| Dev time | A week for the first server, then a standing engineering claim for connector churn | Connect the accounts, point clients at one endpoint; connector maintenance is off your team |
| OAuth refresh | You own the token store, encryption at rest, rotation and the visible failure state | Separate credential service holds per-account secrets; a failed refresh marks the connection needs_reauth |
| Per-org tenancy | An org_id column in a WHERE clause; a dropped clause leaks | One log tenant per organization, enforced in the type system; a wrong tenant returns nothing |
| Tool discovery past 30 tools | Same wall in week two; you build search, ranking and a relevance floor | Meta-tools search_tools and execute_tool, with an IDF-weighted ranking floor already in place |
| Audit ownership | You own the record shape, the error-text firewall and the SIEM export | One record shape across audit and application logs; Datadog export shipped, Splunk HEC and Sentinel accepted |
| Upgrade risk | Each connector update is your merge decision under load | Fork a public connector; a review screen separates safe updates from conflicts and upstream removals |

The build vs buy MCP server decision is settled by the work that arrives after launch, not by the first server. A capable platform engineer can wrap four internal APIs in a working MCP server in a week. MCP (Model Context Protocol) is the standard way an AI assistant calls tools in other apps The demo lands, and the week looks cheap.

Nothing in that week tells you what the next eleven months cost. The recurring line items are predictable, and they are the same five every time. Connector maintenance against APIs you do not control. OAuth refresh, plus the failure states around it. Audit records one organization can read and another cannot. Tool discovery once the advertised list outgrows what a model can hold. An access model with roles, sharing and per-tool restrictions, enforced the same way on every surface.

Price those five honestly and the decision gets easy in either direction. This post prices them against the shape Elaichi ships, because that is an implementation worth describing precisely rather than in the abstract.

## What does connector maintenance cost after the first release?

It costs a standing claim on engineering time, sized by breadth times churn. Every connector is a contract with somebody else's API. Required fields appear. Pagination changes. Endpoints are deprecated on the vendor's schedule, not yours.

Two internal services you own are cheap, because you also write the release notes. Thirty third-party SaaS apps are a different job. Elaichi authors, maintains and serves 450+ from its own infrastructure. Elaichi is the MCP server, not a listing of servers other people run.

If you build, you also build the maintenance surface, not only the connectors. Elaichi connectors are authored from JSON config, and one can be forked from a public connector. A fork pulls upstream changes through a review screen. That screen separates new tools, safe updates, config diffs, conflicts and upstream removals. Conflicts and destructive removals stay unchecked by default, so nobody merges a tool deletion by reflex. A connector cannot be deleted while connections still use it. Each of those is a small decision you make yourself on the build path.

## Who owns OAuth refresh at two in the morning?

Someone owns it. The design question is where refresh lives and what happens when it fails. OAuth is the sign-in handshake that gives one system a token to act in another on a user's behalf. Tokens expire on the third party's clock.

On the build path you own a token store, encryption at rest, rotation, refresh scheduling and a visible failure state. Elaichi does not hold connector credentials at all. A separate credential service holds per-account secrets, AES-256-GCM at rest, and owns refresh. A failed refresh marks the connection `needs_reauth` rather than failing quietly. That is the difference between a user reconnecting and a ticket about an agent that went silent.

Two smaller decisions are worth copying on either path. A connect URL is not a credential. It is a one-time session carrying no token, which is why it is safe to return over MCP. Reading back an account's configuration returns public values plus `secret_paths`, the list of dot-paths that were encrypted, carrying none of their values. Editing one is refused, and the refusal text is identical whichever side produced it, so a caller cannot probe which side said no. Elaichi also accepts a customer's own OAuth app per connector. That is gated on `connector:manage` rather than `connection:manage`, so everyone who can delete a connection does not quietly gain the ability to repoint the organization's OAuth app. Per-organization envelope encryption with a customer-managed key in AWS KMS is available on top.

## Why is audit tenancy harder than adding a log table?

Because the log is where third-party data leaves the system, and because one organization's rows must be unreachable from another organization's query. An `org_id` column in a WHERE clause gets both wrong under pressure.

Elaichi uses one record shape for audit events and application logs, so a single query answers what happened. `actor_kind` is a recorded field, not an inference. Its values include `user`, `system`, `staff`, `scim`, `api_token` and `ai_assistant`. Every tool-call attempt writes one entry, succeeded or failed, and both name the account actually reached, taken from the execution rather than the intent. Which of two Notion workspaces the agent wrote to is the first question after an unexpected change. Recorded per call: the operation and tool, the connection, the classification, whether it was approved, the outcome, and an error code. Argument names and counts are logged. Argument values never are.

Most build plans miss the error-text firewall. Two error strings exist per failed call. The one returned to the caller derives from the third party's response body and is never written anywhere else. The one written to the audit trail is never derived from the request or the response. Audit records are organization-visible, readable by the in-product assistant, and forwarded to whatever SIEM the customer configured. Promoted metadata is a fixed allowlist, because metadata keys can be user-influenced and unbounded flattening grows the field namespace for a whole tenant. Tenancy sits in the type system rather than a database filter: a dropped clause leaks, a wrong tenant returns nothing. Export to Datadog is implemented. Splunk HEC and Microsoft Sentinel are accepted but not yet delivering. The trail is eventually consistent, so a row may take a moment to appear.

## What happens to tool discovery past 30 tools?

Past a threshold of 30 tools, connected tools collapse behind two meta-tools, `search_tools` and `execute_tool`. The threshold counts catalog operations and connected tools together. The control-plane catalog alone is dozens of operations, so one connected app is normally enough to trip it. Collapse is the normal case, and any server you build hits the same wall in week two.

Only the connected half collapses in Elaichi. Control-plane operations stay listed individually, and `search_tools` never returns one. Tools withheld by a restriction are excluded from the count, because they were handed to nobody. `execute_tool` is only a naming indirection. It unwraps to the same name and arguments and falls through the identical gates, so there is no second execution path and no privilege hiding in it.

Then you have to rank. Elaichi ranks lexically over three fields: tool name, description and connector label. Exact name match scores 5, an either-direction name prefix 3, connector 2, description 1. A relevance floor requires a tool to account for at least half the query's own IDF-weighted mass. The incident behind the floor is instructive. A user with only Notion connected searched `list_all_cal_com_schedules` and got back `list_all_notion_users`, because `list` and `all` scored while `cal`, `com` and `schedules` scored nothing. Something from the app you did not ask about is worse than nothing, because the model calls it. Description stop-words and a codepoint tie-break came from the same class of bug. The [reasoning behind the relevance floor](/blog/search-tools-ranking-floor-idf/) is written up separately, and it is roughly the backlog you inherit if you build search yourself.

## What does the access model cost to build?

Three layers kept distinct, plus a resolver every surface agrees on. the code that decides allow or deny Mixing the layers is the usual way a homegrown model becomes unexplainable to an auditor.

Permissions come first. Around 38 action strings are grouped into roles, with exactly one role per member, enforced by a unique index. Every role is therefore a complete persona rather than a bolt-on. The system roles form a strict subset chain from Guest up to Org Owner. Org Admin holds every permission except `billing:manage` and `org:delete`, so an attacker who lands an admin account cannot delete the evidence along with the workspace. Billing Admin and Auditor sit off the chain. Both are free seats, as is Guest, so a compliance reviewer costs no license.

Sharing is one primitive: a grant of view, use or edit on a resource, to a user, a team or the whole organization. A member sees only what they own or what was shared explicitly. No organization-level permission silently widens a listing.

Restrictions target a role or a user. restriction targets are role or user only; the organization default is the absence of a rule, which means allow everything, because the organization default is the absence of a rule. A rule on a user replaces the role rules for that user rather than adding to them Blocks beat allows. the allowlist stage engages on the presence of an allow rule, not its contents, so an allow rule that names nothing denies everything rather than its contents, so an allow rule naming nothing denies everything. Blocks match the tool name or the pinned operation. Allows match the pinned operation only, because [a tool's advertised name is a label the governed party controls](/blog/block-matches-name-allow-matches-operation/). Enforcement runs at browse, connect, advertise and execute, plus a final check on the fully substituted outbound URL.

Freshness is the detail to copy carefully. A role or restriction change takes effect within about two minutes, through a 60 second cache plus edge propagation. Grant revocation, member removal and suspension are effective on the next call. `revoked_at` is re-read on every call, and removing or suspending a member revokes every live grant in the same transaction as the membership change.

## Which arguments should the model never see?

The ones a caller must not be allowed to choose. Elaichi handles that with frozen parameters, a per-entry map over the tool's flattened argument space.

Frozen keys are stripped from the advertised schema, so the model never sees them. Frozen values are merged over caller arguments at execution, so passing the key cannot un-freeze it. The precedence runs entry defaults, then caller or model arguments, then frozen parameters. On the build path that is request parsing, schema rewriting and a merge order you have to defend to a reviewer later. Getting the merge order backwards is the bug that passes every test you thought to write.

Identity plumbing belongs on the same list. Elaichi ships SAML and OIDC SSO built in-house, no third-party auth vendor, plus SCIM v2 and group-to-role mapping. SSO lets people sign in with the identity provider your company already runs. SCIM is the standard that pushes users and groups from that provider into an application, so a leaver is a leaver everywhere.

## What does offboarding cost you to get right?

Offboarding is where a hand-built server usually breaks, because a departing member's credentials are often holding up somebody else's work. Elaichi runs a preflight before removal.

Personal connections referenced by a toolbox entry must be resolved first. A toolbox is a named set of tools and accounts handed to a client. Resolving means transferring the connection to the organization, a team or another member, or deleting it. Otherwise the removal is refused. Unreferenced personal connections are cleaned up. A private connection is not transferable at all: a credential only its owner could ever use does not become someone else's because its owner left. Delegated toolbox entries surface as a non-blocking warning, and re-pinning is the fix rather than a credential decision. The same questions come up for [temporary staff losing AI access](/blog/shadow-ai-contractor-offboarding/).

## What else lands on the build backlog?

Orchestration and residency, both of which are easy to underestimate in a first estimate.

Elaichi supports synthetic tools: a graph of steps with no cycles, where each step calls a connection's tool with templated arguments over inputs and prior outputs. Cycles are rejected at save. Independent steps run in parallel. Every step goes through the same restriction and audit pipeline as any other call. Build it yourself and you own orchestration code plus a second enforcement path, which is where the two paths drift apart over time.

Residency is the other one. Three regions are chosen at organization creation: `eu`, `us` and `apac`. EU and US are hard residency, a Cloudflare Durable Object jurisdiction, so compute and storage stay in that jurisdiction. The `apac` region is a placement hint only. It is best-effort, and not a residency guarantee. Region also selects which regional log instance the audit trail lands in. Building this yourself multiplies deployment and log isolation by the number of regions you promise a customer.

## When writing your own server is the right call

Build when the tools are yours. If you are exposing three internal services, the API contract follows your own release schedule, the audience is one team, and nobody is asking who called what last quarter, then a small server you own is cheaper and clearer. Build also when the target systems are reachable only inside your own network. That constraint decides the shape before cost enters the conversation.

Check one more assumption before you shop. Gateway products generally assume you already run MCP servers. Lunar.dev describes MCPX as a self-hosted enterprise MCP gateway that sits between agents and the MCP servers, APIs and LLM providers they use ([lunar.dev](https://www.lunar.dev/), checked September 2026). Tyk describes an MCP Gateway that proxies and governs remote MCP servers, and that can generate an MCP proxy from a managed REST API ([tyk.io](https://tyk.io/docs/ai-management/mcp-gateway/overview), checked September 2026). If you do not intend to run any servers, that assumption matters. The [full cost comparison for running servers yourself](/blog/self-hosted-mcp-servers-vs-control-plane/) goes further, and there is a real case for [holding off on a gateway entirely](/blog/when-you-dont-need-an-mcp-gateway/).

## How to price the two paths side by side

Write the five obligations down, give each an owner and an on-call rotation, then put one number next to the other. Elaichi Gold is $15 per user per month, or $120 per user per year. The two plans are Gold and Black, with a 14-day trial. The trial is 14 days with no credit card, and checkout sets the paid trial to the remaining days rather than granting a fresh 14. Billable seats are active memberships, with suspended members excluded, along with the free Guest, Billing Admin and Auditor seats.

Be clear about what buying does not remove. The prompt-injection write gate lives in the Elaichi agent window and does not apply to a raw tool call. What does hold on the endpoint is RBAC per operation, the forbidden classification, output redaction, OAuth scope limits and full audit logging. Organization deletion tears down the workspace but has no path to purge the log tenant, and it returns that residue by name rather than claiming everything is gone.

If the answer is buy, check the [connector catalog](/connectors/) for the apps your teams actually use, read the [team setups](/use-cases/) closest to yours, and put your recurring engineering cost next to [plan pricing](/pricing/).

## FAQ

### Is it cheaper to build an internal MCP server or buy a managed one?

It depends on breadth. Wrapping two or three internal services you own is usually cheaper to build, because you control the API contract and the audience is small. Buying gets cheaper as you add third-party SaaS apps, because connector maintenance, OAuth refresh, per-organization audit tenancy, tool discovery past 30 tools and an access model all become recurring work owned by your team. Elaichi Gold is $15 per user per month, or $120 per user per year, with a 14-day trial, which is the figure to place next to that recurring cost.

### Why do MCP tools collapse behind search_tools after 30 tools?

Past a threshold of 30 tools, connected tools in Elaichi collapse behind two meta-tools, search_tools and execute_tool, because long advertised tool lists degrade model selection. The threshold counts catalog operations and connected tools together, and the control-plane catalog alone is dozens of operations, so one connected app is normally enough to trip it. Control-plane operations stay listed individually, and search_tools never returns one. Tools withheld by a restriction are excluded from the count.

### Who handles OAuth token refresh in a managed MCP control plane?

In Elaichi, connector credentials never live in the control plane. A separate credential service holds per-account secrets, encrypted with AES-256-GCM at rest, and owns refresh. When a refresh fails, the connection is marked needs_reauth rather than failing silently, so a user can reconnect instead of debugging an agent that went quiet. An organization can also supply its own OAuth app per connector, gated on the connector:manage permission.

### How quickly does a role or restriction change take effect?

In Elaichi, a role or restriction change takes effect within about two minutes, because it resolves through a 60 second cache plus edge propagation across MCP, console and REST alike. Grant revocation, member removal and suspension are different: they are effective on the next call, since revoked_at is re-read from the organization store on every call and removing or suspending a member revokes every live grant in the same transaction as the membership change.

### What does a hand-built MCP server usually get wrong about audit logs?

Two things. The first is tenancy: an org_id column in a WHERE clause leaks when the clause is dropped, whereas enforcing one log tenant per organization in the type system means a wrong tenant simply returns nothing. The second is error text. An error string derived from a third party's response body must never be written to an organization-visible audit record, because that record is read by assistants and forwarded to a SIEM, which turns the log pipe into an exit for third-party payload.

## Read next

- [Elaichi vs Composio for company-wide MCP access](/blog/elaichi-vs-composio/) — Elaichi vs Composio: one organization-wide MCP endpoint with connectors Elaichi authors, against per-team endpoints over a large third-party app registry.
- [Elaichi vs native AI connectors: one plane, many clients](/blog/elaichi-vs-native-ai-connectors/) — Elaichi vs native AI connectors: what a second AI client costs in repeated authorization, provisioning, audit shapes and role models, and when native is still right.
- [Lunar MCPX alternative: count your servers first](/blog/lunar-mcpx-alternative-governed-access/) — Lunar's MCPX fronts the MCP servers you already run. If you run none, the right Lunar MCPX alternative removes the servers instead of proxying them.
