Why engineers using ChatGPT at work is not a policy question
Engineers using ChatGPT at work is a fact about your organization, not a decision in front of you. The decision is whether the assistant gets scoped access to the systems engineers already query, or whether people keep pasting.
Start with the artifact. An engineer hits a nil dereference in the checkout service at 2am. They copy the stack trace out of the log viewer and paste it into an assistant, asking why the frames are ordered the way they are. The trace carries three things beyond the error. A production customer ID sits in a frame argument. An internal hostname sits in the connection string. A fragment of a query names a table.
None of that was the question. All of it left with the question.
The paste happened for a boring reason. Reading the log needed a VPN session and a saved search. Reading the explanation needed a tool that is not in the log viewer. A clipboard closed the gap between the two.
What does a pasted stack trace actually leak?
It leaks whatever happened to sit on the frame, and nobody picked those fields. The engineer picked the question. The customer ID, the hostname and the table name came along because they were in the buffer.
Three properties make a paste worse than a deliberate disclosure.
It is unreviewable. Data loss tooling sees a paste into a browser tab, if it sees anything at all. It does not see which record went, or which environment the record came from.
It is unbounded. The engineer copies the whole trace because trimming it takes longer than the question is worth. Nobody redacts a trace in the middle of an incident.
It is unrepeatable. The next outage produces a different trace with different fields. You cannot write a rule against a shape that changes every time.
A governed tool call has the opposite properties. It is one record, with a named operation, a named account, and a fixed argument shape you can inspect before you allow it.
Does blocking the assistant work?
Blocking works on the managed browser and the corporate sign-in. It does not work on a phone, a personal account, or the assistant built into the editor. A block changes where the paste happens, not whether it happens. An engineer whose connection drops tethers to a phone and carries on.
There is a second cost, and it is the one that hurts a year later. Once engineers route around the block, you lose the record too. The traffic you could have watched now runs on a device you do not manage (Cyberhaven's Shadow AI report tracks the shift).
A block is still the right call for a short window. If you have no scoped path ready and a regulator is asking questions this month, block, then build. Treat the block as a countdown, not as a control.
What scoped access looks like on one endpoint
Scoped access means one address, per-person identity, and no shared token anywhere in the setup.
Elaichi is a governed MCP control plane. MCP, the Model Context Protocol, is the standard way an AI client calls tools in other systems. Every SaaS account the company uses is connected once. The tools those accounts expose are served through one organization-wide MCP endpoint, POST /mcp, standard MCP over Streamable HTTP, behind OAuth. An endpoint here is a single URL a client points at. OAuth is the sign-in that hands the client a per-person grant instead of a secret somebody pastes into a config file.
There are no per-toolbox URLs and no embedded tokens. Claude, ChatGPT and Cursor point at the same address and sign in as the person using them. The catalog runs to 450+, authored and served by Elaichi rather than listed from servers other people run.
Connector credentials do not live in Elaichi. A separate credential service holds per-account secrets, encrypted at rest, and owns refresh. A failed refresh marks the connection needs_reauth rather than failing quietly. Sign-in to Elaichi itself can run over your own SAML or OIDC, with SCIM v2 provisioning and group-to-role mapping.
Take the nil dereference in the checkout service. The engineer asks the assistant for the error and the deploy that preceded it. The assistant calls the tools. The customer ID never touches the clipboard, because nobody had to move it.
How does the assistant find the right tool?
Past a threshold of 30 tools, the connected tools collapse behind two meta-tools, search_tools and execute_tool. The threshold counts catalog operations and connected tools together. The control-plane catalog alone is dozens of operations, so one connected app is normally enough to trip it. Collapse is the normal case, not an edge case.
Only the connected half collapses. Control-plane operations stay listed individually, and search_tools never returns one. execute_tool is only a naming indirection. It unwraps to the same name and arguments and falls through the identical gates, so there is no privilege in it. Tools withheld by a restriction are excluded from the count, because they were handed to nobody.
Ranking is purely lexical over three fields: tool name, description and connector label. A relevance floor keeps it honest. A tool has to account for at least half the query's own IDF-weighted mass, where IDF means a rare word counts for more than a common one. Without the floor, an engineer with only Notion connected who searches list_all_cal_com_schedules gets back list_all_notion_users, because list and all score while cal, com and schedules score nothing. The full reasoning sits in why tool search needs a relevance floor.
What do you have to write before you turn it on?
Restrictions. A restriction is a rule about which connectors and which individual tools a target may reach, and somebody has to sit down and write them.
This is the trade-off, stated plainly. The organization default is the absence of any rule, which means allow-all within whatever the member can already reach. Opening the endpoint without writing rules gives a broad surface to a fast client.
Targets are role or user only. restriction targets are role or user only; the organization default is the absence of a rule, which means allow everything, so a blanket rule is expressed as a rule on the role everyone holds.
Precedence runs user override, then role rule, then organization default. A user-targeted rule replaces role rules entirely rather than layering on top of them. Within the winning layer, allow rules union, block rules union, and blocks always beat allows.
the allowlist stage engages on the presence of an allow rule, not its contents, so an allow rule that names nothing denies everything. An allow rule naming nothing therefore denies everything. That is the strictest rule you can express, and people write it by accident.
Block matches on the tool name or the pinned operation. Allow matches on the pinned operation only, because a tool's advertised name is a label that whoever edits the connector documentation can change. Governance binds the operation, never the label. The asymmetry is set out in block on the name, allow on the operation.
Frozen parameters handle the case where the tool is fine and one argument is not. Freeze the environment or the project key on an entry. The frozen key is stripped from the advertised schema, so the model never sees it, and the frozen value is merged over caller arguments at execution. Precedence runs entry defaults, then caller arguments, then frozen parameters.
The two minutes between deciding and enforcing
A restriction change takes about two minutes to take effect. So does a role change. Both resolve through a 60 second cache plus edge propagation, on every surface: MCP, the console and REST alike. A runbook that assumes otherwise will be wrong during the one incident that matters.
Three changes are effective on the next call instead. Grant revocation, member removal and suspension. The OAuth grant, the per-person authorization the client holds, is re-read from the organization store on every single call, with no cache. removing or suspending a member revokes every live grant in the same transaction as the membership change.
The operational consequence is simple. If an engineer is mid-incident and reaching something they should not, suspend the member or revoke the grant. Do not edit a restriction and then watch the next call go through anyway.
What the audit log answers afterward
The audit log, the append-only record of what happened, answers the question a paste cannot: which account did the assistant actually reach.
There is one entry per tool-call attempt, succeeded or failed, and both name the account. The recorded connection is the account actually reached, taken from the execution rather than from the intent. Recorded per call: the operation and tool, the connection, the classification, whether it was approved, the outcome, and an error code only. Argument names and counts are logged. Argument values never are.
actor_kind is a field, set at the point of action rather than inferred later from a user agent, and its values include ai_assistant. Whether a person or an assistant made a change is answered by the record, not by a guess.
Two details matter if you forward logs onward. The error string returned to the caller is derived from the third party's response body, and the string written to the audit trail never is, so a remote error body cannot leave through the log pipe. There is one log tenant per organization, enforced in the type system, and customer-visible audit sits in a different tenant from internal application logs. Export to Datadog is implemented.
The seat that reads all of this is free. Auditor is a read-only role with no license cost and no tool execution permission, so a compliance reviewer does not have to be squeezed out of headcount. The audit trail is eventually consistent, so a row may take a moment to appear.
What happens the day an engineer leaves?
Removal runs through a preflight, and it can refuse. Personal connections referenced by a toolbox entry, the saved collection of tools a person or team works from, must be resolved first. Transfer them to the organization, a team or another member, or delete them. Unreferenced personal connections are cleaned up.
A private connection is not transferable at all. A credential only its owner could ever use does not become somebody else's because its owner left. Delegated toolbox entries surface as a non-blocking warning, and re-pinning is the fix rather than a credential decision.
Once the removal goes through, every live grant is revoked in the same transaction as the membership change. That one is effective on the next call, not two minutes later.
Where scoped access does not help
Scoped access reduces the reason to paste. It does not police the chat window, and saying otherwise would be a lie by omission.
The prompt-injection write gate lives in the Elaichi agent window and does not apply to a raw tool call. What does hold on the endpoint: role-based permissions per operation, the forbidden classification that no OAuth scope can reach, output redaction, scope limits and full audit logging.
An engineer who wants to paste a trace into a personal account on a personal phone can still do it. What changes is that the fast path now runs through a tool call, and the fast path is the one people take.
When should a CTO not build this yet?
If you have eight engineers, one internal system worth querying and no contractual obligation to show who touched what, writing restrictions is overhead you will never read again. Use the assistant, keep secrets out of the log lines, and revisit when the second team asks. The case for waiting is laid out in the post on not needing a gateway yet.
If your real exposure is people who are not employees, start with the leaving path rather than the engineering path. Contractor offboarding and AI access covers the day somebody stops working for you.
If engineering is the second team rather than the first, the team playbooks and use cases show the same shape applied to support, finance and legal, and the connector catalog lists what is already authored.