Dekita

Your agent loads every tool schema by default. Decide which ones it should see.

ai agents mcp devtools

Before you change anything about your agent config, answer two questions:

# 1. how many prompt tokens does your harness send on a cold start?
# 2. how many of those tokens are tool schemas?
# measure with your provider tokenizer, not by counting characters.

Most teams cannot answer either question. The reason is not that the information is hard to get. It is that nobody decided who owns it.

On October 1, 2026, Earendil shipped Pi 1.0. The release notes list seven additions. Two of them change a decision you probably thought was made for you:

Both turn tool visibility into a per-tool setting. That is the whole point of this post, and it is a bigger change than it looks.

What changed, and what did not

Pi spent a year saying it did not need MCP. The landing page said it plainly, and the creator wrote a post arguing the protocol was unnecessary. Then 1.0 shipped with native MCP support, and Earendil published an explanation titled "You Said No, MCP!"

The stated reasons were that MCP had matured, and that the tool loading work Pi needed anyway would cover MCP. The technical reason is more useful: Pi had added deferred loading and mid-conversation system messages, so a tool now needs metadata saying whether it is exposed to the model directly, loaded on demand, or callable only from codemode. A regular MCP extension cannot see enough of the tool loadout to make that call.

So the reversal was less a change of heart than a collision between two features that needed the same plumbing.

Why visibility is a budget decision

A tool schema is not free. It sits in the system prompt or the tool block on every single request, including the ones where the tool is irrelevant. Connect enough servers and the schemas alone can eat a large share of your context window before the user has typed anything.

That cost shows up in three places.

Money. Every request pays for those tokens again.

Attention. A long tool list is noise the model has to read past on every turn. More tools means more chances to pick the wrong one.

Reproducibility. When a result changes between runs, the first suspect is often the tool set, not the model.

Pi reports a figure for its own codemode change: the changelog gives an example of a request dropping from roughly 5,300 to 3,300 prompt tokens, about 40 percent. That is a vendor example under one configuration, not a benchmark. Read it as a shape, not a number.

The three exposures, and when each one is right

Pi's metadata makes a tool one of three things. The choice is yours per tool.

Direct. The model sees the schema and calls it itself. Use this for tools used most turns.

Deferred. The tool is declared only when the task calls for it. Use this for the long tail: the specialist API you touch once a week.

Codemode only. The model never sees the schema. It writes code that calls the tool inside the sandbox. Use this when the output needs filtering before the model reads it, or when the tool is only useful in combination with others.

The third mode is the interesting one, because it changes what reaches the context. In Earendil's demo, a script pulls issues from a Linear MCP server, runs a classifier over the comments, and returns a ranked list. Hundreds of tool calls happen. The context window sees the summary.

Do the audit before you tune anything

Do not start by rearranging tools. Start by measuring.

Count the prompt tokens your harness sends on a cold start, with tools loaded and with none. The difference is the tax. Do it with a real request against your provider's tokenizer, not by estimating from character counts, because a schema is mostly punctuation and your estimate will be wrong in the direction that flatters you.

Then sort every tool into one of the three buckets, and write the list down. The exercise is worth doing even if you never change a thing, because the tools nobody can classify are the ones quietly inflating the prompt.

Two questions settle most cases:

codemode or deferred.

What codemode does not fix

Earendil is unusually candid about this, and you should read the caveat before you rewrite your config around it.

The problem is largely on the server side. Many MCP servers were built for harnesses that dump every tool schema into context, so they return text blobs optimised for token count rather than structured data. A harness-side sandbox changes where composition runs. It does not change what the server puts in your context when the harness asks for tools.

Earendil's own framing is that MCP should look closer to OpenAPI with intelligent tool discovery: structured returns, tools discoverable by their documentation. Until that lands, you are measuring your own servers.

So the practical order is: expose less, and where you control the server, return structure instead of prose.

Where I am unsure

I have not run these benchmarks. The 40 percent figure is Pi's own changelog example, and the demo numbers come from the vendor's post. What I have done is treat the three-exposure model as a design question worth asking, because the cost of an unnecessary schema is paid on every request and the cost of a missing tool is paid once.

Two things I would want before trusting any of this in production: the first request's token count from my own harness, and a before-and-after comparison on a fixed task with the model pinned.

Sources

https://earendil.com/posts/pi-1-0/

https://earendil.com/posts/you-said-no-mcp/

Secondary coverage; used for the install commands and the credential-per-server detail.

This post was written with AI assistance. The author is responsible for its content.

The permission half of the same decision

Visibility and authority are not the same thing, and a release can move both at once.

Pi 1.0 also hardened OAuth for MCP: credentials are stored per server name and URL, issuer checks follow RFC 9207, and step-up sign-in preserves the scopes a server was already granted. Those are real improvements, and they are authentication measures. They do not tell you whether a given connector is trustworthy, and they do not stop a model from acting on hostile instructions that arrive inside a tool result.

Treat the two separately in your own review. Ask, per server: what is this allowed to reach, and who approved installing it. A connector you enabled for a narrow read is still a connector the model can call.

This is also why "expose less" is a safety measure and not only a cost

One more thing worth writing down: the audit has a second output. Once the tools are sorted, the shape of your own usage is visible, and that tells you which single tool to fix first. On most teams it is not the model. It is the one connector added for a deadline, still enabled, still loading its schema, rarely used. measure. A tool the model cannot see is a tool it cannot choose to call by mistake.