Appearance

System follows your phone’s setting. Default.

Custom MCP
Date
Read
15 min
Views

My HubSpot agent loaded 88 tools to use five. So I made it search. It picked right every time. The bill went up anyway.

Oct 2, 2026 · 15 min read ·

274 Claude Code sessions on my own HubSpot: tool and skill selection as MCP tools, picked by a decision model with a probability per option. Right 24 of 24 times. Context fell 12x. The bill rose 65%.

01 · The symptom
“Every session loaded 88 tool definitions to use five of them. When the model wanted a skill, it guessed the name.”

I started with a proof of concept: a decision model picking the right sales skill from inside an MCP server, on a mock CRM. This is the real one. I built tool selection and skill selection as MCP tools on my production HubSpot MCP server, put TypeSafe’s Jev in charge of the picking, and ran 274 headless Claude Code sessions against my own portal.

I wanted five answers. Is the pick right? Does it come with a probability I can act on? What does it cost? How long does it take? And does it stop the host from loading every tool and every skill at the start of every session?

Four of the five came back clean. The fifth is the honest part of this post. The host loaded 12 times less context, and the task still cost more. Prompt caching had already made the big tool list cheap, and routing spent its tokens where they’re dear. The arithmetic is below.

02 · The gap

What I set out to prove

An MCP server hands the host every tool it has. Mine has 88, plus HubSpot’s nine sales skills. The host loads the list once per session, and the model reads it on every turn. A rep asking for a morning brief needs one skill and five tools. The session gives the model all of it and lets it pick.

The alternative is to make selection itself a tool. The server exposes a handful of meta-tools. The model says what it wants in the user’s own words. A cheap, fast decision model returns the skill, the charter and the tools that fit, each with a probability. The model loads what it needs and gets on with the job.

That claim has four parts, and I wanted each one measured on real sessions, not on a diff of the tool list:

  • Accurate. The first pick is right nearly always, and when it isn’t, the right answer is still in the list.
  • Scored. A probability per option, so the server can set a threshold, the model can see the runner-up, and I can audit the misses afterwards.
  • Cheap and fast. A rounding error on the session, in money and in seconds.
  • Less context. Four definitions loaded instead of 88 and nine skills’ worth of text, and the model searches for the rest.

The proof of concept measured the tool list and the routing accuracy on labelled prompts. It never measured a session. This time I did.

03 · The fix
Custom MCP

Four tools instead of 88

The server now serves two surfaces from one process. The original endpoint doesn’t change: 88 tools and the write gate. The new one, /mcp/routed, advertises four tools.

find_capabilities takes the request in the user’s words. It returns the matching sales skill, the matching specialist charter and the HubSpot tools they need, each with a probability. The tools come with their input schemas inline, and a one-line next_step tells the model what to load. load_skill returns one of HubSpot’s nine skills, byte for byte, behind a short header that maps the vendor’s connector tool names onto mine and explains the write gate. load_charter returns one of the 44 specialist charters the server already shipped as MCP prompts, which hosts never pull on their own. And call_hubspot runs any of the 88 tools by name. It checks the arguments against the tool’s JSON schema first, then dispatches into the exact function the registered tool uses. Previews, approval tiers, undo and the audit log are identical on both surfaces.

YOU CLAUDE MCP SERVER JEV WRITE GATE Change the amount on one deal to 12,500 find_capabilities hands over the ask 134 questions 2 choices, 132 yes/no Odds per option objects 0.90 Skill, charter, tools odds + input schemas load_charter, then call_hubspot(update) Preview, CONFIRM amount is sensitive You approve or not nothing changed yet Same handlers, same gate: call_hubspot dispatches into the code path the 88 registered tools use. What changed is what the host loads first: 4 tool definitions instead of 88, and a skill or charter the model reads.
01 · You Change the amount · on one deal to 12,500 02 · Claude find_capabilities · hands over the ask 03 · MCP server 134 questions · 2 choices, 132 yes/no 04 · Jev Odds per option · objects 0.90 05 · MCP server Skill, charter, tools · odds + input schemas 06 · Claude load_charter, then · call_hubspot(update) 07 · Write gate Preview, CONFIRM · amount is sensitive 08 · You You approve or not · nothing changed yet Same handlers, same write gate. 4 tool definitions loaded instead of 88.
One routed request. Claude hands the ask to find_capabilities, Jev answers 134 typed questions in one call, and the matches come back with their schemas. Claude loads the skill and writes through call_hubspot, into the same gate as the other 88 tools.

The router is Jev, TypeSafe’s evaluation-only model, reached through Vercel’s AI Gateway. One HTTPS request carries 134 questions: pick one of the nine skills or none, pick one of the 44 charters or none, and a yes or no for every skill, charter and domain tool. Jev returns a probability for each. Anything at 0.5 or above comes back, plus the chosen skill’s and charter’s tools. If the gateway fails or the key is missing, the server falls back to the keyword router it already had and says so in the response. That fallback turned out to matter.

The same selection sits on the full surface too, as two ordinary tools: hubspot_find_skills and hubspot_load_skill. A host that loads all 88 tools still gets routed skills. That let me test the skill idea with and without the small tool list.

04 · The method

274 sessions against my own portal

Every number here comes from headless Claude Code sessions against the PromptMetrics HubSpot portal. Same isolation flags, same operator framing on the prompt, arms interleaved so time of day and HubSpot latency hit them equally. A runner parsed each transcript for exact per-call token usage, the tool-call sequence with timing, turns, duration and cost. It checked completion against the portal, then rejected pending previews and undid every auto-tier write so the next session started clean. Writes only ever touched one test contact and one test deal.

Four runs, 274 sessions in all. A surface A/B: full surface against routed surface against no MCP, 12 tasks three times each, 106 sessions after two hung control prompts were dropped. A skill bench: four skill-driven tasks on both surfaces, 24 sessions. A router comparison on the routed surface alone, Jev against keywords, 20 tasks three times each, 120 sessions. And a Sonnet follow-up: the four skill tasks, Jev against keywords, on Sonnet 5, 24 sessions.

Everything but the last ran on Claude Opus 5.5, the Claude Code default that day. Costs are the API-equivalent at Anthropic’s list prices, as Claude Code reports them per session. The sessions themselves ran on a subscription. Both arms of every comparison used the same model, so repricing moves the totals, not the comparison.

05 · Accuracy

Right every time, with a score next to it

Right pick on the first lookup, live Claude Code sessions on the real portal Jev, skills, Opus 5.5: 24/24 (100%) Keyword, skills, Opus 5.5: 14/24 (58%) Jev, charters, Opus 5.5: 35/42 (83%) Keyword, charters, Opus 5.5: 32/42 (76%) Jev, skills, Sonnet 5: 12/12 (100%) Keyword, skills, Sonnet 5: 9/12 (75%) Jev answered 57 of 57 routes, 651 ms median, 1,079 ms p95, $0.0005 each. Six of its seven charter "misses" picked a skill instead, as designed.
Right pick on the first lookup, live sessions Jev, skills, Opus 5.5: 24/24 Keyword, skills, Opus 5.5: 14/24 Jev, charters, Opus 5.5: 35/42 Keyword, charters, Opus 5.5: 32/42 Jev, skills, Sonnet 5: 12/12 Keyword, skills, Sonnet 5: 9/12 Jev: 651 ms and $0.0005 per route.
Right pick on the first lookup of each live session. Jev: 24 of 24 on Opus, 12 of 12 on Sonnet. Keywords: 14 and 9. On charters the two are close, and six of Jev's seven misses were sessions where it chose a skill instead, which is what the design asks for.

Offline first, because live tasks have to be sparse. On 40 labelled prompts, 16 of them written to sit between two skills, Jev picked the right skill first 36 times, and 13 of the 16 boundary cases. Keywords managed 17 and 4. On the server’s 49-prompt charter corpus, written for the keyword router and full of literal requests like “list workflows”, the two were level: 46 and 47 of 49.

Then live. I ran two copies of the routed server, one on Jev and one with the keyword fallback forced on, and sent 20 tasks through each three times. On the 24 sessions whose task names a skill, Jev’s first pick matched the label 24 times. Keywords matched 14. Every keyword miss had the same shape: a prompt with no trigger phrase in it. “I spoke to Chris this morning, legal is done on their side” never says log a call. “Which deals are closing this month and which have gone quiet” never says pipeline. Keywords returned no skill at all for those, and put call prep above log-call for a prompt that began “just got off a call with”. Jev got every one.

Right skill picked first, 40 labelled prompts (16 on the boundary between two skills) Jev, overall: 36/40 (90%) Jev, boundary cases: 13/16 (81%) Keyword fallback, overall: 17/40 (42%) Keyword fallback, boundary cases: 4/16 (25%) Gate: beat keywords by 15 points on boundary cases. Jev did by 56.
Right skill picked first, 40 labelled prompts Jev, overall: 36/40 Jev, boundary cases: 13/16 Keyword fallback, overall: 17/40 Keyword fallback, boundary cases: 4/16 Gate +15 on boundary; Jev +56.
The offline eval before the live runs: 40 labelled prompts, 16 on the boundary between two skills. Jev 36 of 40 and 13 of 16. Keywords 17 and 4.

On charters Jev was right 35 of 42 times against 32, and the gap is smaller than it looks. Six of Jev’s seven misses were sessions where it chose a skill and no charter, because the prompt was a skill request and the label also allowed a charter. The keyword misses were generic beating specific: “open deals with no next activity” went to lists, and “list our active workflows” went to objects, with workflows second.

The probability is what makes a miss cheap. Here’s what the live server returns today for three requests, placeholder names only:

find_capabilities, live
"Just got off a call with Acme: they agreed the 12-month term, contract goes to their CFO next week. Log it."
  skills:   log-call 1.00, hubspot 0.70     charter: engagements     tools: 26     739 ms, $0.0005
  next_step: Call load_skill(name="log-call") and follow that workflow …

"Which deals are closing this month and which ones have gone quiet?"
  skills:   pipeline-pulse 1.00, hubspot 0.85     charter: objects     tools: 39     559 ms
  next_step: Call load_skill(name="pipeline-pulse") …

"What's the weather in Berlin?"
  skills:   none 1.00     charter: none     tools: 0     713 ms
  next_step: No HubSpot charter matches. … otherwise answer without HubSpot.

A wrong keyword pick usually cost a detour, not the task. The model read past the first result and took the right skill from second place on every log-call rep. Where the right skill wasn’t in the list at all, it failed all three reps of that task, because it worked from a charter without the skill’s “gone quiet” logic. The right answer being in the list, with a score next to it, is the difference between a slower session and a wrong one.

Sonnet 5 repeated the pattern on a smaller model: Jev 12 of 12, keywords 9 of 12, the same wrong pick on the log-call prompt. Sonnet recovered too, but paid more for it. With Jev its sessions ran two turns shorter, 14% cheaper and 16% faster at median, with no malformed tool calls against four. On Opus the keyword arm had been slightly cheaper per session. Twelve sessions per arm is a direction, not a measurement, but it points the way you’d expect. The less capable the model, the more the first pick matters.

06 · Cost

What a route costs

Over 57 live routes on the Opus run, Jev answered every one. Median latency was 651 ms, p95 1,079 ms. Each route cost $0.0005, about 0.2% of the session it sat in. The Sonnet run’s 12 routes came in at 780 ms median and 1,114 ms p95, for the same $0.0005, or 0.3% of the session.

That’s 134 typed questions, answered in one call, for half a tenth of a cent. The routing bill for all 60 Jev sessions in the router comparison was $0.027. The keyword router is free, instant and sends nothing anywhere, which is why it stays as the fallback and as the opt-out for anyone who doesn’t want request text leaving the server. Nothing else leaves either. Jev sees the user’s request and the skill and tool descriptions, never portal data.

One lesson came free. My second attempt at the Sonnet run put 24 sessions through with both arms on keywords, because the Jev instance had been restarted in a shell without the gateway key, and the fallback is silent by design. The runner now probes each arm’s router at startup and refuses to run on a mismatch. A fallback that works too well needs a tripwire.

07 · The catch

Less context, bigger bill

Context in the first request, median over 106 sessions No MCP: 3,304 tokens of Claude Code framing Full surface, 80 tools: 17,944 tokens, +14,640 Routed surface, 3 tools: 4,517 tokens, +1,213 The full surface adds about 12 times more context than the routed one before the model does anything.
Context in the first request, median over 106 sessions No MCP 3,304 tokens of Claude Code framing Full surface, 80 tools 17,944 tokens, +14,640 Routed surface, 3 tools 4,517 tokens, +1,213 About 12 times less context added up front.
Context in the first request, median over 106 sessions. Claude Code itself is 3,304 tokens. The full surface adds 14,640 on top. The routed surface adds 1,213.

The context claim held outright. Before the model does anything, the routed surface adds 1,213 tokens where the full surface adds 14,640. Across a whole session it carries a third less context per API call, 12,906 against 19,242, because the tool list rides along in every request and the routed one is tiny. The tools/list today is 33.7 KB for 88 tools and 2.4 KB for four.

Per task, mediansNo MCPFull, 80 toolsRouted, 3 tools
Context in the first request3,30417,9444,517
Context per API call3,30419,24212,906
API calls2712
HubSpot tool calls, excluding routing044
Cost$0.024$0.067$0.147
Duration8 s20 s29 s
Tasks completedn/a36 of 3636 of 36

Then the bill said the opposite. The routed surface cost 1.65 times more per task and took 1.4 times longer. Both surfaces finished every task, with identical write-gate behaviour, so this isn’t a capability gap. It’s prompt caching.

Context tokens over 36 sessions per arm: cache reads against cache writes Full, 80 tools: 6.37M cache reads, 0.78M cache writes, 5,851 output tokens, $4.29 Routed, 3 tools: 5.19M cache reads, 1.45M cache writes, 7,269 output tokens, $7.07 cache reads, about a tenth of list price cache writes, about 1.25 times list price Routed used 7% fewer context tokens and cost 65% more: it moved tokens from cheap reads to dear writes.
Context tokens over 36 sessions per arm Full, 80 tools reads 6.37M, writes 0.78M, $4.29 Routed, 3 tools reads 5.19M, writes 1.45M, $7.07 cache reads, about a tenth of list price cache writes, about 1.25 times list price Fewer tokens, higher cost: reads became writes.
Context tokens over 36 sessions per arm, split into cache reads and cache writes. The full surface read 6.37M from cache and wrote 0.78M. The routed surface read 5.19M and wrote 1.45M. Fewer tokens, higher bill.

How does context fall 12 times while cost rises? Because “context” above is the first request, and the bill is set by four things the first request doesn’t show.

The big tool list is cached, so it’s nearly free after the first turn. The host sends it with every request, but the API caches it and charges a cache read at a tenth of the price: $0.20 against $4 per million tokens on Opus. In the full arm, 89% of all context tokens were cache reads. The 14,640 tokens of definitions are written once per session and read cheaply after that. The tax I set out to remove was already mostly discounted.

What routing adds isn’t cached when it arrives. The lookup result, with its two dozen schemas and the skill or charter text, lands in the conversation as new context. New context is a cache write, and Claude Code’s one-hour cache writes cost $8 per million, twice list price. Routing removed cheap tokens and added expensive ones.

Then there are more turns. The lookup is a round trip, loading the charter is another, and on write tasks the charter told the model to verify what it wrote. Twelve API calls per task against seven, and every call re-sends the whole conversation. And more output. The routed arm produced 80% more output tokens, at $20 per million the dearest token there is.

Over 36 sessionsFull, 80 toolsRouted, 3 tools
Cache reads, $0.20 per million6.37M tokens5.19M
Cache writes, $8 per million0.78M1.45M
Total context tokens7.15M6.64M
Output tokens, $20 per million5,8517,269
Cost$4.29$7.07

The routed arm used 7% fewer tokens in total and cost 65% more. It moved two thirds of a million tokens from the cheap column to the expensive one and spent more on output. It saved tokens where they were already cheapest and spent them where they were dearest.

So the context claim is true and, on one server with one model and caching on, financially beside the point. The number everyone measures is the one prompt caching already made cheap. Where it isn’t beside the point: a host with several MCP servers attached, where tool lists compete for the window. A smaller model, where the Sonnet run already shows the first pick mattering more. Long sessions, where definitions crowd out the work. And any host or model without prompt caching. The design I measured is also the naive one. It returns 24 tool schemas where three would do, and loads the charter in a separate turn. Both are a day’s work and a re-run.

08 · Behaviour

None of this works unless the model actually searches instead of guessing. Four things the transcripts showed.

It looks up first, once you tell it to. On the full surface, before the server instructions said to call hubspot_find_skills first, the model guessed a skill name, daily brief with a space, failed, then tried daily-brief. One instruction line fixed that in every session after. On the routed surface it called find_capabilities first in every session that had anything to do with HubSpot, and skipped it for the weather.

It reads the whole list, not the top line. Every time the keyword router put call prep first and log-call second, the model loaded log-call. The probability and the runner-up aren’t decoration.

It accepts correction. When the model sent an association lookup with a field missing, the schema check returned the exact missing field, and the model fixed its arguments and carried on. The full surface’s registered tools accept unknown arguments silently. The keyword arm, which started from the wrong charter more often, drew four of these corrections. The Jev arm drew none.

It follows the skill it loads. Call prep saved a prep note. Log-call recorded the call as a note, created the follow-up task and set the lead status, as HubSpot’s workflow says. Sonnet 5 went one further than Opus and set the lifecycle stage too. Every write went through the gate and the runner undid it. On the full surface, with all 88 tools in view, the loaded skill changed what the model did in exactly the same way.

That last one I didn’t expect. Serving a skill through a tool result works as well on a host that loaded everything as on one that loaded four tools. The small tool list and the served skills are separable. You can take the second without the first.

09 · The idea

Skills behind a login

The install problem is gone either way. Nobody copies a folder, and when HubSpot ships version 2.4 of its skills, the server updates them for everyone at once. Once a skill is something a server serves rather than a file a rep installs, three things follow that a folder never had.

The server knows who’s asking. The hosted surface authenticates every request and ties it to a portal. find_capabilities could return only the skills that account is entitled to, and load_skill could refuse the rest, with the same scored lookup for everyone.

A skill is a current version, not a copy. Fix a workflow on Tuesday and every subscriber runs the fix on Tuesday. The mapping header that turns the vendor’s tool names into mine is the same idea applied to compatibility.

Usage is visible. Every routed session leaves a trace of which skill was picked for which kind of request, with its probability. That’s what a product owner needs to know which skills earn their place, and it’s data a folder of markdown never produces.

Put together, that’s a skill pack sold as a connector URL. A team adds the server, signs in, and the skills they pay for show up in the lookup. None of it is built. There’s no entitlement store, no billing, and the routed surface identifies a portal, not yet a person. Two cautions before anyone runs with it. The skill text leaves the server in the tool result, so a subscriber can copy it. What the subscription sells is the current version, the routing, the live mapping onto the tools and the write gate, not a secret. And HubSpot’s own skills are Apache 2.0, so this applies to skills you write, not to theirs.

10 · Open

What isn’t proven yet

  • One model for the surface comparison. Opus 5.5 only. The Sonnet 5 run covered the router question on four tasks, not the surface question, so whether a smaller model benefits from the small tool list itself is still open.
  • One server on the host. The headroom benefit depends on several MCP servers sharing a session, and this bench didn’t set that up.
  • Headless only. Claude Code declares inline confirmation, so confirm-tier writes came back as a request the headless client declined, and the gate recorded a reject. Interactive sessions would show the form and wait. Same code path on both surfaces, but I haven’t sat through it on the routed one.
  • An unoptimised design. Fold the charter into the routing result, return the top three schemas instead of 24, trim the verification boilerplate. Each one is a re-run.
  • Small live samples. 24 skill-labelled sessions per arm on Opus, 12 on Sonnet. The offline eval is 40 prompts. Enough to see the shape, not enough to quote a percentage to two digits.
  • Local runs. The routed surface and the served skills are deployed, and the hosted routed endpoint has its OAuth resource registered, but every measurement here ran against a local instance of the same code.
11 · Lessons

What I’d tell anyone building this

  • Make selection a tool, and give it a probability. The score is what lets the model recover from a wrong first pick and lets you audit the misses afterwards. A router that returns one name is a router you can’t debug. And tell the model to look up first, or it guesses. One sentence in the server instructions does it.
  • Pay for the router. It’s the cheapest part. Half a tenth of a cent and under a second, for the difference between 24 of 24 and 14 of 24.
  • Keep a free fallback, with a tripwire on it. Keyword routing is the right opt-out for anyone who wants nothing to leave the server. It’s also silent when it takes over by accident. Probe for it.
  • Price the tokens you save at what they actually cost. With prompt caching, a static tool list is nearly free after the first turn. Count cache reads and cache writes separately before you claim a saving.
  • Return less than you can. Two dozen schemas inline were most of the context the routed arm spent. Return the skill and the top three tools, and let the model ask for more.
  • Serve skills as a tool result, not a prompt. Hosts don’t pull MCP prompts. A tool that returns the text, behind a header that maps the vendor’s tool names onto yours, works on every host, with a small tool list or without one.
  • Keep the gate shared. The only reason I can say both surfaces behave the same on writes is that they run the same function.

The code is on the main branch of hubspot-mcp: the routed surface, the served skills, the bench, the two evals and the five reports, with the design and every number in docs/routed-mode.md. The runner, the twenty tasks and the report are in bench/. Run it against your own portal, never your live one, and if your numbers point the other way, I want to see them. Tell me where I’m wrong.

12 · What it saved
24 of 24
right skill on the first lookup, against 14 of 24 for keyword matching
$0.0005
per routing decision, in 651 ms, about 0.2% of what a session costs
Reactions
Subscribe

Get the next MCP-gap post in your inbox. Every two weeks, no fluff.

More on the Custom MCP rung
Comments

Comments

Loading comments…

    Leave a comment

    Not shown publicly.
    Work with
    Done-for-you buildsWork with PromptMetrics →