case studies

MCP Case Study: An AI Agent Gets Its Answers in a Few Hundred Tokens Each

An AI agent that needs a number from operational data usually gets it by reading the data: rows from a database through a SQL tool, lines from a log search, an export. Every row it reads costs tokens and room in its context, and most of what it reads is detail the question never needed. A Precomputing file already holds the answers its policy names, so an agent can ask for those instead.

In the MCP demo the files the four other demos write are each served by an MCP server, the one built into precomputing serve. Each file is shown as its policy leaves it once the detail it keeps for a short while has faded. An agent reads what each file keeps, then asks four questions: a p99 from API traffic, a candle from a trading day, an invoice from a month of AI usage, and what went wrong in a web shop’s logs.

Tokens for each of the four questions, on a scale of powers of ten. The p99 of /api/checkout for one hour: 179 tokens in the answer against 283,641 in the raw data behind it. SIM4’s candle at 10:30: 129 against 9,702. Harbor’s September invoice: 126 against 440,661. New log patterns since 12:40: 351 against 4,595,435.

The four answers as the demo measured them: tokens in the text the model reads against tokens in the raw events it summarizes.

The Set-Up

Four files over MCP
Files latency (Demo 1, a day after its three hours), trades (Demo 2, an hour after the close), usage (Demo 3, when September’s dispute window ends), shop (Demo 4, two days after its two hours): 7.6 MB in all
Server One per file: the Go code of precomputing serve, compiled to WebAssembly, on the 2026-07-28 revision of the protocol
Tools describe_file, get_answer, get_windows, get_kept and query, all read only. Answers are short CSV text with a line of context
Agent A script that makes the calls a careful agent would make, so every run is the same. A recorded session with a real agent is on the same page
Tokens Estimated as a quarter of the bytes, for the answers and for the raw events behind them, as CSV
Checks Every figure in every answer against a count of the raw events by separate code

What the Agent Did

It read each file’s description once: 219 to 826 tokens a file, 1,776 in all. A description says what streams the file has, what it keeps ready, and how long each level of detail lives. Then it asked:

  • The p99 of /api/checkout between 10:00 and 11:00 UTC, and the requests that stood out. Two calls: the hour’s window summary and the unusual requests kept whole. 179 tokens back: p99 907 ms over 35,907 requests, the slowest at 2,354 ms, kept whole. The raw requests of that hour would take 283,641 tokens.
  • SIM4’s 1-minute candle at 10:30 New York time. One call, 129 tokens: open 23.36, high 23.37, low 23.33, close 23.36, 522,174 shares in 1,497 trades.
  • What Harbor Legal Drafts owes for September, line by line. Two calls to the precomputed invoice views, 126 tokens against 440,661 for the month’s requests.
  • Which log patterns are new since the payment incident began at 12:40. Two calls: the templates first seen since then and the error lines kept whole. 351 tokens against 4.6 million for the lines.

The invoice, as the model reads it:

invoices (view). Rows with customer = harbor; period 2026-09.
name,plan,requests,input_tokens,output_tokens,list_nano,discount_nano,due_nano,due_cents,cost_nano
Harbor Legal Drafts,pro,39586,99873977,37356271,2374674473500,0,2374674473500,237467,1187337236750

invoice_lines (view). Rows with customer = harbor; period 2026-09.
model,requests,input_tokens,output_tokens,list_nano,cost_nano
large,23730,75589746,27192743,2212328616000,1106164308000
medium,15856,24284231,10163528,162345857500,81172928750

$2,374.67 due, equal to the Meter’s recount to the billionth of a dollar, and the lines add up to it.

A Real Agent

The same page shows a session recorded on 30 September 2026 with Claude, made by Anthropic, connected to four precomputing serve --read-only servers over HTTP with a read token, through the official MCP TypeScript client. Claude built this release and had seen the files while building it; in the session it read them only through the tools. It answered four questions of its own in 18 calls.

The session keeps a mistake. Asked which endpoint got slower, the agent first asked for hourly windows without a time range, and the tool read the last hour, as its description says it will. The answer named the hour it covered, the agent noticed, asked again for all three hours, and found /api/report’s p99 rising from 584 ms to 659 ms, with /api/checkout’s bad hour traced to a two-minute incident at 10:30.

The Numbers

For the four questions Result
Calls 7 for the answers and 4 descriptions: 2,561 tokens in all
The raw data behind the answers 21.3 MB as CSV, about 5.3 million tokens: 2,081 times more
Checks 19 of 19 against the raw events, percentiles within 1%
Native against the browser The same calls to precomputing mcp on the same files: 11 tool results and 4 tool lists, identical to the byte
Official clients The MCP TypeScript client, version 2 on the 2026-07-28 revision and version 1 on the older handshake, over HTTP and stdio: all pass on all four files
Tokens Without a token the server answers 401; a read token that posts events gets 403
In the browser The four files download in 2.1 MB compressed; the server with the Engine is 5.9 MB, 1.6 MB compressed

From a headless run of the demo’s own code with the page’s builds, and from the native binary on a two-core cloud server (Intel Xeon at 2.8 GHz). The data in the files is simulated, as in the demos that wrote them.

What the Demo Revealed

The official client refused the server’s first tool list, which left out how long a client may cache it and sent null for each tool’s annotations. The server now sends both properly. The server is also built for the browser, and its first browser build pulled in Go’s whole HTTP stack, which the page never uses. Moving the HTTP code into a file the browser build leaves out brought it back to size.

Most of the tokens an agent spends here go to the descriptions, read once per file. After that, an answer costs about as much as a short paragraph.

Next: Real Files

The tools only read, and a read token cannot post events, so a file can be opened to an agent without opening it to changes. The server speaks plain HTTP; TLS is on the roadmap for 0.2. The step after that is real files with real agents, in pilots.

Try It Yourself

Open the MCP demo and press Ask the four questions. Each call shows its arguments, the text the model reads and its size in tokens, and each answer is checked against the raw events. Then make calls of your own, mistakes included, and read the recorded session with a real agent.