Live Demo: Traces

A day of AI coding-agent runs goes through Precomputing Traces, right here in your browser. A coding agent sends the whole conversation again with every model call, so a run’s 40th call repeats its first 39. Traces keeps each message and tool list once, under the SHA-256 of its bytes, and each call as the list of its pieces. Secrets are masked before anything is stored, and any call can be rebuilt from the file byte for byte. Every call is also metered: its tokens and its cost by run, repository, model and source, with a daily budget per repository.

Generated runs, real store. The runs are generated by a script in the shape of a public dataset of agent runs, with invented repositories, issues and code. The store, the Engine and SQLite are the real code, compiled to WebAssembly. Prices are examples.

How to Run It

  1. Press Play. The day starts at 07:00 UTC. 100 agent runs on 30 repositories make 2,045 model calls, each reported as its reply comes back. The day passes in about a minute and a half, slowing down when a run gets stuck.
  2. Make trouble. Report a finished run again, the way a tracer that retries an upload would, and watch the meter refuse it. Once a repository spends its budget, raise it.
  3. Rebuild a call. Pick a finished run and a call, and read it back from the file. The page compares it with the request the agent sent, byte for byte and by SHA-256.
  4. Review the results. The day finishes, every call is rebuilt and compared, and every cost is checked against a recount. Then ask the file in SQL and download it.

On a wide screen the demo can use the whole window: open it full screen.

The Controls

Control What it does
Play, Pause, Resume Starts the day, pauses it and carries on. An uninterrupted day takes about a minute and a half
Stop Ends the run and goes back to 07:00
Review the results Finishes the day at full speed, then rebuilds every call and recounts every cost
Report a run again Sends every call of the latest finished run to the store a second time
Raise the budget Once a repository has spent its budget for the day, raises it from $0.20 to $1.00
Rebuild Reads one call back from the file with SQL and compares it with the request as the agent sent it
Policy, Compiled SQL, The store’s tables, Secrets it masks Show the policy, the SQL it compiles to, the tables the store adds and the patterns it masks
Run Runs your SQL against the file while the day is paused or finished

The Day

Each run is an agent fixing one issue in a Python repository: it looks around, reads code, reproduces the bug, edits, runs the tests and finishes. A run sends a system prompt, the task, its own messages and the output of its tools: a shell, a file editor and a scratchpad for thinking. A fourth tool ends the run. Runs make 13 to 28 calls each, in one to two and a half minutes, except one. The script plants four things:

  • 08:47. An agent lists the environment, and the output holds an AWS key id and secret key.
  • 11:10 to 11:22. An agent on tallowcraft/tarnforms gets stuck and runs the whole test suite verbosely again and again: 67 calls, the last ones over 300,000 tokens each.
  • 13:50. An agent prints a git configuration with a GitHub token in the remote’s URL.
  • 17:11. An agent opens a local settings file that holds an API key.

The keys are fake, in the shapes real ones have. Tokens were counted with the o200k_base tokenizer. The provider’s prompt cache is modeled the usual way: a call’s input is served from the cache up to the previous call’s input when that call was sent less than five minutes before.

The Policy

stream calls exact {
  id     call_id refuse repeats 7d            # a call reported twice counts once
  key    repo text
  key    run text
  key    model text
  value  input_tokens integer
  value  cached_tokens integer
  value  output_tokens integer
  derive cost_nano = (input_tokens - cached_tokens) * 400 + cached_tokens * 40 + output_tokens * 1600

  period day close 24h
  raw    until closed + 30d
  rollup 1h keep 400d
  rollup 1d keep forever
}

stream context exact {
  id     part_id refuse repeats 7d            # one event per call and source
  key    repo text
  key    source text                          # system, tools, task, assistant, output, or tool:NAME
  value  tokens integer
  value  cached_tokens integer
  value  output_tokens integer
  derive cost_nano = (tokens - cached_tokens) * 400 + cached_tokens * 40 + output_tokens * 1600
  ...
}

precompute run_cost       = sum(calls.cost_nano) by run
precompute repo_cost_day  = sum(calls.cost_nano) by repo per day
quota repo_budget = repo_cost_day
precompute source_cost_day = sum(context.cost_nano) by source per day

Prices are in billionths of a dollar per token, as examples: $0.40 a million fresh input tokens, $0.04 cached and $1.60 output. Beside the policy’s tables the store adds two of its own and a view: trace_pieces holds each message and tool list once, trace_calls holds each call as the list of its pieces with its size and SHA-256, and trace_requests rebuilds any call with SQL alone.

What the Numbers Mean

  • Model calls kept: calls in the file, one per reply that came back.
  • The calls as sent: the bytes of every request as the agent sent it. A tracer that stores each call whole needs this much.
  • The same calls in the file: the store’s tables and their indexes. The meter’s tables sit beside them in the same file.
  • Spent today: the sum of repo_cost_day, at the example prices.
  • Secrets masked: matches of the store’s patterns, replaced by a label such as [redacted aws-secret] before the piece was hashed and stored.
  • Calls reported twice, refused: calls the store already held. Their events go to the meter again, and the exact streams refuse them by id.
  • The budget bars: the repo_budget view, one row per repository. A gateway in front of the agents would read it before letting a run’s next call through.

Things to Try

  1. Let the day finish untouched. Press Play and wait, or Review the results at once. 2,045 calls, 90.9 MB as sent, kept in 6.5 MB, 14 times smaller. All 2,045 rebuilt from the file are the requests as sent, byte for byte, and every cost matches the recount: $1.46 for the day.
  2. Watch 11:10. The stuck run fills the costliest-runs table and its repository’s budget bar. At 11:17 the bar turns red: call 45 took tallowcraft/tarnforms past $0.20. The run makes 22 more calls after that.
  3. Rebuild the stuck run’s last call. It is 1.46 MB, rebuilt from 135 pieces in a few milliseconds, with the same SHA-256 as the request the agent sent.
  4. Report a run again. Its calls go to the store a second time. The costs do not move, and calls_refused says how many were refused.
  5. Look for the secrets. In Ask the file, “Secrets masked” shows each label in its place. Nothing in the file matches a secret’s pattern.
  6. See which pieces repeat. “Pieces sent most often” shows the system prompt sent with every call, stored once, and the stuck run’s test output sent 45 times over.

Using Real Runs

The generator writes runs with the fields of nebius/SWE-rebench-openhands-trajectories, a public dataset of real agent runs. The prototype includes a script that reads one of its Parquet files, and the demo takes those runs unchanged. A page that shows them credits the dataset, which Nebius publishes under CC BY 4.0.

If It Does Not Start

The demo works in any current Chrome, Edge, Firefox or Safari, on a computer or a phone, with JavaScript switched on. It needs WebAssembly, which all of them support. It downloads SQLite (1.5 MB, 577 KB compressed), the store with the Engine (5.9 MB, 1.6 MB compressed) and the day’s runs (876 KB compressed) from this site. If the demo says it could not start, try another browser, or allow scripts on this page if an extension blocks them.

In testing no check has failed. If one ever does, the results say so in red, and the project would be glad to hear about it through the contact page.

What Is Real Here

Real:

  • The store of agent calls, the secret patterns and the Engine: the same Go code as precomputing traces, compiled to WebAssembly.
  • Every byte kept, every call rebuilt, and every cost in the file.
  • The checks: separate code builds each request as the agent sent it and works out each call’s tokens and cost from the runs. It never reads the file for that.

Generated or simulated:

  • The runs: repositories, issues, code, tool output and the agent’s words are invented, with the fields of the public dataset.
  • Token counts, the prompt cache, the times of day and the prices.
  • The gateway that would read the budgets. The demo shows what it would read.

The native command keeps the same day in about a second and writes the same values in every table: 19 tables, 40,043 rows, 0 differences, with only the demo’s budget limits left out. See the Platform page. The agent traces case study follows the day.