About
Precomputing is a project at an early stage. Its code works on simulated data, and four live demos run it in your browser. The next step is real traffic. This site shows where the project stands, gaps included.
Why the Project Exists
Most systems that collect events ask the same questions of them again and again. How many requests did each endpoint get, and what was the p99? What does each customer owe this month? What does the dashboard show for the last hour? The usual answer is to keep every event and compute the answer each time someone asks. That means storing everything and scanning it on every question. For logs it also means paying a vendor by volume to hold it all.
The questions are known in advance, so the answers can be kept ready. Counts, sums, minimums and maximums, percentile sketches and a few kept examples all update in place as each event arrives, and they merge, so short windows roll up into long ones with no loss. Once the answers are ready, raw detail can fade on a schedule while unusual events stay whole. In Demo 1 that turns 35.2 MB of raw requests into a 4.5 MB file that answers every question in a few milliseconds.
The pieces are well known: materialized views, rollups, quantile sketches, log templates. Precomputing puts them behind one short policy and one SQLite file, so the same idea works inside an application’s own database, in a fast standalone engine, for billing and for logs.
What the Project Is Building
- The language and the file. A short policy says what to keep; any SQLite tool reads the result. The policy language
- Two runtimes. Compiled triggers for any SQLite, and the Engine, a Go binary with SQLite built in, 12 to 17 times faster on trades and safe through crashes. How it works
- Two products on top. The Meter, for usage that has to be counted exactly, and Logs, which keeps a log dashboard’s answers ready and the raw lines on site. The platform
- Next, real traffic. Closing the gaps 0.1 left open, then pilots. The roadmap
How the Project Works
- Measure, then say it. Every figure on this site was measured on the 0.1 code, or is marked as an estimate. The prototype includes the commands that reproduce the native and headless figures, and each demo shows its own figures as it runs.
- Say what is simulated. The demos’ data is invented and follows fixed scripts. The code they run is the real code.
- Check against a recount. Every demo ends by comparing its answers with the same answers recounted from the events by separate code, and the tests compare the two runtimes value by value.
- Small and finished. Version 0.1 covers a stated set of features and documents its limits. It adds more only when real use asks for it.
License and Availability
The engines are not public yet. The license will be chosen before the first public release. The plan is to keep the engines proprietary, with an open-source branch under consideration. Until then the prototype is marked “all rights reserved”, and the live demos are the way to see it run.
Questions People Ask
Can I download it?
Not yet. It will be available once the license is chosen, around the first public release. The live demos run the real code, compiled for the browser.
Is the demo real?
The code is real: the compiler, the Engine, the log reducer and SQLite, built for WebAssembly. The data is simulated from fixed scripts, so a run left alone ends with the published numbers. Each demo checks its own answers against a recount at the end, and you can download its file and open it in any SQLite tool.
How is it different from a time-series database?
A time-series database stores the events and computes answers when asked. Precomputing decides the answers first and keeps them current as events arrive, inside a SQLite file you already know how to use. It is small by design, with no query language of its own and no server you must run. A time-series database is the better choice for exploring data whose questions are not known yet.
How accurate are the percentiles?
Within 1% of the exact value by default, and the policy can ask for more. Counts, sums, averages, minimums and maximums are exact. In Demo 1 every p99 came within 0.65% of the exact p99 computed from every request.
What happens to old data?
Each level of detail lives as long as the policy says: whole events for minutes or days, summaries for months or for good. Unusual events and the ready answers are never deleted. An exact stream keeps every event until its period has been closed long enough, for example 90 days for invoices that might be disputed.
Can I change a policy later?
In 0.1, a file keeps the policy it was made with, so a new policy means a new file. Letting a file gain streams and precomputes is planned for 1.0.
Does the Meter replace a billing system?
No. The Meter counts usage exactly and keeps the totals in SQLite, next to your own price tables if you like. Invoicing, payment and tax stay with a billing system, which can read the meter’s totals.
Does Logs replace Datadog or Splunk?
No. It sits between the services and the log platform, next to the services. The dashboards stay where they are and are fed by answers instead of by every line. The raw lines stay on site for 48 hours, where they can still be searched.
Is the HTTP server secure?
Not yet. In 0.1 precomputing serve has no authentication, so it belongs on localhost or behind a proxy that has it. Authentication comes with 0.2.
Get in Touch
Teams who would like to try Precomputing on their own data, and anyone with a question or a use the project should look at, are welcome to write.