Skip to main content
Introducing packages.sweber.dev
Documentation menuPerformance

Performance

Index size, build time and search time, measured by the benchmarks that run in CI.

The numbers on this page come from packages/core/test/bench.test.ts. The test builds a synthetic site of 1000 pages (3000 sections of about 60 words each) and measures the result. It runs on every pull request with fixed limits, so a change that makes the index or the search several times worse fails the build.

Measured

Measured with npx vitest run test/bench.test.ts --disableConsoleIntercept on a Linux virtual machine with 4 cores, which is about what a CI runner offers. One run, so expect some variation between runs and machines.

WhatMeasuredLimit in CI
Keyword search over 3000 sections, median1.0 ms5 ms
Keyword search over 3000 sections, 95th percentile1.6 ms20 ms
Build of 3000 sections without embeddings193 ms3000 ms
Manifest overhead per section (title, url, headings, without the text)169 bytes400 bytes
Vectors per section (384 dimensions, 8 bit)388 bytes400 bytes

The limits sit a few times above the measurements on purpose. They catch real regressions, not a slow runner.

What it means for your site

The index is the text of your pages plus about 170 bytes per section for the manifest and about 390 bytes per section for the vectors of the default model. Both compress well when your host serves gzip or Brotli.

SiteSectionsVectorsManifest overhead
100 pagesabout 400about 150 KBabout 70 KB
1000 pagesabout 3000about 1.1 MBabout 500 KB

The text of the pages comes on top of that. The table assumes 3 sections per page, as in the benchmark. Real pages differ, so check the count in the cosine build output.

What the benchmark does not measure

  • Embedding time. cosine build with a model takes seconds to minutes depending on the number of sections and the machine. --incremental embeds only changed sections.
  • Model download and start. The English model is about 23 MB (multilingual about 120 MB). It loads once, when the search field gets focus, and comes from the browser cache afterwards. Keyword results do not wait for it.
  • Semantic search time. After the model is ready, each query is embedded in the browser and compared with all vectors. This depends on the device and is not part of the CI limits.
  • Real network and slow phones. The numbers above are for desktop-class CPUs.

For the limits of the approach, see Why Cosine.