# HNSW blog benchmark evidence Benchmark date: October 6, 2026. Results are a single-host experiment, not a general database ranking. ## Environment and versions - Dell XPS 15 9520, Intel Core i7-12700H, approximately 32 GB RAM, Linux 7.0.0-34-generic, ext4 on local NVMe. - Node.js 24.21.0 for Harper, the HTTP client, and the Fastify application in front of PostgreSQL. - Published npm packages: Harper 5.3.0, including `@harperfast/hnsw` 0.4.0; Harper 5.2.15 as the previous release-line baseline. - PostgreSQL 17.11 with pgvector 0.8.7; image digest `pgvector/pgvector@sha256:ac08538c6f8b9904c33c8224c5e5706dbe760aca29db1d096972b4052c22a75d`. Shared buffers 4 GB, maintenance memory 2 GB, up to six parallel maintenance workers. HNSW `m=16`, `ef_construction=200`. Both Harper versions use M=16 and efConstruction=200; 200 is explicitly configured for Harper 5.2. - Harper 5.3.0 tag commit: `726dd1e65b1bda652c8c9db32d33de0f2a45b8f7`. - Harper 5.2.15 tag commit: `0e1da5025d0db90aa127f4437786b18da8712471`. - Six Harper workers and six Fastify workers. Servers restricted to logical CPUs `0,2,4,6,8,10`, one thread on each of six physical performance cores; client restricted to efficiency cores `12-19`. - CPU scaling maximum was already 2 GHz; powersave governor. No machine-wide governor or frequency changes made for this run. Other work remains on the machine, including sibling threads. ## Workload - Primary three-system comparison: first 20,000 SIFT1M base descriptors, 128 dimensions; first 200 independent SIFT query descriptors. - Cosine distance, top 10, no metadata filter. Ground truth recomputed for this exact subset and metric. - Sweep `ef`: 10, 20, 40, 80, 160, 320; concurrency 32; four client workers, each with eight concurrent requests; 50 passes over 200 queries per point, after 50 warm-up queries. - End-to-end localhost HTTP requests including JSON, query execution, and returning IDs. Harper's REST resource versus Fastify plus PostgreSQL/pgvector. - Throughput is specific to concurrency 32 and this client. No claim that each engine reached maximum throughput. - Warm, repeated queries; no cold-cache or out-of-memory workload, replication, concurrent writes, or embedding generation. - Larger two-system comparison: first 200,000 SIFT base descriptors, first 500 query descriptors; Harper 5.3 and pgvector only. Twenty repetitions, 10,000 requests per setting, same four-client/concurrency/core allocation. Harper 5.2 was not measured at this size. - One index build and one sweep per system and corpus size. No repeated-build confidence intervals. Recall is computed from the first pass over unique query vectors; repeated requests improve timing only. ## Validity checks and exclusions - Both databases use HNSW; PostgreSQL's cosine operator matches its index, and `EXPLAIN` must report an Index Scan. - Harper 5.3 must create a native `.hnsw` graph file and successfully complete a causal index-coverage wait after ingest. - Query errors fail the measurement rather than contributing to successful throughput. - The parallel client was verified against mock HTTP endpoints: 13 queries, three repeats, 39 requests, correctly aligned IDs and latency counts, HTTP error accounting, and fatal worker-error propagation. - The integration helper originally forced `THREADS_COUNT=1`. That pilot was stopped and discarded. A benchmark-local copy honors the requested worker count; published Harper packages remain unchanged. - PostgreSQL's CPU set is applied before loading and building, not just before querying. - Do not compare PSS for Harper with container `memory.current` for PostgreSQL. These measure different things. - Do not treat mean sampled clock multiplied by CPU time as measured hardware cycles. - Native graph file apparent size includes sparse preallocation; it is not physical disk use. - Loading differs: Harper uses REST PUTs and indexes incrementally; PostgreSQL uses COPY followed by bulk HNSW construction. Only total time to an indexed, queryable corpus is reported, with that qualification. - SIFT is normally benchmarked with Euclidean distance. These cosine results must not be compared with published Euclidean SIFT recall figures. - Int8 quantization and exact-distance reranking differ from pgvector's float32 index. Compare measured recall, not equal `ef`. ## Reproduction artifacts Original recovered harness: `/home/kzyp/dev/harper/.claude/worktrees/vector-vs-pgvector/benchmarks/vector-vs-pgvector/`. Isolated rerun: `/home/kzyp/dev/harper/.Codex/worktrees/hnsw-blog-bench/`. Run `python3 run-all.py` (20k, three systems) and `python3 run-large.py` (200k, two systems) from that worktree. The runner installs no software, uses the already installed published runtimes, creates a dedicated Docker Compose project `hnswblog20261006`, and runs Harper 5.3, pgvector, and Harper 5.2 sequentially. Results and raw logs live in `results/`. Raw memory and estimated clock counters remain in the machine-readable output for auditability but are excluded from the blog's comparative claims. ## Archived single-client pilot The first completed release comparison used 200,000 descriptors and 500 queries with one HTTP client worker, 20 repeats per point. Harper 5.3 and pgvector completed. Harper 5.2 was stopped during ingest; it has no query results for this size. These files are retained separately under `200k-single-client/` and must not be combined with the primary results. That client's throughput plateaued around 2,500 queries/sec on the faster targets. The primary comparison uses four client workers to reduce this ceiling. The larger run remains useful evidence for recall at a larger corpus, but does not establish comparative maximum throughput or a three-system 200,000-vector result. ## Fastify overhead diagnostics The existing 200,000-row PostgreSQL index was reused, with the same `m=16`, `ef_construction=200`, cosine SQL, `ef_search=160` on every pooled connection, and top 10. PostgreSQL 17.11 / pgvector 0.8.7 and the row count were rechecked. Fastify was 5.12.4 and the Node PostgreSQL driver (`pg`) 8.23.0. Each path used 500 unique queries, one warm-up pass on the actual client-owned connections, and 20 timed passes (10,000 requests). Four client threads on E-cores issued 32 requests concurrently. Each HTTP adapter used six clustered workers on the same six P-cores as PostgreSQL, with a maximum of 24 pooled connections per worker. Three rounds were run in rotated or reversed order; reported QPS is the arithmetic mean of the three round throughputs. | Diagnostic path | Mean QPS | Range across three rounds | | ------------------------------------------------ | -------: | ------------------------: | | Fastify + pgvector, framework comparison | 1,729 | 1,714–1,737 | | Plain Node HTTP + pgvector, framework comparison | 1,775 | 1,750–1,800 | | Fastify + pgvector, adapter comparison | 1,722 | 1,648–1,764 | | Direct `pg` driver, adapter comparison | 2,362 | 2,341–2,381 | | Fastify control without PostgreSQL | 6,107 | 5,773–6,545 | All database paths measured the same 99.82% recall. The plain HTTP implementation uses the same request payload, vector-literal conversion, parameterized cosine SQL, response IDs, `pg` driver, pooling limits, worker count and CPU placement as Fastify. A minimal handler is a diagnostic comparator, not a production server. Mean plain-HTTP throughput was 2.66% higher; paired gains were approximately 2.29%, 0.71%, and 5.03%. This is evidence of a small framework contribution in this setting, not a universal Fastify overhead percentage or a confidence interval. The control parses the full 128-dimensional JSON request, constructs the PostgreSQL vector literal, and serializes ten representative six-digit IDs. Its throughput establishes HTTP headroom; it is not a latency that can be subtracted from full queries under concurrent load. Instrumented HTTP requests record elapsed wall time from Fastify `onRequest` to `onSend` and the awaited `pool.query` duration. The difference includes body parsing, vector conversion, result mapping, serialization and scheduling delays outside that await. It excludes earlier routing/dispatch and later socket I/O; it is not measured Fastify CPU time. The framework comparison averaged 0.134 ms outside the PostgreSQL await in a 16.823 ms server window (0.80%). The PostgreSQL await includes pool/client/transport/backend and event-loop delays, rather than isolating index traversal. Hook metrics are reset only after all client workers finish warm-up, with acknowledgments from all six server workers; each timed run requires exactly 10,000 hook samples. Fresh output directories prevent stale metric files from satisfying later runs' acknowledgment checks. Direct access uses four `pg` pools of eight connections on the client E-cores. It removes the HTTP request/response and places PostgreSQL-client work outside the server's P-core budget. Its 37.1% throughput improvement (HTTP 27.1% lower than direct) is a whole-adapter and resource-placement comparison, not a pure Fastify attribution. The direct diagnostic does not replace the main HTTP comparison with Harper, which was measured through its HTTP endpoint. Research source: `run-overhead.mts`, `overheadDriver.mts`, and `pgvector-app/overhead-{server,worker}.mjs` under the supplied harness. With the existing 200k PostgreSQL data up and pinned, run: ```sh OVERHEAD_PHASE=adapter taskset -c 12-19 node benchmarks/vector-vs-pgvector/run-overhead.mts OVERHEAD_PHASE=framework taskset -c 12-19 node benchmarks/vector-vs-pgvector/run-overhead.mts ``` Detailed samples and per-round summaries are under `benchmark-results/fastify-overhead/` (adapter phase) and `benchmark-results/fastify-framework-06uSzp/` (framework phase). ## Standalone @harperfast/hnsw package (0.4.0) These measurements invoke the published Node package directly, without starting Harper, PostgreSQL, or an HTTP server. The Linux x64 glibc native prebuild was loaded from the installed Harper 5.3 runtime, with no Harper code on the search path. Its SHA-256 is `5b1fa4e068ff127a8588103fe307905f78ce5892f4162507b7ce6d04b997c299`. The binary matches the npm-published 0.4.0 platform tarball, whose SHA-512 integrity was checked independently. `provenance.json` records that check, the package license and Node requirement, and validation that every ground-truth row contains ten distinct neighbors. The package is Apache-2.0 licensed; source/API references: [repository](https://github.com/HarperFast/hnsw), [0.4.0 API](https://github.com/HarperFast/hnsw/blob/v0.4.0/index.d.ts), [builder defaults](https://github.com/HarperFast/hnsw/blob/v0.4.0/src/insert.rs). ### Construction and scope - Same 20,000/200,000 SIFT base prefixes, 128 dimensions, cosine ground truth and 200/500 unique query prefixes as the full-stack benchmarks. - One fresh graph per precision and size: int8 and int16, layer-0 capacity 64, maximum nodes equal to corpus size plus 1,024, no embedded keys (`keyCap=0`). `Plane.create` truncates an existing path; the harness rejects an existing output directory. - Native standalone builder defaults: M=16, efConstruction=200, optimizeRouting=0.5. Sequentially awaited chunks of 10,000 vectors, six native insertion threads. Parallel allocation is nondeterministic: the returned node IDs are inverted to original corpus rows for recall scoring. A rejected record, sentinel/out-of-range ID, reused allocation, or short search result fails the run. - Insertion time starts after file creation and corpus loading. It includes flattening each chunk, Node/native input transfer and native graph construction. Final `flushAsync` is timed separately after every insertion promise settles. - Both graph codecs are quantized, and returned top-10 results are not reranked against original vectors. Int8 uses the asymmetric float-query/int8-stored distance path; int16 also quantizes the query. Standalone int16 is not the Harper default. - Measurements include the Node/native boundary, result allocations and, for async calls, promise/thread-pool scheduling. They do not isolate the bare Rust kernel. | Corpus | Precision | Insert time | Final flush | Allocated graph-file bytes | | ------- | --------- | ----------: | ----------: | -------------------------: | | 20,000 | int8 | 0.803 s | 0.0076 s | 11,747,328 | | 20,000 | int16 | 0.827 s | 0.0096 s | 14,348,288 | | 200,000 | int8 | 13.597 s | 0.0599 s | 117,432,320 | | 200,000 | int16 | 13.596 s | 0.0632 s | 143,437,824 | Allocated bytes are `stat.blocks * 512` after the flush, not apparent sparse capacity or process memory. Corpus arrays remain in the process; these files fit in RAM. Each construction is one sample, not a build-time scaling study. ### Query protocol and validation The process is pinned to logical CPUs `0,2,4,6,8,10`, with `UV_THREADPOOL_SIZE=6` set before Node starts. The event loop and native work share these six physical P-cores; the synchronous mode uses only one calling thread. No separate HTTP clients or client E-core budget are involved. Each graph sweeps `ef=40,80,160,320` and top 10 in two modes: `searchSync` from one caller and six concurrent asynchronous `search` calls. Three timing rounds run 10,000 searches per setting. Mode order is reversed in round two; search-budget order is reversed or rotated across rounds. Each setting warms with one pass over unique queries and calls exposed GC before its timed window. Performance timers, result-length checks and a simple checksum are inside that window; recall scoring and node-to-row mapping are outside. QPS is the arithmetic mean of the three round throughputs. Reported latency is the arithmetic mean of the three round medians (or p99s), not a pooled percentile. Chart bands are the observed minimum and maximum async QPS across three rounds, not confidence intervals. On this shared host async variation is substantial; these repetitions do not independently rebuild the graph or establish a statistical precision-performance difference. After the timed run ends, a fresh process reopens every graph and verifies its precision, node count and durability watermark. At every budget, all unique queries return identical node IDs and distances through synchronous and asynchronous search, and recall is independently rescored against the same full-precision subset ground truth. The sample keyed usage example was also executed, reopened in another process, and verified to return `article:1`, then `article:2`; its existing-file guard was checked. Detailed data: `benchmark-results/standalone-hnsw/`, including one `results.json`, `row-to-id.json`, and `verification.json` per graph, raw logs, `summary.json`, `summary.csv`, and the example verification receipt. Graph binaries and the SIFT corpus are omitted from the download. Harness: `benchmarks/standalone-hnsw/{benchmark.mjs,verify.mjs,example.cjs,run-standalone.py,summarize.py,example-smoke.py}`. On Linux, install the pinned package in an isolated working directory, supply the SIFT `sift_base.fvecs` and `sift_query.fvecs` files. Copy the supplied cosine subset caches from `benchmark-results/standalone-hnsw/ground-truth/` alongside them, or regenerate these caches with `benchmarks/vector-vs-pgvector/dataset.mts`. The cache hashes and SIFT-prefix hashes in each result identify the measured data. Adjust the CPU set for the machine: ```sh npm install @harperfast/hnsw@0.4.0 export HNSW_PACKAGE_ROOT="$PWD/node_modules/@harperfast/hnsw" export VECTOR_DATA_DIR="/path/to/sift" export BENCH_CPU_SET="0,2,4,6,8,10" python3 benchmarks/standalone-hnsw/run-standalone.py ``` The explicit package path prevents an ancestor checkout's different installed version from shadowing the intended runtime. The runner prints a fresh output directory and leaves its graph files available for verification: ```sh UV_THREADPOOL_SIZE=6 taskset -c "$BENCH_CPU_SET" node benchmarks/standalone-hnsw/verify.mjs /path/to/output/20000-int8 /path/to/output/20000-int16 /path/to/output/200000-int8 /path/to/output/200000-int16 python3 benchmarks/standalone-hnsw/summarize.py /path/to/output python3 benchmarks/standalone-hnsw/example-smoke.py ``` Standalone construction differs from Harper's mirrored graph layout and record-key storage, and these searches omit Harper's exact reranking. Their concurrency and client placement also differ from the HTTP tests. Do not divide standalone QPS by HTTP QPS to attribute overhead, or compare raw int8 recall with reranked database recall as if both were the same workload. No standalone pgvector comparison was run.