Skip to content

Benchmarks ​

Benchmarking Your App ​

Use autocannon to load-test your KickJS application:

bash
# Install autocannon
pnpm add -D autocannon

# Start your app
kick dev

# In another terminal — 100 concurrent connections for 30 seconds
npx autocannon http://localhost:3000/api/v1/users -c 100 -d 30

# Quick test — 50 connections, 10 seconds
npx autocannon http://localhost:3000/api/v1/users -c 50 -d 10

Metrics to Watch ​

MetricWhat it tells you
Req/sThroughput — requests handled per second
p50Median latency — half of requests are faster than this
p97.5Tail latency — 97.5% of requests are faster
p99Worst-case latency — only 1% are slower
ErrorsFailed requests under load

Tips for Accurate Results ​

  • Run with NODE_ENV=production — Express disables debugging features
  • Close other applications to reduce CPU contention
  • Run multiple times and compare median results
  • Use a consistent machine for tracking regressions
  • Don't compare across different hardware

Framework Reference Numbers ​

The KickJS monorepo includes a benchmark suite that measures framework overhead. These numbers are from the repo's internal tests, not something you run in your app.

Results vary by machine. Reference from a typical dev machine (Node 24, Linux):

EndpointReq/sp50p97.5p99
Minimal (text)~13,00035ms50ms59ms
JSON object~12,70038ms50ms57ms
JSON array (50 items)~8,60056ms80ms88ms
Middleware stack (3 layers)~12,60038ms51ms58ms
POST + body parse~9,20053ms64ms76ms

Key takeaways:

  • Middleware overhead is minimal (~5% for 3 layers)
  • JSON serialization is the main cost for large responses
  • KickJS's decorator/DI layer adds negligible overhead over raw Express 5

Boot time ​

Time from starting the process to the first HTTP response, for a production build. Measured on a fresh kick new my-api --template rest app (KickJS 8.7, Express, Node 24, Linux), seven runs:

AppMedianRange
kick new --template rest182ms177–191ms
A larger app (auth, DB, 6 modules)~350ms337–371ms

To measure your own app, run kick build, then this script from the project root:

js
// boot.mjs — node boot.mjs
import { spawn } from 'node:child_process'

const runs = []
for (let i = 0; i < 7; i++) {
  const start = performance.now()
  const app = spawn(process.execPath, ['dist/index.js'], {
    env: { ...process.env, PORT: '3311', NODE_ENV: 'production' },
    stdio: 'ignore',
  })
  for (;;) {
    try {
      await fetch('http://127.0.0.1:3311/')
      break
    } catch {
      await new Promise((r) => setTimeout(r, 5))
    }
  }
  runs.push(Math.round(performance.now() - start))
  app.kill()
  await new Promise((r) => app.once('exit', r))
}
console.log(runs.join(' '), 'median', runs.toSorted((a, b) => a - b)[3])

Any response counts, including a 404, so the route doesn't matter. Database connections opened at startup are part of the time.

WebSockets ​

@forinda/kickjs-ws ships a load test in packages/ws/bench/. It starts one or more server instances and several client processes, so clients never share the server's event loop. Each run has two phases:

  • echo — every connection sends at a fixed rate and the server replies; measures round-trip latency.
  • fan-out — one connection broadcasts to a room every other connection has joined; measures delivery latency and reach (deliveries received ÷ deliveries expected).
bash
pnpm --filter @forinda/kickjs-ws build
cd packages/ws

pnpm bench                                             # 1 instance, 2,000 connections
CONNS=10000 WORKERS=8 ECHO_RATE=0 FAN_RATE=5 pnpm bench
INSTANCES=2 CONNS=10000 WORKERS=8 pnpm bench
VariableDefaultMeaning
INSTANCES1Server processes; connections are spread evenly
CONNS2000Total connections
WORKERS4Client processes
SECONDS10Length of the measured phase
ECHO_RATE1Messages per second, per connection
FAN_RATE20Broadcasts per second from the single sender
BASE_PORT4600First server port
REDIS_URLunsetWhen set, instances relay room broadcasts through the Redis broker

The output includes client CPU. If it approaches 100%, the clients are the bottleneck — raise WORKERS before reading the server numbers.

Reference numbers ​

Node 24.20, Linux, 12 cores. Each server process is one event loop, so it saturates at ~100% CPU. Client CPU stayed under 25% in every run.

Scenario (1 instance)Throughputp50p99Memory
10,000 connections, echo9,000 msg/s0.17ms14ms299MB
20,000 connections, echo18,000 msg/s1ms128ms428MB
20,000 connections, echo (overloaded)38,000 msg/s1.0s2.1s481MB
10,000-member room, fan-out50,000 frames/s142ms255ms269MB
10,000-member room, fan-out (overloaded)100,000 frames/s1.8s4.1s270MB

No run lost a message; overloaded runs delivered late. Once a process is past its ceiling, latency grows with the backlog rather than frames being dropped.

Running more than one instance ​

Set REDIS_URL to run the instances with the Redis broker:

bash
docker run -d --rm -p 6379:6379 redis
REDIS_URL=redis://127.0.0.1:6379 INSTANCES=2 CONNS=10000 WORKERS=8 ECHO_RATE=0 FAN_RATE=5 pnpm bench

One 10,000-member room, spread evenly across instances, local Redis. These rows were measured in one session, so they compare with each other rather than with the table above.

ScenarioFrames/sReachp50p99
1 instance50,0001.0125ms178ms
1 instance + Redis50,0001.0141ms184ms
2 instances, no broker50,0000.544ms92ms
2 instances + Redis50,0001.085ms201ms
1 instance (overloaded)100,0001.01.8s2.7s
4 instances + Redis100,0001.058ms203ms
  • Without a broker, each instance reaches only its own sockets, so half the room never gets the message.
  • The broker costs one Redis round trip per broadcast, not per recipient — ~16ms p50 in the single-instance runs here.
  • A load that swamps one process is handled by four with room to spare: server CPU stayed under 70% per instance.

Released under the MIT License. Built with TypeScript — runs on Express, Fastify, or h3.