SCUA

News

SCUA 0.35.0 answers twice as many requests when the handler does real work

October 10, 2026

In SCUA 0.35.0, an http.serve handler runs as compiled code from its first line to its response, and the loops inside it compile to shorter code. A handler that runs a 20,000-step loop for every request answered 46,804 requests a second on one thread of an Intel i7-10700 running Linux, where 0.34.1 answered 23,361 on the same machine. Every request now runs in a partition of its own, so a handler can wait on a database or another service while the server answers other requests, and state that lasts between requests lives in a partition you pass in. With the threaded scheduler, which this release lets you switch on, the same server answered 146,210 requests a second on four threads. The release also adds actors.pool for spreading work over several partitions, a health check the server answers itself, bytes.inflate for gzip, zlib and zip data, and bulk operations on typed buffers. Servers and code that passes typed data across function boundaries may need edits; there is a short guide to upgrading at the end.

#A handler runs compiled from start to finish

Here is a handler that does some arithmetic for each request and answers with JSON:

import http
import json
import bytes

fn score(n)
  let s = 0
  for i in 0:n do s = s + i % 7 end
  return s
end

fn handle(req)
  if req.path == "/score" then
    let body = json.encode({ steps = 20000, score = score(20000) })
    return { status = 200, headers = { "content-type" = "application/json" }, body = bytes.from_string(body) }
  end
  return { status = 404, body = b"not found" }
end

http.serve(handle, { bind = "127.0.0.1:8080", health = "/healthz" })?
$ scua --allow-serve=127.0.0.1:8080 work.scua
$ curl 127.0.0.1:8080/score
{"steps":20000,"score":59997}

In 0.34, compiled code went back to the interpreter at several points a handler like this reaches on every request: a call to a built-in such as json.encode or bytes.from_string, the start of a for loop, a tell to a partition, a string, bytes or table literal, and the return from the handler. In 0.35 none of these leave compiled code, so a typical handler runs compiled the whole way through. The loop itself is shorter too (see below).

These figures come from one sitting on an Intel i7-10700 running Linux. The server ran on four cores and their second hardware threads, and wrk sent load from the other four over 256 connections, three 10-second runs per cell after a 2-second warm-up. Each figure is the median, in requests a second. The work route runs the loop above, s = s + i % 7 for 20,000 steps, written the same way in every runtime; plaintext returns a fixed 13-byte body.

work, one thread work, four threads plaintext, one thread plaintext, four threads
SCUA 0.35 46,804 146,210 118,631 322,541
uWebSockets.js 20.71 (on Node 24) 33,921 113,396 131,080 408,157
Bun 1.4.2 34,020 106,494 102,729 274,463
Node 24 22,382 — 40,048 —

In an earlier run on this machine with the same setup, SCUA 0.34.1 answered 23,361 work requests a second and 111,786 plaintext ones; it served from one thread. The SCUA figures are from a build taken the day before release. Four threads means one process: SCUA with the threaded scheduler at --workers=4, and uWebSockets.js and Bun each running four threads. Node serves from one thread in one process.

#Each request runs on its own

Every request runs as if in a partition of its own. A handler starts from the same place each time, a module's state starts as declared, and nothing one request changes is there for the next. That is what lets the server answer requests side by side without one seeing another half-way through, and what lets the threaded scheduler run them on several cores.

It also means a handler can wait. ask … timeout, wait_for, select, http.get, fs.read and net.read wait for that request alone, and the server answers other requests meanwhile. State that lasts from one request to the next lives in a partition, passed to the handler in context:

import http
import bytes

const greeting = "hello"

partition Hits
  state n = 0
  ask Next()
    n = n + 1
    return n
  end
end

fn handle(req, ctx)
  let n = ask ctx.hits.Next() timeout 1s
  return { status = 200, body = bytes.from_string(`{greeting} #{n}`) }
end

http.serve(handle, { bind = "127.0.0.1:8080", context = { hits = Hits() } })?
$ scua --allow-serve=127.0.0.1:8080 hits.scua
$ curl 127.0.0.1:8080/
hello #1
$ curl 127.0.0.1:8080/
hello #2

--max-mem now applies to each request: one that goes over the limit gets a 500, and the others carry on.

The threaded scheduler, which spreads requests and partitions across cores, is in this release as an option. Start the program with SCUA_SCHEDULER=threaded and a worker count:

$ SCUA_SCHEDULER=threaded scua --workers=4 --allow-serve=127.0.0.1:8080 work.scua

A release build prints a line saying the mode is experimental. It becomes the default in a later release.

#Compiled loops are faster

A counted for loop now checks its operation budget once per block of iterations instead of on every one, and % and // by a constant compile to shorter code, shorter still when the value can't be negative. A while loop that steps by one now compiles to nearly the same code as the for loop that does the same thing. --max-ops still stops a program at the same step and line:

fn checksum(n)
  let s = 0
  for i in 0:n do s = s + i % 7 end
  return s
end

fn checksum_while(n)
  let s = 0
  let i = 0
  while i < n do
    s = s + i % 7
    i = i + 1
  end
  return s
end

print(checksum(20_000), checksum_while(20_000))
print(checksum(50_000_000), checksum_while(50_000_000))
$ scua loops.scua
59997 59997
149999997 149999997
$ scua --max-ops=1000000 loops.scua
59997 59997
scua: loops.scua:3: this turn used its whole CPU budget of 1000000 steps, set by `--max-ops`. Raise it, or remove the budget with `--max-ops=none`

On the Intel i7-10700, the language benchmarks moved as below. Each figure is the whole process's wall-clock time, the best of five runs, with SCUA's JIT on as it is by default:

benchmark SCUA 0.34 SCUA 0.35 LuaJIT 2.1
sum (add up a counter) 55 ms 13 ms 44 ms
orders (% by constants, totals kept with if) 15 ms 11 ms 45 ms
matmul 17 ms 12 ms 15 ms
polyfield 16 ms 9 ms 9 ms
fields 24 ms 19 ms 18 ms
recarray 42 ms 35 ms 42 ms

Across each tier of the suite, the geometric mean of SCUA's time as a fraction of LuaJIT's went from 1.04 to 0.72 on the micro benchmarks, 0.73 to 0.67 on the realistic ones, and 0.96 to 0.87 on the recursion and record shapes. LuaJIT is still ahead on some rows, among them fib (79 ms against SCUA's 119 ms), physics and tail.

#One address for a pool of partitions

A partition handles one message at a time, which keeps its state safe and is a limit when each message waits on something slow, such as a database query. actors.pool(n, make) starts n partitions by calling make and returns one send-only address for all of them. An ask sent to it goes to a free member, so n asks are answered at once, and the reply goes straight back to whoever asked:

import actors

partition Db
  ask Query(id)
    select
      wait_for(Never) -> nil
      wait(100ms) -> nil    -- stands in for a database round trip
    end
    return `row {id}`
  end
end

partition Report
  on Run(label, db)
    let start = now()
    let rows = map_all(range(0, 8), fn(i) return ask db.Query(i) timeout 2s end)
    let ms = (now() - start) // 100 * 100
    print(`{label}: {len(rows)} rows, {rows[0]} to {rows[7]}, about {ms} ms`)
  end
end

tell Report().Run("one Db", Db())
tell Report().Run("pool of 4", actors.pool(4, fn() return Db() end))
$ scua pool.scua
pool of 4: 8 rows, row 0 to row 7, about 200 ms
one Db: 8 rows, row 0 to row 7, about 800 ms

When every member is busy, messages wait in the pool's queue, and each ask keeps its own timeout while it waits. A member that faults is replaced by calling make again, and the ask it was answering gets Error({ kind = "fault" }). Since the address is send-only, it is safe to hand to a web server's handlers: http.serve(handle, { context = { db = actors.pool(8, fn() return Db(URL) end) } }).

#A health check the server answers itself

Give http.serve a health path and the server answers it without running your handler, so a load balancer's check gets through while every handler is busy. The server above has one:

$ curl -i 127.0.0.1:8080/healthz
HTTP/1.1 200 OK
content-length: 3
connection: keep-alive
date: Sat, 10 Oct 2026 12:56:57 GMT
cache-control: no-store

ok

It answers 200 while the server is serving and 503 once a stop has begun, so the balancer stops sending traffic while the server drains. With health_sheds = true it also answers 503 for a short time after the server has turned requests away because it was overloaded.

#Decompress gzip, zlib and zip data

bytes.inflate decompresses DEFLATE data. It reads a zlib stream by default, the format inside PNG images; { format = "gzip" } reads a .gz file and { format = "raw" } a zip entry. It returns a Result, and { max_bytes = N } caps the output as it grows, for input you didn't make yourself:

import bytes
import fs
import str

let packed = fs.read("scores.csv.gz")?
let text = bytes.inflate(packed, { format = "gzip" })?.to_string()?
let lines = str.split(text, "\n")
print(`{len(packed)} bytes in, {len(text)} out, {len(lines) - 2} rows`)
print(lines[1])

match bytes.inflate(packed, { format = "gzip", max_bytes = 4096 })
  Ok(_) -> print("fits")
  Error(e) -> print(e)
end
$ scua --allow-fs=. scores.scua
7855 bytes in, 34292 out, 2000 rows
1,player1,37
bytes.inflate: the output is over the 4096 bytes `max_bytes` allows

#Typed buffers in bulk

Four additions work on whole typed arrays at once. fill(n, v, kind) builds a typed array directly at its final size, without a plain array first. copy_range(dst, at, src, start, end) copies a run of elements in one block between arrays of the same kind, overlapping ranges included, which is the row copy of an image or an audio buffer. bytes.to_array(b, "u8") turns a buffer into a { u8 } in one copy, and elem_kind(xs) tells you a typed array's element kind without converting anything:

import bytes

-- a 4x3 greyscale image, one byte a pixel, and a 2-pixel sprite
let w = 4
let image = fill(w * 3, 0, "u8")
let sprite = bytes.from_hex("ff80")?.to_array("u8")

-- blit the sprite into each row, one block copy a row
for row in 0:3 do
  image.copy_range(row * w + row % 3, sprite)
end
print(elem_kind(image), len(image))
for row in 0:3 do print(image[row * w : row * w + w]) end

let samples = fill(1_000_000, 0.5, "f32")
print(elem_kind(samples), samples)
$ scua buffers.scua
u8 12
[255, 128, 0, 0]
[0, 255, 128, 0]
[0, 0, 255, 128]
f32 [0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, … 999984 more]

The last line shows another change: a typed array longer than 256 elements prints its first 16 elements and a count of the rest, so printing a record that holds an image no longer writes out every pixel. Plain arrays still print in full.

#Smaller additions

  • sys.parse_args fills a list field with every positional argument left over, so a files: { string } field makes convert --quality 80 a.png b.png c.png work. { unknown = "refuse" } makes an undeclared key in a JSON body an error.

  • fs.write(path, data, { exclusive = true }) and fs.open(path, { mode = "create" }) fail if the file already exists, checking and creating in one step, so two programs can't both write it. { sync = true } and fs.sync(f) wait until the data is on the disk.

    $ scua --allow-fs=. report.scua
    Ok(nil)
    Error(report.txt: already exists — `{ exclusive = true }` writes only a file that is not there yet)
    first
  • hash.crc32(data, start, end) checksums part of a buffer without copying it out.

  • fs errors name the path, as in the line above.

  • --jit-stats counts the work done on every worker thread, which is where a server's requests run.

  • bytes.join and bytes.from_array hold their result once while they build it.

#Upgrading to 0.35

Four groups of changes can need edits.

#http.serve handlers run on their own

A handler can't reach the main script's lets, and the server refuses to start until none does. A table the handler only reads becomes a const; anything it changes moves into a partition passed in context, as in the Hits example above. The message lists each variable and what to do:

import http
import bytes

let greeting = "hello"
let hits = 0

fn handle(req)
  hits = hits + 1
  return { status = 200, body = bytes.from_string(`{greeting} #{hits}`) }
end

http.serve(handle, { bind = "127.0.0.1:8080" })?
$ scua --allow-serve=127.0.0.1:8080 hits.scua
scua: hits.scua:12: the handler you pass to `http.serve`, `handle`, reaches 2 variables of the main script. A request runs in a partition of its own, on any core, so it can't reach them:
  - `greeting` (line 4) is read and never changed: make it a `const`
  - `hits` (line 5) is changed by `handle` (line 8): keep it in an actor, e.g. `partition Hits … end`, passed in `context`; see "Keep state in a server"

A module's state starts as declared on every request, so a value a module kept across requests also moves into a partition. --max-mem now bounds each request, and one that goes over it gets a 500 on its own.

#A typed parameter, return or let needs every field of a record

A value that is missing a field with a default is refused at the boundary. Convert loose data, such as decoded JSON, with as(value, Record), which fills the defaults:

import json

record Settings { volume: int = 5, muted: bool = false }

fn apply(s: Settings) return `volume {s.volume}, muted {s.muted}` end

let loaded = json.decode("{\"volume\": 8}")?
print(apply(as(loaded, Settings)))
print(apply(loaded))
$ scua settings.scua
volume 8, muted false
scua: settings.scua:5: field 'muted' of Settings is missing; use `as(value, Settings)` to fill its default
  at apply (settings.scua:5)
  at main (settings.scua:9)

A record literal written straight into a typed let or return still fills its defaults, and a literal passed to a typed parameter with a field missing is a compile error.

json.encode and toml.encode write an enum by its name

json.encode({ kind = Kind.UnsupportedFormat }) gives {"kind":"UnsupportedFormat"}; it used to give the number. A record field declared with the enum still reads the number back, so data you wrote before still loads, but a program on the other end that reads the number needs to read the name. A flags set is still written as its bitmask.

#Typed buffers are checked at function boundaries

A parameter or return declared { u8 }, { f32 } or another element type now checks the value it is given:

fn brighten(p: { u8 }) return p end

let levels: { u16 } = [10, 300]
brighten(levels)
$ scua levels.scua
scua: levels.scua:1: parameter 'p' of 'brighten' is a typed {u8} array: element 1 is 300, and it holds whole numbers from 0 to 255
  at brighten (levels.scua:1)
  at main (levels.scua:4)

Passing nil to one that isn't declared optional ({ u8 }?) faults too. A typed annotation no longer converts an array someone else holds: an array that already has the declared element type is used as it is, and any other array whose values fit is copied. A function that fills a buffer you pass it therefore needs a buffer of the declared type, made with fill(n, 0, "u8") for instance; given a plain array, it fills a copy you won't see. An { f32 } also refuses a number beyond the range of an f32 (about ±3.4e38), where it used to store infinity; inf and nan written as such still store.

#Everything else

Rarer changes, such as fs.rename refusing to replace a read-only file, are in the changelog.

#What's next

The threaded scheduler becomes the default, so a server uses every core without a setting. The next work on serving is a published, reproducible benchmark with the other servers included, and a benchmark of a realistic application with a database, templates and TLS over a real network.

The full changelog has everything in 0.35.0.