In SCUA 0.35.0, an http.serve handler runs as compiled
code from its first line to its response, and the loops inside it
compile to shorter code. A handler that runs a 20,000-step loop for
every request answered 46,804 requests a second on one thread of an
Intel i7-10700 running Linux, where 0.34.1 answered 23,361 on the same
machine. Every request now runs in a partition of its own, so a handler
can wait on a database or another service while the server answers other
requests, and state that lasts between requests lives in a partition you
pass in. With the threaded scheduler, which this release lets you switch
on, the same server answered 146,210 requests a second on four threads.
The release also adds actors.pool for spreading work over
several partitions, a health check the server answers itself,
bytes.inflate for gzip, zlib and zip data, and bulk
operations on typed buffers. Servers and code that passes typed data
across function boundaries may need edits; there is a short guide to
upgrading at the end.
#A handler runs compiled from start to finish
Here is a handler that does some arithmetic for each request and answers with JSON:
import http
import json
import bytes
fn score(n)
let s = 0
for i in 0:n do s = s + i % 7 end
return s
end
fn handle(req)
if req.path == "/score" then
let body = json.encode({ steps = 20000, score = score(20000) })
return { status = 200, headers = { "content-type" = "application/json" }, body = bytes.from_string(body) }
end
return { status = 404, body = b"not found" }
end
http.serve(handle, { bind = "127.0.0.1:8080", health = "/healthz" })?
$ scua --allow-serve=127.0.0.1:8080 work.scua
$ curl 127.0.0.1:8080/score
{"steps":20000,"score":59997}
In 0.34, compiled code went back to the interpreter at several points
a handler like this reaches on every request: a call to a built-in such
as json.encode or bytes.from_string, the start
of a for loop, a tell to a partition, a
string, bytes or table literal, and the return from the handler. In 0.35
none of these leave compiled code, so a typical handler runs compiled
the whole way through. The loop itself is shorter too (see below).
These figures come from one sitting on an Intel i7-10700 running
Linux. The server ran on four cores and their second hardware threads,
and wrk sent load from the other four over 256 connections,
three 10-second runs per cell after a 2-second warm-up. Each figure is
the median, in requests a second. The work route runs the
loop above, s = s + i % 7 for 20,000 steps, written the
same way in every runtime; plaintext returns a fixed
13-byte body.
work, one thread |
work, four threads |
plaintext, one thread |
plaintext, four threads |
|
|---|---|---|---|---|
| SCUA 0.35 | 46,804 | 146,210 | 118,631 | 322,541 |
| uWebSockets.js 20.71 (on Node 24) | 33,921 | 113,396 | 131,080 | 408,157 |
| Bun 1.4.2 | 34,020 | 106,494 | 102,729 | 274,463 |
| Node 24 | 22,382 | — | 40,048 | — |
In an earlier run on this machine with the same setup, SCUA 0.34.1
answered 23,361 work requests a second and 111,786
plaintext ones; it served from one thread. The SCUA figures
are from a build taken the day before release. Four threads means one
process: SCUA with the threaded scheduler at --workers=4,
and uWebSockets.js and Bun each running four threads. Node serves from
one thread in one process.
#Each request runs on its own
Every request runs as if in a partition of its own. A handler starts
from the same place each time, a module's state starts as
declared, and nothing one request changes is there for the next. That is
what lets the server answer requests side by side without one seeing
another half-way through, and what lets the threaded scheduler run them
on several cores.
It also means a handler can wait. ask … timeout,
wait_for, select, http.get,
fs.read and net.read wait for that request
alone, and the server answers other requests meanwhile. State that lasts
from one request to the next lives in a partition, passed to the handler
in context:
import http
import bytes
const greeting = "hello"
partition Hits
state n = 0
ask Next()
n = n + 1
return n
end
end
fn handle(req, ctx)
let n = ask ctx.hits.Next() timeout 1s
return { status = 200, body = bytes.from_string(`{greeting} #{n}`) }
end
http.serve(handle, { bind = "127.0.0.1:8080", context = { hits = Hits() } })?
$ scua --allow-serve=127.0.0.1:8080 hits.scua
$ curl 127.0.0.1:8080/
hello #1
$ curl 127.0.0.1:8080/
hello #2
--max-mem now applies to each request: one that goes
over the limit gets a 500, and the others carry on.
The threaded scheduler, which spreads requests and partitions across
cores, is in this release as an option. Start the program with
SCUA_SCHEDULER=threaded and a worker count:
$ SCUA_SCHEDULER=threaded scua --workers=4 --allow-serve=127.0.0.1:8080 work.scua
A release build prints a line saying the mode is experimental. It becomes the default in a later release.
#Compiled loops are faster
A counted for loop now checks its operation budget once
per block of iterations instead of on every one, and % and
// by a constant compile to shorter code, shorter still
when the value can't be negative. A while loop that steps
by one now compiles to nearly the same code as the for loop
that does the same thing. --max-ops still stops a program
at the same step and line:
fn checksum(n)
let s = 0
for i in 0:n do s = s + i % 7 end
return s
end
fn checksum_while(n)
let s = 0
let i = 0
while i < n do
s = s + i % 7
i = i + 1
end
return s
end
print(checksum(20_000), checksum_while(20_000))
print(checksum(50_000_000), checksum_while(50_000_000))
$ scua loops.scua
59997 59997
149999997 149999997
$ scua --max-ops=1000000 loops.scua
59997 59997
scua: loops.scua:3: this turn used its whole CPU budget of 1000000 steps, set by `--max-ops`. Raise it, or remove the budget with `--max-ops=none`
On the Intel i7-10700, the language benchmarks moved as below. Each figure is the whole process's wall-clock time, the best of five runs, with SCUA's JIT on as it is by default:
| benchmark | SCUA 0.34 | SCUA 0.35 | LuaJIT 2.1 |
|---|---|---|---|
sum (add up a counter) |
55 ms | 13 ms | 44 ms |
orders (% by constants, totals kept with
if) |
15 ms | 11 ms | 45 ms |
matmul |
17 ms | 12 ms | 15 ms |
polyfield |
16 ms | 9 ms | 9 ms |
fields |
24 ms | 19 ms | 18 ms |
recarray |
42 ms | 35 ms | 42 ms |
Across each tier of the suite, the geometric mean of SCUA's time as a
fraction of LuaJIT's went from 1.04 to 0.72 on the micro benchmarks,
0.73 to 0.67 on the realistic ones, and 0.96 to 0.87 on the recursion
and record shapes. LuaJIT is still ahead on some rows, among them
fib (79 ms against SCUA's 119 ms), physics and
tail.
#One address for a pool of partitions
A partition handles one message at a time, which keeps its state safe
and is a limit when each message waits on something slow, such as a
database query. actors.pool(n, make) starts n
partitions by calling make and returns one send-only
address for all of them. An ask sent to it goes to a free
member, so n asks are answered at once, and the reply goes
straight back to whoever asked:
import actors
partition Db
ask Query(id)
select
wait_for(Never) -> nil
wait(100ms) -> nil -- stands in for a database round trip
end
return `row {id}`
end
end
partition Report
on Run(label, db)
let start = now()
let rows = map_all(range(0, 8), fn(i) return ask db.Query(i) timeout 2s end)
let ms = (now() - start) // 100 * 100
print(`{label}: {len(rows)} rows, {rows[0]} to {rows[7]}, about {ms} ms`)
end
end
tell Report().Run("one Db", Db())
tell Report().Run("pool of 4", actors.pool(4, fn() return Db() end))
$ scua pool.scua
pool of 4: 8 rows, row 0 to row 7, about 200 ms
one Db: 8 rows, row 0 to row 7, about 800 ms
When every member is busy, messages wait in the pool's queue, and
each ask keeps its own timeout while it waits. A member
that faults is replaced by calling make again, and the
ask it was answering gets
Error({ kind = "fault" }). Since the address is send-only,
it is safe to hand to a web server's handlers:
http.serve(handle, { context = { db = actors.pool(8, fn() return Db(URL) end) } }).
#A health check the server answers itself
Give http.serve a health path and the
server answers it without running your handler, so a load balancer's
check gets through while every handler is busy. The server above has
one:
$ curl -i 127.0.0.1:8080/healthz
HTTP/1.1 200 OK
content-length: 3
connection: keep-alive
date: Sat, 10 Oct 2026 12:56:57 GMT
cache-control: no-store
ok
It answers 200 while the server is serving and
503 once a stop has begun, so the balancer stops sending
traffic while the server drains. With health_sheds = true
it also answers 503 for a short time after the server has
turned requests away because it was overloaded.
#Decompress gzip, zlib and zip data
bytes.inflate decompresses DEFLATE data. It reads a zlib
stream by default, the format inside PNG images;
{ format = "gzip" } reads a .gz file and
{ format = "raw" } a zip entry. It returns a Result, and
{ max_bytes = N } caps the output as it grows, for input
you didn't make yourself:
import bytes
import fs
import str
let packed = fs.read("scores.csv.gz")?
let text = bytes.inflate(packed, { format = "gzip" })?.to_string()?
let lines = str.split(text, "\n")
print(`{len(packed)} bytes in, {len(text)} out, {len(lines) - 2} rows`)
print(lines[1])
match bytes.inflate(packed, { format = "gzip", max_bytes = 4096 })
Ok(_) -> print("fits")
Error(e) -> print(e)
end
$ scua --allow-fs=. scores.scua
7855 bytes in, 34292 out, 2000 rows
1,player1,37
bytes.inflate: the output is over the 4096 bytes `max_bytes` allows
#Typed buffers in bulk
Four additions work on whole typed arrays at once.
fill(n, v, kind) builds a typed array directly at its final
size, without a plain array first.
copy_range(dst, at, src, start, end) copies a run of
elements in one block between arrays of the same kind, overlapping
ranges included, which is the row copy of an image or an audio buffer.
bytes.to_array(b, "u8") turns a buffer into a
{ u8 } in one copy, and elem_kind(xs) tells
you a typed array's element kind without converting anything:
import bytes
-- a 4x3 greyscale image, one byte a pixel, and a 2-pixel sprite
let w = 4
let image = fill(w * 3, 0, "u8")
let sprite = bytes.from_hex("ff80")?.to_array("u8")
-- blit the sprite into each row, one block copy a row
for row in 0:3 do
image.copy_range(row * w + row % 3, sprite)
end
print(elem_kind(image), len(image))
for row in 0:3 do print(image[row * w : row * w + w]) end
let samples = fill(1_000_000, 0.5, "f32")
print(elem_kind(samples), samples)
$ scua buffers.scua
u8 12
[255, 128, 0, 0]
[0, 255, 128, 0]
[0, 0, 255, 128]
f32 [0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, … 999984 more]
The last line shows another change: a typed array longer than 256 elements prints its first 16 elements and a count of the rest, so printing a record that holds an image no longer writes out every pixel. Plain arrays still print in full.
#Smaller additions
sys.parse_argsfills a list field with every positional argument left over, so afiles: { string }field makesconvert --quality 80 a.png b.png c.pngwork.{ unknown = "refuse" }makes an undeclared key in a JSON body an error.fs.write(path, data, { exclusive = true })andfs.open(path, { mode = "create" })fail if the file already exists, checking and creating in one step, so two programs can't both write it.{ sync = true }andfs.sync(f)wait until the data is on the disk.$ scua --allow-fs=. report.scua Ok(nil) Error(report.txt: already exists — `{ exclusive = true }` writes only a file that is not there yet) firsthash.crc32(data, start, end)checksums part of a buffer without copying it out.fserrors name the path, as in the line above.--jit-statscounts the work done on every worker thread, which is where a server's requests run.bytes.joinandbytes.from_arrayhold their result once while they build it.
#Upgrading to 0.35
Four groups of changes can need edits.
#http.serve
handlers run on their own
A handler can't reach the main script's lets, and the
server refuses to start until none does. A table the handler only reads
becomes a const; anything it changes moves into a partition
passed in context, as in the Hits example
above. The message lists each variable and what to do:
import http
import bytes
let greeting = "hello"
let hits = 0
fn handle(req)
hits = hits + 1
return { status = 200, body = bytes.from_string(`{greeting} #{hits}`) }
end
http.serve(handle, { bind = "127.0.0.1:8080" })?
$ scua --allow-serve=127.0.0.1:8080 hits.scua
scua: hits.scua:12: the handler you pass to `http.serve`, `handle`, reaches 2 variables of the main script. A request runs in a partition of its own, on any core, so it can't reach them:
- `greeting` (line 4) is read and never changed: make it a `const`
- `hits` (line 5) is changed by `handle` (line 8): keep it in an actor, e.g. `partition Hits … end`, passed in `context`; see "Keep state in a server"
A module's state starts as declared on every request, so
a value a module kept across requests also moves into a partition.
--max-mem now bounds each request, and one that goes over
it gets a 500 on its own.
#A
typed parameter, return or let needs every field of a
record
A value that is missing a field with a default is refused at the
boundary. Convert loose data, such as decoded JSON, with
as(value, Record), which fills the defaults:
import json
record Settings { volume: int = 5, muted: bool = false }
fn apply(s: Settings) return `volume {s.volume}, muted {s.muted}` end
let loaded = json.decode("{\"volume\": 8}")?
print(apply(as(loaded, Settings)))
print(apply(loaded))
$ scua settings.scua
volume 8, muted false
scua: settings.scua:5: field 'muted' of Settings is missing; use `as(value, Settings)` to fill its default
at apply (settings.scua:5)
at main (settings.scua:9)
A record literal written straight into a typed let or
return still fills its defaults, and a literal passed to a
typed parameter with a field missing is a compile error.
json.encode
and toml.encode write an enum by its name
json.encode({ kind = Kind.UnsupportedFormat }) gives
{"kind":"UnsupportedFormat"}; it used to give the number. A
record field declared with the enum still reads the number back, so data
you wrote before still loads, but a program on the other end that reads
the number needs to read the name. A flags set is still
written as its bitmask.
#Typed buffers are checked at function boundaries
A parameter or return declared { u8 },
{ f32 } or another element type now checks the value it is
given:
fn brighten(p: { u8 }) return p end
let levels: { u16 } = [10, 300]
brighten(levels)
$ scua levels.scua
scua: levels.scua:1: parameter 'p' of 'brighten' is a typed {u8} array: element 1 is 300, and it holds whole numbers from 0 to 255
at brighten (levels.scua:1)
at main (levels.scua:4)
Passing nil to one that isn't declared optional
({ u8 }?) faults too. A typed annotation no longer converts
an array someone else holds: an array that already has the declared
element type is used as it is, and any other array whose values fit is
copied. A function that fills a buffer you pass it therefore needs a
buffer of the declared type, made with fill(n, 0, "u8") for
instance; given a plain array, it fills a copy you won't see. An
{ f32 } also refuses a number beyond the range of an f32
(about ±3.4e38), where it used to store infinity; inf and
nan written as such still store.
#Everything else
Rarer changes, such as fs.rename refusing to replace a
read-only file, are in the changelog.
#What's next
The threaded scheduler becomes the default, so a server uses every core without a setting. The next work on serving is a published, reproducible benchmark with the other servers included, and a benchmark of a realistic application with a database, templates and TLS over a real network.
The full changelog has everything in 0.35.0.