Is zlib-ng actually faster than stock zlib when NGINX compresses responses? Vendor benchmarks usually answer a different question: one big buffer, one call, a synthetic corpus. NGINX does something else. It feeds the compressor in small chunks, opens a fresh stream per response, and wraps the result in a gzip header. So before recommending zlib-ng as a drop-in zlib replacement, we measured it the way NGINX actually uses it, on both x86_64 and aarch64, from the library call up to requests per second.
This article is the full write-up: the harnesses, the payloads, the NGINX end-to-end runs, the one benchmarking trap that silently invalidates many ab numbers, and the one place where zlib-ng is worse than stock zlib (and the one-line fix).
Results at a glance
zlib-ng 2.3.3 in compat mode versus stock zlib 1.2.11 on Rocky Linux 9:
| What we measured | Result |
|---|---|
| Deflate at level 6, NGINX-shaped streaming, x86_64 | +73% to +82% throughput |
| Deflate at level 6, aarch64 | +92% to +119% throughput |
| Inflate (decompression) | 2.3x to 3.5x throughput |
| Compression ratio at levels 6 and 9 | Parity (within 0.1%) |
| NGINX gzip RPS, CPU-saturated 4-core aarch64 | +30% (HTML), +44% (JSON) |
NGINX gunzip RPS (inflate per request) |
+34% (aarch64), +79% (x86_64) |
| Worker memory (RSS) | Equal |
| Level 1 | Worse ratio. Use level 2 instead |
The rest of this page explains how each number was produced, so you can reproduce it or challenge it.
Test environment
- Library under test: zlib-ng 2.3.3, built with
-DZLIB_COMPAT=ON(the same mode ourzlib-ngpackage ships), so it exports the classic zlib API and thelibz.so.1SONAME. - Baseline: stock zlib
1.2.11-40.el9from Rocky Linux 9. - Distro: Rocky Linux 9 containers, NGINX 1.20 from AppStream,
abfromhttpd-tools. - x86_64 host: Intel Xeon Silver 4410Y, 24 cores, shared with other workloads (roughly 30% background load). Treat x86_64 numbers as relative, not absolute.
- aarch64 host: a 4-core Arm64 server, otherwise idle.
Both libraries ran on the same host, in the same container, back to back. Only the shared library changed between runs.
Payloads
Three files, chosen to represent what a web server actually compresses:
| File | Size | Why |
|---|---|---|
page.html |
~70 KB | A real getpagespeed.com page, fetched live. Typical dynamic HTML |
data.json |
~588 KB | 4,000 API-style JSON records. Typical API response |
random.bin |
1 MB | From /dev/urandom. Incompressible worst case |
The JSON was generated with a one-liner, so it is fully reproducible:
awk 'BEGIN{printf "["; for(i=0;i<4000;i++){printf "{\"id\":%d,\"user\":\"user_%d\",\"email\":\"user%d@example.com\",\"active\":%s,\"score\":%d.%02d,\"tags\":[\"alpha\",\"beta\",\"gamma\"],\"ts\":\"2026-09-01T12:%02d:%02d Z\"},",i,i,i,(i%2?"true":"false"),i%97,i%100,i%60,i%60} printf "{}]"}' > data.json
head -c 1048576 /dev/urandom > random.bin
Harness 1: one-shot gzip deflate and inflate
The first harness is a small C program that compresses the whole file in a single deflate(Z_FINISH) call, repeatedly, for a fixed wall-clock time, and then does the same for inflate. Two details make it NGINX-relevant rather than generic:
windowBits = 15 + 16produces a gzip wrapper, exactly what NGINX sends withContent-Encoding: gzip.memLevel = 8matches the NGINX gzip filter default.
deflateInit2(&s, level, Z_DEFLATED, 15 + 16, 8, Z_DEFAULT_STRATEGY);
s.next_in = in; s.avail_in = insize;
s.next_out = out; s.avail_out = obound;
deflate(&s, Z_FINISH);
Each run loops for 2 seconds on CLOCK_MONOTONIC and reports MB/s of input processed, plus the output size, so ratio and speed come from the same measurement.
Harness 2: chunked streaming, shaped like NGINX
One-shot numbers flatter any library, because a single large call hides per-call overhead. NGINX never compresses that way. Its gzip filter receives the response body as a chain of buffers and calls deflate() repeatedly with Z_NO_FLUSH, finishing with Z_FINISH on the last buffer. The second harness reproduces that pattern with 32 KB chunks:
while (off < insize) {
long n = insize - off < chunk ? insize - off : chunk;
s.next_in = in + off; s.avail_in = n;
off += n;
int flush = off < insize ? Z_NO_FLUSH : Z_FINISH;
deflate(&s, flush);
}
A new stream is initialized for every iteration, the same way NGINX allocates a fresh deflate state per response. We also tried 8 KB chunks. Nothing changed materially, so the 32 KB numbers below stand for both.
Switching libraries without touching the binary
Because zlib-ng compat exports the same libz.so.1 SONAME, the cleanest A/B is to keep one compiled binary and change only which library the dynamic linker loads:
# stock zlib
./bench2 page.html 6
# zlib-ng, process-private, nothing else on the host affected
LD_LIBRARY_PATH=/opt/zng/lib ./bench2 page.html 6
ldd ./bench2 | grep libz was printed for every mode, so each result line is tied to the library that produced it. The same trick switched NGINX: LD_LIBRARY_PATH=/opt/zng/lib nginx versus plain nginx, with the worker’s /proc/<pid>/maps checked in each mode (more on that below).
Library results
Deflate, level 6, chunked streaming (x86_64)
| Payload | Stock zlib MB/s | zlib-ng MB/s | Change | Ratio stock / zlib-ng |
|---|---|---|---|---|
| HTML 70 KB | 41.9 | 76.3 | +82% | 0.2491 / 0.2493 |
| JSON 588 KB | 162.0 | 279.5 | +73% | 0.0888 / 0.0886 |
The one-shot harness agrees: +39% to +88% at level 6. At level 9, zlib-ng is +33% faster on HTML and +52% on JSON, again at ratio parity.
Deflate, level 6 (aarch64)
| Payload | Stock zlib MB/s | zlib-ng MB/s | Change |
|---|---|---|---|
| HTML | 44.1 | 84.8 | +92% |
| JSON | 108.4 | 237.7 | +119% |
The Arm gain is larger than on x86_64. zlib-ng has dedicated NEON and ARMv8 CRC32 code paths, while stock zlib 1.2.11 runs generic C on Arm. If you run NGINX on Graviton or Ampere, this is the bigger story.
Inflate (decompression of level-6 streams)
| Arch | Payload | Stock zlib MB/s | zlib-ng MB/s | Change |
|---|---|---|---|---|
| x86_64 | HTML | 358 | 654 | +83% |
| x86_64 | JSON | 769 | 2,653 | 3.5x |
| aarch64 | HTML | 296 | 678 | 2.3x |
| aarch64 | JSON | 543 | 1,864 | 3.4x |
Inflate matters to NGINX more than people expect: the gunzip filter decompresses for clients that don’t accept gzip, and any proxy setup that unpacks gzipped upstream responses pays the inflate cost on every request.
NGINX end-to-end
Library speed only matters if it survives the trip through NGINX. The end-to-end test served the same three payloads from a minimal server block:
server {
listen 127.0.0.1:8080;
root /bench/www;
gzip on;
gzip_comp_level 6;
gzip_min_length 20;
gzip_http_version 1.0;
gzip_types text/html application/json application/octet-stream;
location /gs/ {
alias /bench/gs/;
gzip_static always;
gunzip on;
}
}
The /gs/ location holds only a pre-compressed page.html.gz. A client that does not send Accept-Encoding: gzip forces NGINX to inflate it on every request, which isolates the decompression path.
Load came from ApacheBench with keep-alive, 8,000 requests per run at concurrency 4:
ab -q -k -n 8000 -c 4 -H 'Accept-Encoding: gzip' http://127.0.0.1:8080/page.html
Worker CPU time was sampled before and after each run with ps -o time= summed over all workers.
The trap: ab speaks HTTP/1.0
Look at gzip_http_version 1.0 in the config above. It is not optional for this test. ab sends HTTP/1.0 requests, and NGINX’s default gzip_http_version is 1.1. Without the override, NGINX silently skips compression for every ab request, still returns 200, and you end up benchmarking plain file serving. Both libraries then look identical, because neither is called.
If you have ever seen “zlib-ng makes no difference to NGINX” from an ab run, this is the first thing to check. Confirm compression is really happening before trusting any number:
curl -s -o /dev/null -w '%{size_download}\n' --http1.0 -H 'Accept-Encoding: gzip' http://127.0.0.1:8080/page.html
If the size equals the uncompressed file, your benchmark is measuring the wrong thing. (wrk and h2load speak HTTP/1.1 or newer and don’t hit this, but the check costs nothing.)
Results: aarch64, 4 cores, CPU-saturated
| Path | Stock RPS | zlib-ng RPS | Change | Worker CPU (stock / zlib-ng) |
|---|---|---|---|---|
| HTML, gzip level 6 | 1,260 | 1,641 | +30% | 15 s / 11 s |
| JSON, gzip level 6 | 614 | 886 | +44% | 49 s / 30 s |
| 1 MB random, gzip level 6 | 123 | 127 | ~flat | 248 s / 238 s |
gunzip (inflate per request) |
4,456 | 5,993 | +34% | 5 s / 4 s |
On a box where compression is the bottleneck, the library speedup shows up almost directly as requests per second. JSON gains the most because the per-response deflate work is largest there. Incompressible data stays roughly flat, as expected: there is almost nothing for either library to compress.
Results: x86_64, 24 cores, not saturated
With 24 cores and only 4 concurrent clients, the x86_64 host was never CPU-bound, so gzip RPS was latency-bound and mostly flat, which is expected. The CPU accounting still showed the saving (HTML worker CPU dropped from 8 s to 7 s per 8,000 requests, in line with the harness prediction). The inflate-heavy gunzip path was the exception: 7,649 to 13,731 RPS, +79%, because inflate is where zlib-ng’s lead is widest.
The takeaway: zlib-ng lowers CPU per compressed response everywhere. It turns into more requests per second only where compression is what limits you. On an idle server it turns into lower load instead.
Correctness and operational checks
Speed is worthless if the output is wrong. In both library modes, on both architectures, we checked:
- Round-trip:
curl --compressedthe gzipped page andcmpit against the original. Identical. - Interoperability: the gzip stream NGINX produced with zlib-ng passes
gzip -tand decompresses with the stockgzipCLI. - Inflate path: the
gunzipmodule correctly inflates.gzfiles on the way out. Identical to the original. - Compressed size at level 6: 17,618 vs 17,637 bytes (x86_64) and 18,077 vs 18,093 bytes (aarch64). A 0.1% difference, which is noise.
- Memory: worker RSS equal within half a megabyte, in both directions across runs.
- Startup: indistinguishable.
- It was really zlib-ng.
lddonly reports what the dynamic linker would resolve. The proof is the running worker’s memory map:
pid=$(pgrep -f 'nginx: worker' | head -1)
grep libz /proc/$pid/maps | awk '{print $6}' | sort -u
In zlib-ng mode this printed the zlib-ng library path and nothing else. We checked it in every mode, before every set of runs.
One more data point that no lab benchmark gives you: this website’s own production NGINX has been running on zlib-ng since April 2026, months before this evaluation, with no compression-related issues.
The level 1 exception: use level 2
Here is the one place where zlib-ng is not a pure win. At level 1 it switches to a strategy called deflate-quick, which trades compression ratio for raw speed:
| Config | HTML ratio | HTML MB/s | JSON ratio | JSON MB/s |
|---|---|---|---|---|
| Stock zlib, level 1 | 0.291 | 125 | 0.105 | 352 |
| zlib-ng, level 1 | 0.344 (worse) | 417 | 0.137 (worse) | 1,189 |
| zlib-ng, level 2 | 0.279 (better) | 226 | 0.093 (better) | 993 |
Lower ratio is better here (compressed size divided by original size). zlib-ng level 1 produces output about 5 percentage points larger than stock zlib level 1, and on incompressible input it can expand the payload by around 5%, where stock zlib stays essentially flat.
zlib-ng level 2 fixes both: a better ratio than stock level 1 and still roughly 2 to 3 times faster. This matters for NGINX more than it sounds, because gzip_comp_level defaults to 1 when you don’t set it. Many configs never set it. If yours is one of them, add this when you install zlib-ng:
gzip_comp_level 2;
Levels 3 and higher need no change. They match or beat stock zlib’s ratio at the same level and run faster.
A note for Rocky Linux 10, RHEL 10, and Fedora
On EL10 and current Fedora, the distro’s own libz.so.1 is already zlib-ng: Red Hat ships zlib-ng-compat as the system zlib (version 2.2.3 on Rocky Linux 10.1). The large stock-zlib-versus-zlib-ng gap above is the EL9-and-older story. On EL10 our zlib-ng package moves you from 2.2.3 to the current 2.3.x.
We checked the level behavior on a Rocky Linux 10.1 aarch64 VM with the chunked harness. The distro’s 2.2.3 at level 1 produced byte-identical output to zlib-ng 2.3.3 at level 2 (21,558 bytes from a 75,838-byte HTML page), while 2.3.3 at level 1 is deflate-quick (26,581 bytes). So on EL10 the level-2 advice is not optional: if you install the newer library and leave gzip_comp_level unset, gzipped responses get about 23% larger. Set level 2 and you get exactly the output you had before. We have not published speed numbers for 2.2.3 versus 2.3.3. On our laptop-hosted VM the run-to-run noise was larger than the difference we were trying to measure.
Where zlib-ng does not help
gzip_staticandbrotli_static. Pre-compressed files are sent as-is. NGINX does no runtime zlib work for them, so the library is irrelevant on that path.- Brotli and Zstandard responses. Those use their own libraries. zlib-ng only affects gzip and deflate.
- Incompressible content. Already-compressed media gives deflate nothing to do.
zlib-ng and Brotli are not competitors. Brotli needs a module and per-site configuration, and it only reaches clients that ask for it. zlib-ng is zero-config: every existing gzip on; site benefits, including non-browser API clients that only send Accept-Encoding: gzip, plus every inflate path. Run both.
Reproduce it
Everything above runs inside a stock Rocky Linux 9 container:
dnf -y install gcc make cmake nginx httpd-tools procps-ng gzip zlib-devel diffutils
curl -sL https://github.com/zlib-ng/zlib-ng/archive/refs/tags/2.3.3.tar.gz | tar xz
cd zlib-ng-2.3.3 && mkdir b && cd b
cmake .. -DZLIB_COMPAT=ON -DCMAKE_BUILD_TYPE=Release -DZLIB_ENABLE_TESTS=OFF -DWITH_GTEST=OFF
make -j"$(nproc)"
mkdir -p /opt/zng/lib && cp -P libz.so* /opt/zng/lib/
Then compile each harness against the system headers with gcc -O2 bench.c -o bench -lz and run it with and without LD_LIBRARY_PATH=/opt/zng/lib. Three practical gotchas we hit along the way:
gzip_http_version 1.0when load-testing withab, as explained above.- Directory permissions on bind mounts. If the payload directory comes from a host path under a
0750home directory, the NGINX worker (running asnginx) gets 403 for everything and you benchmark error pages.chmod 755the mount point and verify a 200 before measuring. ps -o time=has one-second resolution. Worker CPU deltas from short runs are coarse. Use enough requests that the CPU time lands in tens of seconds, or read/proc/<pid>/statfor finer ticks.
Install zlib-ng for your NGINX
You don’t need a custom NGINX build to get these numbers. Our zlib-ng package installs into an isolated directory and is activated through /etc/ld.so.conf.d, so every program on the host that links libz.so.1, NGINX included, picks it up after a restart. It is free and available for RHEL, CentOS Stream, Rocky Linux and AlmaLinux 7 to 10, Amazon Linux 2 and 2023, Fedora 42 and 43, and SLES 16, on x86_64 and aarch64:
sudo dnf -y install https://extras.getpagespeed.com/release-latest.rpm
sudo dnf -y install zlib-ng
sudo systemctl restart nginx
Then confirm all four things: link-time resolution, a healthy restart, the library mapped into the live worker, and a real compressed response:
ldd $(command -v nginx) | grep libz
systemctl is-active nginx
grep libz /proc/$(pgrep -f 'nginx: worker' | head -1)/maps | awk '{print $6}' | sort -u
curl -sS -o /dev/null -D- -H 'Accept-Encoding: gzip' http://localhost/ | grep -i content-encoding
We ran exactly this sequence on a fresh Rocky Linux 10 VM with the GetPageSpeed NGINX package before publishing: ldd and the worker map both pointed at /usr/lib64/zlib-ng/libz.so.1, NGINX restarted cleanly, and responses came back with Content-Encoding: gzip and decompressed byte-for-byte identical to the original. Don’t forget gzip_comp_level 2 if your config doesn’t set a level.
For the full install guide, rollback instructions, and troubleshooting, see Install zlib-ng: faster drop-in zlib for RHEL and CentOS. For the rest of your gzip settings, see the NGINX gzip compression guide.
FAQ
Does zlib-ng change what clients receive? No. It emits standard gzip and deflate streams that any zlib can decode. The output bytes differ slightly from stock zlib at the same level (a different but equally valid encoding), which is why compressed sizes moved by 0.1%.
Do I need to rebuild NGINX? No. zlib-ng compat keeps the libz.so.1 ABI. Install the package and restart NGINX.
Is this worth it on a lightly loaded server? The CPU per gzip response drops either way. On a lightly loaded server you won’t see more requests per second, you’ll see lower CPU use and slightly lower latency on large responses. On a CPU-bound server you’ll see the throughput gains in the tables above.
Why not just use Brotli? Use both. Brotli covers browsers that ask for it. zlib-ng speeds up everything that still uses gzip or needs decompression, without any configuration change.

