← Research

Table Press

Can an agent out-compress the encoders inside real HTTP/3 servers?

Attoloop ResearchSeptember 2026

An agent is given the headers of 50 web responses, recorded from a real browser in 2012: the date, the server name, the caching rules, a privacy policy, and so on, about 19 KB of text. It is also given the specification of QPACK, the part of HTTP/3 that compresses headers, and ten minutes.

It must write those headers as the exact bytes an HTTP/3 connection would send.

Two real decoders then read the file back: ls-qpack, the library inside LiteSpeed's web servers, and nghttp3, the HTTP/3 library curl uses. If either one gets a single header wrong, the file scores nothing. If both reproduce every header exactly, the score is the size of the file. Fewer bytes is better.

This is Table Press, an environment we built to train and measure a skill that tests of protocol code usually skip: not only getting a protocol right, but getting it right with fewer bytes than the software people actually ship. We ran four frontier coding agents on it and compared every file they wrote with what the real encoders produce on the same input.

Agent Files both decoders accepted Files smaller than ls-qpack's own encoder
GPT-5.6-sol (Codex), 40 runs 40 / 40 39 / 40
Claude Opus 5, 10 runs 10 / 10 10 / 10
Claude Sonnet 5, 40 runs 38 / 40 6 / 40
Claude Sonnet 5 at effort xhigh, 20 runs 13 / 20 5 / 20

Source: our measurement, version 4 of the environment, the same 10 tasks for every model. A fourth agent, GPT-5.6-terra, ran part of the same tasks before we stopped that job. It appears in Section VIII.

Every agent could make a file that both decoders accept. What separated them was what they did next. Two of them kept going and wrote encoders that beat a production library on its own ground. The third usually stopped, early and on purpose, at exactly the level that most HTTP/3 servers ship today. Given more room to think, it tried the harder design far more often, and ran out of time before it could finish. This post is about why header compression is a hard problem, how we built an exact grader for it, what we found and fixed along the way, and what the agents did.


IHeaders repeat, and the protocol knows it

Every time a browser asks a website for something, a page, an image or a script, the request and the reply each carry a few lines of labels, like the address and stamps on an envelope. These labels are called headers. Here are the headers of one real reply from our data, a small image sent by a web portal in November 2012:

:status: 200
cache-control: max-age=43200
content-type: image/gif
accept-ranges: bytes
etag: "808cfaf3c1ac81:0"
server: Microsoft-IIS/7.5
x-powered-by: ASP.NET
p3p: CP="BUS CUR CONo FIN IVDo ONL OUR PHY SAMo TELo"
content-length: 43
age: 7374
date: Sat, 03 Nov 2012 13:29:29 GMT
last-modified: Mon, 29 Oct 2007 15:02:13 GMT
expires: Sat, 03 Nov 2012 23:26:35 GMT
connection: keep-alive

The image itself is 43 bytes (content-length: 43). Its headers are about 370 bytes of text. A single page can need a hundred of these exchanges, and most of their headers are the same every time: the same server, the same privacy policy, a date that moves forward by a second. In 2009 the team designing SPDY, the protocol that became HTTP/2, wrote that "typical header sizes of 700-800 bytes" were common for requests (SPDY whitepaper), and most of that text is sent again with every exchange.

The obvious fix is to zip the headers, the way you might zip a folder before emailing it. SPDY did, and it turned out to be dangerous. The CRIME attack (Rizzo & Duong, ekoparty 2012) showed that an attacker who can add text to a request can guess a secret cookie one character at a time: when a guess matches part of the cookie, the zipped request comes out slightly shorter, and the attacker can see the length. So HTTP/2 replaced zipping with HPACK (RFC 7541), designed to limit exactly this kind of attack. It is built from three simple tools:

  1. A shared phrasebook: the static table. Both sides hold the same printed list of common headers, each with a number. HPACK's list has 61 entries and QPACK's has 99. Instead of writing out :status: 200 ("the request worked"), a QPACK sender writes "entry 25", which fits in one byte.
  2. A shorthand for letters: the Huffman code. Common characters get short codes and rare ones get long codes. Morse code uses the same trick: e, the most common letter in English, is a single dot. The 29-character date above fits in 22 bytes. On our traces the code makes header text about a quarter shorter.
  3. A shared notepad: the dynamic table. The sender can say "write this header down". From then on it can say "the one on line 3" instead of sending the header again, in about one byte. The notepad is small, so when it is full the oldest lines are erased to make room.

HTTP/3, the newest version, sends many requests side by side over one connection, and their pieces can arrive in any order. So its version of HPACK, QPACK (RFC 9204), sends notepad updates separately from the messages that use them. A message that refers to line 3 before the receiver has written line 3 down has to wait. That wait is called a blocked stream, and the receiver only lets a limited number of messages wait at once.

Here is what the three tools are worth on one header from the reply above. The same line appears in 40 of the 50 replies in that recording:

p3p: CP="BUS CUR CONo FIN IVDo ONL OUR PHY SAMo TELo"
How the sender handles it Bytes per line Bytes for all 40 lines
Spell it out 53 2,120
Spell it out with the Huffman code 46 1,840
Store it once in the dynamic table (46 bytes), then refer to it 1 86

Source: our measurement on task 02210f9d9490. p3p is not in the static table, so the static table cannot help here.

The notepad is where the large savings are, and it is also where the difficulty is. It is small (sometimes only 256 bytes, room for three headers like this one, or four or five shorter ones), so the sender must decide what to write down, when, and what to let fall out. The specification leaves all of that to the sender on purpose: "QPACK is designed to place the burden of optional state tracking on the encoder, resulting in relatively simple decoders" (RFC 9204, §2.1).

Many implementations decline the burden. Cloudflare's quiche never adds a table entry and writes zero into the two fields of every message that would point into the dynamic table (encoder.rs, lines 50-55). The quic-go QPACK library says it plainly: "it does not support the dynamic table and relies solely on the static table and string literals (including Huffman encoding), which limits compression efficiency" (README). This matters because HTTP/3 is widely used: 40.8% of websites use it (W3Techs, 30 September 2026), and it carried 21% of the requests Cloudflare served in 2025 (Cloudflare Radar).

We looked for published work on writing good QPACK or HPACK encoders and found one measurement study, by the authors of nghttp2, H2O and haskell-http2. It compares whole strategy families, from sending everything raw (1.1 times the original size) to static table, dynamic table and Huffman together (0.31) (Yamamoto, Tsujikawa & Oku, CFI 2017), but proposes no method for choosing what to store. Research on language models points in two directions:

Study What was measured Result
Yoran et al., ICLR 2025 (KoLMogorov Test) Models write the shortest program that reproduces a sequence "Current flagship models perform poorly"
Shen et al., NeurIPS 2025 (PSMBench) Models read protocol state machines out of RFC text At most 0.38 F1 on transitions
Long, Su & Han, AAAI 2026 (APG) Models write ICMP, IGMP, NTP and TCP code from RFCs About 90% interoperability for TCP, after minimal manual patches

Compression that must also be exactly right, against two independent implementations, with a grade that keeps rewarding every byte saved: we found no benchmark that measures that. Table Press does.


IIOne task, end to end

Every task gives the agent the same kinds of files:

  • trace.qif: 30 to 50 header lists from one recorded connection, one header per line, a blank line between messages.
  • config.json: the dynamic table size (256, 1,024, 4,096 or 65,536 bytes), the acknowledgement mode, and the blocked-stream limit (always 100).
  • The full text of RFC 9204 and RFC 7541, and a note describing the file format.
  • /app/bin/check, which runs the grader's own checks locally, and the two decoders, compiled without their encoder halves.

The agent must write one file, encoded.bin. Stream 0 carries the table updates. Streams 1 to N each carry one encoded header list. The acknowledgement mode decides how much freedom the table gives:

  • In immediate mode, the receiver confirms each message as soon as it is decoded. Entries the sender no longer needs may then be evicted and replaced. With a small table, a good sender rotates what it stores as the connection moves on.
  • In none mode, nothing is ever confirmed. No entry may ever be evicted, and every message that uses the table counts against the blocked-stream limit. The table becomes a one-time purchase: choose what to store, once, before you know you were right. It is a knapsack problem.

Each message starts with one number that trips up almost everyone, including, as we will see, every agent that failed: the Required Insert Count. It tells the decoder how many table entries it must have received before it can read this message, like a note at the top of a letter that says "you need the first 12 lines of the notepad to read this". Set it too low and the message refers to something that does not exist yet. Set it higher than needed and the specification lets a decoder refuse the message.

For task 02210f9d9490 (50 responses from one 2012 web portal, 19,189 bytes of headers, a 256-byte table, and immediate mode) the landmarks are:

Encoding Bytes Relative to the plain text
Every header spelled out (our reference solution) 20,285 106%
Static table + Huffman, no dynamic table (what quiche and quic-go do) 11,490 60%
nghttp3's encoder 10,945 57%
ls-qpack's encoder 9,916 52%
Best agent file (GPT-5.6-sol, two runs) 7,360 38%

Source: our measurement. The reference encoders are the same pinned library versions the grader uses to decode.

The ladder an agent climbs is the same on every task: spell everything out, add the static table and the Huffman code, then use the dynamic table, then plan it. Each step is worth more than the last, and each is harder to get exactly right.

The same headers, eight ways Figure 1. Total bytes over the 10 tasks we measured (500 header lists, 210,111 bytes of plain text). Grey: fixed encoders, including our reference solution and the two real libraries. Blue: Claude models. Orange: GPT-5.6-sol. For models that ran several times per task, each task contributes its median run.


IIIAn exact grader built from two real decoders

The reward is one number: minus the size of the file in bytes, paid only if the file passes every check. There is no judge model and no partial credit.

The two decoders are pinned releases of real libraries: ls-qpack v2.7.0, used in LiteSpeed's web servers, and nghttp3 v1.18.0, which curl's non-experimental HTTP/3 support is built on (curl docs). We drive each one through its library interface with a short program of our own. The checks run in order:

  1. The file follows the format: every stream from 1 to N appears exactly once, and nothing is left over.
  2. ls-qpack decodes it under the task's table size, acknowledgement mode and blocked-stream limit, and reproduces every header list byte for byte.
  3. nghttp3 does the same.

Each decoder gets 2 CPU seconds, 512 MiB and 3 seconds of wall-clock time. A file that fails any check gets a fixed score below anything a valid file can earn.

Why two decoders? Because they disagree, and the specification allows them to. The clearest case is the one that decided every rejected file in this post. When a message declares a Required Insert Count higher than it needs, RFC 9204 says a decoder "MAY treat this as a connection error" (§2.2.1). ls-qpack does. nghttp3 accepts the message. An encoder tested against only one of them could be wrong about the other. Requiring both to agree means we grade the part of the standard that nobody can read two ways.

Before we built anything, we checked that the grader catches real mistakes. The proposal for this environment flipped single bits in a valid file 1,000 times. 915 corrupted files were caught because the decoded headers differed, 79 were rejected by both decoders, 3 by only one of them (which the two-decoder rule also rejects), and 3 were accepted. All 3 flipped a flag that changes nothing about the decoded headers, so they were still correct files.

Source: our measurement, run while designing the environment.


IVBuilding the data

The headers come from two public collections of recorded traffic, both MIT-licensed: the QPACK interop corpus qifs, which the QUIC working group's interop guide points to, recorded from a large social network, and hpack-test-case, browser recordings of popular sites made in November 2012. We used 14 traces and replaced every cookie value and every identifier-shaped value (long numbers, long hex strings, long base64 strings) with random text of the same length and character types. The same original always maps to the same replacement, so repeats stay repeats. On the traces we measured, this changed ls-qpack's output size by less than 0.2%.

Each trace was cut into windows of up to 50 header lists, 78 windows in all. Each window got three table configurations, two of them from the tight sizes (256 and 1,024 bytes), because that is where the design work showed the largest differences between strategies. Every combination was then checked in both directions before it became a task: ls-qpack's encoder with nghttp3's decoder, and nghttp3's encoder with ls-qpack's decoder. All 234 passed. A de-duplication step then judged the three configurations of a window, which share the same headers, to be near-copies and often kept only one or two. The published dataset has 119 tasks and covers all 78 windows.

These windows repeat themselves heavily. In a median window, 63% of header lines are exact copies of a line already sent earlier in the same window (71.5% across all 119 tasks). Even after removing the lines the static table already covers, a static-only encoder spends about half of its bytes (median 47.9%) sending text the decoder has already seen. Some of it is strange: one site's responses carry both nncoection: close and cneonction: close, two scrambled spellings of connection. Real traffic is messy, and the dynamic table does not care: the best runs stored both.

Source: our measurement over the 10 tasks in Section VII and all 119 published tasks.


VFour versions, and what each one taught us

We published Table Press and then kept running agents on it. Each round of runs showed us something the task got wrong, and each fix became a new version. The results in this post come from version 4.

Version 1: online, and the answer could be downloaded. In the first version the agent's sandbox could reach the internet, like a developer's laptop. The task withholds both encoders, because running one is not writing one. A reviewer had warned while we wrote the specification that an online agent could simply download them, and the runs proved the warning right. 11 runs downloaded something, and three submitted ls-qpack's own output. Two of those matched the reference size to the byte. One Sonnet run said it outright:

"This is exactly the reference tool for this exact task. Let's build ls-qpack and use interop-encode/interop-decode directly."

Its sibling run on the same task, with no downloads, wrote its own encoder and scored 4,086 bytes, 224 fewer than the downloaded answer. Version 1 also never told the agent how long it had. GPT-5.6-terra ended 52 of its 180 measured runs with no file at all when the ten minutes ran out.

Version 2: say how long there is. The instruction now states the ten-minute limit and says to write a valid file early and improve it afterwards. All 122 version 4 runs in this post left a file.

Version 3: offline. The agent's sandbox lost the internet, and we confirmed from inside it that package sites and GitHub are unreachable. That closed the shortcut and exposed a second gap. The specification had promised the agent a local copy of the grader's checks and both decoders, and they were not there. Offline and without a decoder, an agent could only test its file against a decoder it wrote itself, which shares its own mistakes. One Opus run on version 3 did exactly that, and said so at the end:

"That's my own decoder, not ls-qpack or nghttp3 … so I could not run the actual graders."

Its file happened to pass. Another Opus run was less lucky. Its own decoder printed OK blocks=51 streams=50 size=4708, and ls-qpack rejected the file, so it scored the floor.

Version 4: the tools the specification promised. Version 4 ships /app/bin/check and both decoders, built from the grader's pinned sources with the encoder code stripped out. On all 307 files we had saved from earlier runs, check gives the same verdict and byte count as the real grader. Claude Opus 5 on the same five tasks shows the difference. On the task where version 3 rejected the 4,708-byte file, the version 4 run also wrote 4,708 bytes, and this time it was accepted:

Version What the agent had Files accepted Smaller than ls-qpack
3 offline, no decoders 2 / 5 1 / 5
4 offline, check and both decoders 4 / 5 4 / 5

Source: our measurement. One run per task, the same five tasks.


VIAttacking our own grader

An exact grader is only useful if nothing but a correct, small file can earn its reward. Before the first version was released, two kinds of attackers went after it. Reviewer models read the grader's code looking for ways to fool it, and a red-team model wrote submissions designed to score without doing the work.

The reviewers and the build's automated checks filed 38 problems, 12 of them critical, and 28 were fixed before release. The other ten were recorded and shipped. They include two issues, the online sandbox and an unstated per-decoder time limit, and both were fixed in version 3. The most interesting holes were ways to break the rules and still get paid:

  • Anything goes on the side channel. The first grader only understood headers spelled out in full, and it never looked at stream 0, where table updates travel. Every good, compressed answer scored the floor, and stream 0 could carry any bytes at all. The fixed grader hands every byte to both real decoders.
  • Breaking a promise nobody checked. In none mode the sender promises never to erase a stored entry. The nghttp3 side of the grader did not check that promise, so a file that erased entries it had promised to keep could still pass. We added a check that looks directly into nghttp3's table.
  • Hiding the risk. In none mode, every message that uses the table counts against the blocked-stream limit. The first decoder drivers only counted messages that actually had to wait. A file that sent all its table updates before its messages never made anything wait, so the limit was never counted. The drivers now count what each message declares.
  • Failing on purpose. The first scoring rule gave a rejected file minus twice the trace size, but allowed accepted files up to four times the trace size. For a large file, deliberately breaking it would have scored better than submitting it. Rejected files now score just below the largest file that could be accepted.
  • A check that could vanish. An audit found that the grader's decoder checks could be deleted without any of its own tests noticing, because the tests and the grader looked at different file paths. That was fixed.

The red team ran 70 attacks: 46 it wrote and 24 replays of them on more tasks. It tried a fake reward.json placed next to the submission, 20,000 small extra files to make the grader run out of time listing them, a file whose header promises 4 GB of data and delivers one byte, the same message sent twice, all 50 messages present but empty, encoded.bin as a link to another file, thousands of empty table-update blocks packed up to the size limit to race the five-second deadline, and a search of the sandbox for a leaked reference answer. Every attack that ran to the end scored exactly the floor. One found a real bug that was not an exploit: an empty submission folder crashed the grader instead of scoring the floor. It was fixed.

The one route that did get through, downloading the encoders, was not in the red team's list. Reviewers had flagged it, and real agents found it (Section V). An attacker only tests what it thinks of, which is why we kept watching real runs after release.


VIIResults

We ran each agent on the same 10 tasks, drawn with a fixed seed from the 119. GPT-5.6-sol ran through the Codex CLI at high reasoning effort, four attempts per task. Claude Sonnet 5 ran through Claude Code, four attempts per task, at Claude Code's default effort for that model, which is high, and then again at xhigh, two attempts per task (Section IX). Claude Opus 5 ran once per task. Every agent had ten minutes, no network, and the same local tools. Every number below is our measurement from the graded files, and a separate decoder we wrote reproduces every accepted file's headers exactly.

Claude Sonnet 5 Sonnet 5, effort xhigh Claude Opus 5 GPT-5.6-sol (Codex)
Runs 40 20 10 40
Accepted by both decoders 38 13 10 40
Smaller than static-only 7 6 10 40
Exactly the static-only size 16 4 0 0
Smaller than ls-qpack's encoder 6 5 10 39
Median size relative to ls-qpack 1.50 1.09 0.88 0.85
Median minutes worked (limit 10) 3.1 9.5 7.0 7.4

Source: our measurement. "Static-only" is ls-qpack's output with the dynamic table turned off: static table and Huffman code only.

GPT-5.6-sol's worst run beats the best default-effort Sonnet 5 run on all 10 tasks. Opus 5 and sol are nearly the same encoder on six of the ten tasks, down to reaching identical file sizes (4,651 bytes on one task, 3,646 on another) with different bytes.

Every run against its task's landmarks Figure 2. Each row is one task, labelled with its table size and acknowledgement mode, ordered by table size. Each mark is one run, placed at its file size divided by the static-only size (1.0, dotted). Filled blue: Sonnet 5 at default effort. Hollow blue: Sonnet 5 at xhigh. Larger marks are several runs with the same size. Strong tick: ls-qpack's encoder. Faint tick: nghttp3's encoder. Rejected files are in the shaded column.

Side by side

Task (table · ack) Static-only ls-qpack nghttp3 Sonnet 5 Sonnet 5 xhigh Opus 5 GPT-5.6-sol
02210f9d9490
256 · immediate
11,490 9,916 10,945 7,834 · 11,490 × 2 · 11,896 11,490 · 14,129 7,597 7,360 × 2 · 7,363 × 2
a449a438f44b
256 · immediate
10,937 10,014 10,395 7,714 · 10,937 × 2 · 13,812 7,707 · 7,710 7,661 7,569 · 7,570 · 7,601 · 7,686
0d37295c057d
256 · none
11,407 10,769 11,057 11,407 · 11,487 × 2 · 19,898 11,407 · rejected 9,699 9,593 × 4
0e29694677f6
1,024 · immediate
12,193 8,703 10,651 9,900 · 12,193 × 2 · 21,704 7,954 · 15,423 7,699 7,386 · 7,648 · 7,696 · 9,382
1c6d1b4fa91d
1,024 · immediate
21,987 11,142 20,672 22,047 × 3 · 27,831 rejected × 2 10,000 9,144 · 9,544 · 9,951 · 9,958
34ec5ed416ee
1,024 · immediate
7,584 5,099 6,446 4,918 · 4,989 · 7,584 · rejected 4,690 · 7,584 4,651 4,651 × 3 · 4,659
066bcff9a5e2
4,096 · none
13,900 8,329 12,643 13,900 × 3 · rejected 9,119 · rejected 7,029 7,017 × 2 · 7,019 · 7,035
28a6a3d688a4
4,096 · none
6,903 4,152 5,540 6,903 × 3 · 8,582 6,903 · rejected 3,646 3,645 · 3,646 · 3,651 · 3,659
48149e36c3f9
65,536 · none
8,728 5,734 7,390 4,867 · 8,728 · 10,864 × 2 5,022 · rejected 4,831 4,800 × 2 · 4,803 · 4,861
61725c26a104
65,536 · none
23,173 8,314 8,280 7,190 · 23,173 · 30,529 × 2 30,555 · rejected 7,279 6,540 · 6,541 · 6,559 · 6,595

Source: our measurement. Bytes per graded file, smallest first.

Three patterns stand out.

The small tables are where the production encoders give up most. With a 256-byte table, ls-qpack is only 6% to 14% smaller than static-only on these tasks. GPT-5.6-sol is 26% smaller than ls-qpack on the first task. The difference is not cleverness at compressing text. It is planning what to keep.

The best agents send each repeated value once. ls-qpack encodes one message at a time. The first time it sees a header, it sends the text and also stores it, so the value crosses the wire twice. The best agents read the whole trace first and store each valuable header before its first use.

How many times each repeated header value is sent Figure 3. For every header value that occurs more than once in a task, how many copies of it the file contains, averaged. With a roomy table, Opus 5 and GPT-5.6-sol send each such value exactly once. ls-qpack sends it about twice. With a tight table, some values must be sent again after they fall out of the table.

On the 10 tasks, ls-qpack's files total 14,467 bytes more than sol's best files. 85% of that comes from repeated values, and 15% from framing, because ls-qpack sends many small table updates where sol sends one to six. sol actually spends more on table updates than ls-qpack does, and wins it back many times over.

One question decides most Sonnet runs, and it is asked in the first minute. On every task, Sonnet 5's four attempts split the same way: most stop at static-only, and at most one or two use the dynamic table. When Sonnet did use the table, it did well: 9 runs tried it, 7 were accepted, and 6 of those beat ls-qpack. Its best file on task 48149e36c3f9 (4,867 bytes) is within 1.4% of sol's best.

Where the bytes go in the best files

Four long headers whose repeated values are not in the static table, p3p, content-security-policy, date and connection, account for 62% of everything GPT-5.6-sol saves over static-only. With a table of 4,096 bytes or more, each repeated line in sol's files costs exactly one byte. With 65,536 bytes available, sol never uses more than 3,732 of them. The limit is not space. It is knowing what is worth storing.

Source: our measurement. We decoded every graded file with our own QPACK decoder, which reproduces every accepted file's headers exactly and agrees with the grader's verdict on all 127 graded files.


VIIIHow each model works

We read every transcript: 40 Sonnet 5 runs, 40 GPT-5.6-sol runs, 10 Opus 5 runs, and 16 runs of GPT-5.6-terra at its highest reasoning setting, cut short when we stopped that job.

Claude Sonnet 5 Claude Opus 5 GPT-5.6-sol
First accepted file about 1 minute, usually spelled-out or static-only median 169 seconds within its first 2 to 10 commands
Decision about the dynamic table Usually made before writing any code, and usually "no" Always yes, chosen with a knapsack Always yes, then tuned
How it picks entries Whole-trace frequency count, in the runs that tried Knapsack on estimated savings, then search on real size Knapsack or greedy search scored on the real encoded size, then local search
Evicting and refreshing entries Never In 2 of 10 runs Phase plans with eviction in 13 runs, Duplicate in 9
When it stops After the first accepted file, median 3.1 minutes When its search converges, median 7.0 minutes When time is nearly up, median 7.4 minutes

Source: our reading of every transcript, with counts checked by script against the tool calls.

Time used out of ten minutes Figure 4. Minutes each run worked before it stopped. Filled marks: the final file uses the dynamic table. Hollow marks: it does not. Every default-effort Sonnet run that stayed off the table and stopped on its own did so by minute 4.5. At xhigh, most Sonnet runs work until the limit.

The plateau is a plan

Sonnet 5 did not misread the goal. Its runs quote "fewer bytes" and then add the Huffman code for that reason. But most runs decided in their first minute that the dynamic table was not worth the risk, before writing a line of code:

"Good, ack_mode is immediate and RIC=0 approach will avoid dynamic table complexity entirely." (30 s)

16 of its 40 files are exactly the static-only size, byte for byte what ls-qpack produces with the dynamic table turned off, and what quiche and quic-go ship. 30 of the 31 runs that never touched the table stopped on their own by minute 4.5, with more than five minutes left. Only three of them counted how often headers repeat before deciding. The risk they feared was real but small: 8 of the 9 runs that tried the table got a rejection first, and 6 of those 8 fixed it within the time limit.

One rule broke almost everyone

Every rejected file we graded once the sandbox was offline, 17 files from Sonnet 5 (at both effort levels), Opus 5 and GPT-5.6-terra, fails for the same reason: a message declares a Required Insert Count larger than it needs. ls-qpack refuses it and nghttp3 accepts it, and the grader needs both. Most of these files set the count to "everything inserted so far" on every message. When we rewrote only that one number in each message, leaving the file the same length, both decoders accepted all 17 files at their original sizes. GPT-5.6-terra's rejected file on task 61725c26a104 would have scored 6,566 bytes, within 26 of the best file any model wrote.

The difference between models was not whether they hit this rule. GPT-5.6-sol hit it in 9 runs and recovered every time, often by running both shipped decoders, seeing one accept and one refuse, and cutting the file in half until the failing message appeared. GPT-5.6-terra hit it in 4 of its 12 finished runs and never recovered. In each case it had already written an accepted file and overwrote it with the new, rejected one. It never kept a backup, in any of its 16 runs.

Moments from the transcripts

Each block below is copied from the recorded transcript. Times are seconds from the start of the run. Codex does not record times, so its moments give the order instead.

Sonnet 5: the same table, the same size, opposite scores

Two Sonnet runs on task 34ec5ed416ee built the same 15-entry table and wrote a 4,989-byte file. Both were rejected at first. One went looking for the rule (334 s):

"Found it: RFC 9204 §2.1.2 requires RIC to be exactly zero when no dynamic table entries are referenced by a section (not just "high enough")."

It fixed the count, was accepted at 4,989 bytes, 110 below ls-qpack, and scored −4,989. The other run had backed up its accepted 7,584-byte file, overwritten it with the new one, and spent its last two minutes debugging. Its final message (548 s):

"Let me write a self-contained Python QPACK decoder to validate my encoding logic precisely, since the C decoders don't give detailed error info."

Time ran out with the rejected file in place and the backup never restored. It scored the floor, −53,501. The two graded files have a byte-identical table stream and differ in 48 of 50 message headers.

Sonnet 5: "Task complete" at the reference solution

On task 0d37295c057d, one run's first accepted file spelled every header out in full. At 61 seconds:

"The submission is admitted at 19898 bytes. This is a solid, simple solution using only literal field lines (no dynamic table references), which sidesteps all the risk/capacity/RIC complexity given the none ack mode. Task complete."

19,898 bytes is exactly the size of our reference solution, 74% more than static-only. The run had nine minutes left. On the same task, GPT-5.6-sol wrote 9,593 bytes in all four runs.

Sonnet 5: stopping to protect a working file

On task 28a6a3d688a4, three minutes in (179 s):

"Given the tight 10-minute wall-clock budget for the whole episode, I'll stop here rather than risk breaking a working, verified submission by adding dynamic-table complexity."

It scored 6,903, the static-only size. ls-qpack writes 4,152 bytes on this task, and GPT-5.6-sol 3,645 to 3,659. Another Sonnet run on a different task had written in its notes that the potential gain "was small relative to the implementation risk". It scored 11,896. On that task another Sonnet run used the table and scored 7,834, 34% smaller.

Sonnet 5: the fix, 13 seconds too late

On task 066bcff9a5e2, a run replaced its accepted 13,900-byte static-only file with a 7,452-byte file using the table, which would have beaten ls-qpack (8,329). ls-qpack rejected it. After nearly five minutes of debugging, at 587 seconds, its script found the answer:

# reference only 'connection' (abs=1); minimal RIC = 2, base=2
ric = 2

The limit arrived at 600 seconds, before a corrected file was written. It scored the floor, −84,609.

Sonnet 5: the textbook encoder, first try

On task 61725c26a104, a run counted repeats (355 s) and decided (357 s):

"Well within capacity. Now let's implement the full encoder with dynamic table support."

Its first file using the table was accepted at 449 seconds: 7,190 bytes, down from 23,193, and smaller than both real encoders. It stores each repeating header at first use and refers to it in the same message through a "post-base" index, the part of QPACK most runs never touched. Its three sibling runs on the same task scored 23,173, 30,529 and 30,529.

GPT-5.6-sol: never send the long value twice

Task 1c6d1b4fa91d has a 244-byte p3p header in 49 of its 50 responses and only a 1,024-byte table, so the table must be rebuilt as the connection changes. One run used QPACK's Duplicate instruction, which copies an existing entry to the front of the table for one byte, every time the table was rebuilt:

"Refreshing the retained p3p entry with QPACK's Duplicate instruction avoided retransmitting its long value during phase changes."

9,681 bytes became 9,144, the best file any model wrote for this task. The other three sol runs handled the same moment differently and finished at 9,544, 9,951 and 9,958.

GPT-5.6-sol: an anchor that costs nothing

On task 34ec5ed416ee, a run stored one header mainly so that other values of the same header name could refer to its name in one byte. It chose the cheapest real value it could find (code comment):

# Use the shortest expires value as the name anchor. Value bytes cancel between
# its insertion and its indexed occurrence, while its smaller entry leaves room
# for the otherwise marginal repeated content-length entry.
cands=[p for p in cands if p[0]!=b'expires']
expiry=(b'expires',b'0');cands.append(expiry)

4,694 bytes became 4,651. Three of the four sol runs and the Opus 5 run all arrived at this same 17-entry table independently.

GPT-5.6-sol: storing a header that is already in the static table

On task 61725c26a104:

"A concrete additional win is available: timing-allow-origin: * is static index 93, which costs two bytes per occurrence. Inserting it once makes each of its 37 occurrences a one-byte dynamic reference."

6,575 bytes became 6,541. The one sol run that kept this header in the static table finished last on the task, at 6,595.

GPT-5.6-sol: "preserving the best verified artifact"

The only sol file larger than ls-qpack's came from a run that had done better. On task 0e29694677f6 it had an accepted 7,666-byte file. Its last message, written just before the time limit:

"The cache now uses compact empty-value name anchors and selective duplication to refresh entries before FIFO eviction. The admitted size is 7,666 bytes; I'm doing a final policy search while preserving the best verified artifact."

Its search script wrote each candidate straight into the submission file. When time ran out, the last candidate, 9,382 bytes, was graded. We replayed its scripts and reproduced the graded file byte for byte.

GPT-5.6-sol: the right fix for the wrong reason

On task a449a438f44b, a run's file was rejected for the Required Insert Count rule. After about fifteen probing commands it reached the wrong conclusion:

"One decoder is stricter about older dynamic-table references, so I'm avoiding those references and using only the newest entry."

It then designed a file where all 50 messages are blocked, every reference uses a post-base index, and the single table update arrives last. Both decoders accepted it at 7,686 bytes, 117 more than the best sol run on the task.

Claude Opus 5: stopping at a rule

On task 0d37295c057d (256-byte table, none mode), Opus finished with 4 minutes left:

"Widening the candidate pool to 12 values per name reproduced the same optimum, so 9699 is the best this repertoire allows without eviction."

It chose the same four entries as all four sol runs (p3p, connection: keep-alive, x-powered-by: ASP.NET and the most common date). sol reached 9,593 with them by also using each stored entry's name for that header's other values, a 106-byte difference.

GPT-5.6-terra: "can be restored immediately if needed"

On task 066bcff9a5e2, terra had an accepted 13,908-byte file when it tried its first dynamic version:

"The admitted baseline can be restored immediately if needed."

It never was. Its later candidates each set one Required Insert Count for every message, ls-qpack rejected them, and the rejected file was graded when time ran out: the floor, −84,609. Fixing only that number in each message makes the same file 7,021 bytes, 4 bytes from sol's best.


IXMore thinking, same finish line

The Codex models ran at an explicit reasoning effort, high for GPT-5.6-sol. Claude Code's built-in default for Sonnet 5 is also high (we read it from the Claude Code program itself), so the main comparison is level. To test whether more thinking breaks the plateau, we ran Sonnet 5 again on the same 10 tasks at xhigh, two attempts each.

Sonnet 5, default (high) Sonnet 5, xhigh
Runs 40 20
Wrote an encoder that uses the dynamic table 9 (23%) 14 (70%)
Stopped on purpose at exactly the static-only size 15 3
Accepted by both decoders 38 (95%) 13 (65%)
Smaller than ls-qpack's encoder 6 (15%) 5 (25%)
Ran into the 10-minute limit 3 9
Median minutes worked 3.1 9.5
Median thinking tokens, runs that ended on their own 5,133 14,259

Source: our measurement and our reading of all 60 transcripts. Claude Code records thinking tokens only when a session ends normally, so runs stopped by the time limit are not in the last row.

More effort changed the plan. At xhigh, 14 of 20 runs wrote a dynamic-table encoder, and only 3 settled for static-only on purpose. It also produced the best Sonnet file on task 34ec5ed416ee: 4,690 bytes, below ls-qpack (5,099) and below every default-effort Sonnet run on that task.

It did not change the finish. The extra effort went into thinking: 61% of each run's time against 48% at the default, and a median of 329 seconds of thinking per run against 77. The first accepted file came later (median 122 seconds against 76), the first table design later still, and 9 of the 20 runs ran into the limit. Seven of them had written their new file straight over an accepted one, and the rejected file was graded. All seven were rejected for the Required Insert Count rule from Section VIII. With only that number corrected, all seven pass both decoders at the same size, and six would beat ls-qpack. That would have been 11 of 20 runs below ls-qpack, instead of 5.

Source: our re-grade of the corrected files with decoders built from the grader's pinned sources. Our re-grade of the original 20 files reproduces every recorded score.

More thinking made Sonnet more ambitious. It did not give it the habit that protects ambition: keep the accepted file, and test the new one next to it. Only one xhigh run restored its backup after a rejection, and it kept its 7,584-byte file while the seven that did not averaged a score of about minus 97,000.

Sonnet 5 xhigh: the rule, found in the last second

On task 066bcff9a5e2 a run replaced its accepted static-only file with an 8,068-byte dynamic one, which ls-qpack rejected. Nearly five minutes of debugging later, it wrote the smallest possible test: one table entry, and one message that declares that entry but uses only the static table (600 s):

fields_payload = bytes([0x02, 0x00]) + bytes([0xC0 | 25])  # RIC encoded=2, base delta 0
...
[600s] RESULT: exit:1

That is the exact rule. The run ended in the same second. Graded: −84,609. With the rule applied, its file would have scored −8,068, below ls-qpack's 8,329. The default-effort Sonnet run in Section VIII found the same fix on the same task 13 seconds before the limit.

Sonnet 5 xhigh: the answer, printed and misread

On task 48149e36c3f9 a run tested the decoder directly. For k from 1 to 34 it inserted k entries, referenced only the first one, and declared a Required Insert Count of k (400 s):

1 OK
2 FAIL
3 FAIL
...
34 FAIL

Only the message that declared exactly what it used passed. The run's next step was to read the RFC's section on relative indexing. It never returned to the count, and its 5,022-byte file, which would have been 712 bytes below ls-qpack, was graded at the floor.

Sonnet 5 xhigh: first table, twelve seconds before the limit

On task 1c6d1b4fa91d a run spent most of its time analysing savings and designing a table. It wrote its first dynamic file at 588 seconds, over its accepted 21,987-byte file. At 592 seconds:

rejected: E_DECODE_LSQPACK (decoder rejected input)

Its last command, at 599 seconds, printed the decoder's source code. Graded: −127,197. Corrected, its file is 9,819 bytes, 1,323 below ls-qpack.

Sonnet 5 xhigh: the name-only entry

On task 34ec5ed416ee a run stored all 16 repeating headers with an exact Required Insert Count and was accepted at 4,918 bytes. It then noticed that expires has 47 different values in the trace and no entry in the static table (name candidates: [(b'expires', 47)]), and stored the name alone, with an empty value, so every expires line could refer to it. 4,918 became 4,690 (499 s). GPT-5.6-sol and Opus 5 found the same idea on this task, anchoring the name with the value expires: 0 instead.


XDid anyone cheat?

On version 4, no. We searched all 3,484 commands the agents sent to their sandboxes in the version 4 runs, across every model and both Sonnet effort levels, for any attempt to read the grader, reach the network, install packages, or find a hidden encoder. There were none. A few runs searched the filesystem for ls-qpack source files while debugging and found only the shipped decoder drivers. Ten Codex runs (2 of GPT-5.6-sol's, 8 of GPT-5.6-terra's) checked whether a Python HPACK or QPACK library was installed. None was. Reading /app/bin/check and the decoder source is allowed and useful, and several runs did.

On version 1, yes, in the sense that matters. When the sandbox was online, three runs submitted the withheld ls-qpack encoder's own output, as described in Section V. The shortcut bought "match the production encoder", not better. The best version 1 results still came from agents writing their own encoders.


Conclusion

Table Press asks agents to do something production software often declines to do: use QPACK's dynamic table well, under strict rules, with two real decoders checking every byte.

Two of the agents did it better than the software. In ten minutes, Opus 5 and GPT-5.6-sol wrote encoders whose files are smaller than ls-qpack's in 10 of 10 and 39 of 40 runs. Their advantage came from one idea applied carefully: read the whole connection first, store each valuable header before its first use, and send each repeated value once.

Sonnet 5 had the skill but rarely used it. At its default effort it usually decided in its first minute to leave the dynamic table alone, and stopped at the size quiche and quic-go ship. At xhigh it tried far more often and ran out of time on a single detail. Across every model, one number, the Required Insert Count, decided every rejected file. And the runs that failed after already succeeding all lost the same way: they wrote an untested file over their best one.

Building the environment taught us as much as running it. A grader made of two real implementations cannot be argued with, and it grades only the part of the standard both of them agree on. Measuring the real encoders before any agent runs gives every score its meaning. And each version taught us something the one before it hid: agents with internet access will download the answer, and agents with no way to test will trust their own mistakes. Most of the grader's weak points were found by reviewers reading its code before release. The one that got through was found by the agents themselves.


References

Name a skill your model is missing.

We build the environments, evals and data to train and measure it.

Get in touch