AI-Generated Java Benchmarks for Chronicle-FIX
Chronicle Software CEO Peter Lawrey had AI build a JLBH benchmark for Chronicle-FIX, then corrected its assumptions along the way; from percentile choices to over-engineered code. Round-trip latency landed at 2.4–3.7 microseconds. The real value wasn't the code produced, but what the exercise revealed about the system and the AI itself.
September 7th, 2026
Written by Peter Lawrey, CEO of Chronicle Software.
The value of an AI-assisted experiment is not just the code it produces. It is the assumptions it exposes, the evidence it creates, and the understanding your team retains.
While I am sceptical of using AI for releasing code, it has plenty of uses that previously weren’t practical, such as determining how easy your software is to use. If an AI can “figure it out” with a few hints, then you are on the right track. For me, the value of AI is what you learn using it.
Chronicle-FIX is our low-latency FIX engine, built to decode, encode, and persist FIX messages with predictable microsecond-level latency for trading systems that can't tolerate jitter. It's part of the broader Chronicle stack, sitting on top of Chronicle-Queue for persistence, and is designed so that the messaging layer stays out of the way of your business logic rather than becoming a bottleneck itself.
What AI Does Well and What It Doesn’t
Claude and Codex are effective for producing idiomatic code; for low-latency code, it needs a significant body of example code. In this case, it was able to utilise sample code for benchmarks. If it were used to write business logic, it would need mostly complete code examples, and then it could write variations on them. If you were starting, it would be better to either a) get it to write something functionally correct with the expectation you would rewrite it again manually, or b) write the code yourself and use AI to assist you in improving it.
AI gives us another way to investigate assumptions before asking other developers to spend time on them. The opportunity is to make more exploratory integration attempts practical, not to replace human feedback or engineering judgement.
The AI Benchmark Trial
I gave Codex (GPT-5.5) the task of writing a JLBH benchmark for Chronicle-FIX from documentation and sample code, testing the round-trip latency of W -> D and D -> 8 messages. The throughput is 50K/s each way. The W market data message is ~690 bytes, and the D new order signal and ‘8’ execution reports are a small ~185 bytes. The test is run for 15 minutes each. I verified the benchmark was written but avoided hand-tuning it; then I asked it to trial different GC options, expecting they wouldn’t make much difference, since the application is low GC; however, there might still be some difference. The system is using Java 25.0.3 on a Ryzen 9 9955HX3D, 32 hardware threads, 12 isolated, with 64 GiB of RAM in a laptop running Ubuntu 24.04.4 LTS, kernel 7.0.0-30-generic.
A significant difference between JMH and JLBH benchmark harnesses is that JLBH supports many concurrent asynchronous in-flight actions, whereas JMH tests one action at a time. We chose JLBH to measure completed asynchronous message exchanges at a configured arrival rate, rather than to focus on the execution time of an isolated method.
Some points that needed correcting
The AI agent picked p99 as representative, but I favoured p99.9, which is harder to achieve consistently. In particular, the choice of filesystem made a significant difference at p99.9 and required five 60-second runs. This needed to be explained more than once, as it reverted to running a single 1-second test, which is fine for showing it works, but not for producing a meaningful result.
I asked it to write JLBH tests because they covered asynchronous testing better than JMH.
Tests that wrote more data than fit in the disk cache, i.e., beyond the main memory, impacted the high-percentile latencies, and it needed to consider that as a factor. The agent suggested it was due to the choice of GC configuration, even though no minor collections occurred.
The agent was adding messages without checking the quality of the output. I had to tell it when printing a log message, “if a field is useful, there should be at least one test that expects it” This helped reduce noise.
Often, the agent will over-engineer a solution when a much simpler solution will do the same thing. Boilerplate that is never exercised, overly complex implementations, adding overhead such as logging for low-latency code which doesn’t help, adding code which isn’t justified by any test that requires it, adding mock tests that don’t actually exercise release code, and adding tests that don’t improve coverage or drive a release change.
However, there are more failures in how agents operate than deficiencies in the documentation, examples, or APIs for the application.
NOTE: You can ask the AI to reflect on what went well and where there were sources of friction. i.e., even if it got you the outcome you wanted in the end, could it have done so more efficiently?
The Benchmark Results
The half-round-trip time (RTT/2) was between 2.4 and 3.7 microseconds (< 0.004 milliseconds). For the recommended Parallel GC, the 99.999% was ~11 microseconds. This is a great starting point, considering this doesn’t involve any hand-tuning of the code. The AI can do this because there are many relevant examples to draw on and the code is relatively simple.
When it comes to the crucial business logic, you would either need to assume all the AI-generated code would be rewritten once you have it working, or write most of it yourself and use AI to assist. Having to rewrite AI code might seem like a waste of time; however, it’s much easier to write the release code from scratch once you have comprehensive test coverage and documentation of what you need it to do.
NOTE: The latencies are in microseconds on a logarithmic scale as they vary significantly at the tail. This is the full Round Trip Time, including encoding/decoding/persistence. Chronicle-FIX reads messages to be sent via a Chronicle-Queue, so they are written there first


Caveats
As with all synthetic benchmarks, there are caveats. The most important thing is that there is no business logic, and you can expect that to be an order of magnitude more complex and longer.
Another key assumption is that timings are regularly spaced. However, in reality, there are bursts of activity that disproportionately affect your profit and loss. You can lose the most money when the market is most volatile, i.e., the most active. While you might not need a sustained 50K/s throughput, you are likely to care about a burst of 50 in one millisecond.
Notes
The vertical scale is logarithmic in microseconds.
In the 15-minute run, only a few of the GC options resulted in any GC. Most didn’t.
Within the margin of error, up to the 99.7%ile, different GCs were essentially the same.
Where GC selection varied, it was most likely when the GC woke up to see if anything needed to be done and found not much.
The parallel collector is still the best when you don’t expect to collect.
The absence of collections doesn’t mean no allocations; they were low enough as not to trigger a collection during the benchmark.
The max heap size was 4000 MB and the Eden size was 3000 MB.
"G1+COH" includes Compact Object Headers.
In general, large pages didn’t help. This was clearer in other tests not reported here.
Codex (GPT-5.5, xhigh effort) wrote the benchmark, and a second Codex reviewed the work, as did Claude.
Conclusion
There are sufficient resources for AI to write a benchmark that is close to what you would write yourself, and it can be a good starting point. However, you should expect to rewrite the release code once it's working, and to spend significantly more time on documentation and testing than you would without AI in your workflow.
I am comfortable using disposable code to explore a problem. I apply a different standard to code we intend to support in production. That may require a rewrite, a sufficiently deep review, or both. The important transition is from “the AI produced this” to “we understand this, can justify its choices and accept responsibility for maintaining it”. A useful AI experiment can leave you with less code and more understanding.
Before AI, we would have to guess and rely on user feedback to determine whether the software was easy to use. Now we can use AI to trial new functionality, and if it can figure it out with a few hints, then we are on the right track. The value of AI is what you learn using it, not the code it produces.
In terms of garbage collectors, it’s not so important as the code is so low-garbage that it doesn’t rely on it, and if anything, you want to avoid disturbing the application when it’s running. It is still needed to avoid an out-of-memory error if something goes wrong.