CKKS vs. GL Benchmark: From 38 Minutes to 12.67 Seconds
At a glance (TL;DR)
- Ciphertext-ciphertext matrix multiplication (CCMM) is the core operation of encrypted AI inference and one of its main bottlenecks. We compared two implementations of it on the same machine: the CKKS implementation (built on OpenFHE) that won the matrix multiplication challenge on FHERMA, an open international competition, and the commercial implementation of the GL scheme in the DESILO FHE Library. The task was 256 pairs of 64×64 real-valued matrix multiplications.
- On a single CPU core, GL was up to about 229× faster than the CKKS implementation. In the higher-level setting (level 16), 256 pairs took the CKKS implementation about 6 hours 58 minutes and GL 109.44 seconds. At the minimum level, it was about 38 minutes 5 seconds against 12.67 seconds, a gap of about 180×. CKKS times are the mean of 10 runs for one pair, scaled by 256.
- One matrix multiplication consumes two levels in CKKS and one in GL. A level is roughly the remaining battery charge of a homomorphic ciphertext. The deeper the circuit, the more this difference accumulates into a difference in the number of bootstrapping operations, the most expensive operation in FHE.
- GL's advantage grows under three conditions: large AI models with many matrix multiplications and deep circuits, workloads that multiply many small matrices at once, and environments where a GPU is available.
Contents
- Why matrix multiplication between ciphertexts matters
- The baseline: FHERMA's winning CKKS implementation
- Measurement setup
- Level consumption
- Batching
- Result ①: minimum-level setting
- Result ②: high-level setting
- Summary
FAQ · References · Appendix (experimental setup)
1. Why matrix multiplication between ciphertexts matters
Most of the computation an AI model performs is matrix multiplication. To run AI on data that stays encrypted, those matrix multiplications also have to happen on ciphertexts.
There are two kinds of matrix multiplication on encrypted data: one where only one side is encrypted, and one where both sides are. When one side is plaintext, as with model weights, the cost is relatively low. When both sides are encrypted, the cost rises sharply. In this post, we call a matrix multiplication where both sides are encrypted CCMM (ciphertext-ciphertext matrix multiplication).
CCMM cannot be avoided. Attention in a transformer multiplies the query matrix Q by the transpose of the key matrix, Kᵀ. Both come from the user's input, so once the input is encrypted, both are ciphertexts. A large language model repeats this multiplication in every layer. The speed of CCMM therefore decides whether encrypted AI inference can run fast enough to be practical.
In "Understanding the GL Scheme," we used the figures from the paper to show how GL addresses this problem at the level of scheme design. This post is the next step: a direct comparison between GL and a third-party CKKS implementation that won an open challenge, on the same problem and the same machine.
2. The baseline: FHERMA's winning CKKS implementation
FHERMA is an international FHE challenge platform run jointly by Fair Math and the OpenFHE team. Participants compete on implementation performance for the same problem, and winning solutions are incorporated into open-source libraries. IBM researchers have also set problems on the platform.
Our baseline is the winning implementation of the platform's encrypted matrix multiplication challenge. It was written by a researcher at Graz University of Technology (TU Graz) in Austria and runs on OpenFHE, a widely used open-source FHE library. The original code is public in FHERMA's challenge repository (github.com/Fherma-challenges/matrix-mult), and the challenge results are on the FHERMA challenge page (login required).
For this measurement we used the submitted code from that repository. We kept the matrix multiplication algorithm as it is and only moved the preparation of the mask plaintexts outside the timed section. Repeated runs and timing were handled by a separate benchmark driver (see the appendix). The ring dimension is 2¹⁵, the same as in the submission, and we ran it with OpenFHE's 128-bit security setting (HEStd_128_classic) enabled. GL was set to the same ring dimension and the same security level.
3. Measurement setup
| Item | Details |
|---|---|
| Task | 256 pairs of ciphertext-ciphertext multiplications of 64×64 real-valued matrices |
| CKKS implementation | One run processes one pair. 256 pairs = mean of 10 runs for one pair × 256 |
| GL implementation | 256 pairs processed in one batched operation, using a 3D ciphertext of shape (256, 64, 64) |
| CPU | AMD Ryzen Threadripper PRO 3975WX, single core |
| Security level | 128-bit for both schemes, same ring dimension of 2¹⁵ |
| Measurement | Wall-clock time of the matrix multiplication itself, mean of 10 runs (encoding, encryption/decryption and key generation excluded) |
We measured under two conditions. One sets both implementations to level 2, the minimum starting level the CKKS implementation needs (Section 6). The other starts both at level 16 to mimic a point in the middle of a deep computation pipeline (Section 7).
4. Level consumption
Two pieces of background help in reading the results. The first is levels.
A homomorphic ciphertext has a level. It is standard terminology in schemes such as BGV, BFV, CKKS and GL, and it indicates how much room is left for further multiplications. It is related to the noise budget described in "Understanding Fully Homomorphic Encryption (FHE)," but it is not the same concept. In the CKKS and GL implementations compared here, the level drops with each multiplication, and when it runs out it has to be recharged through bootstrapping. It can be likened to battery charge. Bootstrapping is the most expensive operation in FHE.
How many levels one matrix multiplication consumes depends on the scheme. The winning CKKS implementation consumes two levels per matrix multiplication, because its algorithm needs one step of plaintext multiplication for masking and one step of ciphertext multiplication (the winner's published paper also states a multiplicative depth of 2). In GL, matrix multiplication is a native operation of the scheme, so it consumes only one.
This difference has three consequences.
- If only matrix multiplications are chained, leaving out other operations, the same number of levels supports twice as many of them.
- More matrix multiplications fit between bootstraps. How much bootstrapping actually drops depends on the level consumption of other operations and on the overall circuit.
- The same work can start from smaller parameters.
The more matrix multiplications stack up in a circuit, as in transformer inference, the larger this difference becomes.
5. Batching
The second is how many matrix multiplications can be processed at once.
The CKKS reference implementation used here processes one matrix multiplication pair per run, so we obtained the time for 256 pairs by multiplying the time for one pair by 256. CKKS can also pack several matrices into one ciphertext and process them together, but this reference implementation is not written that way. As described in "Understanding the GL Scheme," a GL ciphertext has a three-dimensional structure of batch × rows × columns from the start, so a single (256, 64, 64) tensor handles 256 matrix multiplications in one batched operation.
This batch structure matches the shape of AI workloads. Multi-head attention in a transformer performs a separate matrix multiplication for each batch and head. BERT-base and BERT-large have a head dimension of 64, so when the input length is also 64, the product of Q and Kᵀ is a 64×64 matrix multiplication. The (256, 64, 64) task is a simplified example of this kind of batched computation.
6. Result ①: minimum-level setting
First, the setting that puts both implementations at level 2, the minimum starting level the CKKS implementation needs, because it uses two levels per matrix multiplication. GL uses only one of them.
| Basis: 256 pairs | CKKS, 1 pair (measured) | CKKS, 256 pairs (scaled) | GL, 256 pairs (measured in one batch) |
|---|---|---|---|
| CPU (Threadripper PRO 3975WX, 1 core) | 8.93 s | 2,285.31 s ≈ 38 min 5 s | 12.67 s |
For 256 pairs, that is 2,285.31 seconds against 12.67 seconds, or 8.93 seconds against about 0.05 seconds per pair (for GL, the batch time divided by 256): a gap of about 180× on the same single CPU core.
The CKKS figure of 2,285.31 seconds is the mean of 10 runs for one pair, multiplied by 256. The figures in the table are rounded to two decimal places; the scaled totals and ratios were calculated from the unrounded measurements. Because the reference implementation processes only one pair at a time, we counted 256 pairs as 256 independent runs. Cache and memory effects that could arise when actually processing 256 pairs in a row are not reflected. A CKKS implementation written to pack several pairs into one ciphertext could come in below this scaled figure.
7. Result ②: high-level setting
The second measurement starts both implementations at level 16. In the middle of a deep circuit, as in neural network inference, a ciphertext has to keep a high level for the operations that remain, and the higher the level, the more each operation costs. This setting mimics that situation.
| Basis: 256 pairs | CKKS, 1 pair (measured) | CKKS, 256 pairs (scaled) | GL, 256 pairs (measured in one batch) |
|---|---|---|---|
| CPU (Threadripper PRO 3975WX, 1 core) | 98.04 s | 25,097.97 s ≈ 6 h 58 min | 109.44 s |
In absolute terms, GL's 109.44 seconds is larger than CKKS's 98.04 seconds, but 98.04 seconds is the time for one pair and 109.44 seconds is the time for 256. On a 256-pair basis, that is 25,097.97 seconds against 109.44 seconds, or 98.04 seconds against about 0.43 seconds per pair (for GL, the batch time divided by 256): about 229×.
The gap is wider than the 180× in the minimum-level setting. In this measurement, raising the level increased CKKS's time more than GL's (about 11× for CKKS, about 8.6× for GL).
8. Summary
In this measurement, the gap between FHERMA's winning CKKS implementation and GL was about 180× under the same conditions on a single CPU core, and it widened to about 229× in the higher-level setting.
The gap grows as more of the following conditions apply.
- Workloads dominated by matrix multiplication. AI inference and training fall into this category.
- Workloads that multiply many small matrices at once. Multi-head attention and batched serving are typical examples, and GL's 3D ciphertext holds this shape directly.
- Deep circuits. The fewer levels matrix multiplication consumes, the less bootstrapping is needed.
- Environments where a GPU is available. Because GL reduces the computation to plaintext matrix multiplication, it benefits fully from GPU acceleration (this measurement was on a single CPU core).
The reason lies in the scheme design. GL holds matrices as a 3D data type that includes the batch dimension, reduces encrypted matrix multiplication to plaintext matrix multiplication, and spends only one level on that multiplication. DESILO implements this in a commercial library and provides a GPU backend.
Where THOR reduced encrypted BERT inference time by changing the algorithms on top of existing CKKS, GL lowers the cost of matrix multiplication by changing the scheme itself.
Frequently asked questions (FAQ)
Q1. Does this result mean GL replaces CKKS?
No. CKKS is the standard fourth-generation FHE scheme that first made real-number arithmetic practical, and it remains a good choice for many tasks, such as element-wise operations. What this comparison shows is that, for one specific operation, matrix multiplication, GL's matrix-oriented structure and batching produced a large performance difference. The goal is to add GL as an option for workloads with repeated large-scale matrix operations, and the DESILO FHE Library supports both CKKS and GL.
Q2. Why compare against the challenge-winning implementation?
To avoid any suspicion that we compared against a weak baseline of our own making. FHERMA's winning implementation is a CKKS matrix multiplication implementation that went through open competition, and its algorithm and code are public on GitHub (Fherma-challenges/matrix-mult) for anyone to check.
Q3. Are the two schemes at the same security level?
Yes. Both settings were configured for 128-bit security at a ring dimension of 2¹⁵, and the CKKS side ran with OpenFHE's HEStd_128_classic setting enabled. Parameter details are in the appendix.
References
- FHERMA encrypted matrix multiplication challenge page (login required to view results): https://fherma.io/kernels/matrix-multiplication/challenges/matrix-multiplication-2024
- Repository of the winning CKKS implementation (the code used in this measurement): https://github.com/Fherma-challenges/matrix-mult
- Aikata, Sujoy Sinha Roy, "Secure and Efficient Outsourced Matrix Multiplication with Homomorphic Encryption," IACR ePrint 2024/1730 (paper describing the winning implementation; states a multiplicative depth of 2): https://eprint.iacr.org/2024/1730
- Craig Gentry, Yongwoo Lee, "Fully Homomorphic Encryption for Matrix Arithmetic," CRYPTO 2026 (IACR ePrint 2025/1935): https://eprint.iacr.org/2025/1935
- Eric Crockett, Craig Gentry, Hyojun Kim, Yeongmin Lee, Yongwoo Lee, "Efficient Bootstrapping in Fully Homomorphic Encryption for Matrix Arithmetic," CRYPTO 2026 (IACR ePrint 2026/956): https://eprint.iacr.org/2026/956
- Related post: Understanding the GL Scheme: Homomorphic Encryption Redesigned from Scratch for Matrices
- Related post: Understanding THOR: Running LLM Inference on Ciphertext
Appendix: experimental setup
The reference CKKS implementation uses the matrix multiplication algorithm from the FHERMA winning submission (github.com/Fherma-challenges/matrix-mult) as it is, with only the mask plaintext generation and encoding moved out of the computation function. Repeated runs and timing were handled by a separate benchmark driver, which completes mask plaintext generation, encoding, encryption and key generation outside the timer and then times only the matrix multiplication. Fair Math publishes the winning solutions under the Apache-2.0 license (fairmath/components).
The CKKS parameters are as follows.
- Ring dimension 2¹⁵, batch size 2·64·64
- Scaling modulus 40 bits, first modulus 40 bits
- Security level flag: HEStd_128_classic (both settings)
- Multiplicative depth: 2 (minimum-level setting) / 16 (high-level setting)
- Key switching HYBRID, scaling technique FLEXIBLEAUTO, 2 special primes
- Single core, mean of 10 runs of steady_clock wall-clock time
On the GL side, we created a GLEngine with shape=(256, 64, 64) in the DESILO FHE Library 1.16 (public release) and timed ciphertext-ciphertext matrix_multiply, with the ring dimension and security level matched to the CKKS side.
The measurements were run by Yeongmin Lee of the DESILO library team.