Benchmark results¶
Disturbed run. Other work shared the CPU. After re-timing, 1 sentinel checkpoint remained slow, so some times below are too long. Before quoting them, regenerate the page on a quiet machine with
pixi run --locked report --publish.
This report compares APN Mojo with GMP, MPFR and MPC through Rug, and with FLINT and Arb through python-flint. For the workloads, measurement method, and commands, see Benchmarks.
How these were measured¶
- Recorded: 2026-10-06T21:15:38-04:00. This is a snapshot of the stated build, not a measurement of the current checkout.
- Machine: Intel(R) Core(TM) i9-9980HK CPU @ 2.40GHz, CPU 10 only, governor
powersave, turbo on; Linux 7.0.12-1-t2-noble. - Software: apn_mojo 1.0.0, Mojo 1.1.0 (8189361e); GMP 6.3.0, MPFR 4.2.2, MPC 1.4.1 through Rug 1.30.0; FLINT 3.6.0 through python-flint 0.9.0 on Python 3.13.15.
- Correctness: each case ran once on every backend before timing. Integer, Rational, Float and Complex results must be equal; Float and Complex results are correctly rounded in both libraries. Balls must overlap. 555 of 555 cases agree.
- Timing: following the timeit protocol, calls increase through 1, 2, 5, 10, ... until a sample takes 10 ms, then 7 samples are collected. The full report gives medians per call; ± is half the interquartile range relative to the median. Inputs are prepared before timing, and each timed call creates and destroys its result without checks. Backends take turns case by case.
- python-flint overhead: an empty call in the same loop takes 19.7 ns with one argument and 22.8 ns with two. Comparisons with FLINT/Arb subtract it. † marks a call less than four times the overhead; its net time is uncertain.
- Drift: a sentinel case, Rug's 1024-bit Integer product, was timed every 25 cases; its median varied by 6.5% over the run's 3 min. 2 checkpoints ran more than 15% slow, a sign of other work on the CPU; the 75 cases next to them were timed again. Checkpoints still slow afterwards: 1.
Summary¶
Each ratio is a geometric mean of APN Mojo / reference times for matching cases whose results agree. 0.5 means half the time; 2 means twice the time. Averages span the tested sizes and operations, so they are not a prediction for every call. A dash means no comparison; Text conversion has FLINT results only for Integers.
| Area | Cases | APN / Rug | APN / FLINT or Arb |
|---|---|---|---|
| Integer arithmetic | 52 | 1.83 | 0.94 |
| Integer number theory | 42 | 2.50 | 2.40 |
| Rational arithmetic | 18 | 1.08 | 0.72 |
| Float arithmetic | 40 | 2.43 | – |
| Float functions | 102 | 1.21 | – |
| Complex arithmetic | 21 | 1.89 | – |
| Complex functions | 45 | 0.80 | – |
| Ball arithmetic | 20 | – | 1.82 |
| Ball functions | 85 | – | 2.94 |
| ComplexBall | 30 | – | 2.41 |
| Text conversion | 12 | 4.15 | 2.75 |
| Batches of 1000 | 18 | 1.92 | – |
| Hypergeometric functions | 70 | – | 1.31 |
Rug supplies GMP for exact arithmetic, MPFR for Float, and MPC for Complex. The other column uses FLINT for exact arithmetic and Arb for balls. Ball overlap checks consistency of enclosures; it does not mean their radii are equal.
Detailed results¶
Download the full report for every operation and precision, absolute timings, sample variation, ball radii, and batch timings on multiple CPUs. It is a Markdown file you can search or render locally.
Short python-flint calls marked † in that report are sensitive to overhead subtraction; their ratios also contribute to the averages above. Check those rows before interpreting small differences. All summary ratios use the single-CPU measurements, including batches.
To measure a particular workload, follow the benchmark guide.
Local runs retain the full report and the individual samples in build/report/.