Development¶
Start with the design principles before changing arithmetic or storage. They define the exactness and ownership guarantees your changes must preserve. The sections below cover setup, focused checks, and the Mojo practices used in the implementation.
Environment¶
Set up your development environment using Pixi as described in the
installation guide. Run the commands below from
the project directory containing pixi.toml.
- The default Pixi environment pins Mojo 1.1.0.
- The
comparisonenvironment adds Python reference libraries (mpmath,gmpy2with MPFR and MPC, andpython-flintwith Arb) and requires Cargo to build Rug. - The
docsenvironment provides MkDocs and Pygments.
Library programs and documentation builds run on Linux and macOS. The
test, docs-test, and report runners require Linux: their compiler
memory guards use systemd or prlimit to enforce the 12 GB cap.
Repository layout¶
| Path | Purpose and Contents |
|---|---|
src/apn_mojo/ |
Core library source code, organized by numeric family plus batch/ and common/ (see the Module map). |
tests/test_apn.mojo |
Functional test suite: cross-feature test scenarios using std.testing assertions. |
tests/test_batch_refinements.mojo |
Shared batch contracts: borrowed storage, retained views, failure cleanup, broadcasting policies and budgeted interchange. |
tests/test_batch_construction.mojo, tests/test_batch_json.mojo, tests/test_batch_promotions.mojo, tests/test_mask_layout.mojo |
Focused construction, interchange, operator-type and Mask layout contracts. |
tests/test_lift_inputs.mojo |
Mixed numeric inputs, Ball contexts and enclosures, fold seeding, scalar broadcasting, and failure cleanup. |
tests/test_vmap_results.mojo |
Tuple and Optional result borrowing, partial-result cleanup, and nested Ball outputs. |
tests/test_execution.mojo |
Shared elementwise execution: chunk readers, failure ordering, cleanup and frontend coordinates. |
tests/test_ball.mojo |
Ball suite: inclusion of exact results, tightness, kinds, comparisons, sets, text and hashes. |
tests/test_float_utilities.mojo |
Float utilities suite, with shortest decimals checked against tests/fixtures/shortest_decimal.txt (written from Python's repr by tests/fixtures/shortest_decimal.py), and decimal conversions at large exponents against tests/fixtures/large_exponent_decimal.txt (written from Arb's balls by tests/fixtures/large_exponent_decimal.py). |
tests/test_interning.mojo |
Interning suite: representation identity and order, FloatKey and ComplexKey, and stable_hash vectors. |
tests/test_number_theory.mojo |
Number-theory suite, checked against tests/fixtures/number_theory.txt (written by gmpy2 through tests/fixtures/number_theory.py). |
tests/run.py |
Bounded build and test runner managing memory caps and timeouts; it builds each tests/test_*.mojo as its own executable, and --suite NAME runs one. |
tests/benchmarks/ |
The consolidated report: its case catalog, the workers for apn_mojo, Rug and python-flint, and the differential check. |
scripts/ |
Tooling for documentation audits, reference generation, and validation; generate_function_tables.py (comparison environment) writes src/apn_mojo/ball/_tables.mojo, certified by mpmath's interval arithmetic. |
docs/content/ |
Markdown source files for this documentation site (configured in mkdocs.yml). |
docs/examples/, docs/examples.json |
Maintained code examples and their verified standard outputs. |
docs/theme/, docs/hooks.py |
Custom MkDocs site theme and build hooks. |
Temporary build artifacts and micro-benchmarks belong in build/ or .cache/,
both of which are git-ignored.
Verification gates¶
Run the suites affected by the changed components. Use focused AddressSanitizer checks when ownership changes; reserve full sweeps for an explicit request. The commands below include both focused and full-suite forms:
| Command | What it checks |
|---|---|
pixi run --locked test --timeout 1500 |
All tests/test_*.mojo suites, compiled separately at -O3 |
pixi run --locked test --suite ball --timeout 1500 |
One suite; replace ball with another suite name |
pixi run --locked python3 tests/run.py --asan --timeout 3400 |
The suites under AddressSanitizer at -O1 |
pixi run --locked report --check-only --area <area> |
Results against GMP, MPFR and MPC (Rug) and FLINT and Arb (python-flint), once per case |
pixi run --locked -e docs docs-check |
Documentation inventory, strict site build, links, anchors and search |
pixi run --locked python3 scripts/api_docs.py --check |
docs/api.json against docstrings generated by mojo doc |
pixi run --locked docs-test |
Every maintained example, compiled with --Werror and checked for expected output |
Build times depend on the suite, compiler cache and machine. A cached test executable starts quickly; a cold build can take several minutes.
The harness caches binaries by source hash and compiler identity. If neither
changes, the next run executes the cached binary immediately. AddressSanitizer
builds use -O1 because Mojo 1.1.0's runner can crash with -O0 --sanitize=address.
Before a commit, measure only the speedup of the functions the change is
about, on the scalar benchmark or by cycle counts, and run nothing else: the
suites, AddressSanitizer, compare, the differential check and the corpus
are not gates before a commit. Changes affecting the public API still update
documentation, docstrings and verified examples.
Iterating on a kernel¶
Choose the suites that exercise the area you changed. Each is a separate executable, typically taking about a minute to compile and seconds to run. For elementary functions, use:
- Correctness:
pixi run --locked test --suite elementary(every rounding mode against MPFR's fixtures, ball inclusion and Arb's tightness), and--suite constantsafter touchingball/_tables.mojo. - Speed:
pixi run --locked report --area float_function --area ball_functiontimes our functions against MPFR and Arb; add--baseline HEADfor a before-and-after column. For where the time goes, run the worker'scallscommand under callgrind (see Benchmarks).
Before a commit, measure the changed functions' speedup and nothing else (see above).
When submitting changes:
- Arithmetic modifications require passing the affected functional suites.
- Changes to pointer manipulation or memory ownership require focused
AddressSanitizer checks of the affected ownership paths
(
--asan --suite NAME). - Numerical changes require running
compareagainst reference libraries. - Changes affecting the public API require updating documentation, docstrings, and verified examples.
Mojo 1.1 practices¶
- Lambdas: Use the syntax
lambda (x: T, y: T) raises -> R: expr. Lambdas that capture closures cannot be passed as compile-time parameters tovmaporlift. Because lambda bodies must be single expressions, in-place updates are written asacc.__iadd__(x). - Struct return costs: Returning large structs (such as multi-limb results or
Optionaltypes) across non-inlined function boundaries incurs a 15–20 ns overhead. Inline small fast paths, and prefer sentinel values overOptionalwrappers in internal kernels. - Reference counting: Copying heap-allocated structs triggers atomic reference count increments and decrements. Pass arguments by borrowing (
ref), and preferrebind_varoverrebindfor owned variables. Note that evaluatingx and yon arbitrary-precision types creates a full copy. - Tuples of heap values: Unpacking a returned tuple (
var q, r = f()) copies each element, and the tuple's release undoes the copy: two atomic reference-count updates per heap element. Keep the tuple and move parts out with a swap (_takenininteger/value.mojo), or return it whole (return parts^). A conditional expression (a if c else b) with a heap result also copies it once more; use statements. Optionalof a large record: Building anOptional[ComplexContext](orOptional[FloatFormat]) can compile to an out-of-line initializer that stores the payload a byte at a time, hundreds of instructions. Argument records use a plain slot instead: the value, a presence flag and a placeholder (_FormatSlot,_ComplexContextArgument). Taking a large value out of anOptionalhas the same problem:Optional.take()is inlined only while few call sites use it, and otherwise moves the payload byte by byte, about 400 instructions for aBall. A new call site elsewhere can push it out of line everywhere: adding the derived Ball functions made the 53-bit Ballsin12% slower that way. Where the Optional is dropped next, copy the value out with.value()instead; check hot paths' instruction counts after addingtake()calls.- Fast paths beside slow ones: The compiler hoists cheap computations out of the branches of an inlined function, so an inlined fast path can pay for its alternatives' setup. Keep the alternatives behind one
@no_inlinecall, and choose an operation at compile time (_float_operation_of[operation]) rather than at run time in one function. A large function with many live values spills: a small hot routine can run faster out of line, with its operands in registers. Measure both. - Constant tables: A
comptimeArrayorSIMDtable indexed at run time is copied to the stack on every call (anArrayneedsmaterialize). A string literal is static data: encode entries as printable characters and read them throughunsafe_ptr()(see_RECIPROCAL_SEEDS). - Native wide integers: Use native
UInt128andUInt256types along withcount_leading_zerosandcount_trailing_zeros. Variable 256-bit divisions lower to bit-serial routines; for inner loops, prefer reciprocal multiplication. - Compiler warnings: Compiling with
--Werrorflags unused variables; declare variables only in the code paths where they are initialized. - Trait declarations: Traits added via
__extensionare invisible to modules compiled before the extension is imported. All library types must declare trait conformances directly on their definitions. - Constraints on types: A
whereclause cannot evaluate a call to a library function, even a compile-time one, so overload resolution cannot search a type list. Give the types a marker trait and testconforms_to(T, Marker); acomptime assertin a body may call functions. - Unreachable code:
abort()returnsNever, and code aftercomptime assert Falsecounts as unreachable; under--Werrorboth are errors. Assert the condition itself (comptime assert T == A or T == B) and write no code afterabort(). - Owned generic values: A
varargument whose type is only known to beIterableOwnedcannot be abandoned or rebound; consume it withiter(value^), whose iterator is movable. - Converting from
None: The literalNonehas typetype_of(None), notNoneType; an implicit constructor that accepts the literal takestype_of(None). One takingNoneTypeis not found for it. - Float64 elementary functions:
std.math'sexp2,log2and**onFloat64are good to about \(2^{-31}\) relatively (measured over 200,000 arguments),cbrtto about \(2^{-42}\), andexp2is not exact at integers. Build powers of two from their bits (bitcast), refine an estimate with a Newton step inFloat64or on an exact residual, and never let a result's correctness rest on these functions' accuracy. - Stack scratch behind a pointer: Once a stack array is written through a pointer cast to
MutUntrackedOrigin, read and write it only through that pointer, and keep the array alive until the last access (_ = scratch^), as_long_dividedoes. A native quotient that wrote its limbs through cast pointers and read them back by array name got the arrays' initial zero fill: the compiler did not connect the two.
Performance discipline¶
Before optimizing, estimate how much time the change could save. Measure it against the same workload and an unchanged baseline, then record the method and results in the relevant architecture chapter. The benchmark guide describes the measurement protocol. Keep the change only when its measured gain justifies the added complexity.