Batches and execution¶
Batch[T] adds shape and layout to scalar values. Views share their storage,
and common executors handle mapping and reductions. This keeps the numerical
rules in scalar functions while the batch layer manages traversal and work
distribution. The Batch reference defines public
behavior; private layouts and scheduling thresholds may change.
Batch records¶
A batch holds a shared list of scalar elements, a shared layout descriptor (shape, strides and offsets), and default format metadata where needed. Rank and shape are runtime properties, so the same type represents rank-0 values, vectors, matrices and higher-rank arrays.
Copying a batch shares its storage. An update installs a new value for the
destination without changing earlier copies or views. Each stored element is
a scalar handle: Integer, Rational, Float, Complex or Ball. Mapped
functions can borrow these handles directly and move results into output lists.
Float and Complex elements retain their own formats and special values.
Slicing, reversal, transposition, broadcasting and contiguous reshaping share
element storage and describe the new layout. Creating a view may allocate a
layout descriptor; it does not copy every scalar. Operations walk logical
indices through that layout, gathering values only where their implementation
needs a contiguous buffer. Internal interfaces that erase the element type use
_ArrayData, which carries a family tag, element storage, layout and formats.
Construction from iterators¶
batch/_iterator.mojo consumes an iterator once. For each element, it calls
the family's converter, checks the next list size, runs any failure checkpoint,
and moves the converted value into the list. The tensor becomes available
only after the iterator finishes. An exception destroys the partial list.
Converters keep their own accepted types and diagnostics. Same-family values move into storage; native integers keep their full width. Constructors still handle batch shapes and default formats, including the unread part of a batch iterator. Ball uses the same driver with a same-family converter.
Element families¶
batch/_families.mojo lists the element families once. That list drives
the buffer variants in _ArrayData, runtime tag dispatch, and the types that
vmap can collect. Construction, views, indexing, mask selection, assignment,
iteration and printing share one implementation across the families.
Arithmetic operators, comparisons, JSON and conversions between families have
separate support requirements. Integer, Rational, Float and Complex provide
these operations; Ball provides container, mapping and lifting support.
Use scalar Ball functions through vmap or lift for ball arithmetic.
A scalar family declares _BatchElement and supplies a zero that needs no
heap storage. The mapper uses that zero for missing Optional results and for
slots whose tuple fields have been moved out. Functions accepting several
numeric families use an argument type implementing _MapAdapter; this keeps
the element conversion with the scalar function's interface.
To add a container family, add it to the family list, implement _BatchElement
and its zero, and register any new context type with vmap. Its function
argument type, if needed, implements _MapAdapter. Compile-time checks keep
the marker trait and family list consistent.
A Mask stores packed bits and a runtime shape. Masks require matching shapes,
never broadcast, and select values in logical row-major order. Views and owned
batches use the same public type, so selections such as x[mask] += 1 and
x[1:3] *= 2 use the same assignment rules.
Masks and numeric batches use the same checked shape product, transpose calculation, and axis-selection plan. A Mask transpose gathers values from the permuted layout. Axis selection uses the shared plan with a direct block gather, avoiding temporary full layouts. Both operations produce fresh packed storage. Mask-specific shape and permutation diagnostics stay at the Mask boundary.
The axis-selection plan accepts a validated shape. Selecting an empty axis fails before computing a size. Otherwise, any empty axis remains in the result; a nonempty result's dimension product cannot exceed the source product. This lets the plan reuse the source's addressability guarantee.
Operator rules¶
Elementwise arithmetic, comparisons and updates call scalar kernels through
batch/_parallel.mojo. Two rank-one batches must have equal lengths. Higher-rank
operations align trailing dimensions, which must be equal or one; broadcasting
uses views over the input storage.
_prepare_binary_shape in batch/value.mojo checks shaped operands once
and selects a shared shape or a broadcast target before entering a family
kernel. It carries
the result shape through arithmetic and comparison paths. Rank-one operators
keep their equal-length rule. lift uses the same trailing-dimension and
stride helpers, but permits singleton broadcasting for vectors too; vmap
requires equal mapped extents.
Operator result types come from _Sum and _Quotient, the aliases used by the
operators themselves. Tests check the types of public arithmetic expressions;
there is no separate promotion model maintained just for the tests.
Mixed operations borrow their operands without first constructing a converted
batch. A shared scalar operand is prepared once and borrowed by each call.
Results follow logical row-major order, even when workers evaluate calls in
parallel. Operators do not fuse: a * b + c creates an intermediate product
batch. Mapping a scalar function that computes the whole expression avoids
that intermediate batch, while still allocating the output and any scalar
arithmetic temporaries.
Mapping functions¶
vmap[f] maps a pure function with a supported signature. f is a compile-time
parameter; Mojo 1.1 cannot pass runtime function values across worker threads.
The mapping reference gives the
supported argument and result types.
in_axeschooses an axis for each argument, orNoneto share it.out_axeschooses where to place the result's new dimension.- Numeric results form batches, Bool results form masks, and batch results stack into a larger batch. The first result supplies any inferred shape.
- Flat tuples map field by field, with independent axis settings.
.vmap()adds another mapping level.- Empty mappings do not call the function on invented inputs. Supply
out_shapewhen a returned batch's shape cannot otherwise be inferred.
Mapped functions borrow stored inputs and move results into output lists. Measure mapping overhead against a scalar loop using the same values; the result depends on the function, layout and batch size. See the benchmark guide.
The signature wrappers share nesting logic and execution runners. Calls with
one, two or three arguments use one runner per arity, whether they return a
single leaf or a tuple. Four-argument calls, such as hyp2f1, have a runner
on the general path only, since each call costs far more than the flat loops
save. A small adapter supplies the unused context for plain
callbacks. Functions taking a tuple use one runner for both result forms.
Typed readers and specialized flat loops still borrow ordinary arguments;
only a callback that takes a tuple needs its argument tuple assembled.
_MappingResults describes result handling at compile time. _LeafResults
collects one leaf; _TupleResults collects each field in order. The runners
share output validation, error handling and traversal, then ask the adapter
to finish the outputs. Both adapters reuse the existing collectors and flat
output builders, including their ownership records and shape rules. There
is no runtime result-type dispatch.
A callback returning an empty tuple still runs at every mapped position and can fail. Empty mappings keep their existing shape-hint requirements. If a later tuple field fails validation, earlier fields are discarded with it; the caller receives a complete result or an error.
Flat and general execution¶
The flat path handles numeric or Bool results, including flat tuples, when the
arguments and axes can be described by one plan. It supports transposed,
strided and reversed inputs, shared arguments, nonzero in_axes, and reordered
out_axes. The planner in batch/mapping.mojo resolves extents and strides at
each mapping level.
A flat run:
- reads a single strided run directly, or advances a cursor through a more general layout;
- passes a row or column argument as a view, using a remaining-axis layout prepared once per chunk;
- expands a flat tuple argument of up to three fields into argument slots;
- moves numeric results into lists and packs Bool results into masks;
- collects supported Optional results as a value batch and a validity mask, filling missing values with the family's zero;
- places output axes using views and sends sufficiently long work to the pool.
The general path runs on the calling thread. It handles batch results, Bool
or Mask arguments, tuple arguments with more than three fields, and mappings
using axis_size or out_shape. It also defines validation diagnostics: the
planner sends unsupported flat plans to this path, and both paths use the same
messages for equivalent failures.
Lifting binary functions¶
lift[f] supplies broadcasting calls, outer, reduce and accumulate for
supported binary functions on every batch element family. It also supports Bool
functions on masks. Calls and outer products use the same _FlatInput readers
as flat vmap runs: matching elements are borrowed, and mixed operands go
through _scalar_argument. Scalar inputs have rank zero. Broadcasting and
outer products supply their own storage positions to those readers.
An identity supplies the result of an empty reduction when no initial
value is given. associative=True permits chunked reductions whose partials
combine in operand order; it does not require commutativity. The caller must
ensure that this regrouping preserves the function's result.
Float and Complex folds use exact working intermediates by default. A reduction
rounds its final result once. An accumulation rounds each output prefix once,
without feeding that rounded output into the next step. Intermediates must be
finite binary fractions: a result such as 1/3 or sqrt(2) raises because
it cannot be represented exactly. Pass
exact=False to round each step with the caller's context instead.
Same-family reductions retain their existing kernels. Mixed reductions and
accumulations share _InputPart, which holds either an untouched source
operand or a callback result. The first source operand stays exact until the
first callback, including when it is the only element of a parallel chunk.
Only a final singleton result or an emitted prefix converts that operand to
the result family. This prevents early rounding of Rational inputs. Float and
Complex adapters select exact working formats for callbacks, while the fold
tracks source formats separately for output rounding.
Ball folds call the scalar function at each step with BallContext. The
resulting enclosure becomes the next accumulator; there is no exact-working
Ball mode. exact affects only Float and Complex results. As with any lifted
function, callers must justify associative=True before allowing regrouping.
Functions with a context parameter receive the appropriate context through an adapter; plain functions are wrapped. Bool results become masks. Mask folds use the same lane machinery with internal zero and one values.
An in-place accumulator such as lift[update], where update calls
acc.__iadd__(x), can reuse Integer storage across steps. It may still need to
allocate when the result outgrows that storage.
Shared elementwise execution¶
Flat vmap calls, lifted calls and outer products, and the Integer, Rational,
Float and Complex arithmetic runners use _execute_values in
batch/_parallel.mojo.
The executor checks the output allocation size, chooses the execution path,
stages results, and handles cleanup and failure replay. Existing indexed
kernels reach it through _map_tensor or _map_values.
Packed comparisons use the same executor with one mask tile per result.
Each frontend supplies two small contracts:
- An
_ElementKernelowns one chunk's readers. It starts at a logical index and computes successive elements from a borrowed job. Flat mapping keeps its cursors and lent Float or Complex argument records here; lifted calls supply broadcasting or outer-product positions. An indexed arithmetic kernel needs no reader state. - An
_ElementErrorspolicy formats the failing index and the error raised by the scalar call. This preserves nestedvmapcoordinates, lifted coordinates, and each arithmetic family's messages.
Reader state belongs to a chunk, never to the shared job. The job outlives every chunk and any replay. The pool receives an immutable pointer whose origin keeps that borrow tied to the calling scope. Jobs need not be copyable, and dispatch does not clone their plans or retained inputs. A failed parallel run releases all written values before the executor constructs a fresh kernel at the lowest failing index and repeats that call. Both reader setup and scalar failures pass through the frontend's error policy. A successful run transfers the complete output and its ownership record to the caller.
Error policies borrow their shape metadata on the calling thread and build coordinates only after a failure. They do not add a shape copy to successful calls or travel to worker threads.
Each adapter owns shape checks, exact input conversion, output formats, and scalar arithmetic. Complex format checks and the speculative stored-pair path keep their error precedence. The executor leaves promotion and rounding to these family rules.
This contract covers runs with one stored result per call, including flat tuple
results. General vmap result collectors still validate dynamic result shapes
on the calling thread. Reductions and accumulations keep their lane kernels:
their ordering and rounding rules differ from independent elementwise calls.
They share the worker pool and storage machinery below the executor.
Parallel execution¶
common/_threads.mojo and batch/_parallel.mojo provide a persistent worker
pool whose default size follows usable physical cores, or a positive
APN_MOJO_NUM_THREADS setting. set_num_threads changes the participation
count, including the caller; growing the pool restarts the workers between
jobs. See thread controls. Workers take chunks
through atomic counters and write results into separate slots. If several
calls fail, the executor reports the lowest failing logical index.
The failure counter has its own shared allocation, retained until all workers finish. With an inline atomic field in the mutable slot record, a Mojo 1.1 run could lose a worker's recorded failure: a lane wrote 997 of 1,000 outputs, recorded its error, and the caller later read the no-error sentinel. Keeping the atomic outside that record prevents an incomplete result from being published and lets cleanup destroy only the outputs that were written.
Most runners consider the pool only for at least 64 elements; flat vmap
keeps its initial probe for shorter runs too. For eligible runs, the calling
thread times the first 16 elements. It sends chunks to workers when the
remaining work is estimated to take at least 100 µs. If the pool is busy,
including during a nested call, the calling thread continues the work.
Elementwise operations, reductions, flat mappings and unlimited batch JSON writes can use the pool. Unlimited JSON reads parallelize numeric conversion for Integer, Rational and sufficiently wide Float or Complex values. The general mapping path and conversions with limits stay on the calling thread.
Owner-affine memory release¶
Freeing an allocation on a different CPU can incur cache and allocator costs. Parallel runs record which worker owns each chunk. When the last reference to a large result (at least 1,024 elements) is released, its chunks return to their owning workers for destruction. Multi-CPU measurements are included in the full report linked from Benchmark results.
Sharing and lifetime¶
Views retain the underlying element storage and a layout descriptor. Updating a parent batch leaves existing views intact. Reversed, strided and transposed views expose logical indices, which also determine serialization order and error indices. Mask selections gather the selected handles into a new contiguous list.
These ownership rules protect saved values during ordinary updates. They do not make concurrent mutation of the same batch variable safe.
Element readers use _borrow_values in batch/_values.mojo to obtain an
immutable Span. Storage, strided views, mapping readers and reduction lanes
share this access path. The span does not copy elements or acquire another
storage reference. Its origin is the reader that already retains the storage;
even the pointer used by a linear lifted fold is read-only and carries that
origin.
Mojo 1.1 cannot infer this relationship through the retained owner and list, so the helper contains an explicit origin cast. Each caller must keep the list alive and leave its storage and borrowed elements unchanged until the borrow ends. Mapping readers may advance their cursor or replace a converted argument only after the callback returns. Their return types expose an immutable borrow even though preparing the next argument mutates the reader.
A Float input can hold a Float run, a Complex component run, or a broadcast
scalar. Its element method borrows each representation through one private
helper that ties the reference to the input. This helper may only receive
values retained by that input; it handles the compiler's origin restriction
on optional payloads and component fields. Format inspection uses the same
borrow. The value method makes an owning copy when one is needed.
Float dot products use _borrow_values for native runs. Their pointers are
immutable and cannot outlive the input. The starting position includes the
selection offset, and the step combines the tensor and selection strides,
including negative strides. Callers must establish a nonempty native run
before requesting its starting pointer.
The separate mechanism that lends Float and Complex significands to argument records still clears those records before destruction.
Updates and transactional safety¶
An update checks shapes and indices, reads the original inputs, and builds
a private result. Overlapping assignments such as x += x[::-1] therefore
use the original values throughout. Scalar kernels check domains and normalize
results as they run. The wrapper replaces the destination only after every
step succeeds, preserving its shape and default formats.
A recoverable error discards the private result. The destination, earlier copies and views keep their previous values. General arithmetic has no allocation quota; conversion limits apply only to conversion calls.
Reductions¶
Built-in reductions and lift share batch/_lanes.mojo. It traverses selected
axes and preserves dimensions when keepdims=True. The
reductions chapter explains exact accumulation, ordered
rounding and the arithmetic adapters.
Resource boundaries¶
Live elements, spare capacity, layouts and retained views all consume memory.
ConversionLimits bounds newly requested storage within one conversion; it
does not cap retained batches, later arithmetic or execution time.