Known limitations¶
numpy counting is an explicit non-goal¶
The counting model prices scalar code as a compiled port would execute it; array operations are bulk vectorized routines outside that model, and no part of numpy's semantics is adopted. Concretely:
np.float64scalars work and count correctly on either side of an operator — not because numpy is supported, but becausenp.float64is a plain C double subclassingfloat, so it flows through the ordinary float path.- mixing
CountedFloatwith numpy arrays, or with numpy scalar dtypes that do not subclassfloat(np.float32,np.int64, ...), raisesTypeError.CountedFloatrefuses numpy's ufunc protocol (__array_ufunc__ = None) on purpose: the alternative is an operation that silently returns an uncounted result whose type may later recover toCountedFloat, hiding that flops went missing — the loud boundary is the honest one.
Keep counted algorithms in scalar float/CountedFloat code; hand values to
numpy only after converting to plain float (e.g. float(x)), outside the
counted region.
Constant folding keys on the operand's value, not on it being a literal¶
The cost model presents a plain float operand as a
compile-time constant — something the imaginary
compiled program knows while being compiled. The implementation cannot see the
source, so it applies that rule to the operand's runtime value: any plain
float is treated as a constant, whether it was written as a literal or
arrived from somewhere else.
The two readings coincide for a value that really is fixed. They come apart when one call site meets different plain floats over its lifetime:
That single division counts 2 MUL + 2 DIV: the powers of two fold to exact reciprocal multiplication, the others do not. No compiled port produces that mix — a real program has one instruction there, chosen once at compile time. The model's central premise is what quietly breaks, not merely a weight.
The discipline that avoids it: anything that varies while the algorithm runs
should be a CountedFloat, so it is priced as dynamic input rather than
folded on whatever value it happens to hold. Reserve plain floats for values
that are genuinely fixed.
This is named rather than fixed: detecting it would mean inspecting the caller's syntax tree to see whether the operand was a literal, which is out of proportion to a modelling assumption that correct usage already avoids. It is observable, though — a counting context asked to report what it counts logs each fold with the reason it was applied, so a site being folded inconsistently shows up in that output.
Loops the library cannot count as loops¶
math.prod and math.fsum count n−1 operations for n elements, whatever the
elements hold (math.prod counts n when start is not 1). The built-in sum,
min and max, and any loop written by hand, cannot be counted that way: the
library cannot patch them, so their counts come from the individual operator
calls. Two things then cost counts that a loop would charge — the ordinary
constant folds apply to each arithmetic operation, and an operation between two
plain operands counts nothing at all.
Counting through the operator calls goes wrong in two directions:
- one addition too many in
sum(data)over counted values:sumseeds with an integer0, and adding zero is not a no-op for signed zero (-0.0 + 0.0is+0.0), so+ 0is counted exactly as a strict compiler would emit it. An explicit0.0seed changes nothing, for the same reason. - too few operations wherever an operation meets no counted operand:
sum([1.0, 2.0, cf])counts one addition where a loop runs two, the plain prefix costing nothing. Forminandmaxa plain winner also turns the running result plain again, so the shortfall is value-dependent:min([1.0, 2.0, CountedFloat(5.0), 3.0])counts one comparison of three, while the same call withCountedFloat(0.5)— which wins — counts two.
Two fixes, each with its own reach:
math.fsum(data)ormath.prod(data)— n−1 operations for n elements, every element counted whatever its value. Counts nothing at all when the sequence holds noCountedFloat, and there is nomin/maxequivalent.- seed a counted accumulator —
sum(data[1:], start=CountedFloat(data[0])), and the same shape for a+or*loop written by hand: contagion carries the accumulator, so every later operation has a counted operand and counts, even over otherwise-plain data. Elements that fold against that accumulator are still dropped, by the identity folds of the cost model as everywhere else — a-0.0added or a1.0multiplied costs nothing.
Neither fix reaches min and max: they take no start, and seeding a loop
does not help, because a plain winner turns the accumulator plain again and the
remaining comparisons stop counting.
Other limitations¶
- the uncounted built-in operations are cataloged per surface: the
mathmodule's not-instrumented set (frexp,ldexp,modf,nextafter,ulp— the port emits no floating-point instruction for them; see themathcoverage table) andfloat's own uncounted members (hex(),as_integer_ratio(), formatting, truthiness — see the float surface). A counting context can be asked to report the contagion-relevant ones as it meets them, so a count that is quietly missing them says so — see watching what gets counted - operator-level fused multiply-add is not modeled:
a*b + ccounts as separate MUL + ADD, matching the contraction-off reference semantics the cost model pins. Real builds routinely contract: on aarch64, FP contraction is the compiler default at plain-O2(no fast-math flag involved), and x86 FMA3 targets fuse under the same defaults — so on such builds, flop counts over-estimate fusable multiply-add sequences (dot products, Horner evaluation) by up to the fused sequences' MUL share. Python cannot observe operator-level fusion, so the library cannot count it.math.fma(Python 3.13+) is the one place a fusion is observable, and it is counted as a singleFMA; expressing a multiply-add that way is therefore the only means of having one counted as fused — and there is none on older interpreters, wheremath.fmadoes not exist - mixed operations with non-float numeric types are outside the counting
model:
CountedFloatdelegates to the other operand exactly likefloatdoes, so e.g.CountedFloat(x) * Fraction(1, 2)yields a correct but plain — and uncounted —floatresult (downstream counting stops). The reverse order counts normally:Fractionhands the operation back, and theCountedFloatperforms it. Adecimal.Decimaloperand raisesTypeErrorjust as with plainfloat, in either order. Comparisons with aFractionregister 1 ABS + 2 COMP even though the comparison is delegated:Fraction's own float handling guards withmath.isnanandmath.isinf, and the classifiers count for a counted operand whoever calls them — the same honest-count stance as the dict-membership note below. Counting properly in the presence ofFractionorDecimalvalues is a non-goal (a compiled port has no such types to price); numerical algorithms should usefloat/CountedFloatvalues throughout, rather than relying on which side an operand happens to sit. - counting state is per OS thread (created lazily, freed with the thread):
a
FlopCountingContextmeasures only the thread that opened it, is confined to that thread while open (cross-thread use raisesRuntimeError), andPauseFlopCountingpauses the calling thread only. To measure a multi-threaded computation, open one context per worker and sum the results:
def worker(job):
with FlopCountingContext() as ctx:
run(job)
return ctx.flop_counts()
with ThreadPoolExecutor() as ex:
total = sum(ex.map(worker, jobs), FlopCounts())
All asyncio tasks running on one thread share that thread's counter — there
is no per-task isolation. Note that while any thread has a context open,
all threads see the patched math.* functions (patching is inherently
process-wide); plain-float calls still take the fast path, and counts always
land in the thread performing the operation. Free-threaded builds (3.14t) are
supported and covered by CI; counting there needs no lock, since each thread
mutates only its own state. Note that the benchmarking extra requires numba
0.65 or newer on a free-threaded build — earlier versions ship no
free-threaded wheels
- builtin min/max return the winning operand object, so with mixed
counted/plain arguments the result is a plain float whenever a plain constant
wins — countedness ends silently, and value-dependently: the same call site
keeps or drops it depending on the data. Each comparison with a counted
operand counts COMP; those between two plain operands do not, which is what
makes an n-argument call cost fewer than n−1 — see
Loops the library cannot count as loops.
No dunder exists
through which CountedFloat could intercept the returned object. When the
result feeds counted computation, re-wrap it — CountedFloat(min(...))
counts nothing (a float-source construction) and is correct whether or not
the wrap was needed
- dict/set membership of CountedFloat keys inflates COMP: hash-bucket
equality checks count as comparisons. This is consistent with the model
(those comparisons really execute) but can surprise when a dict is used as
bookkeeping rather than algorithm — pause counting or use plain-float keys
for bookkeeping structures
- while a FlopCountingContext is open, the patched math.fsum, math.prod,
math.sumprod and math.dist materialize their iterable inputs (they call
list(...) on them) even when no CountedFloat is involved, so the argument
can be inspected after the value is computed. The stdlib versions consume some
of these lazily, so passing a very large one-shot iterator to one of these
functions inside a context holds the whole sequence in memory — O(n) space
where the unpatched call would stream. The computed value is unchanged; only
peak memory differs, and only while a context is active. (math.hypot takes
its coordinates as separate positional arguments, already materialized as a
tuple by the interpreter, so it does not diverge.)
- flop weights should be taken with a grain of salt and should only provide
relative ballpark estimates w.r.t. computational complexity. Production
implementations in a compiled language could have vastly differing
performance depending on cpu cache sizes, branch prediction misses, compiler
optimizations using vector operations (AVX etc...), etc...
See also:
- The counting model — the contract defining what is (and isn't) counted.
- Math patching semantics — which
mathfunctions are instrumented, and the third-party-patching contract. - FLOP types reference — per-operation lists of what is counted and what is not.