# aadc — AADC for Python (agent guide)

Adjoint algorithmic differentiation. Record a computation once, replay it many
times with exact derivatives — vectorised over AVX lanes, optionally JIT-
compiled.

This file ships INSIDE the wheel. `aadc.llms_txt()` returns it, so it always
describes the version actually imported. Check `aadc.__version__` (wheel) and
`aadc.__engine_version__` (C++ engine) — they are different facts and neither
implies the other.

## Install

    pip install aadc

Linux x86-64/aarch64, macOS arm64, Windows x86-64; CPython 3.10–3.14.
Unlicensed use runs in Community Edition (non-commercial) and prints one line
to stderr at import. `aadc.license_status()` reports the state; a licence is
supplied out-of-band via `$AADC_NG_LICENSE`, `./aadc-ng.lic` or
`~/.aadc-ng/license`. `aadc.license(key)` is a 1.x API that no longer exists
and raises NotImplementedError saying so.

## The core loop

    import aadc

    f = aadc.Functions()
    f.start_recording()
    x  = aadc.idouble(2.0)
    ax = x.mark_as_input()
    y  = x * x + 3.0
    ay = y.mark_as_output()
    f.stop_recording()

    ws = f.create_workspace(kind="scalar")   # 1 lane: floats in, floats out
    ws.set_val(ax, 5.0)
    ws.forward()                  # ws.val(ay) -> 28.0
    ws.set_diff(ay, 1.0)
    ws.reverse()
    ws.diff(ax)                   # -> 10.0   (dy/dx = 2x at x=5)

Four rules that explain most confusion:

1. `mark_as_input()` / `mark_as_output()` return HANDLES (`Argument`,
   `Result`). Replay is addressed through those, never through the Python
   variable `x` — after `stop_recording()`, `x` is a stale value.
2. The recorded graph is fixed. Changing an input and replaying is free;
   changing the *structure* means recording again.
3. Reverse mode needs a seed: `ws.set_diff(<output>, 1.0)` before
   `ws.reverse()`, then read `ws.diff(<input>)`. Seeding nothing gives zeros —
   that is the commonest "my gradients are zero" cause. Adjoints do NOT
   accumulate across independently seeded passes on this engine: a fresh
   `set_diff()` wipes the stale plane (measured). `reset_diff()` is for
   clearing state deliberately.
4. `Kernel` is an alias of `Functions`. Same class, two names.

`record_kernel` is the same thing with less ceremony:

    with aadc.record_kernel() as kernel:
        x  = aadc.idouble(2.0)
        ax = x.mark_as_input()
        ay = (x * x + 3.0).mark_as_output()

## The trap: branching on an active value

`if x > 0:` needs a plain `bool`, so the comparison COLLAPSES — the direction
it took while recording is frozen into the tape. Replay with a different input
still follows the recorded branch, producing a plausible wrong number and a
wrong derivative.

    y = aadc.iif(x > strike, x - strike, aadc.idouble(0.0))   # correct
    if x > strike: ...                                        # frozen

This is REPORTED, not silent:

    aadc.recording_passive_warnings()   # during/after recording
    kernel.passive_warnings()           # per kernel
    kernel.num_passive_warnings()

Treat a non-zero count as a finding to explain, not noise. Related helpers:
`aadc.where_arr` (array form), `aadc.masked_assign`, `aadc.smart_assign`,
`aadc.branching_function` + `aadc.aadc_early_return` for whole functions,
`aadc.lower_bound` for searching a sorted knot array (it returns the
INSERTION index -- count of knots < x -- so the containing interval is
one less).

## iint and ibool: replay-varying integers and predicates

`aadc.iint` / `aadc.ibool` are tape types with their own banks. They can be
inputs and outputs; there is NO adjoint with respect to them (the derivative
w.r.t. an index does not exist -- it flows into the element the index
selects). `evaluate()` feeds an iint input as int64 (a float is refused) and
an ibool as bool, and returns iint / ibool outputs as int64 / bool arrays;
`Argument.kind()` / `Result.kind()` say which. `aadc.array(table)[i]` with an
iint `i` is a RECORDED lookup (engine arrayOps::get); an out-of-range index
CLAMPS to [0, len-1]. `np.searchsorted(knots, x)` on an active x records
`lower_bound` and returns an active iint. Since aadc-ng #896 (v2.16.0) the
engine's lower_bound accepts every element layout (const, Random, Diff,
mixed) with ONE precondition: the array must be sorted ascending AT REPLAY.
The Python binding still emulates ACTIVE knots with an O(N) recorded
comparison-sum (issue #140); constant knots take the native O(log N) op.

The trap: `int(np.searchsorted(knots, float(x)))` runs in host arithmetic,
is frozen at record time, and a varying input changes nothing on replay --
counted as a passive conversion. Keep the lookup on the tape.

`aadc.idouble(i)` on an active iint RECORDS the int -> double conversion
(engine >= 2.20.0), and mixed arithmetic promotes the same way: `i * 2.5`,
`x * i`, `i < 2.5` return an active idouble / ibool that replays with the fed
int. There is no adjoint w.r.t. the int (exactly 0); the adjoint w.r.t. the
double operands is the usual one. `float(i)` / `int(i)` are still counted
passive extractions. On an ACTIVE iint these raise TypeError instead of
freezing: `//`, `%`, `divmod()` (no integer-division op on the tape,
aadc-ng #958), `aadc.math.max/min` / `np.maximum` / `np.minimum` / `np.fmax` /
`np.fmin` / `np.clip` with a float or idouble operand, and
`aadc.iif(c, i, 2.5)` -- write `aadc.idouble(i)` to say you want the double
surface. `math.floor(i)` / `ceil` / `trunc` / `round(i)` are the IDENTITY on
an integer and return the ACTIVE iint unchanged.

SCALAR ONLY. With an `ArrayNG`/`AADCArray` operand the same spellings
PROMOTE rather than refuse: `np.maximum(a, i)`, `np.minimum(i, a)`,
`np.clip(a, i, 5.0)` record a double max/min through the same iint -> double
op and replay exactly. The refusal is about Python's `max` returning the
winning OPERAND OBJECT -- `max(iint, 2.5)` is an int for one recorded value
and a float for another, a result TYPE no tape can carry. An array operand
fixes the result to a double array, so that ambiguity does not exist.
`aadc.iif` / `aadc.where_arr` still refuse an active iint branch in every
shape. Details, examples and history: wiki page "Python iint and ibool".

## Arrays

`aadc.array(...)` gives an `AADCArray`: a NumPy-compatible array of active
values. NumPy ufuncs work on it; `mark_as_input()` / `mark_as_output()` apply
elementwise and return arrays of handles.

    import numpy as np
    a  = aadc.array(np.array([1.0, 2.0, 3.0]))
    aa = a.mark_as_input()
    s  = np.sum(a * a)

Use `aadc.math` (not `math`) for scalar transcendentals on active values;
`np.exp` and friends already dispatch correctly on `idouble` and `AADCArray`,
and so does `np.where` on active arrays.

`isnan()`, `isfinite()` and `isinf()` all RECORD (scalar and array): they
return an `ibool` / mask that re-evaluates per replay, and cost no passive
conversion. Collapsing one with `if` does — use `aadc.iif(x.isnan(), a, b)`.
`isinf()` needs aadc-ng >= 2.15.0; older engines had no IsInf opcode and
returned a frozen, *uncounted* bool.

Rounding RECORDS from engine 2.21.0 on (aadc-ng issue #965, ops
CAADNG_G_Floort / _Ceilt / _Trunct). `x.floor()`, `x.ceil()`, `x.trunc()`,
`aadc.math.floor(x)`, `np.floor(x)` and `np.floor(array)` return an ACTIVE
`idouble`; `math.floor(x)` / `ceil` / `trunc` are the INTEGER spelling and
return an ACTIVE `iint` (bindings issue #157); `x // y`, `x % y` and
`divmod(x, y)` are built on the recorded floor. None of them costs a passive
warning, and the adjoint through all of them is 0 -- the correct derivative
of a step function almost everywhere, and what the engine returns AT an
integer too, where the derivative does not exist.

On engine 2.20.0 and older every one of those EXTRACTED the value: frozen at
the recording point, counted but still wrong on replay. `x % y` was the worst
of them -- an ACTIVE result carrying a frozen quotient.

Still off the tape: `x.to_int()` (a plain Python `int`, the deliberate
uncounted opt-out), `int(x)`/`float(x)`, and `round(x, ndigits)` -- there is
no recorded decimal rounding, so it is a COUNTED extraction.

`round(x)` (no ndigits) RECORDS (engine op CAADNG_iIntRoundFromDouble)
and returns an ACTIVE `iint` that re-evaluates per replay, with a zero adjoint
-- the correct derivative of round almost everywhere. Two deltas from CPython:
ties go HALF AWAY FROM ZERO (std::llround), not half-to-even, and the result
is an `iint` rather than an `int`. A PASSIVE idouble keeps CPython's exact
behaviour. `round(x, ndigits)` has no recorded form and is counted.

## Speed: compile, then batch

    kernel.compile("scalar")                 # JIT for the 1-lane workspace
    ws = kernel.create_workspace(kind="scalar")  # the kind those kernels run on
    ...                                      # set_val / forward / reverse as above
    kernel.last_forward_used_compiled()      # True iff a compiled segment ran

compile() and the workspace must agree on the workspace KIND, or the replay
interprets (same numbers, slower than not compiling). The defaults always
agree: compile() with no argument is "scalar" -- free on every platform, the
same script runs compiled on x86-64, Apple Silicon and Linux AArch64 -- and
create_workspace() / evaluate() with no workspace argument follow whatever
compile() installed. For the host's vector lanes, compile("auto") ('jit' on
x86-64, 'arm64-jit' on AArch64; licensed on Linux AArch64): a following
create_workspace() is then multi-lane, and set_val()/val() take and return
one value PER LANE. Avoid 'jit-scalar' in portable code: it is x86-64 only
and installs nothing on Apple Silicon. Pick the backend at the FIRST compile()
of a recording: compile() then compile("auto") keeps the scalar kernels (a
compiled segment is never recompiled) and reports no error.

`aadc.evaluate(...)` replays a kernel over many input rows, across threads,
returning values and derivatives together. For batches call
compile("auto") first: evaluate() then packs scenarios into the vector lanes
(measured 3-4x faster than after plain compile(), whose scalar kernels
evaluate() replays one scenario at a time; on transcendental-heavy tapes the
scalar JIT can even trail the uncompiled vector interpreter). This is where AADC pays
for itself: per-row cost after recording is a fraction of re-running the
original Python. `aadc.ThreadPool` sizes the workers; one `Workspace` per
thread — a Workspace is NOT shareable.

## Source tracing: find the code that computed a number

Record with `trace=True` and the tape carries file/line/function for every
operation:

    f.start_recording(trace=True)
    ...
    f.stop_recording()
    print(f.trace_report(outputs={ay: "npv"}))
    f.trace_toc()                 # per-function index, [begin, end) op ranges
    f.trace_range(begin, end)     # operations in a range

For an agent this is the difference between guessing which code ran and
knowing. Attribution requires the code to have been COMPILED with the
instrumenting toolchain (`aadcpp -faadc-debug`): Python-level arithmetic and
the engine are always attributed; third-party C++ (e.g. QuantLib) only if you
installed an instrumented build. A report with 0% attribution looks exactly
like a healthy one — check the counters.

## Verifying a derivative (do this before trusting one)

    ws.set_val(ax, 5.0); ws.forward(); base = ws.val(ay)
    h = 1e-6
    ws.set_val(ax, 5.0 + h); ws.forward(); up = ws.val(ay)
    fd = (up - base) / h                  # ~10.0, compare with ws.diff(ax)

If AAD and finite differences disagree by more than FD noise, suspect a frozen
branch first (check the passive-warning count), then a missing
`mark_as_diff()`, then a genuine bug — in that order.

## Error messages you will actually hit

* `ImportError: AADC_BACKEND=...` — the 1.x engine was removed. Unset the
  variable, or `pip install "aadc<2"` for the old line.
* Zero gradients everywhere — no `set_diff` seed, or the value was never
  connected to an input (it was a plain float, not an `idouble`), or it passed
  through something passive: `float()`, `.val()`, `.to_int()`,
  `round(x, ndigits)`. Rounding itself is NOT on that list from engine 2.21.0
  on -- `x.floor()`, `aadc.math.floor(x)`, `np.floor(x)`, `math.floor(x)` and
  `//` / `%` all record (aadc-ng #965, bindings #157) -- but their adjoint is
  0 by construction, which is correct, not a leak.
* Two copies of `libaadc-ng` in one process — values can be right and
  gradients silently wrong. Never vendor the library into a second wheel; all
  AADC packages must resolve to one image.
* `dir(aadc)` is the declared API (`aadc.__all__`). Names outside it exist but
  are not interface.

## Map of the API

    idouble ibool iint          active scalars
    Functions / Kernel          recording + replay driver
    Workspace                   one replay's values and adjoints (per thread)
    Argument Result             input/output handles
    array fromiter ArrayNG      NumPy-compatible active arrays
    record_kernel               record a block
    record                      record a function -- NOTE its Jacobian is
                                FINITE-DIFFERENCE, not adjoint
    iif where_arr               branchless selection
    masked_assign(target, value) branch-mask assign; returns the new value
    evaluate                    batched replay (evaluate_sums raises
                                NotImplementedError on this engine)
    optimize least_squares      solvers over recorded kernels
    root_scalar                 scalar root find on the tape
    aadc_assert                 precondition recorded on the tape
    trace_report trace_toc      source attribution (record with trace=True)

Full documentation: https://github.com/matlogica — see the `aadc-docs`
Python book. Contact: info@matlogica.com
