ed25519: documents RFC 8032 byte-equal Metal verify kernel (100 vectors)
+ N_threshold = 256 on M1 Max with 26.7x speedup at N=4096 vs equivalent-
shape CPU port. dGPU CUDA port pending in lux/crypto/ed25519/gpu/cuda/.
mldsa: documents structural-skeleton M1 kernel + dGPU residual. Per-thread
serial work on Apple Silicon (~122 us SHAKE256-dominated) vs ~42 us NEON
narrows expected M1 wall-clock speedup to ~7x at N=4096; CUDA H100/Ada
closes that gap (~33x ceiling).
mlkem: same shape as mldsa, ~75 us per-thread Metal vs ~26 us NEON.
SHA3/SHAKE256 chains are the M1 ceiling; dGPU port closes it.
luxfi/crypto becomes the single Go entry point for ALL Lux-family crypto.
Every public function in this module now dispatches between three
implementations through a runtime-selectable backend:
- vanilla: pure-Go reference (always available)
- cgo: native binding (blst, libsecp256k1, ckzg) where present
- gpu: batch acceleration via github.com/luxfi/accel
The dispatcher reads LUX_CRYPTO_BACKEND (auto|vanilla|cgo|gpu); auto
picks the most capable backend the binary was compiled and linked with.
New canonical packages:
backend/ runtime backend selector (env + programmatic)
internal/gpuhost/ accel session lifecycle, single per-process
keccak/ Keccak-256 with batch GPU dispatch
sha256/ SHA-256 with batch GPU dispatch
sha3/ SHA3 / SHAKE family
ripemd160/ RIPEMD-160 (Bitcoin/Lux address derivation)
ed25519/ Ed25519 with batch GPU verify
bn254/ canonical alias for bn256 (matches FIPS naming)
modexp/ canonical alias for bigmodexp
evm256/ EIP-196/197 precompile ABI wrappers
poseidon/ Poseidon2 hash via gnark-crypto
pedersen/ Pedersen commitments over BN254
ntt/ Number-Theoretic Transform reference
polymul/ negacyclic polynomial multiplication
Extended existing packages with batch GPU paths:
bls/batch.go BatchVerify routes through accel.BLSVerifyBatch
mldsa/batch.go BatchVerify (ML-DSA-65) via accel.DilithiumVerifyBatch
mlkem/batch.go BatchEncapsulate / BatchDecapsulate via Kyber kernels
secp256k1/batch.go BatchVerifySignature via accel.ECDSAVerifyBatch
GPU dispatch is gated on (a) backend.Default(), (b) batch size threshold,
and (c) accel.Available(). When any gate fails the call falls through to
the vanilla CPU path; output is byte-identical.
The legacy gpu/ stub is replaced with a thin probe surface (Available,
Backend, Devices, Version) that delegates to the same gpuhost session.
Tests show vanilla and gpu backends produce identical outputs across all
batch entry points (-race clean).
See AUDIT.md for the per-algorithm state matrix and honest gaps.