Files
Hanzo AI ecaca10cdb canonical Go entry: backend selector + batch GPU paths via lux/accel
luxfi/crypto becomes the single Go entry point for ALL Lux-family crypto.
Every public function in this module now dispatches between three
implementations through a runtime-selectable backend:

  - vanilla: pure-Go reference (always available)
  - cgo:     native binding (blst, libsecp256k1, ckzg) where present
  - gpu:     batch acceleration via github.com/luxfi/accel

The dispatcher reads LUX_CRYPTO_BACKEND (auto|vanilla|cgo|gpu); auto
picks the most capable backend the binary was compiled and linked with.

New canonical packages:
  backend/             runtime backend selector (env + programmatic)
  internal/gpuhost/    accel session lifecycle, single per-process
  keccak/              Keccak-256 with batch GPU dispatch
  sha256/              SHA-256 with batch GPU dispatch
  sha3/                SHA3 / SHAKE family
  ripemd160/           RIPEMD-160 (Bitcoin/Lux address derivation)
  ed25519/             Ed25519 with batch GPU verify
  bn254/               canonical alias for bn256 (matches FIPS naming)
  modexp/              canonical alias for bigmodexp
  evm256/              EIP-196/197 precompile ABI wrappers
  poseidon/            Poseidon2 hash via gnark-crypto
  pedersen/            Pedersen commitments over BN254
  ntt/                 Number-Theoretic Transform reference
  polymul/             negacyclic polynomial multiplication

Extended existing packages with batch GPU paths:
  bls/batch.go         BatchVerify routes through accel.BLSVerifyBatch
  mldsa/batch.go       BatchVerify (ML-DSA-65) via accel.DilithiumVerifyBatch
  mlkem/batch.go       BatchEncapsulate / BatchDecapsulate via Kyber kernels
  secp256k1/batch.go   BatchVerifySignature via accel.ECDSAVerifyBatch

GPU dispatch is gated on (a) backend.Default(), (b) batch size threshold,
and (c) accel.Available(). When any gate fails the call falls through to
the vanilla CPU path; output is byte-identical.

The legacy gpu/ stub is replaced with a thin probe surface (Available,
Backend, Devices, Version) that delegates to the same gpuhost session.

Tests show vanilla and gpu backends produce identical outputs across all
batch entry points (-race clean).

See AUDIT.md for the per-algorithm state matrix and honest gaps.
2025-12-27 19:30:33 -08:00

79 lines
2.0 KiB
Go

// Copyright (C) 2020-2026, Lux Industries Inc. All rights reserved.
// See the file LICENSE for licensing terms.
package mlkem
// BatchThreshold is the minimum batch length at which BatchEncapsulate /
// BatchDecapsulate will try to dispatch through github.com/luxfi/accel.
var BatchThreshold = 64
// BatchEncapsulate runs Encapsulate for a slice of public keys, returning the
// resulting (ciphertext, sharedSecret) pairs. All keys must be the same mode.
//
// When the batch is large enough and a GPU backend is available the
// computation runs on the GPU; otherwise it falls back to per-key
// Encapsulate calls. The output is byte-identical for the deterministic
// path; the streaming path uses fresh randomness either way.
func BatchEncapsulate(pubs []*PublicKey) (cts [][]byte, sss [][]byte, err error) {
n := len(pubs)
if n == 0 {
return nil, nil, nil
}
mode := pubs[0].mode
for i := 1; i < n; i++ {
if pubs[i].mode != mode {
return nil, nil, ErrInvalidKeySize
}
}
cts = make([][]byte, n)
sss = make([][]byte, n)
if n >= BatchThreshold && mode == MLKEM768 {
if ok, gerr := batchEncapsulateGPU(pubs, cts, sss); ok && gerr == nil {
return cts, sss, nil
}
}
for i, pk := range pubs {
ct, ss, e := pk.Encapsulate()
if e != nil {
return nil, nil, e
}
cts[i] = ct
sss[i] = ss
}
return cts, sss, nil
}
// BatchDecapsulate runs Decapsulate over a slice of (sk, ct) pairs.
func BatchDecapsulate(sks []*PrivateKey, cts [][]byte) (sss [][]byte, err error) {
n := len(sks)
if n != len(cts) {
return nil, ErrInvalidKeySize
}
if n == 0 {
return nil, nil
}
mode := sks[0].mode
for i := 1; i < n; i++ {
if sks[i].mode != mode {
return nil, ErrInvalidKeySize
}
}
sss = make([][]byte, n)
if n >= BatchThreshold && mode == MLKEM768 {
if ok, gerr := batchDecapsulateGPU(sks, cts, sss); ok && gerr == nil {
return sss, nil
}
}
for i, sk := range sks {
ss, e := sk.Decapsulate(cts[i])
if e != nil {
return nil, e
}
sss[i] = ss
}
return sss, nil
}